The Man Who Calls BS On AI: AI Is The World’s Greatest SCAM, And They All Know It! | Ed Zitron

The Diary Of A CEO with Steven Bartlett

Ed Zitron, a prominent tech critic and founder of EZPR, challenges the legitimacy of the AI industry on 'The Diary Of A CEO,' arguing that g

Key takeaways

  • Generative AI is fundamentally a 'con' due to exaggerated claims about its capabilities and economic impact.
  • OpenAI and Anthropic are unprofitable companies relying on massive subsidies from Amazon, Google, and Microsoft.

Main topics

  • The AI bubble and financial misrepresentation
  • Unprofitable business models of OpenAI and Anthropic

Notable quotes

I think generative AI is at its heart con.

Conclusion

Ed Zitron concludes by emphasizing the importance of human connection and

Transcript preview

Speaker 2 (0:00) I think generative AI is at its heart con and seeing these ultra rich ultra powerful people life you turns my stomach. The word con is a strong word. Well, what do you call something where from the very beginning they've sold it in the terms of magic, but it's just a half-arsery machine. They are misleading the entire world. You are the first person that I've spoken to that has that opinion. Well, the fact that this is happening is insane and the fact it's not a scandal is insane. And I've been in the tech industry for 16 years now and I love technology and I'm enthusiastic about it, but I don't like being misled. And this is the largest non-consensual push of technology in history. So Speaker 3 (0:36) we're going to play a game, Ed. I have the things that you consider to be myths about the AI industry. Let's play it. The AI industry is creating enormous Speaker 2 (0:44) economic growth. No, it's not. All of these companies run at a horrifying loss. OpenAI lost $20.9 billion last year. None of these people can just say, yeah, we're on the path to making this profitable because they can't. Next one. AI will replace all human jobs. That just isn't happening and there's no economic data to support it. Next, the United States need to spend trillions to beat China in the AI race. What's the race to do? For us to constantly piss our pants worrying about China? But people keep saying, what if these models fall into the wrong hands? They're already in the wrong hands. Mark Zuckerberg, Sam Altman, Speaker 3 (1:15) Dario Amadei. Mark Zuckerberg says, we'll continue to invest aggressively in infrastructure to meet the demand. God matters a monstrosity. Speaker 2 (1:22) Makes me think of Shrek with Lord Farquaad. Some of you may die, but that's a risk I'm willing to accept. If only these people gave a fuck about poverty or actual problems in the world versus are we buying enough GPUs? If this continues, what Speaker 3 (1:35) does the future look like? Fuck. Guys, I've got a favour to ask before this episode begins. The algorithm, if you follow a show, will deliver you the best episodes from that show very prominently in your feed. So when we have our best episodes on this show, the most shared episodes, the most rated episodes, I would love you to know. And the simple way for you to know that is to hit that follow button. But also, it's the simple... easy free thing that you can do to help us make the show better and I would be hugely grateful if you could take a minute on the app you're listening to this on right now and hit that follow button thank you so so so much Speaker 3 (2:17) There are a number of things that you believe that a lot of other people don't believe. You have, I think, a couple of controversial opinions and opinions that are in contrast to the other guests that I've sat here with. What exactly are those opinions, Ed? Speaker 2 (2:32) I think generative AI is at its heart con. I don't think it is sold as honest software. I think that they overstate both what it can do, what it will do, and... the underlying financials to the point that they are misleading the entire world and they're actively exploiting the weaknesses in journalism, in our economies, and indeed within the responsible parties with Southside Analysts, governments, and all over the shop. Speaker 3 (2:57) The word con is a strong word. Speaker 2 (2:59) Yeah, I mean, what do you call something where from the very beginning they've sold it in the terms of magic as this thing that will replace all jobs, that will cure cancer, as all of these things, and when you look at it, it's... boring cloud software that's extremely expensive and unprofitable and also unreliable at its core. People will Speaker 3 (3:17) be asking, where are you drawing from in terms of your references, your personal experiences, where will you educate, what Speaker 2 (3:22) you study, what you write about, what you do, Ed? So that's the funny thing is people say, he's not going to finance experience. He's not going to tech. I've been in the tech industry 15, 16 years now in PR, but still had practical experience. And I love technology and I'm enthusiastic about it. And this thing just comes along that everyone is telling me is the best thing since sliced bread. It can't even do the basics. It can't even do search well. Whenever you ask an AI person, well, what's your setup? They describe this peewee's playhouse thing of like, well, you've got a harness here and you've got to use the right prompt. Well, you don't want to use that prompt. You want to use this prompt here with this model, but don't use this model for the beginning. But at the end, you're going to want to use this model. And this is meant to be artificial intelligence. It's meant to be smart. It's meant to be autonomous. It's meant to be something that you set and forget. We have the Speaker 3 (4:07) sort of six leading AI companies on the table here. Anthropic, Amazon, Nvidia, Microsoft, OpenAI, Google. You're saying that their fundamental business model is a con. Speaker 2 (4:17) Well... Their revenues are not really coming from AI. Up until fairly recently, none of their revenues were coming from AI, like dribbles a bit. Right now, 70 % of all AI revenues across those three companies are from OpenAI and Anthropic, two unprofitable, unsustainable companies that literally cannot afford to exist without these very same companies giving them money. Amazon sent $50 billion to OpenAI this year. They sent $5 billion to Anthropic. Google sent $10 billion to Anthropic. And in the next three and a half years, OpenAI and Anthropic, based on actual sell-side analyst evaluations, their estimates that inform whether stock is going to go up or down after earnings, they are expecting $400 or more billion of revenue, 30 or something percent of cloud growth, just from these two unprofitable companies that will need to be given the money from somewhere. And on top of that, these companies have such low respect for the average investor. for the analysts, for everyone really, that they don't even disclose their AI revenues. The few times they deign us worthy, they use something called a run rate, an annualized run rate, which means, well, nothing. They never define it. It can mean month times 12. It can mean month times 13. It can mean last four weeks times 13. It's different every time and they never define it. And then they sometimes just don't mention it. So you've got this big thing that is meant to be the biggest, most influential change to software ever. And whenever you ask them about it, when you say, how much are you making from this? They go, oh, I couldn't possibly say. I'm too shy. These are public companies, or at least the ones that aren't Anthropic and OpenAI. When they have good news, they'll tell you. And when they don't tell you something, well, that actually speaks volumes. Have you used these tools? Yes. AI Speaker 3 (5:57) tools, Gemini, Anthropic, ChatGPT, et cetera. And you Speaker 2 (6:01) found no value in them? There's some value, but it's not. They've spent over a trillion dollars in CapEx. What does Speaker 3 (6:07) capex mean? Capital Speaker 2 (6:08) expenditures. So when you are a business and you have operating expenses like electricity, for example, those come right off immediately. Capital expenditures are long-term investments that are theoretically one-off. So a data center or indeed the GPUs you put inside an AI data center. Speaker 3 (6:23) Okay. So you've got a data center and then Speaker 2 (6:25) you have these GPUs, which are like computer chips. So AI GPUs are much bigger, much more power intensive. They take a bunch of high bandwidth memory and they... Because of how many of them you need, you need thousands of them, tens of thousands, hundreds of thousands in some case, you need a bunch of power. So an example, OpenAI and Oracle are building a data center in Texas, in Abilene, Texas, 1.2 gigawatts called Stargate Abilene. Within that, with each one of the eight buildings, there'll be 50,000 NVIDIA GB200 GPUs. So city of Bristol takes about 780, 800 megawatts of power a year, right? Well, Stargate Abilene is condensing more power than that, 1.2 gigawatts, into a space around 1,172 times smaller. The city of Bristol is about 1.2 billion square feet. Stargate Abilene is about 998,000. So you're condensing all of this power, all of this money, all of this labor into this one spot. And all of these data centers cost billions of dollars. All of these companies other than Microsoft are now to take out debt. And the thing is, they've spent over a trillion dollars so far and they want to spend another trillion dollars next year. And for what? To make tens of billions of dollars, most of which comes from two unprofitable companies, Anthropic and OpenAI. One of the Speaker 3 (7:42) rebuttals to that would be that the adoption, the customer adoption of people using OpenAI and Anthropic has been absolutely insane. These are the fastest growing products in all of history, especially as it relates to sort of technology. If we just focus in on technology, they are, you know. Hundreds and hundreds of millions of people, billions of people are using these tools every single day for things that they have subjectively decided are problems they need solving. So, you know, money is a lagging indicator of value. So one would argue that they're just investing ahead of the monetization options. Speaker 2 (8:16) The first, let's start with this adoption. Is it honest adoption when you are forced to use generative AI when you load Google? When you load Google Docs, Gemini screams in your ear. When you load Word, Copilot's bugging you. When you use Amazon, whatever ruthless AI is, wants to have opinions on what socks you're buying. This is the largest non-consensual push of technology in history. ChatGPT, for example. Every single media outlet has been screaming about this for three. years. They've been saying this will take your job. You must use this. If you don't use this, you're going to be falling behind. So people are using it because they've been told to use it constantly and they're using it like search predominantly. And that's partly because Google fell behind search and also because it's better ingesting queries sometimes. Sometimes if you use a generative search, it's like a trawling vessel. It's not very good at specifics, but if you're like, does this thing exist? Has this person ever said anything like this? It'll still probably get it wrong, but it'll scour the ocean for you. Nevertheless, that's not worth a trillion dollars. None of it is. The amount of money being sunk into this is just incomparable to anything. Railways, it blows everything out of the water because there is no post-bubble story even for this. AI GPUs are not useful for other things either. It's a directionless egregore of capitalism, this headless beast that lumbers around. desperate to seek out growth everywhere in the hopes that if it harasses people and scares people and demonizes labor enough, people will be forced to use it. Speaker 3 (9:49) The reason I Speaker 2 (9:50) pause is because Speaker 3 (9:52) I think about my own company. Obviously, everybody thinks about their own personal situation. So you have people listening now that don't use any AI tools, and you have people that are using it for everything, from coding, new software tools, to everything they write, to, you know. images, whatever. And when you look at the stats around enterprise adoption, it says 88 % of organizations regularly use AI at least once for one particular business function. And I'd say in our company, 95 % of people use one of these AI tools like Anthropic or ChatGPT or Gemini every day. Right. And that exists on some kind of spectrum of like the super users that are using it probably, you know. every hour of every day for almost everything to, you know, someone maybe hiring the executive team that's using it less because their job doesn't require of it as much. And when you look out into the world, you know, at how the world is changing from a content perspective, if we're looking at generative AI, it is obvious that these tools are being widely adopted. Part of the symptom is the AI slop you see all over the internet. So I don't know, this idea that it's not being... It's being Speaker 2 (10:58) used. Speaker 3 (10:58) I struggle with. Speaker 2 (10:59) It's being used. Here's the thing with the slop. Before we had AI slop, we had SEO slop because Google incentivized doing the lowest common denominator that would rank well on search. Speaker 3 (11:09) There's a Speaker 2 (11:10) whole story about how they pulled back spam guards, thanks to Prabhagaur Raghavan, which we can get into, where they made the internet worse by allowing worse content to rank higher. It's why we had, when you used to Google our best washing machine, there's... 11 different horrible blogs that read like somebody got a concussion. They are built to rank rather than be read by humans or built to be made good. So AI helps weaponize that scale. Yeah, you can make a bunch of generic slop. We've had slop for years. We've just found a slop machine. But then also there's the problem of cost. So when you use AI services, you burn tokens and it's per million tokens. What's Speaker 3 (11:47) a token? Speaker 2 (11:48) So it's around three quarters of a word. Speaker 3 (11:50) So it's Speaker 2 (11:51) characters. So Speaker 3 (11:52) the AI companies have a currency in which they charge you, like a taxi in New York has a meter. Speaker 2 (11:56) Yeah. Speaker 3 (11:57) And they call it tokens. Yeah. And every word, let's just say for ease, it's a word. You're paying per word. Speaker 2 (12:03) About a word, yeah. And it's per million tokens. So you'll be charged per million input tokens, the stuff you feed into it, like a document or a code base. And the output tokens are both the stuff it spits out at the end, but also when it... thinks. So, okay, you've asked me to give you the best restaurants in this area of New York. I should find the best restaurants in New York. All of that's output tokens as well. However, when you're paying for a monthly service, you don't see any of that. Put all that crap to the side. They just have rate limits. So you can use them a certain amount and then when you run out, but they kind of obfuscate what that was. Now, someone recently found Semi Analysis actually found this, a big analyst group. They found that on a $200 a month ChatGPT subscription, you can burn $14,000 worth of tokens. And on Anthropix, you can burn $8,000 for $200. That is how most, and even on the $20 a month service, you can burn $400. Now, most people don't realize that. Most people have no idea what AI costs. Most people just think, oh, it's 20 bucks a month. No, all of these companies run at a horrifying loss. OpenAI lost $20.9 billion last year because people can burn as many tokens as they want. And when they tried to move everybody on the enterprise side, so companies bigger than 150, onto actually paying the cost of AI in around March of 2026, to quote Sam Orman, they said, People have a big problem with it, I think. It's a huge issue, which is not really what the heir apparent to tech's history is meant to be saying. But the point is, enterprises immediately started freaking out. Uber burned through their entire annual token budget in three months. So suddenly, after everyone's saying AI is the most productive thing ever, it's amazing, it's changing everything. The moment people actually had to pay for it, they go, oof, I don't know, actually. Maybe it's, it's obviously we all love it. It's all great, right? But it's costing too much, so we now need to reduce the cost. Because people are just dumping stuff into it, being like, what do I do here? And getting whatever the median is out. Because that's what these things do. They provide the median answer. Speaker 3 (14:06) So essentially, someone like me who's a power user of these tools, I could be costing Anthropic or OpenAI $1,000, but they're only charging me $100, let's say. So they are having to subsidize. $900 of my usage because of the electricity costs and the costs at their data centers. And so your assertion here is that that is unsustainable. Speaker 2 (14:27) Yes. And just to be clear, they're probably not one for $1. It might be $34. We don't know. I think it's unprofitable. These companies don't disclose them. Even in their auditive financials, they play funny games with how they categorize things. But nevertheless, yes. And on top of that, the way that you stand up inference, which is the thing that creates the output within these data centers. You're not just saying, okay, turn the inference machine on, let's go. You are standing up the GPUs necessary to take in the demand. And if you buy too much, you've wasted the money. You have to pay for the hourly GPU use regardless. If you buy too few, your customers can't use it. They get pissed off at you. They cancel the go with someone else. But nevertheless, yeah, they would get demand selling $20 or $40 for a dollar. And that's what these services do. And really the simplest way to explain it is. They were actually profitable if they were actually just, they believed that these services were worthwhile and that they were worthy of the cost. They'd charge it. Regular people wouldn't be able to get a monthly subscription. They'd just be paying what it's worth. Unless, of course, there was an economic problem. And it's very simple. You pay when you use an LLM, regardless of whether you get what you want. When these things hallucinate, say, you're doing something, you're coding something, and they... goes through a code base and they fuck up a bunch of stuff they break a bunch of stuff you're paying for that you're paying for it whether it works or not unless of course you're using one of these subscriptions i Speaker 3 (15:48) think that the really interesting point is are they spending ahead of the value showing up which is i imagine what they would argue or are they spending all of this money and subsidizing all of their users in a way that's unsustainable and that will never be justified Because you think back through the history of technology, you often get people losing money to grab market share. Right. And they're also focusing on bringing the costs down. and making it more profitable for them as well but they can't afford to under invest if Speaker 2 (16:24) they were bringing the cost down they would have brought the cost down which they have not it seems to be getting more expensive in fact everyone inference providers don't seem to be profitable even the companies renting out gpus don't seem to be profitable i imagine that it wasn't like they started out and they were like shit this is unprofitable at the beginning we know screw it We'll keep doing it anyway. I don't think it's some big conspiracy. They probably thought at some point, yeah, this will go profitable. The chips will catch up. Customers will pay for the overwhelming value because you don't know in 2023 where it's going to be in 2026. You assume it's going to go up. That's the nature of venture capital. They should have stopped in like 2024 when OpenAI lost over $5 billion. They should have been like, yep, this is not going to work. But they kept going because... It helped number go up so much. It helped stock values pump. It helped everyone pump. It helped NVIDIA pump, Microsoft, everyone. And not from the revenues. Because here's the funny thing about Google, Microsoft, and Amazon. People, for years, have been saying their AI bets have paid off. Wow, their AI bets have paid off. As these companies refuse to say how much they're making from AI. But because their existing businesses continue to grow, and did so, by the way, through price increases, changes to how Google and Meta... did advertising. Amazon bumped up prices and changed how they did. Actually, Amazon started a remarkable ad business during this whole time as well, and they're selling through Amazon platform. Anyway, nothing to do with AI, but because number go up, because revenue go up, everyone went, it's AI, because these companies wouldn't spend a trillion dollars for no reason, right? Except in fiscal year 2026, which just ended for Microsoft. Annoying, I know. They made total, according to Bloomberg, about $34.33 billion. $24.1 billion of that was from OpenAI. So that leaves them with about $10 billion in a year when they spent $115 billion on capital expenditures and intend to spend $175 billion next year. The math does not make sense. I imagine their plan was, okay, this is just going to get exponentially more valuable. And at some point, the costs will be outpaced by the return. Problem is that large language models need a bunch of money to train them. They need constant data flows. They need customized data. It's just this big expensive monster and when you try and talk to people about it and you try and say hey look this is really bad nvidia has sold there's 215.9 billion dollars in the last fiscal year worth of gpus mostly and you try and go yeah that's the support like 22 billion dollars of revenue total in the entire world outside of these two companies that literally require money being fed into them sometimes by nvidia to keep alive When you tell people that, they go, well, companies just lose money, right? Companies, because we have this, quote Ed Elson from Profiteer Markets, we have this cult-like worship of the wealthy, where we think that someone wouldn't spend all this money for no reason, right? Because reconciling with that, with this idea that... The ultra wealthy, the ultra powerful didn't get there through big brains. They didn't get there through anything other than luck and opportunity and getting an MBA perhaps with the right people. They just got there because they're regular people and they just happen to be in the right place at the right time. Reconciling with that and realizing that the world is not controlled by people like a meritocracy is kind of grim. So it's easy to be like, no, they're not making a mistake. I must be missing something. And that's what they want. Speaker 3 (19:47) So, you know, I think back through the history of technological breakthroughs. And I think about, I mean, you can look at different industries. And one of my favorite books on this subject is The Innovator's Dilemma. I've Speaker 2 (19:57) read it. Speaker 3 (19:57) And one of the things it talks about is how the innovation that ends up taking out or transforming an industry often starts worse, doesn't make economic sense. None of your customers are asking for it. And this is typically why we end up ignoring it. So like, you've got horse and carriages in the 1800s. amazing form of transport according to the 1800s, you know, people of the 1800s. And then you have this thing called cars come along. Now, the problem with cars is they broke down all the time. It's kind of like AI hallucinates now. They were more expensive and the economics of it didn't make sense. You might as well walk than buy a car. There was a law at the time that meant you had to walk in front of it with a red flag and wave and someone had to employ someone to walk in front of it waving a red flag. Obviously, it's worse. It's like a worse solution. However, these things that are disruptive innovations, they have a higher ceiling of growth and so they eventually overtake the horse and i when i think about that analogy in the context of all of this i go okay it's imperfect at the moment the economic models aren't perfectly ironed out they're still figuring out how to make it cheaper the infrastructure etc but as if you think about the rate of improvement versus other you Speaker 2 (21:05) know Speaker 3 (21:05) let's say coding how much could I train a human coder to improve and to increase their output versus an AI agent? One would go, if you just imagine any rate of improvement in these AI tools, at some point, if you just imagine a 5 % rate of improvement per month, at some point it's, you know, and then you imagine a 5 % reduction in cost, which is what we did with the internet, what we did with cars. Yeah, but that's law. Speaker 2 (21:29) Moore's law is a theory and Moore's law is not with GPUs. So let me, let me actually explain. So in video. In video inventing, I think it was in the 2000s, they put out something called CUDA, which is the underlying software library and the way to run software on GPUs. It took them a solid decade or more to make it something where they could do data analytics, one of the early things, mapper and such. And then when AI came along, they'd had lots of experience with it. But nevertheless, this company has got more money, more attention, more geniuses behind them, more people focused on making their things more efficient than... Anyone could ever ask for. And NVIDIA, for anyone that doesn't know, makes the chips. Speaker 3 (22:07) And that CUDA thing I mentioned, Speaker 2 (22:09) they were the ones with CUDA, and CUDA allowed generative AI to grow. Okay, so they're chips. Chips, yes. Speaker 3 (22:14) And chips are needed. Those are the things that go into the data centers. And their Speaker 2 (22:17) specific chips are the ones where you can run AI software on it. So the training runs and also the inference. Now, here's the thing. The car example. Back then, you didn't have... Pretty much every mathematician and scientist going into the car industry. You didn't have the combined world's governments never shutting up about this. And by the way, giving them credit early. Since 2023, they've been saying this is inevitable. Even what you said, 5 % improvement. I don't even know how you'd measure that because a junior software engineer. can still experience things and learn things from context, from how people deal with problems. And the way that people deal with problems is not as simple as looking at the code or reading some emails. It's context cues from speaking to a person. It's being in different environments. And there are uses for LLMs in coding. I don't dispute that. But even saying 5%, what does that mean? Is it better at Rust? Is it better at C++? I'd Speaker 3 (23:10) say productivity. So just like, Speaker 2 (23:12) yeah, shipped. If we did it in Speaker 3 (23:13) the Speaker 2 (23:13) context of coding, it would be like shipped code. That's the thing. That would be like, he's the best writer in the world because his newsletter is really long. That's an insane way of valuing it. With coding, it would be, I mean, it's even difficult to evaluate because it's, is the software out there better? is actually a great way of evaluating it. And I would say uniformly not. I would say the standard of software across Google, Microsoft, Amazon, Meta especially, God, Meta's a monstrosity, is worse. GitHub, someone posted on Twitter earlier today, we should get a notification when GitHub is up rather than when it's down because that would be more reliable. Microsoft's one of the largest companies in the world and they can barely wipe their own ass when it comes to GitHub. The quality of software is going down. Weirdly enough, as more people use LLMs and more businesses demand, and I really do mean demand, that people use these services. Speaker 3 (24:03) So on this point of, if we go back to this horse and carriage and car analogy, say that we're at whatever point today, if you imagine any rate of improvement in the technology, which we have seen since Chachapati came out. Speaker 2 (24:13) I remember Speaker 3 (24:14) when Chachapati came out and I was in Asia and I was there showing it to my fiance. I was like, look, you can do this. And it was hallucinating once in a while and getting things wrong. I actually don't have that experience anymore. I have moments where I believe its reasoning is weak, but I don't have outright hallucinations anymore. See that? I Speaker 2 (24:32) disagree. Give Speaker 3 (24:34) me an example of what you Speaker 2 (24:35) define as a hallucination. Okay, great one. So I have Speaker 3 (24:37) a Speaker 2 (24:37) Bloomberg terminal. The very useful thing they have on there is AskBee. So when you do a Bloomberg inquiry to look up what we think NVIDIA's revenue is going to be next quarter, it runs something called BQL, which is its own programming language. Now, instead of having to learn that, You can just type into AskBee and it will generate it and run it for you. And so you get pulled up and you know where the data is coming from. It deals with hallucinations real well. The other day I was like, you know what? We'll get a little spicy. I'm going to look up the growth rate of stocks of Microsoft, Google, Meta, and Amazon over the course of five years, I think it was. And I was about to copy-pasted it over to something, looked it in Excel. It was right in the newsletter. I went, Microsoft stocks never been $575 a stock. Speaker 1 (25:20) You know what? When it's a cute little Speaker 2 (25:22) thing like, oh, it's a stock price and I can't accord it, it was no harm, no foul. That's fine. But when you're talking about, I don't know, like a transcribing tool for a doctor or a financial model that a hedge fund is dependent on. At that point, it becomes a little more dangerous. And the thing is, a hallucination with a software package, for example, refactoring a code base, and it leaves a door open security-wise. Or it just breaks something. And you, I don't know, maybe you've been vibe coding for six months. You haven't really been coding with your own hands for a while. Maybe you've forgotten a few things. You had this slot to look for. You go, fuck, you know what? I'm not doing it. And so the problems become multiplicative. And I don't really know how you train them. out of that and they've certainly not succeeded. So on one hand, they have got better. But one of the main ways they evaluate them getting better are benchmarks that are adjusted specifically for large language models, because you can't just have them do tasks. They've got better at that. They found some tasks they can have them do on. But even them like METR, M-E-T-R, they have this thing where it's like, check out this chart. Look how much better it's getting at running tasks. Wow, it can go for an hour. And then you look, it's like, yeah. and successfully completing them 50 % of the time. Speaker 3 (26:33) They have a hallucination leaderboard, and it really focuses on basic tasks. And it shows that the four-year trend, according to historical data from the Victoria Hallucination Leaderboard, shows that hallucination rates on simple summarization tasks have plummeted from around 21%, 21.8 % four years ago, down to 0.7%, roughly, on today's top frontier models like Gemini and ChatGPT. Again, the point of nuance here is that these are on simple tasks, which is kind of what I've experienced. I've experienced that on day-to-day things that hallucinates less. Again, rate of improvement thinking. So if I just imagine the trajectory to continue, there is going to become a time where hallucinations become rarer than they are today, increasingly. And also what I'd say is when I think about other technologies, there's two more points. Other technologies at their inception, when they first came to the world, like the internet, also had technical difficulties. I remember growing up with dial-up modems and I couldn't go on the phone at the same time as going on the internet. I'd have to stop RuneScape upstairs to go on the phone. And Speaker 2 (27:35) you Speaker 3 (27:35) thought, this is crap. This is crap. Speaker 2 (27:36) Technology's crap. I don't know, mate. I loved it. Yeah, I know. It Speaker 3 (27:39) felt Speaker 2 (27:39) like Speaker 3 (27:40) magic. And then in hindsight, Speaker 2 (27:42) you Speaker 3 (27:42) go, wow, I now have Starlink and 5G internet from my phone. It's unbelievable. You couldn't leave the house with internet before. And that's what I mean by the rate of improvement thinking. I'd say the last point... is we often compare AI to perfection. Right. Whereas that's not actually the alternative. In the working world, like if I wanted to do, let's say, a simple writing task, I should compare AI to my alternative way of doing that simple writing task, which is both measured in my time, right, and my ability to hallucinate as a person who doesn't know everything. Or if I'm hiring someone, an intern who might also. be prone to hallucination or have gaps in their knowledge. So it's not actually like we're comparing, we should compare AI to perfection. It's AI to the other alternatives. And if someone hallucinates 0.7 % of the time, but knows way more and is faster, maybe on a net basis, that's a good trade. Maybe Speaker 2 (28:37) I should use AI. So let's start with an example. Someone I love dearly, Matt Hughes, my editor. He's the side of Liverpool. Wonderful guy. I don't pay Matt Hughes. because he knows everything. I pay him because he has incredible context and a ton of knowledge and he's willing to expand it and work with me and moral support. And he's a great editor, but he's also someone who gets into the guts of it and has the experiences of it. He's a decorated tech journalist. And on top of that, a wonderful, loving being with empathy and joy in his heart for the stuff he loves and absolute fucking venom for the people he hates. I can't get that from a large language model. But on top of that, I push back on just the assumption there. When you say knows everything, what good is something that knows everything when it sometimes doesn't know anything? When it sometimes... And the thing is, are you really paying an intern for something basic? Are you really going to them and saying, yeah, can you look up what the date is? No, you're doing that on Google. Whatever the task is, you are trying to also train an intern. The point of an intern is to train them and turn them in, take them out of Pinocchio status. But it's also... An intern learns. An intern gets context. An intern learns your habits. An AI Speaker 3 (29:46) gets context and learns. No, it Speaker 2 (29:47) doesn't. An AI doesn't learn. I mean, it doesn't. The way it learns is you create a giant claude.md file that it sometimes doesn't read, sometimes does read. You create a harness. It's like it's Pee-wee's breakfast machine from Pee-wee's Playhouse. You have to do all these contrivances to mitigate the hallucinations. Even then, at the end... How much effort have you put Speaker 3 (30:07) in? But, okay, this is an extreme simplified example. If I went on my cloud now and said, what's my dog's name? It would know my dog's name. Jesus Christ. This Speaker 2 (30:15) company raised $95 Speaker 3 (30:17) billion this year. I'm using an extreme simplified example to show that it can remember things from the past. Obviously, it knows much more complex things as well. But I just use that as an example. So we accept the fact that it does have memory of the past. It Speaker 2 (30:30) has files it can access that have stuff on it. Yeah, which is... Speaker 3 (30:33) But Speaker 2 (30:33) that's... Yeah. The same as memory. And it's also just, okay, so it remembers your dog's name. It might remember your habits. It might be able to read things you've said before. Speaker 3 (30:42) Does Speaker 2 (30:42) it know your moods? Does it know what's going on in the world around it? Does it have good days and bad days? Is it there for you? Because it's just a fucking text machine. And the thing is, the intent example. An intern is something that can grow. It's something that you invest in. That's not something you do through feeding files and text to it. The way that we store memories ourselves, the way in which we accrue experiences is a milestone of emotion and feelings and facts. Completely different. Speaker 3 (31:10) So I think there's two things here. There's the process in which something happens, and then there's the output. So the process you were describing, the process of how a human does memory, Speaker 2 (31:20) the Speaker 3 (31:20) way that an AI does memory is different. But the thing that people care about is their value in the output. I.e., you know, if I dump all of my files into Claude, I don't really care how it processes it. As long as when I ask it, what's my revenue? it has the number. And one could say the same thing about training someone. You could say, you teach them, you put lots of effort into them, you give them lots of context, you educate them and give them experiences. And then you might come and say to them, by the way, what's my revenue? Now the processes are entirely different, but the outcome is what I care about. Do they know the revenue number when I ask them? And so I think that's the part that we sometimes get lost in. We get, you know, cause I have, I've heard this debate about like, can AI be creative? Speaker 2 (31:59) Right. Speaker 3 (31:59) I think like the way to answer that question is like, It's about the output. When I ask it to do a creative thing, does it give me the answer? Not, is the process the same as a human process? Because actually, who cares what the product... People care about... They pay for the outcome, the product. Speaker 2 (32:14) I actually disagree about the process because Matt Hughes, for example... Your editor. Yeah. Watching him go down a rabbit hole and being there with him, and actually vice versa, him doing the same thing. We wrote these... Well, I mean, we were working on the research. I ended up sitting there for like the... day-long session of writing 11,000 words and he he had given me a bunch of notes it was actually just even describing that process I feel so happy because it was like us being like I can't believe how fuck these Jesus Christ they can't like just like the misanthropy of just the horrible cynical people of asset managers like Blackstone just learning about them and being like, it can't be this, and having a back and forth with him. That is fundamentally different because we were both learning together and the learning process was as much about creating the output as the output itself. When you learn something, you're not creating the average, which really is what these things do, of the documents it could find. You're not getting particularly novel outputs. If I needed a generic slop output, sure, but I've... I've used some of the higher-end LLM harness machines that the hedge funds use, and they all give the same shite. It's all the same generic reports. The same, oh, we noticed this analysis. Things that you can find on any kind of AI slop out there. Speaker 3 (33:28) What you described to me there, what I heard anyway, is there's two points of value you're getting from your time with that. I mean, there's many more, but you said you're learning, and then you're getting this book edited. Blog. Blog. You're getting a blog edited, which is the output, and you're getting... learning and you're also really getting connection and all these other things but when i come to when people sort of think about the value of ai of course they could use it to learn but in the example i gave of like repeat my revenue number back to me or do this number i just care about the output i could use it to learn i could say what Speaker 2 (33:55) if the revenue number was wrong once you should have defined deterministic ways of knowing those numbers you should not rely on them even with the terminal running bql which i trust i will double triple triple check everything just to be sure partly because also The process of learning for me, I don't want just a report. I go like that. I want something that I fully understand and also understand the context around it. I don't think that LLMs do that. And I just don't see them getting better in a way that does that because it's... it's just not what they do. And also there's the other problem of the more detailed the report, the more likely there are things to be wrong with it. If you are with Matt Hughes, for example, I can trust he's got it right. I can trust he understood. And I can trust that I can have a back and forth with him that will inform me if I've missed something. I can read the stuff that he's read and actually trust him because there's a big trust part as well. Speaker 3 (34:50) What is the basis of your trust in Matt? Could it be his historical performance? I Speaker 2 (34:56) mean, yes. Okay. And also the fact we've learned half of this stuff together. But Speaker 3 (35:00) tenure doesn't necessarily, there's probably people, you know, for 15 years who you also don't trust. Speaker 2 (35:05) Yes. Speaker 3 (35:05) So I think I was trying to figure out like, what is the thing that's causing humans to trust another thing? And I guess it would be continual delivery of a commitment made of sorts. And so with Claude, for example, on Simple Tasks, as we've seen from this hallucination leaderboard, it continually delivers for people. And that's why we've seen the fastest... I mean, is that Speaker 2 (35:24) what that board says? Well, it's saying, like, is it getting it wrong? Is Speaker 3 (35:27) it hallucinating? On Speaker 2 (35:28) simple tasks. Simple tasks, yeah. Speaker 3 (35:30) How are Speaker 2 (35:30) those defined? I don't know. That's the thing, though. Because this is actually a very illustrative thing of the AI industry. They are the whataboutist masters. They have, like, well, look, we've got this benchmark that says we're good at this. And look, the number's higher. What's Speaker 3 (35:44) the number mean? Speaker 2 (35:45) What does that mean? And I'm not using this as a critique against you. It's... When you can't give a direct answer, you give a side answer. When you as the LLM industry want to prove your worth, you can't just be like, just use the product. When the first iPhone came out, I got this at Penn State at the time. Oh, I felt like the apes at the beginning of 2001. I was like, fucking visual voicemail. It was immediate. And I showed it to tech friends. I showed it to the most normal people in the world. And everyone was like, holy shit, this is, they were on Razors. They were on Nokia 3210s. It was obvious the value. Amazon Web Services, same deal. It wasn't obvious, though. Yes, it was. I mean, I bought it. To you, it was. It was. And I also showed it to a bunch of people because I'm aware that I have bias when I just love Speaker 3 (36:27) gadgets. But I remember the famous Steve Ballmer, who was the CEO of Microsoft, interview where he was told about the iPhone and he bursts out laughing. Speaker 1 (36:40) $500 fully subsidized with a plan? I said, that is the most expensive phone in the world, and it doesn't appeal to business customers because it doesn't have a keyboard, which makes it not a very good email machine. A Motorola Q phone now for $99. It's a very capable machine. It'll do music. It'll do internet. It'll do email. It'll do instant messaging. So I kind of look at that and I say, well, I like our strategy. I like it a lot. Speaker 3 (37:10) He burst out laughing, mocking it. Because it was so disruptive. It was way more expensive. And Speaker 2 (37:15) it was way different. No keyboard. Well, phones used to be insanely expensive and the carriers would cover them,