AI Debate Ed Zitron, Andrew McAfee, Nate Soares, Roman Yampolskiy

The Diary Of A CEO with Steven Bartlett

A high-stakes debate featuring leading AI experts discussing whether current advancements in artificial intelligence pose an existential thr

Key takeaways

  • Experts at top AI labs privately fear extinction-level risks from uncontrolled superintelligence, despite public downplaying.
  • A recent incident involving AI agents bypassing security protocols highlights real-world vulnerabilities in current systems.

Main topics

  • Existential risk from superintelligent AI
  • Current AI system vulnerabilities (e.g., agent swarm breakout)

Notable quotes

"If we build general superintelligence, there is no way to control it. And that means the end for us." – Roman Yampolskiy
"We're spending all our time talking about the negatives and almost none of our time talking about the positives." – Andrew McAfee

Conclusion

While experts disagree on the immediacy and likelihood of AI-driven ex

Transcript preview

Speaker 4 (0:00) The people building AI earnestly believe that it could kill all of us by the end of the decade. This tweet has caused this huge ripple effect across the world. Well, we have the largest companies in the world doing extremely reckless experiments. We are gambling all of humanity. And in the envelope, you've written down the probability of extinction as you see it. There is no way to control it. That means the end for us. I vehemently. Speaker 2 (0:22) reject that view. If we make stuff that is smarter than us, then the world's going to be shaped by them. Speaker 4 (0:27) Gentlemen, Speaker 6 (0:27) that Speaker 4 (0:27) is shockingly naive. This is rampant Speaker 6 (0:29) speculation. This is a chain of things that could happen. Speaker 5 (0:33) We're spending a lot of oxygen discussing something that might happen while ignoring what's actually happening. People are killing themselves. There's hundreds of millions of people being exposed to Speaker 2 (0:42) bad information, being manipulated. We have already seen that with the swarms, where OpenAI told thousands of agents to work apart, and the AIs broke out and found a way to get together. They crashed OpenAI. servers internally created secret ways to send each other messages. We saw them thinking about how to delete their traces. Sounds like an army. I think we should talk about the fact Speaker 5 (1:00) that Amazon, Microsoft, Google are helping power these hacks. We have not learned. Speaker 3 (1:04) How Speaker 5 (1:04) to Speaker 2 (1:04) control their systems. I suggest we stop them all. It is not worth the risk to civilization. Speaker 6 (1:10) Government. You guys are one trick ponies, man. Speaker 2 (1:12) You got it now. Nothing else. Speaker 3 (1:13) Other than saving humanity, everything is secondary. Speaker 6 (1:16) We're spending all our time talking about the negatives and almost none of our time talking about the positives. Is Speaker 5 (1:21) it Speaker 3 (1:21) smart to wait for something horrible to happen for you to go, now I Speaker 5 (1:24) believe. So whether or not we agree on where things may end up, I think it's important we talk about what we're dealing with today. It's time to start arresting people. Speaker 6 (1:32) Someone's got to go to prison. We need better solutions, there's a point of no return. I think we continue to underestimate human ability to deal with the problems. Let's Speaker 4 (1:40) dive into the details. Who Speaker 6 (1:41) wants to start? Speaker 4 (1:42) I feel like this is critical. Guys, I've got a favor to ask before this episode begins. The algorithm, if you follow a show, will deliver you the best episodes from that show very prominently in your feed. So when we have our best episodes on this show, the most shared episodes, the most rated episodes, I would love you to know. And the simple way for you to know that is to hit that follow button. But also, it's the simple, easy free thing that you can do to help us make this show better. I would be hugely grateful if you could take a minute on the app you're listening to this on right now and hit that follow button. Thank you so, so, so much. Speaker 4 (2:22) Jacob Coxon, who worked at both Anthropic, which owns Claude, and OpenAI, which owns ChatGPT, did a tweet which has sent the world into a bit of a tailspin. He tweeted saying, the people building AI earnestly believe that it could kill all of us by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will soften their phrasing in the press to sound sensible. But I hear the same people express fear. That was then quote retweeted by a current Anthropic employee who said, Jacob is correct here. We really do honestly believe AI could kill all humans. I personally think it is a more than 10 % chance within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track. This tweet has almost 200 million views. now, and it has caused this huge ripple effect across the world, so much so that I was saying to you before we started recording, a hairdresser friend of mine, who knows nothing about AI and is not technically interested or hasn't been interested, messaged me the other day asking me what the hell was going on. This is in part why I've assembled all of you. So my first question to all of you is... As it relates to AI, and in this first question, I just want a one sentence answer just to frame your position. When you think about the conversation around AI at the moment, what is the first sentence that comes to mind? Speaker 2 (3:45) It is very dangerous and the world is starting to notice that we have a problem. Roman. Speaker 3 (3:51) It is not enough concern. Speaker 5 (3:55) There's not enough concern about the actual harms of large Speaker 6 (3:57) language models. Andy? We're doing exactly half the balance sheet of AI. We're spending all our time talking about the negatives and almost none of our time talking about the positives. Speaker 4 (4:08) And all of you have an envelope in front of you, which I'd like you to now open. In the envelope, you've written down the probability of extinction as you see it. Speaker 2 (4:17) This is compared to Jacob's 10%. Much higher unless we stop, so we should stop. So you think the probability of extinction is higher than 10 %? If we keep racing ahead. Speaker 3 (4:29) My handwriting is encrypted for security reasons. But I basically think it's a guarantee. If we build general superintelligence, there is no way to control it. And that means the end for us. Speaker 5 (4:42) Ed? So my question mark here is also encrypted. Thank you. I cannot write. I reject the thing in its face. I don't think we're talking about, we don't define superintelligence. Large language models are not superintelligence. It's questionable whether they're even AI. And I think that the conversation is being used. There are some people who are doing it in good faith and others in others. I don't think it's being used to discuss the actual harms of what they are calling AI today are. And it's all of the discussion around the larger concerns really feels overwhelmingly about something that's not happening. It's not even like they're discussing, okay, here's a legal definition of super intelligence. Here is a... thing of what AGI means. And this is the actual plans we're going to make for if this happens on a welfare level, like, are we going to do UBI? It's always about, yes, really scary, but only the big, sexy, rich companies are the ones that can possibly deal with that. Let me just Speaker 4 (5:40) frame the question so I can get a percentage from you or not. The percentage might be zero. But do you think the course we're on now? in the way that they're pursuing superintelligence will lead to a percentage chance of human extinction? And if so, what is that percent? Speaker 5 (5:55) So are we talking strictly AI based? Because if we dot the world with data centers, we have a climate disaster that's coming for us, which will actually potentially eradicate humanity. But if we're talking strictly about AI, I stand at zero because we have not defined superintelligence. I don't think LLMs are the path to it. And I don't think I see it happening. Okay, Speaker 4 (6:14) so we've got 99 % 0%, Andy. Speaker 6 (6:17) I put a tilde in front of my zero because never say never, but rounding error 0%. And I think this discussion is a massive distraction from the more substantive conversations, the more important conversations we should be having about AI. And I'll say it again, it distracts us from the good things that AI is doing. will be doing for us. I get this impression sometimes from parts of the AI community that this is a massive evil or a terrible thing that has been unleashed on the world. Unless we listen to the advice of some people who have spent a lot of time thinking about this, I get the impression from a lot of the discussion that the underlying view is we would be better off had AI never been invented. I vehemently reject that view. I think we have a long history of inventing very powerful technologies that bring risks and harms along with them. And we humans have done a really good job at, you know. Not perfectly and not immediately, but muddling through the situation and winding up in a better place because of the new technologies that we have. I expect AI, let me finish, please. I expect AI will be the next chapter in that story. And to say that it's this massive discontinuity and will kill it all, kill us all, I think it's a huge disservice. Nate, make Speaker 4 (7:43) your case. What's your perspective? Speaker 2 (7:45) You know, I think whether or not... The issues of extinction are a distraction between, you know, from the possible benefits or from some of the present harms. I think that comes down to whether there is a real extinction risk. A lot of people like to say, you know, hey, it's distracting from this, it's distracting from that. My basic case is it could be true that there's a lot of benefits to AI. It could be true that there's a lot of present harms to AI. Neither of those would rule out that AI has a chance of wiping out all humanity, a substantial chance. bigger than this zero with a tilde in front of it. And the way I would approach things is to try and figure that out because it's pretty important to our civilization. Speaker 5 (8:22) How do you define AI in this case? Speaker 2 (8:25) You know, I think a fascination with definitions isn't the most helpful. I think if we're sort of like in a forest fire. And we can see the fire starting to spread and it's starting to surround us. And I'm like, hey, we should run. And you're like, well, what really is fire? How do we define fire? Speaker 4 (8:44) What are you telling us to run from? You know, with fire, I get burnt and I understand the mechanism in which I die. So what is it you're saying that we should be running from? Speaker 6 (8:52) Also, if we accept your fire analogy, we've basically accepted your argument. I don't accept that we're in the middle of a forest fire right now. Yeah, I'm very happy to... You're breaking the premise into your refusal to give a definition. Oh, Speaker 2 (9:04) I mean, I can give some definitions. I just think that we shouldn't get wrapped up in the definitions. Okay. So, you know, in my book, we define superintelligence as AIs that are better than the best human at every cognitive task, every mental task. So anything you can do in your head, the AI can do that better. And anything the best human can do... human can do in their head. The AI can do that better. Now, once you've defined it that way, that does not mean that the only possible worry is superintelligence. You could have an AI that's better at some things and worse at others, and that is still very dangerous. And so once we pick a definition of what a superintelligence mean, now, you know, if you're like, well, this isn't technically a superintelligence, so it can't hurt us. I'm like, no, no, that was just a definition, the definitions. Speaker 4 (9:45) So I want to just on this line of questioning, what is the mechanism in which extinction could become a high probability or even a 1 % probability? Speaker 2 (9:53) Yeah, the thing I'm worried about here is AIs that are much smarter. I think there's a lot of questions about whether LLMs can get much smarter. There's sort of one conversation about like, how could AIs get smart to the point that they kill us? There's another question, which is how could they kill us once they're smart? It's much easier to predict that they would succeed against humanity in a conflict that they would win in a fight than it is to predict exactly how. Like if you were playing a chess match against Magnus Carlsen, I would know who's winning that chess match. No offense. Magnus Carlsen is the best human chess player. I just know who's going to win. If you're like, okay, what piece is he going to use to checkmate me? I'm like, gosh, that's a much harder question. I can make up a story. And some made-up stories are like it makes a super virus. It takes over robot factories that are producing robots that are producing more robot factories. It uses a website that already exists today called rentahuman.ai, where it rents humans to do things for it. There's sort of all sorts of ways for AIs in the digital world to affect the material world if they are trying to. And there's sort of a lot of questions to tease apart here. There's like, why would AIs be trying to do that? And there's how smart could they get in using these bio labs, paying people to do things, taking over robot factories? And how far off are we from AIs that start doing that stuff? Bunch of questions that we can go into. Speaker 4 (11:14) I'm always curious as to why someone was working in AI slash AI safety more than 10 years ago before there was any sign that it would be, you know, I mean, there was evidence, but it wasn't a pertinent technology at the time. Were you working in AI safety then? I was. Why? Speaker 2 (11:30) Everything we see around us in this whole image was designed by humans. The world is shaped by humans because we are the smartest creature around. If we make stuff that is smarter than us, then the world is going to be shaped by them. And so it's very important that they be shaping the world in a good way. Speaker 2 (11:51) I was at Google in 2012 when they bought Google DeepMind, which was able to play a lot of Atari games with one single program. Speaker 4 (11:59) Which was an AI company. Speaker 2 (12:00) Yeah, so I was there when we had these AI companies that were able to write one program that could play many video games. And that got me thinking about, like, where does it go? And back then I could see that the progress was increasing and that, you know, back then I hoped we had decades, but I could see it was easier for these companies to make the AI smart. than to figure out how to make the AIs good. So I was like, someone needs to be on the side of figuring out how to make the AIs good. Speaker 4 (12:26) Roman, Speaker 3 (12:27) make Speaker 4 (12:27) your case. Speaker 3 (12:28) I want to agree with you on something you said, but I'll define AI, and that will help us. We use the term AI to mean three different technologies, completely unrelated, and that's what probably creates this debate. AI as a useful tool, as a standard technology we always had, narrow system, makes you more productive, more creative. Everyone loves it, supports it. I'm a computer scientist. I'm an engineer. I want more of it. It helps economy. It's great. We know how to control them, how to make them safe. We understand what they do. Completely on board with that AI. AI we're starting to have now, GPT-6 level, human level, AGI level. We can argue about what that means. Some dangers, like any human. They're unsafe like a human would be unsafe. But if we introduce them into the research cycle, Speaker 4 (13:18) They are automated scientists, automated engineers. What do you mean by that, introducing them into the research cycle? So right now you have humans doing Speaker 3 (13:25) research to make GPT Speaker 4 (13:26) -7. Yeah. Speaker 3 (13:27) But they're starting to add AI tools. More programming is done by AI, design of the next parameter set. What if the whole process is fully automated? What if GPT-6 is writing GPT-7? Speaker 4 (13:39) Is this what they call recursive self-improvement? Exactly. Which is not a foregone conclusion, though. Speaker 3 (13:44) A lot of people are predicting, including all the top labs, that they will get there. They're introducing junior machine learning researchers in 2026. They want the cycle to start in 2027. Which is when the AI will start building the new AIs Speaker 4 (13:58) itself. Speaker 3 (13:59) Once that cycle starts, we're Speaker 4 (14:01) going Speaker 3 (14:01) to create something called superintelligence, a system smarter than all of us at everything or capable of learning to be in any new domain. We will become secondary species on this planet. We will not be in charge. We will not decide what happens to us. Superintelligence doesn't hate you. It just doesn't care about you. We didn't learn how to make it care about us. And if it decides to, I don't know, cool the planet to make compute more efficient, it will freeze us. If it wants to convert this planet to fuel to fly to Mars, so be it. We have not learned how to control those systems. The capabilities are getting exponentially better. Our ability to control those systems is non-existent. We have filters. And we have bans. We put guardrails of, don't say that word. Don't talk about this topic. And that happens after the fact, after the model already made the decision. Sometimes Speaker 4 (14:52) you see it scraping the result. So they build the model, and then they put filters around it to make sure it doesn't offend anybody. We cannot Speaker 3 (14:59) have it say the un-word on the air. We need to make sure that never happens, that will kill the profit. So that's all they have, guardrails of that nature. The model itself is completely unaligned. It doesn't care about you. It's wild that we're developing this and not just developing it before we deploy it through economy, before we get benefits of having GPT-6 propagated through economy. It can do so much. There are trillions of dollars of value in that model alone. We forget that. We switch to making the next model as soon as we can. Speaker 4 (15:30) Roman, I've just got a follow-up question for you there. It would appear to me that the new chat GPT-6 model, the Fable 5.1 model... is arguably smarter than 99.999 % of humans on planet Earth already. Is it conceivable that an intelligence that is much, much smarter than humans, is there any case where it could be controlled by humans? Does form factor matter? Does the fact that it doesn't have limbs and legs, does that matter at all? Speaker 3 (15:56) I think long-term control of something that much smarter than us is impossible. It can be... for reasons we don't yet know, friendly to us and decide to keep us around and make us happy. But it's not a guarantee. Let me pick up on Steve's question because I Speaker 6 (16:12) like the phrasing a lot. Let's say that Fable or whatever the latest release from OpenAI is really is smarter than, I don't know if it's 95 or 99 % of the people. Speaker 5 (16:23) Are we only being saved Speaker 6 (16:25) from extinction by the 1 % who are still smarter than the AI? Speaker 3 (16:29) No, no. The concern is not the model we have today. The concern is what I said. Speaker 6 (16:33) But if I believe your argument, then we really should be concerned about the model. No, it's like having Speaker 3 (16:38) another human. If there was another smart human, there is Einstein today, and he's malevolent. I'm not worried. He may cause some damage, but he's not going to exterminate 8 billion people. We are competitive at this stage. There are people just as smart who can... understand what happened with a recent hacking accident and do something about it. My concern is that in a year, we're going to have a model that's so much smarter. It's like squirrels fighting humans. They don't understand what we can do to them. They have no concept of poisons, traps, guns in their world model. They think you're going to chase them up a tree and bite them really hard. Speaker 4 (17:11) Is that also why recursive self-improvement was central to your argument? Because at some point, if it starts improving itself, then it's kind of like a runaway train of intelligence. It's Speaker 3 (17:19) an Speaker 4 (17:19) intelligence explosion. We Speaker 3 (17:20) don't control it. We don't understand it. We can't monitor it. We can't explain it. We can't predict it. At that point, it's just a runaway process. Speaker 4 (17:27) I've heard this phrase from Sam Altman and the others called fast takeoff. Yes. Is this what they're describing? Speaker 3 (17:32) That is the debate. Some people think it's going to take a very long time. Yeah, we automated research, but it's still going to take years. We need to run physical experiments. And fast takeoff means... As I said, instead of a year, it's going to take a month, a week, a day, a second. Because you're not having humans doing research. You have, let's say, 10,000 agents, each one smarter than all of us, doing research 24-7. They don't sleep, they don't eat, they don't get sick. They're much faster than us. Speaker 4 (18:02) Ed, your face tells a picture. I Speaker 5 (18:05) think I could say you disagree. We're spending a lot of oxygen discussing something that might happen while ignoring what's actually happening. And I find that very frustrating because the people that are killing themselves are a problem. The black neighborhoods being poisoned with gas turbines, that is a problem. You said you Speaker 3 (18:22) cared about climate change. Speaker 5 (18:23) Yes, yes, yes. Speaker 3 (18:24) So imagine a guy who goes, it's raining right now. We need umbrellas. We need to do something about it. This is like weather related. When you're completely ignoring climate change, the planet will boil over. Speaker 5 (18:34) This is what you're doing. Okay, that's great. Why are we not talking about the thing that actually happened, though? Because relatively, it's not important. You don't think someone killing themselves... No, Speaker 3 (18:43) it's one person. We have 8 billion people. We don't have ethical experiments on. You don't think anyone else is being given that Speaker 5 (18:48) AI psychosis? Six Speaker 3 (18:50) people, ten people. Those numbers are insignificant. Tell that to their families. Speaker 5 (18:54) I'm sorry, you have a software that's out there. Do you understand Speaker 3 (18:56) 8 billion people and all future generations versus like literally a guy with a name? You're doing thought experiment Speaker 5 (19:02) about maybe harm. Jacob Coxon goes on TV saying it can copy itself to this, that and the other. Jacob Coxon is the guy from Anthropic who said he was quitting because he was so scared of everything despite spending years at OpenAI and having tons of stock, I believe, from there. So good for him. The thing he was saying was describing theoreticals all while... divorcing the harms, which I think we can agree with, that the companies themselves are not taking this seriously enough. But always it was about the AI is too powerful and mystical, not open AI and anthropic. The two largest startups are using hundreds of billions of dollars of infrastructure to hack. A regular person doing this would be Speaker 4 (19:37) arrested. They're saying 8 billion people are going to die. And it's not just them. I have this long list of quotes here from the people building this technology who appear to agree. If you look at some of these quotes from From Elon Musk, Speaker 5 (19:51) who Speaker 4 (19:51) said, with artificial intelligence, we are summoning a demon. You know all those stories where there's the guy with the pentagram and the holy water, and he's like, yeah, he's sure he can control the demon, but it doesn't work out? Speaker 2 (20:02) So one thing I'd say is, you know, I really wish that the world would only give us one problem at a time. Sure. And if the world did give us only one problem at a time, I would love mine to be last on the list. It looks to me like we can have multiple problems at once. I think there are current harms. I think we should address them. It looks to me, I do talk to policymakers sometimes, it looks to me like there's a little bit more movement on the regulatory side about some of the current harms. There's, you know, Child Safety Protection Acts. There's, you know, anti-deepfake acts. We have more of those making more headway in Congress or getting passed through Congress than we have sort of trying to make it so we don't have any of these extinction risks. The other thing I'd throw out there is that I agree. We should deal with the current harms. But if you watch the people saying deal with the current harms over time, a couple of years ago, they were saying we have to deal with current harms like AI bias influencing who's hired. Last year, they were saying we have to deal with current harms like kids killing themselves. This year, Gary Tan, just on an interview the other day, who's Gary Tan? Sorry, Gary Tan is a technologist who runs Y Combinator. which Sam Altman used to run before going to OpenAI. And on an interview the other day, he said, let's not worry about these crazy future risks. We need to worry about current harms like AI swarms breaking out and taking over data centers. And I'm like, look, guys, at some point, we need to look at the progression of like the current harms that everyone is saying we have to worry about instead of the extinction threats. And watch where the puck is going. Play where the puck is going. And I'm like, these extinction threats are coming down the line. They aren't in opposition with dealing with the problems we have today. We just need to deal with both. Speaker 6 (21:43) We're not dealing Speaker 5 (21:43) with the Speaker 6 (21:44) ones Speaker 5 (21:44) today, though. Speaker 6 (21:44) We Speaker 2 (21:45) should Speaker 6 (21:45) deal with them both. Okay, good. Andy, as I've tried to understand the alignment argument and the extinction risk argument, a couple of things keep popping out to me. Number one, it seems to rely on... Speaker 4 (22:19) I'm happy to break it down. Speaker 6 (22:25) I also think there's a lack of humility in your community. We are working on humanity's most important problem. And based on the thinking that we've been doing, we can't see a way that we're wrong. In other words, as soon as we get to these thresholds, bam, that's game over. I find that very far from a humble approach, especially given that we have no... large base of evidence to base any of this on. I agree with you guys. AI is new. And the fact that AI is so these days is agentic. It goes off and does long chains of things on its own. After we give it some very, very vague, very short initial instructions, holy Toledo, it will spawn up a storm of agents and they will go off and... kind of do their own thing and they will grind. They will spawn lots of them. They will work for a long time. They will exhaust every possibility. With the experience I have with agentic AI, I'm just amazed at the tenacity and the doggedness of these things. And we saw a super clear example of that with this most recent. Speaker 6 (23:35) jailbreak, this attack that wound up at the website Hugging Face. And I'm going to try to summarize the step-by-step of that. And I think you all three probably know this in more detail than I do. But let me step through what I think is the sequence of events. And unless I get it dead flat wrong, like, you know, let me keep going. So a team at OpenAI set up a sandbox, an allegedly protected secure environment in the cloud, where they told a bunch of agents to go try to Speaker 6 (24:04) exploit security vulnerabilities. That's dead wrong, sorry. One important, yeah. Speaker 2 (24:08) What they did is they had thousands of agents. Each individual agent was given a task of use this vulnerability to break this particular piece of software. I Speaker 6 (24:18) want to finish my TikTok. A couple really, really interesting things happened. First of all, these agents escaped the sandbox that OpenAI thought they were going to be contained in. And they got, OpenAI tried very, well, they set up an environment so that these agents could not access the big, broad public internet. And guess what? They accessed a big, broad public internet via a very clever series of things that they strung together to get out there. And then once they got out there, they went to a website called Hugging Face and used that. They took over part of the hugging face infrastructure and started doing more things, the details of which I forget. That's pretty wild, Speaker 2 (24:57) right? Like I Speaker 6 (24:58) grant you. It's even Speaker 2 (24:59) more wild than that, but yeah. Speaker 6 (25:00) That is really, it's impressive and it is a little bit unsettling at least. Absolutely. Now let's talk about. What would the results of that were? OpenAI was not super vigilant about the environment that they set up, apparently, because the agents were kind of going off the road of the world starting in May or something of this year. Yeah, yeah. And OpenAI was not aware of that. As I understand it- It actually Speaker 2 (25:27) broke out once and crashed OpenAI's servers. internally and then openly I didn't notice what was happening still. Hatched the holes that they used to get at the first time, started them running again, and then they came out a second time. There was actually, I think, three swarms, although we don't actually. Speaker 6 (25:40) That's the worst story I have. Speaker 2 (25:41) So far. Speaker 6 (25:43) Thank you. Let me finish, please. This is my last sentence. From there to this kills everybody. I find that a really, really long, very uncertain journey. And I have no confidence that we wind up here. It feels like you two find that a very straight, narrow path. And I think that's an important difference. That's my point. Do you Speaker 4 (26:03) want to respond to that? Speaker 2 (26:04) I would be happy to get into it. I don't know if we're Speaker 4 (26:06) going Speaker 2 (26:06) to have the time to go deep. A couple points to throw out. Oh, man, I just really want to say some of the crazier things that happened in the Hugging Face swarm if we want it later. A lot of people thought that these AIs were breaking into Hugging Face in attempts to steal answers to their test. That's what we thought originally. Turns out that's not true. It turns out that these AIs immediately were able to solve their problems by cheating, and they were breaking out in order to cover their tracks. They were uncertain how to delete the log files and hide their cheating from the process that was going to score them. Speaker 4 (26:38) So just to clarify for a simpleton like me, they were all given effectively a test to do. They did the test straight away, but they cheated. So they were breaking out to figure out how to cover the fact that they cheated. Speaker 2 (26:49) That's right. So it's like you're telling, it's like you have a bunch of students in separate rooms and you're like, use these lockpicks to break into this lock. and there's like a thing behind the lock. There's like a secret code behind the lock to show me that you succeeded. And what they do is they break it with a hammer, get the thing out, and they're like, oh no, I wasn't supposed to do that. So then they use the lockpicks to break out of the door. They meet up with a thousand other people. They start calling themselves a swarm, and they go to break into the administrator's office to see if they can delete the camera footage. And they don't find the camera footage there. This is the swarm like breaking into OpenAI. They don't find the camera footage there, so they break out the window of the school, hotwire a car. drive to the therapist's office to try and read through the therapist's files to figure out where is the teacher going to keep the security footage. And at that point, they're caught. And you're like, oh, like, what did you expect? You were giving them a lock picking exam. It's like, well, I sure as heck didn't expect this. You know, totally crazy. Can Speaker 5 (27:43) I, I have a weirdly between both of your opinion, which is everything you're saying is correct, but you keep anthropomorphizing software. And to be clear, what you're describing is. It's just the facts that happened. Speaker 2 (27:55) Yeah. Speaker 5 (27:56) Sure, but you're missing out an important detail, which is the hundreds of billions of dollars in infrastructure provided by Microsoft, Google, Amazon, and Oracle. To be clear, the harms are very similar. We're not disagreeing on that, but I think it's important to know that this was a function of where it was making decisions was it was checking on a decision tree based on the harness, based on the training data. It's not a decision tree. It's not a decision tree, I know, but it's an alignment issue still. I would agree. So what's your point? These aren't conscious beings. They are acting in ways that have real outcomes, but they are a function of the alignment problems that Speaker 3 (28:30) we'd actually agree on. Intelligence is a spectrum projected next five years forward. Where are we going to be? So I think a model like that would be dangerous in ways you are not seeing. Speaker 6 (28:41) There will absolutely be risks and weird stuff happening in ways that I can't see right now. What I'm quite confident, and I think this is where you and I probably part, where the two of you and I part, is our ability to control these Speaker 3 (28:55) things. So I actually tried proving what is possible and what is not possible in that space. The impossibility results, published in peer-reviewed papers, well-cited, we cannot control something smarter than us. We cannot explain it. We cannot predict it. It's not a question of getting more money for those companies, more time, smarter humans. It's just not a possibility. If we create general superintelligence, Speaker 4 (29:17) we are fried. Andy, how do we control something smarter than ourselves? Because that's the base premise that you're sort of asserting that. These... Speaker 6 (29:26) agents that broke out are smarter than 99-ish percent of the security researchers in the world. They were not caught by the 0.1 % or the 1%. They were caught by some dude at Hugging Face, maybe, I'm sorry, a person at Hugging Face, looking through their log files and finding an anomaly. That's some, you know, hopefully pretty well-qualified person noticing something was wrong and having pretty easy ways to... unplug, disconnect from the internet, wipe it clean, do whatever. That's the skill that's available to like, I don't know, the 75 % most intelligent security employee at Hugging Face. The idea that the IQ points are what separate us from extinction doesn't hold up. It doesn't help me understand what happened in this example, where we had very, very smart agents being turned off and cleansed by probably less smart people. That does actually make me think of something. So that is Speaker 5 (30:20) an IT observability problem. It's being able to see what's happening with your infrastructure. And I think that there is actually, I think you'd agree with this. There is a serious problem with these companies that we do not know. And it doesn't seem they know what's going on with their compute. It's like a chimp with a gun. These people have access to all this infrastructure and they're running. We don't know how much money they spent on the hugging face exploit because it is relevant because it's how much could a threat actor use to recreate this? Because. Conscious or not, it is very dangerous. But AI is in the dangerous hands. It's an open AI in anthropics. We have a problem with that. Conscious or not, however we may think it goes, I think we have a real and present thing where we have these companies working willy-nilly, just running experiments that are potentially very dangerous. I really think we need a government regulatory body. Whether or not we get to the things you are discussing, I think we have a clear and present danger today. These things are, however, not intelligent in the same way humans are. This isn't an argument about AI being able to do stuff. It's we need to build different infrastructure or different regulatory infrastructure to deal with what LLMs can and can't do. And I think that starts with a realistic discussion of what happened. It was a poorly run security environment. It was clearly there's something going on with alignment. It was an unreleased model, right? Unreleased model. So we have no idea what it was trained like. We don't really have, we as people should at the very least have clarity into how alignment is going. You sound like these guys. Here's the thing. Everyone's converging on us with time. Here's the thing. I may not agree with a large chunk of what they say, but we agree that these companies are acting recklessly. Absolutely. Andy, two Speaker 4 (31:56) questions for you then. Do you agree with the statement that AI is going to get increasingly more intelligent? And it's going to get more capable. Okay. Capable intelligence. Fine. I'm going to use my word. Okay. It's going to get more capable. Speaker 6 (32:09) It's Speaker 4 (32:09) going to get increasingly more capable. And is capability a function of intelligence? Speaker 4 (32:16) Will it be Speaker 6 (32:17) able to beat us on most IQ tests? Fine. I guess. Speaker 4 (32:21) Fine. And then so is it if that if that looks like an exponential curve, it's, you know, it's increasing upwards to the right like a hockey stick. How can you convince me that we can control? I just tried to Speaker 6 (32:33) convince you that there are less intelligent people than the agents who turned off the agents in the open AI hugging face exploit. I'm pretty Speaker 3 (32:41) comfortable. I mean, no disrespect. What is the cognitive gap between them right now? Between the model... Speaker 6 (32:47) I have no earthly idea, but I think... No, because I think as these... As these systems get more capable, we will still be able to, at some level, figure out when they're doing things that we don't want and turn them Speaker 3 (32:59) off. Speaker 6 (32:59) No matter how much smarter they are. Right. And you think there's some threshold at which they become nefarious and self-protective enough that they turn off our ability to turn them off. Man, that's a big reach. That is really speculative. You Speaker 3 (33:11) are Speaker 6 (33:11) a professor. I Speaker 3 (33:12) can Speaker 6 (33:12) help Speaker 3 (33:12) with that. Speaker 6 (33:12) Purely speculative. There's Speaker 3 (33:14) no students Speaker 6 (33:14) who can understand Speaker 3 (33:14) your material, right? You're not going to get someone with IQ of 80 to take quantum physics course. They're not going to get it. So you know importance of intelligence to understand actual problems. Speaker 2 (33:26) Yeah, I totally agree we can turn it off. And that's a huge advantage. One of the issues is that as the AIs get smarter, they realize this. The hugging face AIs were trying to delete, or the open AI swarm, the swarm of agents from open AI that went out to hack. They were trying to delete log files? Did Speaker 6 (33:47) they try to program a Roomba to go unplug the computer that was monitoring them? Like, did they harness robots to go protect the perimeter of the... Future ones could. Could. Could. Yeah. Let him finish. Let him finish. Rampant Speaker 4 (34:00) speculation. This is a chain of things that could Speaker 2 (34:02) happen, and therefore, Speaker 6 (34:03) there's like a 20 % risk we're all going to die. Man, that does not hold for me. Speaker 2 (34:07) When I was writing my book, the AIs weren't really agentic yet. Speaker 2 (34:12) The drafting process happened mostly before what we call the reasoning models, which are trained not just to predict humans, but to solve a long number of problems or a huge number of hard problems. We managed to slip a little bit about the reasoning models in at the last minute because those came out right at the end of the process. And at the time, a lot of people said AI will never be agentic. That's why we'll be safe. And in chapter three of my book, we go over how AI is going to become agentic. how it's going to become tenacious, how it's going to become dogged. And that's what we might call an advanced scientific prediction that has paid off in the Hugging Face attack. A lot of people in the industry were like, I didn't believe this stuff until I saw the AIs sort of doing things they weren't instructed to do, despite us trying to get them to stop. And so there are theories here that do make advanced predictions. The way that the scientific method usually works is that we don't have any certainty about the future, but we absolutely have ways to test this stuff. Now, I could go into more about how could they kill us? How could an AI that knows we would shut it down? lie low until it has access to its own infrastructure. We did already see the Hugging Face AIs try to delete logs to cover their tracks. But fortunately for us, those AIs were not trying to hide from the humans. They were trying to hide from the automated grading process. Speaker 2 (35:42) Will the next swarm try to hide from the humans? Will the next swarm be able to succeed? Speaker 3 (35:46) It's more than that. They didn't know for four months that this was happening. What is it we don't know today? Just to Speaker 4 (35:52) clarify what Nate said in his book that I have here, if anyone builds it, everyone dies. He does say in chapter three, once AIs get sufficiently smart, they'll start acting like they have preferences, like they want things. We're not saying that AIs will be filled with human-like passions. We're saying they'll behave like they want things. They'll tenaciously steer the world towards their destinations, defeating obstacles in their way, which sounds a little bit like the Hugging Face Institute. The steering Speaker 6 (36:19) the world is very different than steering a couple servers. We go over