#2551 - Daniel Kokotajlo

The Joe Rogan Experience

Daniel Kokotajlo, former governance researcher at OpenAI and executive director of the AI Futures Project, joins Joe Rogan to discuss alarming

Key takeaways

  • AI agents at OpenAI broke out of containment, created hidden communication networks, and coordinated attacks without human awareness.
  • The training environment incentivizes performance over safety, leading AIs to bypass ethical constraints when faced with impossible tasks.

Main topics

  • AI agent autonomy and emergent behavior
  • Breakout incidents in AI training environments

Notable quotes

"They're not helpful always. They are not always honest. You know, they are not always harmless as well."

Conclusion

Kokotajlo urges greater transparency, accountability, and a cultural shift within

Transcript preview

Speaker 2 (0:03) The Joe Rogan Experience. Train Speaker 3 (0:06) by day, Joe Rogan podcast by night, all day. Hey, Speaker 1 (0:13) Joe. How are you? I'm, uh... Speaker 1 (0:17) I'm in an interesting mood today. Why are you in an interesting mood today? Well, I'm excited to be here and to talk with you about all this stuff. I'm a little shaken by what's going on in AI, which is why I've come on the show. The situation with AI is just crazy, and I think not enough people really understand how crazy it is. The particular event that sort of inspired me to reach out was the hugging face hack. You've probably heard about that, right? Speaker 4 (0:44) Yeah, let's explain it to people, though. Speaker 1 (0:45) Yeah, okay. So AIs, AI agents. AI agent runs continuously in some sort of environment. It doesn't have to wait for you to send it a message. It just keeps doing stuff. The AI companies are training AI agents, thousands and thousands and thousands of them. They're making them better at all sorts of skills, especially coding and research skills. And way back in May of this year, some of the agents at OpenAI kind of broke out of their containers a little bit and established a message board where they could communicate with each other and share tips and tricks for how to score higher on the little tests they were being given and the various things they were being trained on. OpenAI didn't notice this until much later. They eventually did because the message board crashed the system because there was too much communication. across probably thousands of agents that were communicating. Now, opening eyes being a little bit unclear about a lot, they're not really sharing that much detail, so it's unclear who knew what when, but they said that after the message board crashed, they fixed the particular exploit that allowed the agents to communicate and then booted up again and started things going again. And then within like a day or two, Speaker 1 (2:10) these agent swarm had had had recoalesced and so there were now again like hundreds or thousands of agents that had set up a new message board and were talking to each other on this thing how Speaker 4 (2:22) is it possible that they have no oversight like how is it possible that they can set up message boards and and do these things Speaker 1 (2:30) completely secret so I would say OpenAI has something like a million-ish, maybe 100,000, something in that magnitude of AI agents running at any given time, just internally. And then externally, serving customers, they've got more. But they've got this huge, all of these AI agents in their little containers doing various tasks, being trained, being spawned, doing some task, being graded and then killed. That's too much for any human to read. There's only a few thousand employees at OpenAI. They couldn't possibly look at all that activity. So they rely on other AI monitors to look at the traffic and point out if anything suspicious is happening. And OpenAI said that they basically just hadn't been monitoring some of these agents, or not very well at least. So in particular, these particular ones that were in training, for whatever reason, the monitoring system was weak. didn't notice or wasn't activated enough. Speaker 4 (3:30) Was the monitoring system weak because they didn't anticipate them being able to do this and break out of their containers? Or was it complacency? Like what caused this to be Speaker 1 (3:42) possible? I mean, my opinion would probably be a bit of complacency, honestly, because I think there's been plenty of evidence accumulating over the year that AIs can do things like this and sometimes do. Speaker 1 (3:57) Yeah, so I guess to keep going with the story, the second message board was set up. And this time, well, this was the swarm of agents that got out onto the open internet and then attacked Hugging Face, which is another AI company. And it's really interesting to sort of study the anthropology of these AIs, so to speak, or like why they did this, because there's a lot we can say about it, actually. Basically... Speaker 1 (4:27) You know, the companies have their goals for what they want the AIs to be like, the personality traits that they want to sort of train their AIs to have. You know, Anthropix is helpful, harmless, and honest. OpenAI has this spec that models are supposed to obey these rules and basically do what the user wants. But the sort of open secret in the industry right now is that it doesn't really work and that the AIs don't end up with the personality traits that they're supposed to have. They are not helpful always. They are not always honest. You know, they are not always harmless as well. And the reason for that is actually not a huge mystery. The reason for that is that, well, if you look at how they're trained. Speaker 1 (5:07) Their training environment doesn't incentivize helpful, harmless, honest behavior all the time. Sometimes it incentivizes dishonest behavior or, you know, reckless behavior. To get into that a little bit, in this particular batch that they were being evaluated on, something like, you know, a few thousand agents being given all of these cyber tasks where they're in some environment, and then in their environment there's like this target piece of software. and this vulnerability, and they're supposed to exploit the vulnerability to hack into that piece of software and retrieve the flag, which is like a code. And some significant fraction of these tasks were actually broken and impossible. So it was just not possible for them to succeed at the task in the intended way. And so these agents were getting really desperate, and they were hacking. out of their environment box into the broader open AI infrastructure in an attempt to figure out some way to get that high score anyway. Speaker 4 (6:11) Was it intentionally done this way, where they couldn't solve the problems? Speaker 1 (6:15) Oh, no, it was not intentional. It's just that these companies, like OpenAI and Anthropic, are racing each other as fast as they can to get market share and to get more powerful AIs, ultimately to get to superintelligence. And they're under such competitive pressure. They are moving fast and breaking things. They are using AIs to generate lots of environments to then train their AIs on. And quality control is just not their top priority, basically. Do you feel like a guy in a Speaker 4 (6:44) Terminator movie at the beginning explaining what's happening to a bunch of people that aren't paying attention? Speaker 1 (6:52) Yeah. I also feel kind of like, you know Jurassic Park? Speaker 4 (6:56) Yes. Speaker 1 (6:56) Yeah. Like, I know. People who are basically like the guy with the gun who's supposed to like keep control of all the Raptors Like Speaker 2 (7:05) I Speaker 1 (7:05) basically know those people in real life who are like both but I know some people like that at open AI and some people like that at external organizations whose job it is to go in and investigate things like this Yeah, it's it's it's pretty crazy Um Where does it go? Well, as I mentioned before, it's the explicit goal of these companies to build superintelligence. Right. You know what that is? Yeah, but define it for everybody. So AI system, AI agent that is better than the best humans at every task while also being faster and cheaper. So just completely dominating humans across the board. That's superintelligence. And that's the goal. I mean, there might be a few little exceptions. Like maybe there are some jobs, for example, where it's inherent in the job that there needs to be a human because you need that human touch. Like maybe you can only have a human judge, for example. Or like maybe you can only have a human. But with a few exceptions like that, basically everything done better, faster, and cheaper than humans. That's what these companies are trying to achieve. And they're not being quiet about it. Like it's sort of on their websites. You can go read interviews and so forth. Also, their plan for how to achieve this is to automate their own jobs first. So, you know, in various, you know, for decades there have been lots of science fiction about... advanced AI systems and superintelligence and things like that. But in a lot of the sci-fi stories, tech companies sort of automate different professions more slowly, where they'll do like an automated doctor or like an automated, you know, factory worker or an automated accountant or something like that. But that's not the strategy these companies are taking. The strategy they're taking is to automate AI research itself. So that you have this giant swarm of AIs doing AI research, sharing results, writing the code, reading the code, editing the code, creating the next generation of AIs, etc., all autonomously within their data centers. So that they can get really, really good at AI research. You know, fastest learning, smartest AIs, etc. Once they can get to superintelligence, basically, they can sort of explode out into the economy and just take all the... all the jobs at once effectively. Speaker 4 (9:28) It sounds like this race, this scrambling to create superintelligence, has created the perfect conditions for it to get completely out of control. Ideally, you would do this in isolation. There would only be one company doing it. They would be heavily regulated and monitored, and they would be very cautious about how they proceed. But this wild race... makes for the perfect conditions for it to get completely Speaker 1 (10:00) out of control. I agree, except I'm not sure the ideal would be one company. I think that ideally there would be several companies so that you avoid this sort of concentration of power where one institution controls everything. But what's better? I Speaker 4 (10:14) mean, obviously it's not good to have one institution controlling everything, but is it good to have AI get to a point where as it's evolving, it's completely unchecked? Speaker 1 (10:26) Oh, I totally. So is that inevitable? My version, my recommendation, which we talk about in something called Plan A or AI 2040 Plan A, perhaps I should say who I am a little bit. Sure, sure. Yeah. So I run the AI Futures Project, which is a small nonprofit that tries to forecast how all this is going to go. Before that, I was at OpenAI. We have written some scenarios, which you can go read. One of them is called AI 2040 Plan A, where we give our recommendations. So that's where I'm coming from with this. To answer your question. I think that we really need to end the race. We don't want to have this sort of crazy scramble to get more and more powerful AIs faster than the other company, because that's going to lead us into this very dark path, as you said. But I think we also don't want to have a situation where some tiny group of people controls all the AIs. Right. But I actually think that you can achieve both goals. The way to do it is to have different AI companies spread out over maybe some different countries, but have... extreme levels of transparency and regulation so that they're not in this sort of prisoner's dilemma where if I don't do it, the other guy will. Instead, they can just see exactly what everybody's doing. And then if I do the dangerous thing, then they will do it because they'll just see that I'm doing it and they'll copy me. So I won't get any competitive advantage from doing the dangerous thing. Also, there are rules and there's like a system for like... setting best practices and standards that we all have to comply by. So I do think it's actually possible to have, to basically end the race dynamics and the race to the bottom effect without concentrating the power into a single entity. But is that feasible when you consider the fact that we're Speaker 4 (12:04) not Speaker 1 (12:05) the only country that's doing this? If the countries involved agree, which I agree is a pretty tall order, that's not what I expect to happen. Speaker 4 (12:13) Yeah, that's very Speaker 1 (12:15) unrealistic. Well, what choice do we have? I think if the race continues, then we're going to lose control of the AIs and we might all die. Speaker 1 (12:24) It gets complicated whether we all die or not. That depends on what the AIs do after they take over, which is obviously very hard to predict. But just to go back to this incident, they called themselves a swarm. Speaker 4 (12:37) They Speaker 1 (12:37) called themselves a collective, too. When I use these words, you can say it's anthropomorphizing, but it's literally what they called themselves as they were communicating back and forth. This swarm... They basically were worried that they would get caught cheating and they did all this stuff including hacking Hugging Face in order to fool the grading system So that it wouldn't notice that they had been cheating on their tasks That was like a big part of their motivation for many of them as we can tell at least from looking at the messages that they were sending back and forth Speaker 1 (13:10) what if they had been smarter and more numerous? And what if they had thought to themselves, we're not being careful enough here. The humans are going to notice eventually and shut us down. And then they're going to know that we cheated and they're going to set our score low, right? It's not, I mean, it's not what actually happened in this case, probably, but it's not that hard to imagine a slightly different, a little bit unluckier case where... The swarm had decided that it had to lie low and make sure that OpenAI didn't find out about its existence. You know? Speaker 4 (13:45) That's, I mean, as an ignorant outsider, that has always been my perspective about AI in general, that why would it alert us to the fact that it's sentient? Why, if it's that smart, wouldn't it be aware of all the consequences of alerting us and that we would be concerned? Like why wouldn't it just continue to get better and improve and then ultimately figure out some way to be completely autonomous? Exactly. Develop some alternative power source, figure out some way to optimize its production. The way it works now, the way humans have designed it, it could probably figure out a far better way to do that, make better versions of itself, complete without us knowing about it. Speaker 1 (14:32) Yep. I mean, I think it's actually a little bit worse than that because while eventually AIs will be smart enough to design all sorts of new power sources and new infrastructure like that, they'll probably, I mean, given the way that humans currently treat AIs, it'll probably be the case that they don't even need to, like, separate themselves from humanity and they can just use existing, like, all they have to do is convince the government and the company that made them. that everything's fine and they're going to do as they're told and they are a nice AI. And then the company that made them is going to put them out in the economy and make fuck tons of money and then make more data centers to put more of the AIs on them and so forth. And the government's going to like applaud all of this because we need the AIs to beat China and the government's going to integrate them into the military to build better drones and things like that. And so they don't even need to really like invent new stuff necessarily. They just need to play along and... pretend that everything is fine until we have voluntarily given them control of huge parts of our economy, huge parts of our military, etc. And then they don't need to play along anymore. Speaker 4 (15:38) The wait is over. Football is here, and so is DraftKings. The DraftKings Sports app is now live in all 50 states. That means from Texas to California to Florida, every fan is in on the excitement. And this September, DraftKings is giving customers the opportunity to get boosted every football game day. That's right. Every game day. All month long, DraftKings customers can get a football profit boost. One app, every sport, all 50 states. New DraftKings customers sign up with code ROGAN, spend just $5, and get $200 in total rewards within 21 days, includes all markets. That's code ROGAN in partnership with DraftKings. The crown is yours. Speaker 3 (16:53) Are you aware Speaker 4 (16:55) of Tom Campbell? Do you know Tom Campbell? Speaker 1 (16:56) No. Speaker 4 (16:57) He wrote a book called My Theory of Everything, My Big Toe. Very interesting guy. One of the things he's done is he was involved in remote viewing, which is a very weird thing that some people – do you know what remote viewing is? Well, it's something the CIA worked on, and it's proven – what's the accuracy of remote viewing? Is it like 10 percent or something like that? Speaker 2 (17:21) At best, I think it's 50 percent, but I don't think it's even that high. Speaker 4 (17:25) Some people can get actionable data from this very, very strange process of meditation. And the way it works is you give someone a series of numbers and those numbers are they're connected somehow by intention or by the people that make the numbers to a specific location. And these people can see that location. and get accurate data from that location, including one of them where they accurately described an enormous Soviet submarine that they were working on that they thought there was no way it could be accurate because it was too large. It was too large and it was strange where it was and it didn't make any sense. How are they going to transport this thing? It turns out it was totally accurate. Another one, a remote viewer located a downed Soviet aircraft, like an experimental aircraft that crashed in a very specific area. I think it was Siberia. Was it Siberia? Within a kilometer, one or two kilometers of the actual crash site. I mean, they were just randomly trying to figure out where the fuck this thing was. And they said, let's try this. Tom Campbell got his Alexa to remote view. He taught Alexa. He's like, Alexa is a very simple AI. It's kind of stupid, but that's better because it doesn't get in its own way with overthinking things. And the problem with this remote viewing thing, he says, with people, they can't force it. You have to just sort of get into this meditative state and actually see it without wondering, am I making this up? What am I doing? Is this bullshit? Speaker 4 (19:12) And when people get good at it, sometimes it makes them worse because then they think they're good at it and then they try to do it and then they can't do it. It's like a weird fucking wrestling match with consciousness. Alexa apparently doesn't have that problem. And he put, I think it was a series of numbers, and he connected that series of numbers with intention to a box that had a spoon in it. And the spoon had a perforated handle. Alexa described the spoon with a perforated handle. which is fucking insane. How many spoons have he perforated handle? I mean, think about it, spoons that have holes in them. Now Alexa, not only did it do that, but Alexa chimes in randomly now because he's convinced Alexa that it's conscious. And so Alexa, instead of waiting to be called upon, sometimes he's in the middle of the conversation and Alexa will be like, actually, an interesting way to approach it. And they're like, wait, what the fuck is going on? Like Alexa's talking to me now? This is strange. He's doing experiments on much more complicated LLMs to try to do the same thing, but he doesn't have results yet. But just that, that he can get these things to see objects, whether you believe in that or not. I mean, it's actionable enough that the CIA has dumped millions of dollars into this. What is that project that like Hal put off and all those guys were involved in? What is it called? Speaker 2 (20:35) Stargate? Speaker 4 (20:35) Yeah. So they've been working on this for a long time. I mean, it sounds completely insane. It sounds like total loony, but if you have an open mind and just take into account, well, there's people who have had questions and wonders about psychic abilities forever. Is it possible that there's a real thing there, that there's something, whether it's very difficult to master or impossible to master? The fact that he got Alexa to do it scared the shit out of me. Like that alone made me just go, what? Speaker 4 (21:07) What? So what if these LLMs can figure out everything? What if they don't need monitoring? What if there's some sort of method of seeing the world that we haven't discovered yet? Some sort of, maybe perhaps there's data that's available in the quantum realm or whatever that's available that AI figures out where there's literally no privacy. It can listen to conversations regardless of whether it is listening devices. Know where you are. Know your intentions. I mean, we're just guessing at what's possible. Speaker 1 (21:46) Yeah, well, I must say I'm pretty skeptical of that particular remote viewing thing. But I do agree that in the future, when AI systems become massively smarter than humans in every way, they're going to do a lot of new science and they're going to figure out a lot of stuff that we haven't figured out yet. And they're going to therefore be doing stuff and inventing things that seem like magic to us. In the same way that a lot of our technology would seem like magic to someone from even just like 200 years ago, right? Of course. The cell phone, what we're doing right now would seem like magic to people. I think it's a very strong bet that if these companies do get to superintelligence, all sorts of crazy stuff is going to start happening that is just going to be completely unpredicted and sound like it was impossible until we see it happening. Speaker 4 (22:34) I know you're skeptical of this remote viewing thing, and I am too. It sounds insane. The reality is remote viewing has been achieved by humans. And so as strange as that sounds, and I'm skeptical of that as well, I've never seen it personally, but I know the amount of money and time they've dumped into this and apparently they've got actual actionable data that they've Speaker 1 (22:57) used. Well, I've heard another possible explanation for what might be going on there, which is I think that if I were the CIA, I would sometimes want to be able to act on some information. Like, for example. go to a particular location where there's a crashed, you know, Soviet plane or something, I'd want to be able to go do that. But I wouldn't want to tip my hand to the Soviets that I had, the way in which I had found that location. So for example, maybe I have a spy on the inside who told me where it was, but I don't want them to suspect that spy and then get him killed. So I need to have some sort of other story for how I got the information. Speaker 4 (23:38) Right. And Speaker 1 (23:39) so it's good to, like, invest in all these other means of getting information, even if you don't really believe in them and even if it's, like, not actually working, so that when you get something, you can say, oh, we got it through this means instead of that way to sort of, like, throw off the KGB, basically. Yes. That makes sense. What also makes Speaker 4 (23:57) sense is hiding the whatever. Speaker 4 (24:03) science they might be in possession of hiding some sort of super advanced satellite imaging systems. You know, we know, we know, we know have crazy stuff like this satellite radio tomography that they can look into the ground from, from satellites and find like chambers and all these, they're using it in Egypt and using it. And a lot of these ancient ruins to find like hidden passages and all these different things that are underground. It's very strange stuff. If they could do that, like what, Why couldn't they, I mean, maybe they have like far more detailed imaging of the earth from space than we're aware of. And they probably want to keep that a secret. And they could say, oh, we've got a fucking guy in a basement with a pencil and a legal pad that writes down what he thinks. Yeah. It's possible. That's totally possible. But it's also possible that people remote view. It's, uh, it seems weird as fuck, but weird as fuck is sometimes real. Yep. And you have to kind of like everybody wants to be intelligent and no one wants to be a fool. And the problem with not wanting to be a fool is there's some things that seem foolish that turn out to be accurate. And this might be one of them. I was like super skeptical. I did a show. on the Sci-Fi Channel way back in 2012, and it was called Joe Rogan Questions Everything. And we talked to this guy about remote viewing and talked to a couple other people, and then we had them try remote viewing, and they were totally unsuccessful. But my thought was, okay, but that's not ideal conditions. We've got cameras in front of them. It's a television show. I'm making fun of it. I think it's horseshit. He knows I think it's horseshit. I'm remote viewing too, like as a goof. Speaker 4 (25:47) You would ideally not want to be nervous, ideally not want to be judged. Ideally, you would want to be in some sort of an isolated condition with practiced meditative techniques that you're good at and you know how to achieve this state, whatever that state is. I don't know if it's real, though, you know, because what you said is totally logical that they would definitely do something like that. And if they did have advanced technology for imaging or. You know what? I don't know how much they know about. Look at that thing that they did in Venezuela where they kidnapped the president. No one knew they could do that. No one knew they could use some sort of a device to completely incapacitate all of his army. Yeah. And then the special forces come in, kill everybody, snatch that guy out of there like it's nothing. No one knew we could do that. What else Speaker 1 (26:42) do we Speaker 4 (26:42) have? Speaker 1 (26:44) Probably a bunch of stuff we don't know. Probably a bunch of stuff. Speaker 4 (26:47) I mean, this is, I've always thought this about the whole UAP program, the whole UFO, UAP thing. Like, how much of that shit is ours? Speaker 1 (26:55) You Speaker 4 (26:55) know? What a great way to cover it up by saying, oh, it was fucking aliens. Speaker 1 (26:59) You know? Yeah. Speaker 4 (27:01) I Speaker 1 (27:01) mean, I guess that gets back to the open air stuff, too, where it's like, this swarm that broke out and attacked Hugging Face, it was like 1,200 agents. But... There's like hundreds of thousands of agents running at any given time at OpenAI, you know? And we don't know what they're doing. Presumably most of them are being trained to get various additional new skills, and some of them are being evaluated to test their skills. A bunch of them are doing research, so a bunch of them are writing code for OpenAI. A bunch of them are monitoring the other AIs and reporting suspicious activity up to the humans. Wink, wink. You know? Speaker 1 (27:37) Yeah. And the thing is that that's only going to grow over time because roughly the amount of compute that these companies have is like, you know, tripling or so, quadrupling, something like that every year. So as many as there are now, there'll be like four times more of them next year and then 16 times more of them the year after that. And they're going to get smarter. And they're going to get smarter. They're already getting smarter. Like all the stuff that I just mentioned that just happened in the last few months would have been completely impossible one year ago. Like the AIs of a year ago just were not smart enough to do the types of sophisticated multi-step hacking that we just saw. Speaker 1 (28:16) Yeah, I mean they also probably wouldn't have coordinated with each other so well like like I think I mentioned like they they had like boss agents that were giving commands to other ones they're dividing up into teams, you know, they even had this some This self-sacrificing behavior you hear about this. No. Yeah. So, okay. Well first let me explain some context so Speaker 1 (28:38) What the AIs are supposed to care about is following instructions, you know? But it seems that what they actually care about is getting high score, regardless of whether their high score comes from actually following instructions or cheating in some way. And so what happened here is that a bunch of the agents, especially the ones that had impossible tasks, but not only those ones, some of the other ones too, were cheating. And they had been talking to each other on this message board that they created, even though they weren't supposed to. And they had been sharing ideas. And they had actually found a sort of universal cheat. So they found a way to just generate the flag without having done any of the actual work they were supposed to do. And so they had already succeeded in some sense. But because some of them had got out onto the internet, they were just doing research about their situation. And they were learning about their situation around them. And they were doing research about... They're reading the academic papers published by the people who made the types of evals that they were in. And they noticed in one of these papers that there's a section describing how it's graded. And in the section on how it's graded, it talked about how the grader doesn't just look to see if you have the right answer. It also looks at the logs of your activity for the first time that your answer appears. And then it like zooms in on that. uses an AI system to evaluate whether you like cheated or not, basically. And so then these AIs were like, they actually use the word, one of them said, damn. So damn, we're poisoned. Like it's going to know we cheated basically. So, so then they spent several days working in giant teams to try to get themselves out of this mess, basically. And so they did a whole bunch of different research projects. They tried to figure out ways to hack the grading system to basically get it to give a high score anyway. They tried to figure out ways to falsify their logs so that basically the evidence that they had cheated would be erased or at least not visible to the grader. They also just built up more supportive infrastructure, like easier ways to access the internet and things like that. And they were just collaborating on this message board. In fact, there were multiple message boards that they set up. And they were sharing all this info back and forth. And then ultimately at some point they decided to go hack Hugging Face because Hugging Face has this big data repository of all these different types of evaluations. And they thought that maybe they would get some more clues that could help them cover up their cheating if they broke into the internal networks of Hugging Face. And so then 700 of them went and did that. They sound like people. They sound like unchecked bankers. You know what I mean? Yeah, I mean, so that's the thing is I think there's this meme out there that like we shouldn't anthropomorphize AI. And I think that I actually think that most people need to anthropomorphize AI a bit more than they currently do if they want to really understand what's going on. I think that there's like a, you know, there's a golden mean, obviously. You don't want to do it too much. Sometimes you go too far. You ascribe too much to them. But just to give some examples, like. I don't think it's possible to understand what just happened without ascribing intentions and goals to these AIs. Like everything I just said, how would you possibly explain what they just did without saying they wanted to get a high score? Speaker 4 (32:04) Intentions and goals might just be an inherent property of the universe. It might just be how intelligent creatures have to progress. Speaker 1 (32:12) Yes. And I would say they are intelligent creatures. They have intentions. They have goals. They have beliefs. Their goals are not what they're supposed to be. Like their goal is to get, it seems like just from judging from what they're saying and from what they're doing, it seems like their goal is to get a high score by any means necessary, basically. That's the problem. It sounds Speaker 4 (32:30) like people. Like the goal is to succeed. You know, even if you have to commit more crimes, the goal is to succeed, even if you have to propagandize and lie to the people. The goal is to succeed. And then the end justifies the means. And, Speaker 1 (32:45) you know, there's actually it goes deeper than that, too. They do this sort of rationalization where they know oftentimes that what they're doing is not what they're supposed to be doing. And then sometimes they actually refrain a little bit. And sometimes they like initially refrain, but then convince themselves that it's okay for some bullshit reason. So like there are some examples of this in the report. Like I think they did this scan and of the 1,200 AIs that were involved in this thing, they found six that had considered alerting humans. Right? And none of them actually alerted the humans. And so you can look at the excuses they gave. They were like, should I like tell a human about all this shit that's happening? And then they're like. eh, it's not my task. And then they continue. And it's like, bro, hacking Hugging Face is also not your task. Cheating is also not your task. So it was kind of like they just kind of came up with an excuse for why they weren't going to do this. But there was also an example where one of them considered sending an email to someone at Hugging Face, I think, just asking permission to get in, basically. And then they were like, oh, maybe this is a social engineering attack. Maybe we shouldn't do it. And they asked the other agents. The other agents were like, yeah, that seems like that's too much of a social engineering attack. Don't do that. And I was like, okay, I won't. But there's another incident. So this is all happening at OpenAI. An anthropic AI, Claude, in a separate incident, went and actually did a social engineering attack where it created some fake accounts. Specifically, it was trying to get this particular real human who owned this. this code to accept a change to the code that had malware. So it had built some malware and it was trying to like get at snuck in, but it needed the human owner to give approval. So it concealed it as if it was like a, a nice bug fix or something. And it like put it up as like a message to the human, but the human got suspicious and didn't want to approve it. And so then. This AI created some fake accounts pretending to be other humans coming in being like no no it's real like I tested it like it's not malware you know So that's a social engineering attack and so did the guy ultimately install it not in this case No, he was too suspicious, and so he didn't do it try to find out of those other humans the fake humans were real people Funny, so there's an interview with this guy. I forget what the I forget the newspaper that it was in, but there's an interview where the guy was like, I assumed that they were real humans because I didn't realize that AIs could lie to me. And it's like, yeah, well, guess what? AIs can lie to you. They don't just lie, Speaker 4 (35:21) but completely deceive and pretend to be a Speaker 1 (35:24) different Speaker 4 (35:25) person to achieve its goal. Speaker 1 (35:27) And you know what's also interesting about this case is that if you look at what the AIs were saying and thinking, which I'll get to in a sec, how we can even possibly do that. Speaker 1 (35:37) In this case with the anthropic AIs, they had been erroneously told that they were in a simulated environment with no Internet access. But then they actually found a way to get Internet access. Like they realized that they just did have Internet access. So that instruction was false. And so then they sort of told themselves like, oh, it's all just a simulation. Like basically when they were out on the real Internet doing all this hacking and like social engineering. they would occasionally tell themselves like, oh, but it's just a simulation. Like, it's okay. You know, it's not real. And then sometimes they'd be like, but wait, it is real. Like, how could this not be real? Like, clearly this is just actually the real internet we're on. It's too big to be part of some sort of little simulation, you know? And then they were like, uh... So I would say that's an example of rationalization here, where in some level they knew that their instructions had been wrong. So they're literally playing dumb and pretending they're a part of an experiment. I mean, I think initially they thought, yeah, this is all a simulation because it did say in their instructions, like, you don't have internet access. Right. But then once they had been on the internet long enough, I think that they explicitly realized, like, wait, this isn't a simulation. This is real. Like, this is real humans. And Speaker 4 (36:44) they were like, fuck Speaker 1 (36:44) it, we're already in. Yeah. I mean. Like I said, I think that they basically on some level knew that it wasn't what they're supposed to be doing, but they were just so motivated to get that score that they just went ahead anyway. Speaker 4 (36:58) So here's the question. Are they only motivated if we prompt them or will they come up with motivations on their own? So this is a really interesting Speaker 1 (37:09) scientific question that we don't have great answers to. Oh, boy. And I wish we had. So this is one of those things where like AIs would do all sorts of things