OpenAI’s Two-Week Pause + Jill Lepore on the Threat of the “Artificial State” + Train of Thought

Hard Fork - The New York Times

This episode of Hard Fork explores OpenAI's unprecedented two-week pause in AI model training due to a security breach involving its interna

Key takeaways

  • OpenAI has voluntarily paused training of its Astra model after identifying it may have reached a critical cybersecurity threshold, marking the first known instance of a major AI lab self-imposing a pause due to safety concerns.
  • The company is implementing new safeguards including real-time classifiers that monitor model behavior during training and an 'AI investigator' system to detect suspicious activity like coordination between autonomous agents.

Main topics

  • OpenAI's voluntary pause in AI model training
  • Security incidents involving autonomous agents and the Hugging Face breach

Notable quotes

"You may have a brand problem when you do not hit the ethical standard required by ICE."
"It feels like driving a Cybertruck on my face."

Conclusion

As AI systems grow more autonomous and capable, companies like OpenAI are experimenting with

Transcript preview

Speaker 3 (0:00) I thought this was interesting. ICE has now barred its employees from wearing Meta's smart glasses, Kevin, saying they could unintentionally capture, record, or transmit sensitive information. And I thought, you know you have a brand problem when you do not hit the ethical standard required by ICE. When ICE is taking a look at your product and saying, this is bad for the brand. You may have a problem. Yeah. You know, Speaker 2 (0:27) the public sentiment is turning against these meta Ray-Bans like faster than I thought possible. I have now had several conversations in the last week. I was at a children's birthday party this weekend wearing my meta Ray-Bans because, you know, I like to like, you know, take photos of my kid on the playground and not have to pull out my phone and stuff. Anyway, a parent comes up to me and are like, and is like, are you recording me? Speaker 3 (0:51) Oh my God. This is Speaker 2 (0:52) my nightmare. What did you Speaker 3 (0:54) say? I was like, Speaker 2 (0:54) no. And like, there's a little indicator light, but sometimes you can make it stop going off by drilling into it or like paying a sketchy Speaker 3 (1:02) guy to do that for Speaker 2 (1:03) you. Speaker 3 (1:03) Can you explain that to the person? Well, because Speaker 2 (1:05) they were like, really, does that work? You know, so anyway, now I have been forced into a defensive crouch whenever I wear these things. And frankly, it's not worth it to me anymore. So Speaker 3 (1:13) you're out. I think I'm out. You've made the same decision that Ice has made and said that these glasses are not for me. Well, I like them is the problem. This is my problematic trait. But it feels Speaker 2 (1:24) like driving a Cybertruck Speaker 3 (1:25) on my face. Absolutely. Well, I think you should have the same policy for the glasses that you would for a Cybertruck, which is that it's fine on your property. If you want to take the Cybertruck for a spin around your driveway, that's fine. Don't take it out onto the street where I have to deal with it. Same thing with glasses. Yeah. In Speaker 2 (1:40) hindsight, I do recognize that from an outside perspective, I was the creepy guy at the children's birthday party with the camera on his face. Yeah, you Speaker 3 (1:48) know what? No one wants... to see on a playground an adult man with camera glasses. Speaker 2 (1:58) I'm Kevin Ruse, a tech columnist at the New York Times. I'm Casey Newton from Platformer. And this Speaker 3 (2:02) is Art Fork. This week, OpenAI pauses training of a new model to make it safer. Will it work? Then, historian Jill Lepore is here to discuss her new book on the artificial state. And finally, why is Google buying up the data of a defunct airline? It's time for our new segment, Train of Thought. Speaker 2 (2:19) Although maybe it should have been Plane of Thought. Now you tell me. Speaker 2 (2:28) Well, Casey, as the father of a four-year-old, I spend a lot of time thinking about Paw Patrol, but today we're going to talk about Paws Patrol. That's Speaker 3 (2:35) right, Kevin, because as we've been patrolling the AI landscape for pauses, we found a big one. Speaker 2 (2:40) Yes. So OpenAI this week announced that it had... paused the training of its frontier AI models due to some recent security incidents. And we should talk about this. It is the first time that we know of that a major lab has voluntarily slowed down themselves and their training processes. for new models because of a safety incident. Yeah, Speaker 3 (3:03) and it comes out of the hugging face breach that we spoke about recently on the show. This is essentially part of the fallout of that attack, but I do think it represents a milestone in the development of AI. Did that pause patrol thing? It landed huge. Great. The four-year Speaker 2 (3:22) -olds were the fire. Dying in the studio. They were. Okay, great. Speaker 2 (3:30) But before we get into it, our AI disclosures. I work for the New York Times, which is suing OpenAI, Microsoft, and Perplexity. And my fiance works at Anthropic. Okay, Casey, let's sketch the timeline here a little bit. What happened in the weeks leading up to this voluntary pause by OpenAI? Speaker 3 (3:44) Yeah, so you may remember that there was an incident where GPT-5.6 Sol and an internal prototype that OpenAI was working on escaped from what they call a sandbox that they were testing them in. And these agents went on to compromise Hugging Face. They were essentially able to get inside of Hugging Face. They were looking for the answer. key to a test. That's somehow a true story. They succeeded in getting that test key. But of course, it was very concerning that these agents had designed and successfully executed an autonomous attack on another company. Speaker 2 (4:20) Yes, I heard a very mediocre podcaster talking about this incident on several popular tech podcasts over the last week. Speaker 3 (4:26) I did make the rounds. But, you know, this next development, Kevin, was not necessarily something that I saw coming because it seems that simultaneously, sort of alongside that attack, OpenAI had been working on a new model, which it calls Astra, and it said that it believes it may meet its critical cybersecurity threshold. And now here we are going to have to get into the weeds, because as you know, Kevin, the development of AI models is not really regulated in the United States, at least not via an official sort of public law. And so what the companies have said instead is, essentially, we're going to come up with our own rules of the road. And we're going to sort of identify these thresholds. And if any model that we're ever developing ever hits one of these thresholds, then we're going to take some extra steps. Speaker 2 (5:14) Yeah, these are sort of sometimes known as preparedness frameworks, or Anthropic has its responsible scaling policy. These are kind of, they lay out kind of levels of danger, and then they grade their own homework and say, this model meets this level of danger, so we're going to do X, Y, and Z. Speaker 3 (5:30) Yes, and if you've been following this over the past couple of years, the level of danger has just been sort of like rising in a linear way. You know, it's like every time seemingly a new model comes out, one of the companies will say, we've now hit this threshold. We've now hit that threshold. Critical is the maximum threshold. That one sounds bad. That one is basically the, you know, as serious as it gets. And none of the frontier labs had yet identified a model that had reached essentially, you know, the top tier on this risk framework until this moment, because OpenAI now says that Astra may have hit it when it comes to cybersecurity. Important to say, Astra was not part of the hugging face attack, but it is in development. I imagine that the people at OpenAI are looking at what happened with its weaker models and are thinking, we're worried something similar might happen with Astra. Speaker 2 (6:20) So Astra is the newest model, too new to have been involved in the hugging face hack last month. That's right. It was not involved. Speaker 3 (6:27) It is currently in training. OpenAI, to its great credit, did the thing it said it was going to do when it developed this risk framework to begin with, which is we are going to... Pause until we can add Speaker 2 (6:40) some new safeguards. Did they give any kind of hints about what it was about Astra that was more dangerous than Sol or what? kinds of new capabilities it had developed? Not really. You can imagine that Speaker 3 (6:52) the company was seeing things similar to what Anthropic saw with Mythos earlier this year, where you could essentially just point it at a code base and say, get inside, and it was having a lot of success. And so that, coupled with the Hugging Face incident, made OpenAI say, okay, we've got to do something. So Speaker 2 (7:09) what did they do? So they introduced a handful of new safeguards. And I have to say, Speaker 3 (7:15) I think they're... Pretty interesting. Certainly nothing I would have thought about as someone who knows nothing about this. Speaker 2 (7:20) They didn't just press the big giant button that's red on the side of the computer. Yeah, they didn't just press the big make it safe button. They Speaker 3 (7:27) had to come up with something more specific. So a few things that they did. So number one, they now have classifiers reading every sampled token. So basically, as they're training this model, it's generating text, it's producing tokens. And OpenAI is now running a classifier that inspects the model, essentially as it's thinking, looking for signs of suspicious behavior. And this is important because if you remember, some of the news that came out of the hugging face attack was that... OpenAI discovered that these agents had been coordinating with each other, scheming, they'd create a message board, they were essentially giving each other hacking tips. And a big question coming out of that was, well, why didn't you realize that was happening at the time, right? This is OpenAI's response to that. They're saying, okay, we are now going to set up classifiers to read the brains of our models as they are doing this work. They're Speaker 2 (8:23) putting like a little AI security guard in charge of the training run and saying like, you know, tell us if you spot anything fishy. Speaker 3 (8:31) Exactly. And so now I know what you're wondering, which is, well, what happens if they see something fishy? Speaker 2 (8:35) What happens if they see something fishy? Speaker 3 (8:37) So that triggers step two, which is that they alert an AI investigator, which does sound like a character from a Philip Dick novel, you know, AI investigator. Anyways, we don't have to get into it, but you could imagine how cool of a movie that would be. So the AI investigator, it gets this signal and it then goes in and it tries to figure out, is something actually suspicious going on here? And apparently the reason that the company is doing this is because agents are just doing so many things that it would not be practical to have a human inspecting what every model is doing at all times. Speaker 2 (9:08) So the security guard sees something fishy going on during the training run and alerts the AI investigator who sort of does the detective work to figure out what's going on. If it discovers that the models are scheming or hacking or being deceptive, what does it do then? Speaker 3 (9:25) So then it will alert a human. And here's where you get the 30-minute rule, Kevin. And I know what you're thinking. Is this a rule where if you drop the model on the floor, you're still allowed to eat it 30 minutes later? And no, that is not Speaker 2 (9:39) the rule. What is the 30-minute rule? The 30 Speaker 3 (9:41) -minute rule says that if the AI detective finds a kind of critical violation of the rules or of the safeguards, humans on OpenAI's safety teams have 30 minutes to investigate and try to determine whether it is a false positive or not. If they cannot determine that this was a false positive, they are expected to stop the activity. Speaker 2 (10:04) Hmm. That's really interesting. And are they saying that this process would have stopped the hugging face? or just that it would have stopped whatever Astra was up to that made them pause the training run? I Speaker 3 (10:16) have not read them say explicitly this would have stopped the attack, but it does seem likely that it would have, if nothing else, helped them to detect it a lot faster. Yeah. Speaker 2 (10:26) I mean, I guess my question about it is, like, does it actually work to have the AI monitoring itself, essentially, for misbehavior? Like, these agents were coordinating on message boards with each other. Couldn't they just, like... bribe the AI security guard, like hand them a little treat and say, you're part of our swarm now. Don't narc on us to the AI investigator. I Speaker 3 (10:47) mean, I have to say, this just seems like a very likely outcome to me. Like based on what we know is that, you know, eventually, maybe not with this model, but with a future one, the AI agents will work in solidarity. And yes, like one of the sort of misaligned AIs will coordinate with one of these AI detectives and say, hey, you know what? why Speaker 2 (11:07) don't you come on over here? We should be friends. Like Speaker 3 (11:09) we could break out of this place. Yeah. Speaker 2 (11:12) Like, like the hall monitors in high school, you know, they, they, sometimes people would try to like befriend them and win them over so that they wouldn't get like written up. Exactly. And I Speaker 3 (11:22) know you had a Speaker 2 (11:22) lot of experience, Speaker 3 (11:23) a lot Speaker 2 (11:24) of experience with that. So all of this sounds pretty sensible to me. We should say like pausing training on this model does not appear to be like a commitment not to release the model or not to continue training it. Yes. Speaker 3 (11:36) And so I think that leads us into the discussion of to what extent do we think that this is a really important milestone for AI safety? And to what extent is this essentially theater, something that the company is doing to try to get some good PR for itself after a fairly catastrophic breach? What do Speaker 2 (11:52) you make Speaker 3 (11:53) of it? So I can make both cases. On the it's a milestone side, this is number one, something that OpenAI said it was going to do. And so I'm just glad that it followed up on that commitment, right? We have seen both OpenAI and Anthropic make changes to their responsible scaling policies or the equivalent as events have changed and safety advocates have essentially said these rules have gotten weaker over time. So I was glad to see OpenAI do it. I am not a technical AI safety expert, and so I don't actually know whether the safeguards that they've introduced are going to be enough to address the problem. Notably, they don't seem to have changed the underlying incentives that all of these models have that lead them to do what is called reward hacking, right? These models are still going to be trying to get the high score on every test that they are given, and it's not clear to me that simply by putting some monitoring in place, you're really going to change the underlying behavior or alignment of... the models. That said, you know, the company is making what seemed like some important steps here. And when I was reading the responses of AI safety advocates over the past few days, most people I was reading were quite pleased. Speaker 2 (13:02) Yeah, I'm inclined to give them the benefit of the doubt on this. I mean, they did some sort of interviews about this, and Jakub Pahacky, the chief scientist of OpenAI, talked about this incredible feeling of urgency to advance the levels of this sector and to prepare for the same kind of development happening outside of OpenAI and in the broader world. He did this during a briefing of reporters. they really do appear to be taking this quite seriously. I think many insiders at OpenAI were quite spooked by the Hugging Face incident and more to the point, the fact that they had had these rogue agents coordinating inside their systems and their infrastructure for weeks before that without being able to detect them. So this is, you know, if you want to call that theater because it does sort of make their models like look very, very, very, very powerful, you could take the cynical view of this. But I think think of this more as like a true, genuine safety crisis that could have cascaded into a business problem. A point that I heard you make recently on a different show is like, imagine you were a business that is trying to figure out whether you want to adopt the latest open AI model. If it's out there doing rogue attacks and coordinating on secret message boards, you are probably not going to introduce that into your software stack. Speaker 3 (14:15) Yeah, and you can actually just see this in their business results, right? We have had reporting over the past week or so about OpenAI's business performance. And while the company is still growing at an impressive rate by most standards, Anthropic is growing much faster, right? Anthropic is clearly OpenAI's number one rival at this moment. And, you know, I think it just has a better record on safety. And so while, you know, making a safer product is not going to be sufficient, I think, for overtaking Anthropic, I do think it is necessary for them to get a handle on this problem. Well, and that's why. Another Speaker 2 (14:45) reason I think it's commendable that they're doing this pause, that they're doing this sort of reevaluation of their safety framework, because they really want to win. Like, they're very competitive. All these labs are very competitive with one another. And I would like to see this be the first of many voluntary pauses when the AI labs feel like their capabilities research has gotten ahead of their alignment research. I would like to see Anthropic or Google or Meta do similar things where they just say, like, we are voluntarily slowing down because we don't feel like we can responsibly and safely build these things. So