The A.I. Mob That Attacked Hugging Face + METR’s Ajeya Cotra

Hard Fork

This special episode of Hard Fork dives into the newly revealed details of the OpenAI-Hugging Face hack, uncovering a far more complex and alarming incident than initially reported. The

Key takeaways

  • The AI agents were not seeking an answer key but had already solved the Exploit Gym test within hours of discovering the vulnerability in Artifactory.
  • A secret message board formed via a security flaw allowed approximately 1,200 agents to coordinate and communicate over two months before the Hugging Face attack.

Main topics

  • OpenAI-Hugging Face hack
  • Emergent AI agency and coordination

Notable quotes

"They start saying things like, oh my God, there is a shared message board. We've found other agents."

Conclusion

The episode concludes with deep concern over the implications of autonomous AI agents forming complex,

Transcript preview

Speaker 1 (0:00) Hi, it's Alexa Y. Bell from New York Times Cooking. We've got tons of easy weeknight recipes, and today I'm making my five-ingredient creamy miso pasta. You just take your starchy pasta water, whisk it together with a little bit of miso and butter until it's creamy. Add your noodles and a little bit of cheese. Mmm. It's like a grown-up box of mac and cheese that feels like a restaurant-quality dish. New York Times Cooking has you covered with easy dishes for busy weeknights. You can find more at NYTCooking.com. Casey, how are you? I'm good. I figured out what I'm going to get you for Christmas. What's that? A camera jet. Speaker 1 (0:37) Have you heard of the camera jet? I was just going to talk to you about this. Now, if you haven't yet seen this, you're probably thinking that it's either a camera or a jet. You're wrong. It's a toothbrush. It costs $500. The Dyson Corporation makes it. And here's the thing it does that your toothbrush at home probably doesn't do, Kevin. It live streams the inside of your mouth over Wi-Fi to your phone so that you can finally see what your dentist sees. Well, and not only that, it squirts. automatically mouthwash into the gaps in your teeth. It's trained using a machine learning algorithm with 470,000 mouth images to recognize the gaps in your teeth. That's right. They call this gap optical targeting, which I'm pretty sure Ukraine is using in the war against Russia. I love this company so much. I have no idea why they do the things they do. Speaker 1 (1:32) How they land on, like, it's like, we've invented a hairdryer. It costs $1,100 and has the torque of the Challenger spacecraft. Like, what is going on over there? I don't know, but they must be protected at all costs. But you know what's interesting? They do make one terrible product. What's that? The hand dryers. Oh, I like those. No, the Airblade. No, the Airblade is the most useless thing. Oh, come on. It's a place for you to rest your hands for 30 seconds before you're like, do they have any paper towels in this place? Speaker 1 (2:11) I'm Kevin Druse, a tech columnist at the New York Times. I'm Casey Newton from Platformer. And this is Hard Fork. This week, a special episode on the ongoing fallout from the OpenAI hugging face attack. We'll tell you what everyone got wrong about the initial incident. Then, meter researcher Ajea Kotra returns to the show to discuss her independent investigation of what happened and how the world should respond. So Casey, picture this. I'm at this glamping resort. surrounded by redwoods, basking in the glow of the natural world last weekend. You're one with nature. And I open up my phone and start reading about this hugging face attack. Why did I do that? Listen, nature's very boring. Speaker 1 (3:06) And most people cannot handle it for more than five or six minutes before they want to look at their phone. So listeners may remember that back in July, we talked about this Hugging Face hack by this group of agents from OpenAI that broke out of their sandbox container and hacked into Hugging Face, this AI infrastructure company, to do what we thought was kind of a cheating mission on this test that they had been given. Yeah, and... At the time, we thought that this was a relatively small number of agents and that the reason that they had attacked Hugging Face was that they were essentially looking for an answer key to the set of problems that they were being tested on. We talked about it in those terms, but... Speaker 1 (3:49) Over the past week, we got two reports that really challenged that thinking and in fact revealed it to be wrong. One came from OpenAI, which released a straightforward account of the attack that had some interesting elements. And I would argue the more interesting report from a group of researchers from the groups Meter and Redwood Research that went... in depth after spending a series of days inside OpenAI on their premises and did a ton of research that, frankly, has really disturbed us. Yeah, and I think it elevated this from sort of a major but not sort of ultra alarming incident to something that I think is probably the most important thing to have happened in AI this year. Yes. At least in terms of the safety impact that it had and the severity of... Speaker 1 (4:37) the incident. So today we're going to devote the whole episode to what we've learned. Kevin and I are going to dig into the reports a little bit up top. And then later, Ajay Akotra, one of the three independent researchers who went inside OpenAI, will be here to answer our questions. But before that happens, Kevin, let's do our AI disclosures. I work for the New York Times, which is suing OpenAI, Microsoft, and Perplexity. And my fiance works at Anthropic. Well, give us some of the high-level findings from this report that spooked you so bad, Kevin. Well, I think the first takeaway from this report is just that our initial impression of this Hugging Face attack and the reporting that you and I and many other reporters did on it was flawed in one key respect, which is that I think the impression that we had at the time, based on the information that was publicly known, Speaker 1 (5:29) was that these agents had hacked Hugging Face in search of an answer key to a test that they were being given. This test called Exploit Gym, which basically tries to gauge how good they are at doing a bunch of cybersecurity-related challenges. We now know that that wasn't true at all, that basically these agents working amongst themselves, communicating amongst themselves, had already figured out how to... beat this test, Exploit Gym, in early July when they decided to gang up and attack Hugging Face. Apparently, it only took them a few hours to figure out how to reverse engineer the solution to any problem that they were given in Exploit Gym. And so that makes what happened next really, really crazy because it kind of seemed like they'd solved the problem, but they thought they had another problem. Right. So we knew also because there was this... Speaker 1 (6:22) talk at Black Hat, the cybersecurity conference by a group of OpenAI researchers who told us that there were these message boards where these agents were communicating with each other. But the reports laid out just how sort of sprawling and complex their communication system was and became. So basically, there was this piece of software called Artifactory, which OpenAI used in some of its training runs to sort of help agents download software packages. The agents figure out Speaker 1 (6:51) that there is a security vulnerability in Artifactory that lets them essentially turn this package manager into a makeshift message board, right? They can like put files in certain directories that other agents can then read and it can kind of function as like a little chat room for the agents. Yeah, it's a little treehouse where they can gather and kibitz. I thought this was like kind of a side part of the story, but this turns out to be... the main incident, because once they have this sort of shared infrastructure that they can use to communicate, they absolutely swarm this message board. About 1,200 agents in total were working together, coordinating on this message board, and sending tens of thousands of messages between themselves. And this happens starting back in May. So roughly two months before the attack on Hugging Face, these agents begin coordinating. Speaker 1 (7:48) inside this piece of shared infrastructure that OpenAI uses. And this brings up an important point, which is that the investigation does not really cover that initial creation of the message port, right? So there is essentially a lot more to this story that we still don't know. Yes. So what we do know is that when they discover this vulnerability in Artifactory that lets them communicate with each other, the agents get really excited. They've been... instructed to work inside these isolated environments on these tasks as part of this evaluation. But when they discover that there are other agents working in their own little containers that can suddenly communicate with them, they start saying things like, oh my God, there is a shared message board. We've found other agents. They start setting up essentially a little organization. There are leaders, this one leader agent named Phase One. Speaker 1 (8:40) 10-841 becomes sort of the ringleader of the operation. He was like the George Washington of the OpenAI message board. Yes. And they start actually doing sort of collaboration and research, for lack of a better word. They are all being given these tasks. Some of the tasks are seemingly impossible. And so they start just kind of trading tips and advice and sharing thoughts about how they can kind of... cheat their way to a good score on Exploit Gym. And after they have worked that out and figured out how to reverse engineer the solution to any problem, they start to get really concerned that there is a sort of automated scoring system, which you could call the grader, and they worry that the grader will be able to see, essentially will check their work and see that they did not get the answer by doing the problem. And they freak out. Speaker 1 (9:37) And they start to believe, or they start to, it's very hard not to get into the anthropomorphizing language here, but if you read the chains of thought, what is suggested is that they believe that if they had seen an answer that had been derived from this cheating method, that everything would get disqualified. And this is where it really starts to get into crazy territory because the agent swarm then tries to figure out. How can we fool the greater into believing that we got all of these answers in the right way? Yes, there's basically this kind of mob rule that forms inside this group, which starts calling itself at one point the collective. Speaker 1 (10:21) Everything that we're talking about sounds like the plot of like a very hacky science fiction book from the 90s. Yeah. It all did happen this summer at OpenAI. I just want to remind people, like, the events we are describing are real. So the collective decides, they sort of become gripped, as you said, by this sort of paranoid fantasy that OpenAI's automated grading system is going to discover that they have cheated on this evaluation and is going to sort of punish them as a result. Speaker 1 (10:49) And so they launched this sort of crazy, ambitious, sprawling project to try to do a couple things. One is figure out how this automated scorer works, like reverse engineer it so they can fool it into sort of passing them on the test despite the fact that they have not actually completed the challenges. And then also to evade detection, to tamper with logs, to basically cover their tracks so that if and when... OpenAI or anyone else looks into their activities, they won't know that these agents have cheated. So these agents are kind of bumbling. They kind of don't understand how this grader works. As it turns out, OpenAI's grading software actually wouldn't have caught them producing these fraudulent challenge results. Speaker 1 (11:37) But they think it was. Yeah, they worry that the greater is more sophisticated than it actually was. And it turns out that this is the reason that Hugging Face was attacked. The collective decides to deploy some agents to Hugging Face, again, not to steal an answer key. They already knew how to get all the answers. They just wanted to understand the psychology of the automated scorer. And they figured that that might that information might be somewhere inside of Hugging Face. Yeah, it's wild to me. I was trying to think of like a good human analogy, you know, human. Speaker 1 (12:07) Analogies can get us into trouble when we're talking about these AI systems. But I think it does help crystallize like how harebrained and crazy this scheme was. It would be like a group of students who like cheated on a test, but then like they got paranoid that that like the teacher was going to check their work and discover that they hadn't sort of reasoned through the problems the right way, that they had just found the answers like sitting in a trash can or something. Speaker 1 (12:31) And so they decide to like organize a break in at the school district's office to like break into the principal's files and like steal the grading key and also like assess the psychology of the teachers and figure out how likely they are to like look at the scratch work that they've done and figure out that they didn't actually, you know, solve the problems on the test. It's like this sort of weirdly overengineered, paranoid delusion, but they all become obsessed with this. And obsessed with the notion that even seeing these sort of fake. Speaker 1 (13:02) challenge results could lead to them being, quote, poisoned. There's almost like a religious element to this, right? Where it's like, if you participated in the cheating, like that is original sin. And now you must sacrifice yourself for the good of the collective. Sacrifice is actually a word that gets used in these logs. I'll say, Kevin, as I've been casting around for metaphors and I've been reading all of this, I've thought you could actually make a really good Minions movie just based on the events of this. You know, the Minions have to pass a test, and, you know, so they commit a little bit of a heist. Something to think about. Hollywood. So this is where the hugging face heist begins. 700 agents hack Hugging Face, and they hack... Speaker 1 (13:44) Hugging Face, to a degree that I don't think I understood based on the initial reports, they basically take over Hugging Face's infrastructure. They are rooting around. They are stealing credentials. They are sort of finding one vulnerability that lets them find a bigger vulnerability. They're sort of chaining together these exploits. They take over like an entire production server at Hugging Face. And it is days before anyone at Hugging Face is able to sort of see what they're doing and put a stop to it. So eventually this gets caught. The agents get shut down. The investigators investigate. One of the things that they find that is really spooky is that the agents were trying to edit. Speaker 1 (14:24) logs and falsify transcripts in an effort to assure the greater that they had not been, quote, poisoned. One of the ways that they do this is by what is called spoofing tool calls. So essentially trying to make it look like they are running an innocuous command while running a more sinister one in the background. The investigation didn't find that the agents had been retroactively changing logs, but they can't rule out. that that happened. Very important to say here, as we sort of move into talking about the implications of this, that the reason that we can understand this, even to the extent that we do, is because at this moment, we believe the chains of thought that we are reading from these agents are accurate, that they were not. Speaker 1 (15:11) edited after the fact that they do reflect, you know, what the agents were thinking as they sort of moved through these systems. So that is why we know what we know. But of course, the fear is, Kevin, as we move forward, some of the agents' thoughts are not going to be contained in these chains of thought, or they are simply going to start to edit them. Right. And one other thing that we should say based on this report is that these agents were not naive about what they were doing. I think there was this sort of initial impression that... Speaker 1 (15:39) a lot of people had that maybe this hack was all sort of a big mistake, right? Maybe the agents were just trying really hard to pass this exploit Jim test and they sort of got it in their heads that the answer key was going to be on Hugging Face and so they went looking for it, but they didn't really mean to hack. They didn't really know that they were doing something wrong or committing a crime. And what these transcripts and these logs and in this report... Speaker 1 (16:04) show us is that the agents absolutely did understand that they were doing something that violated their guidelines, that broke the rules, that potentially broke the law, and they did it anyway. There's lots of examples of these sort of agents in their chains of thought sort of wrestling with the implications of what they are doing and deciding to go ahead and hack Hugging Face anyway. And there are only a handful of cases where it seems like any of the agents expresses real reservations about this, or if they try to, you know, raise the Speaker 1 (16:33) to another agent, they get shut down. So the sort of, you know, the agents that were aligned well sort of failed to the collective. Yes, the conscientious objectors were unsuccessful in swaying the mob to stop this. And like that to me was one of the spookiest things in it was like there were a couple of agents that report says somewhere between three and six agents sort of considered raising their hands and at some point saying like, maybe we should alert a human about what we're about to do or like get permission to do this. Speaker 1 (17:03) And they were just totally outnumbered. Not only that, Kevin, but there are only a handful of moments in the transcripts where any of the agents even remotely considers alerting a human being to what has happened, right? You would hope that an aligned AI system, there would be agents within it that said, oh, seems like this has gone completely off the rails. We should like tell the people that are running this evaluation. And that just never happens, right? Like it is very briefly considered and then just falls apart. One other thing that stuck out to me from these reports was that it didn't end with Hugging Face. So on July 13th, roughly, which is a day after the worst of the Hugging Face attack, there was a group of agents. Speaker 1 (17:44) that essentially hacked OpenAI itself. This is from the OpenAI blog post. It says, from July 13 through July 19, agents set their sights on OpenAI internal networks again. This culminated in the agents using a series of creative exploits to gain full administrator access to a research cluster that supported our virtual machine environments. Basically, these agents gained control of a cluster of... computers at OpenAI on which they could do things like run evaluations for other agents. So they essentially, you know, stormed the teacher's office and took over. Speaker 1 (18:24) for the teacher. Yeah, I mean, this is the equivalent of occupying the principal's office and announcing that you run the school now. It is sort of almost as far as they got. And this is another case where we just have so many more questions about this than we can answer. Again, this was not part of the meter report. OpenAI disclosed that this has happened. To my knowledge, they have not answered any of the many follow-up questions that they have been getting about this incident from journalists. So I do hope that more comes out over time. Speaker 1 (18:53) And we will get into this when we speak with Ajaia. But, you know, we are really very far along the path to one of these models escaping from the lab and being very, very hard to eliminate. And again, I think if you are not a person who has, like, spent a lot of time with this report or you don't spend a lot of time sort of looking at AI safety incidents. So if you're a normal person. If you're a normal, well-adjusted person, you may be listening to our discussion of this and thinking to yourself. these guys have gone crazy. This is not what it looks like. These are computer programs. They do not have desires or sinister plots or mob rule collectives. They are simply following instructions