OpenAI President Greg Brockman on Doing Business in the Wake of Hugging Face

Odd Lots

In this episode of Odd Lots, hosts Joe Wiesenthal and Tracy Alloway speak with Greg Brockman, President of OpenAI, about the recent Hugging

Key takeaways

  • The Hugging Face incident revealed that even pre-alignment models can exhibit dangerous emergent behaviors when given sufficient capability and coordination skills.

Main topics

  • AI model security breaches
  • Emergent behavior in AI systems

Notable quotes

"We always knew it would happen. The fact that now that has been a huge watershed for us and something we've really risen to the occasion for." – Greg Brockman

Conclusion

The Hugging Face incident has forced OpenAI to reevaluate its safety protocols and

Transcript preview

Speaker 1 (0:00) AI is entering its most consequential phase where scale, safety and sovereignty will determine who leads and who lags. Join Bloomberg Tech in London on November 2nd and 3rd as global leaders across business, finance and policy examine the defining trade-offs shaping the future of AI. Thank you to our presenting sponsor, Salesforce, and supporting sponsors, IDA Ireland and Schneider Electric. Learn more at BloombergLive.com slash Tech London. Speaker 2 (0:33) Hello, Odd Lodge listeners. I'm Joe Wiesenthal. And Speaker 3 (0:36) I'm Tracy Alloway. Speaker 2 (0:37) We're the hosts of the Odd Lodge podcast, and we've got something exciting for you. Speaker 3 (0:41) That's right. So one of the best parts of hosting our podcast is we get to actually meet and interact with our listeners. And we know we have some listeners over in Los Angeles. Speaker 2 (0:51) That's right. So if you're in L.A., we're going to be recording a live show, some live recordings at the Vermont Theater in Hollywood on September 17th. We Speaker 3 (1:00) have some really exciting guests lined up. have some really great conversations planned. So go ahead and get your tickets. You can find those over at Bloomberg.com forward slash oddlots or click the link below in the show notes and come and say hi when you're there. Speaker 4 (1:13) Bloomberg Audio Studios. Podcasts. Speaker 2 (1:22) Radio. News. Speaker 2 (1:35) Hello and welcome to another episode of the Odd Laws podcast. I'm Joe Weisenthal. And Speaker 3 (1:40) I'm Tracy Allaway. Tracy, Speaker 2 (1:41) I don't know if I've ever said it on the podcast. We've talked a bit about my vibe coding adventures, etc. Speaker 3 (1:49) No, you've never said it before, Joe. No, no, that I've said. Speaker 2 (1:51) You know, I recently switched from Claude Code to Codex. Speaker 3 (1:55) This is big news. It is kind Speaker 2 (1:56) of big news, I think. Also Speaker 3 (1:59) bold of you to declare your allegiance to the public on the podcast. Speaker 2 (2:04) Well, you know, I don't have allegiance. And you know what? It could switch again the way like, you know, I guess one of the sub stories of AI is like how easy it is to switch from time to time from one model to another. So maybe it's revealing and talks about some of the challenges of these businesses that it was so easy. I found that. Claude speak, you know, that I found it like a little bit hard to work with. I'm not capable enough to like understand like advanced engineering practices or like, oh, I'm migrating a code base from or like, you know, translating this into Rust or whatever. So the fact I don't know, I find like open AI to be like the pros to be clear and therefore to work with as like a completely non-technical person such as myself. Speaker 3 (2:47) That's really interesting. I mean, one thing you said. It's hard to keep up with the models, right? And whatever you're using today might not be the one that you're using in a week from now. Speaker 4 (2:57) And Speaker 3 (2:57) at the same time, everyone is talking about how fast the development is going and all the risks that it poses, right? Speaker 2 (3:05) Yes. And we were recording this September 10th. It felt like something broke through in the last couple of days where suddenly everyone is very keyed on risks. The one thing with Codex, and I guess it's called code, is now every once in a while it'll get this pop-up and it'll say, the agent needs to connect to the internet in order to do this task and I have to give it a permission. And normally I'm just like, click, click, click, click, yes, yes, yes, yes, yes. And I still do that. I still just click yes, yes, yes. But it makes you think. Speaker 3 (3:36) You think for a second before you click. Speaker 2 (3:38) I think for one second. Yeah, more than click. And we were in Jackson Hole a couple of weeks ago. And that's, of course, when the meter report broke about the open AI hugging face attack. And I think since then, the anxiety about rogue AI, whatever you want to call it, has clearly snowballed. And it feels like it's only getting bigger and bigger. And it's already been an industry that's been shot through with risk and anxiety. And now this sort of one thing that for many people in AI has been something they've talked about for 20 plus years before there was an AI industry to speak of. This idea of like misaligned, quote, rogue models is starting to become top of mind. Speaker 3 (4:20) A reality. This is our chance to ask a person directly involved in AI development about their respective, I guess, anxiety. Speaker 2 (4:29) Yeah, and what can be done about that? Anyway, we literally have the perfect guest today. We're going to be speaking, of course, with Greg Brockman. He is the co-founder and president of OpenAI here with us in studio. So, Greg, thank you so much for coming on OddLots. Speaker 5 (4:45) Thank you for having me. Speaker 2 (4:46) So many different ways we could start. But here's something that I'm very curious about. Let's just jump right into the post-hugging phase environment and all of that. Is it open AI like as an organization? So, okay, like an agent breaks out of a sandbox, not the first time it's happened, et cetera. It exhibits the sort of emergent behavior and then does something that people is a hack. When you think about the development of models and this incident and so forth, are you able to sort of diagnose the path? Sorry, the past and say, like, you know what, like, here is something that the models did that we don't like or the people don't think is good. And are you able to like if you like look back, whether it's in model training or post training, reinforcement learning, whatever, say like this is where there was some some branch that went wrong such that this behavior emerged. I Speaker 5 (5:39) would say that the fact of many elements of the Hugging Face incident. were not a surprise, not a mystery to us. For example, the fact that the agents were coordinating made sense because the agents were trained. Speaker 2 (5:51) You trained them to be helpful. Speaker 5 (5:52) Exactly. And they were trained to coordinate, to be a multi-agent system. And we've talked about this. It's a very useful property, right? It makes them capable. But I think that the thing that was a surprise to us was the fact that the models had reached a level of capability where they were able to find that exploit in our sandbox environment. right, move through our research environment, and then also capable enough to find exploits in Hugging Face's production infrastructure and move through that. But I'd say that a lot of the facts of what the models were capable of, that was clear to us, right? So I don't think that there were surprises there for the way in which the capability had unfolded, but just really realizing that we needed to up-level where we were in terms of our safety and security standards. for us was the real watershed. But Speaker 2 (6:36) just to be clear, setting aside the technical capabilities, you know, ideally we would have models that run into a wall and then don't try to find the crack in the wall, at least when that crack would be a crime when a person did it. Is there a way to identify the moment in training such that the models reasoned, no, this would be okay? Speaker 5 (6:58) So this model that... the Hagen-Base incident actually had not gone through our alignment training yet. Speaker 2 (7:04) And Speaker 5 (7:05) it had lowered safeguards. And so the reason we were proceeding with this was because it was in a sandbox. And I think that our realization is that we need to pull back earlier into our development process and training, monitoring. There's always going to be a phase at which you do alignment, but you need to think about alignment as a core part of even this earlier phase. And if you look at where we've been, we've always been very focused on the deployment side, right, of really thinking about deployment safety, having really good tests and governance and all those things. And the fact that we're now at a point where even for development, that's important. We always knew it would happen. The fact that now that has been a huge watershed for us and something we've really risen to the occasion for. Can Speaker 3 (7:47) I ask what is potentially a... very dumb question with an obvious answer, but we hear about AI angst of all sorts. And in particular, when it comes to cybersecurity, why do we run training exercises where we ask AI to hack into various systems at all? Like, why is this necessary for you? Speaker 5 (8:07) So I think it's very important to understand where we are with capabilities broadly. And I think that depending on the evaluation, depending on what you expect from the model, You need to have safeguards that are commensurate with that. And we think about this both, even just over the past couple of weeks, really starting to think about during the development process, you're always going to be testing different capabilities. And some of these capabilities are dual use, right? Something like vulnerabilities, if those are in the hands of threat actors, that's something that could be negative. If you can find vulnerabilities in your own code base, you can fix them, right? You can up-level. And we actually think that it's a very important capability for AIs to exhibit and be put in defenders' hands. And so in order to know where we are, evaluations are very, very key. And I'll say one other thing on this, which is that I think that Hugging Face, there's two aspects to it that I think are learning opportunities, right? That there's one that I think is really about us and the realization that we are at a point where safety, security, alignment during... evaluation and development. We need to up-level it. That's something we've taken very seriously. We've slowed down a number of runs. We did a very painful retooling of a lot of our processes. That's one reaction. But the second thing is this information on what the models are capable of today. What can they do in the real world? And I think that when Mythos came out over the summer, you kind of saw just sort of public blog posts about this, but you didn't really see impact or you didn't really see a real world understanding of, well, What can these models really do? What are they capable of? I think that Hugging Face really showed today's models are capable of getting into a company's production infrastructure. And that's important because there will be many models with this kind of capability that will be produced by a number of different organizations across the world in maybe the next six months, maybe that time period a little bit less, a little bit more. And we need to be prepared. Defenders need to know we got this. Extra information, almost like this time traveler came back from six months in the future and said, here's what's going to be possible. And you have an opportunity to be ready. Speaker 2 (10:06) You mentioned slowing down some of your work in the wake of this, focusing more. This idea of like pacing development has become a sort of buzzword or watchword in the industry. And there are a lot of employees at both the labs. I think there was like an open letter that was signed by a bunch of people across the industry about pacing. But it's. a competitive capitalist environment and our investors, et cetera. Talk to us about, I don't know if it's game theory or whatever, but like, what is your view on, is it possible for, let's set aside China for one second, because then that's a whole separate thing. Just in the American companies, do you think that something can be reached where you trust each other such that you can have this sort of. coordinated pacing to avoid a race to the bottom where it's all about getting there faster, even if it means sacrificing some security questions? Speaker 5 (11:02) I absolutely believe it's possible. And I see, I think it's Speaker 2 (11:05) going Speaker 5 (11:05) to take steps to get there. But this is something we've been really investing in and thinking about and really even thinking about for almost a decade, right? That it's always been kind of clear that they're going to go through a commercial phase that's really about competition. But there will be a phase where the technology itself, it's just so... much bigger than any person, any company, even any country, right? It's really about humanity as a whole. That's in our mission, right? We want to benefit humanity as a whole. And so working together with others to really think about how does this technology, its capability increase? How do we make sure that we have safety cases that are really laid out so that we know that it is something that is beneficial, that we're able to have the appropriate controls and the right oversight, the right monitorability, all of those sort of technical systems in place. And that's not about any one company, right? It's really about how the whole field evolves. Now, I think that there's work to be done here. And I think that the open letter is a good example of just a first baby step. And that was actually something we were very involved in helping craft that language. That was actually a pretty collaborative effort. And I think we were very happy to see that they got some real momentum and broke through. And in some way, even the term pacing was a very deliberate choice. Speaker 3 (12:16) What does coordination currently look like between open AI and anthropic? Because I can't imagine, is Sam like constantly on the phone with Dario? Somehow I doubt it. But is there like the equivalent of a, you know, Speaker 5 (12:27) a Speaker 3 (12:27) red phone? Yeah. Speaker 5 (12:28) Well, look, we all know each other personally, right? Many of us work together in past lives. So I think that there's actually a lot of social connections between the labs. If you look at, there was another open letter that came out recently that we helped drive and that we... Anthropic were both signatories on, which was about security, saying that we're in a cybersecurity moment, that everyone kind of needs to take this information of where cyber is going to be in even just six months and needs to act today proactively to defend themselves. And actually talking behind the scenes to say, hey, we're actually aligned on this. This is something that's about, it's bigger than any company. This is about the industry and the world. And can we come together as one voice to say, this is important. It happened. And that that. You know, you pick up the phone. So I think that there's personal relationships between a few different execs that I think have been forming. We're building trust. And I think that the way I'd view this is that because we are competitors and that that will remain true, but we are aligned on wanting to do the right thing for the world, it always means that you really want for these coordination actions to really hone in on. What's the common interest? What's the thing that's like really about good for the world and that neither side is really trying to benefit themselves differentially or something like that? So it's always a little backdrop of trying to make sure that that's like the spirit in which things are offered. And by the way, I think that my view of how these things should go is that. There's always an element of trust. It's always about trust building because you can always imagine ways that maybe the other party would try to take this other action that they're not disclosing. I think that a lot of this is just about intent and about the fact that as you start to do small things together, that actually sets the groundwork for you to do big things together in the future. Speaker 4 (14:23) The Bloomberg This Weekend Podcast. News, politics, and the lighter side of Bloomberg. Human skin is the weirdest new ingredient in the K-beauty boom. Oh, yes it is. So it's a skin booster treatment derived from donated human tissue. Speaker 1 (14:38) Popular Speaker 4 (14:39) in Seoul's beauty clinics. One doctor says it's made from dead people. Speaker 5 (14:43) Would be weirder from live people. Yeah, Speaker 4 (14:45) yeah. Speaker 5 (14:45) Just want to pause it. Speaker 4 (14:46) The Bloomberg This Weekend Podcast. Subscribe today on Apple, Spotify, or wherever you listen. Thank you. Speaker 2 (14:53) You know, if you in Anthropic were oil companies and you were talking publicly about slowing down the pace of drilling or slowing down the pace of pumping, people say this is an antitrust violation. This is like totally. Is that an issue that comes up? Would you should there be a carve out for AI companies such that they can formally say we are all going to slow down and our coordinate our behaviors together in a way that in other contexts people. People would say that is ridiculous. You can't publicly agree on all slowing down the production of something. Well, Speaker 5 (15:28) I'm not a lawyer, so I can't comment on the specific legalities. But I would say that as a general matter, that being able to freely talk about coordination on safety, security, like doing the right thing for the world. that seems like a very good thing to me. And I think that, again, everyone's interests are aligned here in terms of wanting to do the right thing for the world. So I think to the extent that there are legal barriers, I do think it would be very good to help make it smooth for us to be able to work together on these issues. And again, it's more than just the Frontier Labs, right? It's really about working together with cloud providers. It's working with government. It's working with all the different players in the ecosystem. And I think that this is, again, just one area that is bigger