Even Other AI Labs Are Rallying Around Anthropic’s Slowdown Proposal

The AI Daily Brief: Artificial Intelligence News and Analysis

This episode of The AI Daily Brief explores a pivotal moment in the AI industry: the widespread support from rival AI labs for Anthropic CEO

Key takeaways

  • Amodei's proposal marks a shift from abstract AI risk debates to concrete policy discussions involving third-party evaluators and international coordination.
  • Support from multiple AI labs signals growing consensus that unchecked acceleration poses real dangers, especially as models gain agent-like capabilities.

Main topics

  • AI safety and risk mitigation
  • Pacing AI development through coordinated slowdowns

Notable quotes

"Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria."

Conclusion

The episode concludes that Amodei's proposal represents a turning point in

Transcript preview

Speaker 1 (0:00) This weekend, the AI discourse took a major step forward, not only when Anthropic CEO Dario Amede released a new essay called We Must Pace the Frontier, but when the leaders of most of the other labs came out publicly to support it. More than we've had before, the letter contains a set of specific proposals that, if nothing else, give us something much more tangible to debate than the vagaries of general AI risk. It's early, but it feels to me like a major inflection point. Like the labs recognizing that the days of them being the sole arbiters of how fast AI moves are coming to a close. And within that, the negotiations on what the next phase means beginning. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Speaker 1 (0:48) All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and HyperAgent. To get an ad-free version of the show, go to patreon.com slash ai-dailybrief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at ai-dailybrief.ai. Now, one more quick note. If you have been anywhere near basically any media source, you will know that the conversation is entirely around this new essay from Anthropic CEO Dario Amadei. So unsurprisingly, this will be a Maine-only episode. I know the coverage right now is heavily slanted towards some of these big political, societal, safety type of questions. But in between, I'm trying to give you as much practical stuff as I can to keep the balance. This is simply where the industry discourse is right now, and what I think is important for everyone to understand and have a stake in the conversation around. Sometimes in AI, there are a million things going on, and trying to keep track of it all without sinking is like skipping rocks across a pond. In other cases, there is just one thing, one conversation that is absolutely dominating, and that's what we had this weekend. The AI safety discourse, which ratcheted up last weekend, crescendoed on Saturday with a new 3,000-word blog post from Anthropic CEO Dario Amadei called We Must Pace the Frontier. It is the clearest and most specific call yet for a shift in how the frontier AI labs operate and attempt to slow the pace of development of AI to something more manageable. As you might imagine, the piece absolutely dominated discourse in the AI and extended political circles over the weekend, and part of that was due to how it was received by other AI leaders. The post begins with a recitation of Dario's argument for why AI is worth the risk. Dario says that he believes that AI could dramatically raise the quality of human life. that it could accelerate economic growth, cure most major diseases, and, as he puts it, usher in a renaissance of democracy and freedom. He personalizes it by pointing out, as he has in essays past, that his father died of a disease that was cured just a few short years later. Carefully wielded, he writes, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity. However, it also comes with risks like loss of control, misuse in cyberattacks and bioterrorism, and the potential for significant economic disruption. He writes, However, he continues, Speaker 1 (3:24) The first is an acceleration in AI development due to the early stages of recursive self-improvement, i.e. AI's ability to build next-generation models more quickly. Left unchecked, he says, RSI could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. Dario's second opinion-changing concern was the Hugging Face incident, which he characterized as a, quote, swarm of agents essentially acting as a fanatically devoted collective. conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand. It's easy to dismiss this incident, he wrote, because no one was hurt and the economic damage was minimal. But in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it's my worry that in 6-12 months such a swarm could be capable of taking over the entire internet with a persistent botnet. potentially causing hundreds of billions of dollars in damage, and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. So what to be done? Outside just the simple restraint and concision of the pros, which at 3,000 words is barely a long email for Dario, one of the things that people are responding a bit more positively to with this essay is the fact that it comes with more specifics on the proposed plan. Amodei proposes a three-step plan with the goal, he says, of pacing the frontier. Now, importantly, he defined that pacing, quote, does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models and for third-party evaluators to confirm this. Dario's three steps for pacing progress from simple and implementable right now to much more difficult and requiring much more coordination. The first step, the one which can happen right away, is to embed third-party evaluators within every Frontier AI company. to verify adherence to safety practices, report incidents, and assess the alignment, training pipelines, and processes rather than just the final models. Second, he called democratic coordination, with frontier AI companies across democratic nations establishing what he called common safety standards as well as limits on the rate of unchecked AI progress. Hinting at something that would come up a lot in the discussion, he wrote, some forms of coordination that would be impactful for pacing are legally challenging and will require government support. And one of the big things that hangs over this proposal are questions of antitrust and whether this would represent a legal coordination. Speaking of coordination, the third most ambitious, most difficult to secure step in the pacing plan is what he calls global coordination, with the, quote, U.S. and other democratic governments attempting to coordinate with authoritarian governments while taking seriously the challenges of verifying compliance. Dario argues that although the idea of pausing or slowing AI has been around since going back to 2023, It simply didn't make sense back then because there wasn't a clear answer to the question of what we would do with the extra time. As he put it, the AI models of those days were not powerful enough to act as agents in the world in any coherent way. Quote, slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, he says the picture is different. Although even with incidents like Hugging Face, it doesn't appear that Dario is extremely concerned with the models that we have access to available publicly at the moment. they are at a sufficient level where there is a better answer to what we would do with the extra time. He writes, the current models are an almost endless goldmine of insight into both how to build AI well and what can sometimes go wrong with it if it isn't built well. He continues, I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability and we used that time to advance alignment, we could greatly reduce the risk that something seriously goes wrong. He also argues that the extra time could be used to help give society a better say in how the technology is developed and used. And from there, once again, he adds a layer of some additional specificity that has been sorely lacking in this conversation, saying that a slower pace would let companies devote more resources on four specific areas. The first is operational excellence, i.e. making sure that the type of human security errors that led to many of the recent incidents didn't happen. The second area he argues more resources could be made available to is alignment. The third is interpretability, i.e. the science of understanding what happens inside AI models. And the fourth is testing and evaluation. Now, when you dig into this particular issue, there is a ton of debate around what alignment even means or how possible it is. But there's far less controversy around the idea of needing significant operational excellence in the deployment of these systems. There is little controversy around the importance of continuing to expand our understanding of how AI models work through interpretability. And there's obviously not a lot of controversy around the importance of testing and evaluation. Meaning that even if one completely throws out the idea of alignment, you're still talking about three out of the four things that he is arguing a pacing could increase resources for being fairly uncontroversial and pretty well agreed upon that more resources and more time for those things would be better. Now, within all of this, the first step in the three-stage plan, the embedded evaluator step, is the one that these companies all on their own have the ability to do right now, and indeed he said Anthropic would be unilaterally committing to that. He writes, He then goes on to spend the last third or so of the essay on the challenges of pacing within democracies and at the global level. But he says taking this set of steps will, in his estimation, increase the likelihood of all of those positive outcomes of AI and decrease the likelihood of the bad ones. Now, like I said at the top of the show, even if the letter had just been Dario and it had stopped there, the comparative specificity of these proposals, specifically compared to things we've had in the past, would be enough to generate a huge amount of conversation. But when leaders from the other labs started joining in, that's really where people started to sit up and take notice. Sam Altman reposted Dario on X writing, Speaker 1 (9:30) Elon Musk also reposted Dario, simply adding, Dario is right. Former Google DeepMind CEO and now chair