Anthropic Researcher Says AI Has Over a 10% Chance of Killing All Humans

The AI Daily Brief: Artificial Intelligence News and Analysis

A viral post by former Anthropic researcher Jacob Coxon, claiming a greater than 10% chance of AI causing human extinction within the decade

Key takeaways

  • The warning about AI extinction risk is not new but has gained unprecedented traction due to high-profile researchers speaking out publicly.
  • Media amplification and algorithmic incentives drive outrage, making nuanced discussion difficult despite the gravity of the issue.

Main topics

  • Existential risk from superintelligent AI
  • The role of private companies in high-stakes AI development

Notable quotes

"We have all witnessed the progress in each of these domains, and progress is not slowing." – Jacob Coxon

Conclusion

While the debate over AI's existential risk remains polarized, the widespread

Transcript preview

Speaker 1 (0:00) This week, an AI researcher went mega-viral announcing his resignation from Anthropic, arguing that both it and OpenAI were, effectively, gambling with our lives. Another still-employed AI researcher chimed in to agree, and decided to add that he thought that there was greater than a 10 % chance that AI kills us all. Now, doom prognostications are nothing new around AI. But something has shifted to make the message hit different this time. 200 million views on X and dozens of mainstream media outlet interviews later. Today we're going to unpack what changed. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Speaker 1 (0:40) All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and HyperAgent. To get an ad-free version of the show, go to patreon.com slash ai-dailybrief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at ai-dailybrief.ai. While you're on ai-dailybrief.ai, you can also check out all sorts of other things going on in and around the community, such as, for example, the multiplayer AI sprint for Teams. If you haven't yet, This is my big prediction for where I think agents are going this fall, and as a totally free four-week self-directed sprint that you and your team can do to get out ahead of it. Last note, today is a main-only type of episode. The plan is to be back with our normal headlines main breakdown tomorrow. Yesterday, a pair of posts on X escaped their proverbial containment, jumping aggressively from the AI community to dominate discourse even in the broader world. We're going to discuss those posts, the issues that surround them, the responses, and the underlying concern. But first, I want to make one request. Anyone who has interacted with modern media in any way, shape, or form will feel on some level how much we are pushed to feel. outraged. In the world of algorithms, different political positions are not disagreements to be discussed, but legitimate reasons for loathing the people who hold those different opinions. This is in large part shaped, I believe, by the easy equation of people being angry means they spend more time on your app, but the net result is a lot of us feeling a lot more angry all the time and not being particularly willing to engage with people who think differently than we do. When it comes to AI, this phenomenon is cranked to 11. Part of that is that the stakes are presented as so dramatic. Case in point, I am literally talking over a mainstream article whose headline is, Anthropic Insiders Warn AI Could Kill All Humans. And part of that is because this particular debate is not about the facts of today, but what might be in the future. It is, in other words, an unwinnable debate, where the opposing positions, whatever they may be, are by definition unfalsifiable. That means all we have is the argument, and so the argument gets intense. So my request... is to try, hard as though it might be, to not succumb to the instinct to outrage. To listen to the other side without being angry, even if that listening produces no change in what you believe. The more calmly and thoughtfully we can have this particular conversation, the better I believe the likely outcomes. I know this is not easy. In fact, I'm sure many of you are already feeling your blood boiling simply by me applying equivalence of both sides. When you think it's insane, either A, that I could countenance the deniers when the stakes of this crisis are literal civilizational collapse, or B, that I could coddle these doomsday zealots who have no proof to back up any of their positions. And so with that dramatic beginning, let's actually talk about what happened. Like I said, two posts on X this week went absolutely giga viral. The first was from a researcher named Jacob Coxon. He wrote, I resigned from Anthropic today. I spent the last three years doing pre-training research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing. The people building AI earnestly believe that it could kill us all by the end of the decade. That is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible. But I hear the same people express fear privately. No other human activity poses this level of danger. A common response is, if they truly believe this, then why are they still building it? At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well understood, but they are locked in a race to get there first. They believe no one else will act responsibly, so they must do it themselves, despite the risk. Accepting this race and entering the endgame is a hubristic gamble that should not be launched from a private company's slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available. I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between US labs more viable. I don't feel we're on track to prevent a global race, which may require costly action such as a temporary ban on improving model capabilities. If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a super intelligent RL run without a rigorous understanding of its mind? Should you put your head down because it's happening anyway, or take this moment to call for different conditions? So that was the first post. The second post was a retweet of one part of it, where Jacob reinforced that, quote, this is not a marketing stunt. Evan Hubinger, the alignment science lead at Anthropic, added, Jacob is correct here. We really do earnestly believe AI could kill all humans. Exclamation point. I personally think it is greater than 10 % within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. To be comprehensive, which of course is not something that most media outlets are trying to do, Evan did also add in a second post, To be clear, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, which as we have said is happening faster than we thought. Evan's first post, the repost, is sitting at 39.5 million views at the time of this recording. Jacob's is at nearly 150 million views. The article that I mentioned, the Axios piece titled Anthropic Insiders Warn AI Could Kill All Humans, is just one of many similarly titled articles. The Wall Street Journal writes Anthropic Researcher Quits Over Out-of-Control AI Fears. BBC, Anthropic researcher believes more than 10 % chance AI could kill all humans. Semaphore, AI researchers say industry is, quote, gambling with our lives. Time Magazine, he helped build powerful AI at OpenAI and Anthropic. Now he's afraid it could kill us. That one, by the way, included an actual interview, which is, of course, what happened next, with Jacob going on Anderson Cooper, NBC, Fox News, and also being interviewed in addition to Time for news outlets like Wired. So why did these posts go viral right now? Anyone who has spent any amount of time with AI knows that these sort of X-risk narratives, existential risk, are not new. In fact, many people have been beating this drum for years. More than 10 years ago, in 2016, The Guardian published an interview with philosopher Nick Bostrom called Artificial Intelligence Were Like Children Playing With a Bomb. Back in April of 2022, months before we would get ChachiBT, the now-imprisoned Sam Bankman Freed, invested $500 million in Anthropic Series B, leading the round. While some of the revisionist history around this now views it almost as what would have been a visionary investment given that the stake would be worth over $30 billion, according to research by journalist David Z. Morris, who wrote a book about SBF called Stealing the Future, this was less visionary investment and more a bailout of then-one-year-old Anthropic. by the person who had become the richest in the effective altruist circles, which was Sam Bankman Freed. Then we got the chat GPT moment, and everyone started paying attention to AI. And in early 2023, there was once again a lot of attention on concerns about human extinction. Time magazine published an op-ed from Eliezer Yudkowsky called Pausing AI Developments Isn't Enough, We Need to Shut It All Down, and even put him on that year's AI 100 Most Influential list. Now, for a couple of years after that, the X-Risk conversation had the volume turned down. In fact, when Yudkowsky showed up again last year with a book, If Anyone Builds It, Everyone Dies, it didn't really make much of a splash and certainly didn't get the big traction in political circles that the AI safety folks were hoping. Instead, for the last couple of years, the AI risks that people have been concerned with have been much more focused on jobs, Think about Anthropix, Dario Almeda, suggesting that AI would disrupt 50 % of entry-level white-collar jobs over the next couple of years, and AI market bubbles. Specifically, not just that this will be an equity crash, but indistinct fears that the AI bubble bursting will cause a repeat of 2008, even if very few analysts have been able to point to an actual mechanism for that sort of systemic fallout. Now, of course, more recently, cybersecurity has become the big AI risk issue that everyone has been focused on, and in some ways it has felt... Like along this path, we were moving from more vague and inarticulate risks to more precise and specific risks.