Mythos Comes Back But Not for Everyone

The AI Daily Brief: Artificial Intelligence News and Analysis - Nathaniel Whittemore

Mythos returns under restricted access, with GPT 5.6 also limited to trusted partners. The U.S. government now controls frontier AI distribution via executive discretion, raising concerns about equity and transparency.

Key takeaways

  • Frontier AI access is now gated by government approval

Transcript preview

Today on the AI Daily Brief, the return of mythos begins, but the bigger questions remain. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. subscribe on Apple Podcasts, to learn more about sponsoring the show, send us a note at sponsors at AI daily brief. AI. Of all of the sins of this particular administration when it comes to artificial intelligence, the one that is personally most disruptive to my life at this point might be the fact that important news keeps breaking late on Friday afternoon after I've finished recordings for the weekend. And this Friday it was a big one. In a letter to anthropic, Commerce Secretary Howard Lutnik set the terms for a narrow reintroduction of Mythos. Notably. notably, the letter the letter was letter was letter was the letter was addressed not to Darrio but to Chief Compute Officer Tom Brown, who has become increasingly the main point of contact between this White House and Anthropic, and in the letter, Letnik starts to craft a path forward. Since the issuance of my June 12th letter, he writes, Anthropic has worked with the US government to address risks associated with Claude Mythos 5 and Claude Fable 5. These efforts have yielded significant progress. In addition, Anthropic has committed to work with the US government on protocols and standards and releases for these models. In light. In light of this progress, as progress, as Department of Commerce's evaluation of the diversion risks currently presented by the covered models, I have determined that appropriate safeguards are in place to permit certain trusted partners to access the Claude Mythos model. Basically, Letnet goes on to say that a certain selected handful of partners, including presumably both companies and US government agencies, could once again have access to Mythos. No one has seen the full list provided by the Commerce Department, but reports suggest that around a hundred organizations will regain that access. Still, what's clear from the letter from the letter is that, is that Interior AI models, if you were in any doubt, are now subject to a licensing regime. It's a licensing regime that hasn't been passed by Congress, established in an executive order, or even fully articulated in public. At this moment, it is a licensing model based on the whims of Howard Lutnik. Indeed in that same letter, he says, I reserve the right to re-evaluate and adjust the scope of license requirements on the covered models should circumstances change. So, presuming this is the beginning of the end, people should be excited, right? Fable 5 can't be all that far behind it. And yet excitement is not the word that I would use to describe the tone. Future forwards Matthew Berman was very upset about this all weekend. Writing, Anthropic just struck a deal with the government to allow 100 select companies and governmental agencies to use Mythos. The government in Anthropic are now deciding who uses Frontier Intelligence. Hopefully this is just Mythos and not the standard for all frontier models going forward. Well, sorry Matthew, but it appears that it is not just Mythos and not just Anthropic models going forward. As, as the, as the other, as the other, as the other, as the other, as the from Friday was the release surrounded by the biggest air quotes possible of gPT 5.6. GPT 5.6 is actually three models, Sol, which they call their next generation frontier model, as well as 56 Terra, a balanced model for efficient everyday work, and 56 Luna, a fast and affordable model for high volume work. Now as I mentioned in the addendum to the weekly recap, at the request of the US government these models will once again only be available to a small group of trusted partners. In their announcement post open AI wrote, We believe in broad access and we plan to make geep. Seoul, Tara and Luna generally available in the coming weeks. As part of our ongoing engagement with the US government, we previewed our plans and the model's capabilities ahead of today's launch. At their request we are starting with a limited preview for a small group of trusted partners whose participation has been shared with the government before releasing more broadly. During this preview we will continue testing and coordinating closely with partners as we work toward broader availability. We don't believe this kind of government become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them. We are taking this short-term step because we believe it is the strongest path to broader availability in the coming weeks, while we work with the administration to develop the Cyber Executive Order Framework and a repeatable process for future model releases. In additional comments, Sam Altman once again supported the premise of a limited rollout but disagreed with how it's being executed. He wrote, I think it's quite reasonable to roll out models, especially as they reach they reach significant new ability in this way. It fits with our long-held strategy of iterative deployment. But this isn't quite the process that we think is optimal. Now we will with the government attempt to get a transparent reliable reliable reliable reliable reliable reliable, reliable, reliable, reliable partner that works with all stakeholders, and we also want to live by our mission of benefiting all of humanity. I believe the government shares most of our goals and that they are overall doing a good job in a very difficult situation. We will work as quickly as we can to get this model as we can to get this model in your we hope you will love it. Now as for the actual model open AI has introduced new nomenclature for the family. Again the model will be available in three different sizes, Luna the cost, the cost, Terra the medium version which they say will deliver gpity 5.5 level performance at half the cost and soul which is the new flagship that open AI says will be a step function better than gpity 5.5. A p.5. Which cost for soul remain the same as gp. 5.5. at $5 per million per million input tokens and 30 per million output tokens, which, which, which was, which was, which was, the $10 and $50 pricing for fable. Open AI will also introduce a new max reasoning setting for soul as well as an even heavier setting called Ultra. When working in Ultra mode, Soul will spin up multiple subagents to allow the completion of more complex work. Now at this stage, none of the three model variants are available for public release, making it impossible to know exactly how strong they are. Based on the benchmarks released by Open AI, 56 Sol on Ultra settings is the new state of the art in agentic coding. It scored 91. It scored 91.9. 2.0, beating mythos by almost 4 percentage points. Sol on max settings is also slightly ahead of mythos. Tara matched fables score, which is slightly behind mythos, while Luna is slightly less performance than geep. 5.5. On exploit bench, a cyber security benchmark that tests a model's ability to autonomously find, code and execute an exploit, open AI claims that sol pushes the performance efficiency frontier. It appears that its performance on max settings is roughly in line with mythos, but using around one third of the token. Terra's performance on this benchmark is slightly better than GPT 5-5 or Opus 4-8, while Luna is roughly in line with Opus 4-8. Open AI also released a handful of other benchmarks showing strong performance in biological analysis and cyber security. On model safety, open AI is taking a layered approach similar to anthropic with the fable release. Some guardrails are trained into the model layer, others are present as prompt refusals. And open AI also plans to continually monitor for prolonged misuse. In addition, open AI will be feeding outputs outputs to check for misuse before delivering the output to the user. Now one thing that you might be scratching your head about is that it's not obvious that Tara or Luna are significantly more advanced than GPT 5-5 or Opus 48, making it theoretically puzzling on why the less powerful model variants are being held back. It could imply that the government has halted all model releases for the time being not just the largest and most capable versions. Or even if they haven't explicitly that open AI is just being extra careful not to tread on any toes. Or finally, of, of, of finally, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, of, want to release an incomplete set of these models without the flagship available. Now when it comes to external benchmarks one group that did apparently have early access to gp. 5.6 soul was meter. They wrote, with our access meter conducted a pre-deployment evaluation of gp. 5.6 soul including an attempted measurement of its 50% time horizon. This is of course meter's well-known test to see in human equivalent terms how long the most complex task that a model can accomplish is. Just as a reminder, if the 50% time, if the 50% time horizon measure is 10 hours. That does not mean that the model worked continuously for 10 hours, but that the equivalent task that it accomplished at a 50% success rate would take a human 10 hours. Meter benchmarks this