AI Model Month Is Off to a Blistering Start
The AI Daily Brief: Artificial Intelligence News and Analysis
The AI Daily Brief explores a flurry of new AI model releases in September, including Gemini 3.8 Flash, Meta's MuSpark 1.3, the Muse personal agent, and
Key takeaways
- The shift from single-model dominance to dynamic model selection is accelerating as specialized AI tools emerge.
- OpenAI's alleged breakthrough on the Navier-Stokes problem has ignited controversy over data access, academic ethics, and potential scooping of prior work.
Main topics
- New AI model releases in September
- OpenAI's Navier-Stokes controversy
Notable quotes
"The development of a singularity would have to deepen despite the presence of viscosity which tends to smooth out motion."
Conclusion
The rapid pace of AI model innovation underscores a new era where tool selection and ethical
Transcript preview
Speaker 1 (0:00) Throughout the summer, the big thing we've been exploring at the AI Daily Brief is all about the move from a single model paradigm where you pick the best model overall and that's the one you stick with, to a more complex model architecture where we are, both as individuals and as teams, able to navigate nimbly between different models and even different harnesses to get the most out of AI based on whatever particular use case we might have. And what's more, this summer, we got really clear on the fact that getting the most out of AI is not just a question of model or harness capability, but also a question of efficiency and cost, especially as we move to more complex agentic workloads. And so it's fitting that the beginning of September has been just a cavalcade of new models. From Fable 5.1 to GPT-6 Astra to the models that we're looking at today, including Muse Spark 1.3 and ChatGPT Images 2.5. All of these add up to way more diversity in the tools we have access to for you to design the perfect AI stack for your actual life and work. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Speaker 1 (1:07) All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Section, and HyperAgent. To get an ad-free version of the show, go to patreon.com slash ai-dailybrief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at ai-dailybrief.ai. And finally, if you haven't yet, you can check out our latest free self-directed training program. It is called the Multiplayer AI Sprint for Teams. And basically, the idea is to shepherd you through a process of figuring out how to build agents that don't just help you, but actually sit at the intersection of work that is shared across your teams. I'm pretty convinced that this is the next big paradigm for AI inside companies, and so I wanted to build a sprint that could help you guys fully embrace that. There's, of course, a link to that on the AIdailybrief.ai website, but you can also find it at multiplayerai.ai. We kick off today with a story that very easily could have been the main episode. given how much drama is surrounding it. On Tuesday, OpenAI published a solution to the Navier-Stokes problem, one of the seven problems selected for the Millennium Prize in the year 2000. The Wall Street Journal characterized these problems as the, quote, holy grail of math, and that's fairly accurate. Each Millennium Prize problem has a million-dollar reward attached, and only one has been solved in the 26 years since the prize was established. The other problems include the most famous unsolved problems in math, such as the Riemann hypothesis and P versus NP. You know, the things we all talk about when we get together for dinner. Now, for the purposes of this particular episode, I'm actually not going to get into the details of the problem itself or debates around whether it has any significant real-world applications. I'll read OpenAI's description of the problem just to give you a flavor. They write, The Navier-Stokes equations use Newton's second law of motion, F equals ma, to describe how fluids move. Importantly, they treat a fluid as a continuous medium rather than tracking individual molecules. These equations are used for aircraft design, weather forecasting, and the study of blood flow. A fundamental open question for these dynamical equations has been whether the continuum approximation of the fluid can break down. Specifically, can the Navier-Stokes equations for a three-dimensional, incompressible fluid with constant density develop a singularity even when the motion starts smoothly? Here, a singularity means the dynamics lead to speeds in the fluid growing without bound within a finite amount of time. The development of a singularity would have to deepen despite the presence of viscosity which tends to smooth out motion. Because a real fluid cannot move infinitely fast, this would mark a breakdown in how the equations model the fluid. To continue modeling the system, one would then need to track the behavior of each particle individually. So that's the problem they're addressing here, and I think again for context. The important note is that this represents a huge step up from something like the Airdos problems that made news last year. OpenAI claims to have solved this problem, and they did so using an internal model that is significantly more capable than GPT-6 Astra. Noam Brown said that the result cost several million dollars to find, and it seems to have taken a week or two. However, the big controversy surrounded exactly how OpenAI had arrived at this result. Shortly after the result was published, New York University professor Tristan Buckmaster published his version of the events. According to Buckmaster, he and an anthropic employee named Levent Alpagi had been working on the Navier-Stokes problem together for more than a year. This was an outside project for Levent and the pair had used a range of different AI models, including GPT-56 Sol in the Codex Harness. Crucially, Levent and Buckmaster did not find a solution to Navier-Stokes, but they did find novel solutions to related problems that could be viewed as a stepping stone to the Millennium Prize problem. They were also using extremely novel methodology that few in the mathematics world were pursuing. Rumors of their work spread through AI circles in recent weeks, incorrectly claiming that Anthropic had solved a Millennium Prize problem. Buckmaster says he reached out to OpenAI last week to clarify the situation. According to Buckmaster's telling, Sebastian Bubeck from OpenAI informed him on Sunday that they had solved Navier-Stokes and wanted to discuss publication. Buckmaster wrote, I asked whether the model had been trained on or had access to our sessions in Codex, into which we had been putting all of our drafts for the whole of the project. I was told the model did not look up user data. I asked again about training and I did not get an answer. He said he was offered two proposals, for either OpenAI to publish separately or Buckmaster to join the publication if he agreed to remove Levin from the authorship because of his ties to Anthropic. Buckmaster declined both options and threatened to go public if OpenAI published. Continuing his account, Buckmaster wrote, The reply was, Why would you ruin your career? I replied that I am an academic and asked why he thought going public would ruin my career. The reply was, If you don't want me to be nice, then I don't have to be nice. OpenAI leaders responded with a series of statements. Sebastian Bubik called the allegations false and inflammatory, then revealed part of his text message chain, which he claims contradicts Buckmaster's version of events. Sam Altman gave a series of explanations for how this went sideways, but claimed that OpenAI's approach was different than that of Buckmaster and Levin. The OpenAI account claimed, We, the researchers and the agents, did not see any of their work through any means until they released it publicly. In particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from the usage of our products helped improve our models. After Levin characterized this statement as quote-unquote coming clean, OpenAI Chief Research Officer Mark Chen responded, Now, as for the controversy, there are two distinct strains of conversation. Firstly, academics are up in arms over what they see as unethical behavior. Assistant Professor Talia Ringer of Illinois University wrote, Rushing to get a result after you hear someone else has a result is messed up. That is AI scooping culture and goes against every academic norm that exists in reasonable fields like mathematics. This is how AI culture rots entire fields. Thomas Wolfe, the co-founder of Hugging Face, suggested this might just be a preview of accelerated AI science, commenting, Hope this is not a glimpse of the future we'll get in science research, with these dominating players playing marketing games hurtful for the real scientific community. The second, and likely far more relevant criticism, at least for the AI Daily Brief audience, was questions of trust in OpenAI. From Buckmaster's account, we can assume they were using some consumer version of Codex, but it's unclear whether they agreed to share data to improve OpenAI's models. For some, it's a wake-up call for anyone who is using AI to work on proprietary tasks. Former DeepMind employee Susan Zhang wrote, Everyone getting sniped by the personal drama, but missed the more interesting unanswered question. Can these labs see all your work and scoop you when the stakes are high enough? Seeking clarification from OpenAI leaders, mathematician Tryon Zyloris asked, Important question. If I opt out from training then paste a trade secret using my paid subscription, do you de-identify my personal details but keep the trade secret and may add it to your training data? At the time of recording, that has not received a response. So taking a step back, there are a few reasons that this whole episode is having such resonance. First is honestly the voyeurism of it. People love drama. Fighting against that is like trying to fight the tides. But the question of ethics around advanced AI and what these companies can do with data is a question that, while it has been present basically since the beginning of LLMs, has gotten a lot louder in consideration more recently. Especially as model leadership starts to consolidate around a couple of companies, it brings up a