AI researchers debate how close we are to recursive self-improvement

Dwarkesh Podcast

In this episode of the Dwarkesh Podcast, host Dwarkesh Patel convenes leading AI researchers John Schulman, Beren Millidge, and Charlie O'Ne

Key takeaways

  • The current trajectory of AI development may hit asymptotic limits due to persistent challenges in generalization, meta-learning, and continual learning.
  • Even with massive compute and scaling, models may remain bottlenecked by their inability to perform open-ended scientific discovery or paradigm shifts without human-defined objectives.

Main topics

  • Recursive self-improvement and AI takeoff
  • Scaling laws and diminishing returns in deep learning

Notable quotes

"If it requires another one of those discontinuities to solve, I'm not sure that the current method of training LLMs with these RL environments would be able to discover that discontinuity." – Charlie O'Neill

Conclusion

While the researchers agree that AI could achieve recursive self-improvement and surpass human capabilities within

Transcript preview

Speaker 5 (0:00) Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk, and who also happen to be at somewhat open-ish labs and companies, so you guys can actually say things on the record. I'm joined by Baron Milic, who is the CTO of Zyphra, which is developing open source models. John Shulman, who is the chief scientist at Thinking Machines, previously the co-founder of OpenAI, led the RLHF work that led to Chet GPT. And Charlie O'Neill, who is head of model training at Base10. The first question I have, if we're in 2036, it's been 10 years, and we don't have billions of crazy super intelligences that are running around that have radically transformed the world, what is the most likely reason that that doesn't end up being the case? Other than sort of exogenous political shocks, or like there's a war, or they ban AI or something. But what is the most likely technical reason that 2036 isn't like a crazy alien super intelligence world? Speaker 3 (0:55) I mean, like, my reason would just be, like, it's got to be this sort of, like, there's been a classic thing, almost like Marv X Paradox, right? Where, like, we see, like, you know, we think of the AI being like, if it can do this, it's going to be amazing, right? Like, if it can solve these hard math problems, if it can win a chess, blah, blah, blah, and then it solves these things, and then it's, like, not that impactful. Obviously, it's somewhat impactful, but, like, not everything. It's like, if somehow that continues, and, like, there's never, like, the true, like, spark of generalization that occurs, I think that could lead to, like, the AI just being, like, extremely good at kind of everything that people, like, put into a benchmark, put into an environment, but, like, there's still some persistent, like, sim to real, which is somehow blocking everything. I think this is kind of unlikely. I think we... actually see this kind of generalization even from our LLM practice already. But if it is just ridiculously hard to generalize meta-learning, plus we don't solve continual learning, it's just super hard and impossible. This would be my default scenario in that case. Speaker 1 (1:47) Yeah, I agree with that. Humans have a lot of advantages over models now, and each time a new model comes out, it'll catch up in some of these areas. Like you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment or the models can't check themselves well enough. Yeah, so there's this cycle that keeps repeating where people think, where a new model comes out and people are blown away and they're like, this is it, this is AGI, but then they use it a bit and then it starts to feel dumb after a month or so. So that cycle just might keep going and it's hard to predict how many times it's going to repeat. Like right now, you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you like a hundred times more productive. But yeah, so maybe there are just more of these cycles than we would expect. Speaker 4 (2:52) For me, it's like a question of how far off like this global optimum of a learner you could have on a chip. is the transformer plus RL, basically the current recipe. So I think people imagine that once you have an agent which is better than all humans at AI research, even if it's 0.1 % better than all humans, then the fact that you can run hundreds of thousands, if not millions of these in parallel, you can run them much faster, chips going to speed up, that's going to outweigh every other bottleneck and you're eventually just going to hit this very fast takeoff with recursive self-improvement. I could imagine that if we continue along the trajectory that we're currently on with that paradigm where, you know, it's basically just like self-attention, RL, scaling up RL environments. I guess, like, if you think about what happened with Moore's Law, right? Like, we had this very, like, nice straight line and that held for a really, really long time. But there were so many, like, discrete, like, discontinuities and innovations that had to happen to keep that scaling law going. And the same thing has kind of happened with LLMs. Like, we... had this pre-training scaling law, and then that was kind of hitting the diminishing returns. And then we came up with RL and solved that. And then we got this new diminishing returns curve to hit that made it keep looking like a straight line going up. And so if it requires another one of those discontinuities to solve, I'm not sure that the current method of training LLMs with these RL environments, even RSI-targeted RL environments, would be able to discover that discontinuity. And if not, like... we're probably going to hit this asymptotic curve. Speaker 5 (4:25) Sorry, but do you think discontinuity will be harder than anything that's come since 2012? Speaker 4 (4:30) If we had the answer to that, we'd kind of have the ability to implement it. But maybe we should distinguish between discontinuity, which adds to the current paradigm. Again, it's cumulative. There's something beyond the RL that we have to discover, and maybe they're capable of connecting the dots in that straight line. But again, how far off the global optimum are we? Do we have to go back and throw out gradient descent and neural nets in general? And I don't think if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away. Speaker 5 (5:05) Yeah, the only hope really is if deep learning just can't get us to an AI, which is at least... Speaker 5 (5:14) can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or like, I don't know, maybe humans would also never have discovered the next learning architecture, but to the extent humans could have discovered it eventually. But it just seems like, I don't know. if you just look at the progress that's happened in 2012 till now, and you just continue that on, I mean, I know it's been powered by huge amounts of compute scaling and so forth, but it would be weird if, like, it just didn't get to the point where it could, like, dominate humans, at least in R &D. Especially over the next few years, there's going to be... Ryan Greenblatt was on the podcast recently, and he made this point that you could imagine as the AIs get more and more capable and are capable of making progress on... simulations which incentivize getting better at not only AI R &D, but generally at science. So this is a thing that all the labs are targeting, many startups are targeting. Or another intuition pump is if you look at the ELO score of chess bots since the 80s, there's just like a very linear increase in ELO over time. But there's this huge discontinuity as they cross the human range of human experts always win against AIs to like human experts never win against AIs as this linear increase in ELO happened. And you could think, I agree with your point that So far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world. But that's because they're slowly rising in ELO relative to humans. Yeah, Speaker 3 (6:39) I mean, I agree. I mean, the only way for this to not happen is if like... as you said, somehow asymptotes just before basically, because we're already pretty close, in my opinion, to where we'll start crossing the human ELO score. And so we'll need to asymptote before that. And that's the only way, in the scenario you posed, where somehow we're sitting here in 2035 and everything is normal for this to happen, I think. I mean, the only other way is there's some dramatic regulation on AI. This is kind of what I see as the most likely way for this scenario to happen, actually, rather than the technical thing. Speaker 4 (7:07) Yeah, I think there's different kinds of research. There's research where it's like the auto research style where the objective is already specified very cleanly and you're optimizing that objective. And I think everyone is picturing like... If we continue along this path of, like, you know, making pre-training loss go down, making our own environments go up, that's going to lead to, like, improvement. But, like, you know, maybe what Ryan is talking about is, like, this much more open-ended type of science, which is required for, like, paradigm shifts, where we can't specify the objective, and the AIs are definitely not able to specify that objective either. Like, we have to be really, really careful about how we specify objectives for any of these things. Speaker 5 (7:39) And maybe your point is that, like, the nature of the breakthroughs that have happened since 2012 is that we have found, like, in 2012, people weren't saying, I'm assuming, I don't know, you guys were there. Or at least, John, you were there. I was in primary school. Actually, John, I'm curious for your wisdom of the ages or wisdom of being in the trenches way back when. But presumably a big breakthrough was realizing that next token prediction is the... You wouldn't have thought that nano GPT speedrun is the thing to be optimizing for in 2014. But now that we have come to this new paradigm, you wouldn't think to do a speed run on that and have AIs get really good at that. But maybe there's like a next inner loop to optimize that the AIs wouldn't anticipate. And there's an outer loop of like revenue or something that eventually should be strong, but it's a very slow outer loop. Speaker 1 (8:29) Yeah. In fact, I remember in the early OpenAI days having the intuition that actually just do like... Minimizing log loss wasn't going to get you to intelligence because the important bits are accounting for such a small fraction of the loss that it was going to be overwhelmed by noise. So just training a language model on next-token prediction just wasn't going to learn the interesting things you wanted to learn. And we needed to craft better objectives that would put more emphasis on the important things. Speaker 1 (9:07) You can make all sorts of arguments for this and you could say, oh, humans probably don't learn how to, like, we don't learn how to model everything in our environment. We can't, most, like, people can't create a photorealistic reproduction of some kind of scene they've looked at. So there must be, we must need a better objective. But then it turned out that it just worked anyway. Speaker 5 (9:32) As you were pointing out, the inner loop, even in current AI research of like post-training benchmarks or whatever, it doesn't necessarily translate into what users like. Speaker 1 (9:41) Oh yeah, I mean, the whole field relies a lot on generalization and it's very hard to predict when you're going to get generalization or when you're going to get some kind of out-of-distribution generalization. So we know that if you train on the task you care about, you're going to do better. But like the most important advances are often the opposite. Speaker 1 (10:03) types of generalization that we have no right to expect. So for example, from just pre-training on this very naive next token prediction objective to like various tasks of interest where some very, that require understanding of the input in some deep way or learning some skill from pre-training that's like very rare and like not very like heavily represented. And then also generalization from these verifiable tasks to less verifiable ones. This is also a type of generalization that there's no reason a priori to expect it. Speaker 5 (10:43) So this is an interesting question because one intuition pump that you could have for why you would see some sort of singularity very rapidly without even scaling up the inputs to AI progress that are not just AI labor is that before every single experiment you run, That's, you know, like a seven figure experiment. You spend an equivalent amount of compute on AI labor. And so you just have automated versions of you guys spending a century thinking about like, what is the optimal experiment to run? Do you like small scale ablations? Developing literally like. a century's worth of theory. So going back even before like deep learning, before you decide what experiment to run, doing extremely optimal, like setting up of the experiment, then you do a century of thinking after the experiment is over, where you're like analyzing what happened and what the next experiment to run is. Well, Speaker 1 (11:33) I think if you think hard enough, you probably could have expected some of these things beforehand. Like, there is probably some very clever way to do a small-scale experiment that'll let you build the theory that then will generalize to the large-scale experiment. So I would expect that, like, we're nowhere near the ceiling of how well you can do research. And, like, I would imagine a future where AI is, like, is doing a lot of, like, analysis and theory building. Speaker 1 (12:04) Spending a comparable amount of compute to the amount that you're spending on the experiments themselves, doing various kinds of analysis and building a theory around what Speaker 2 (12:14) we've Speaker 1 (12:14) seen so far. Speaker 4 (12:15) I think there's really concrete examples of this when the objective is well-specified. So again, all thinking can do is update your posterior based on the bits that you've gotten since you formed your prior. You can't gain any new bits from just thinking. But when the objective is well specified and there is this data sitting around, I imagine there will be this big speed up in the current paradigm we're in. And a good example of this is if you've got an AI to think about the Kaplan scaling laws, an AI at this point would have noticed that they've just taken these intermediate checkpoints and didn't account for the annealing. And so this is wrong. And that would have caught that years earlier. We would have made progress. We would have cut off a year or two of progress just from that observation from an AI. And again, once the objective is well specified, which is lower pre-training loss or whatever, there's many, many good examples where if you just thought about it a bit more, you would have been able to cut down a significant amount on things that you've done. So like mu p and how learning rate scales with model size and realizing that model width is important in that as well. I feel like you can really back out a lot of these things and cut off a lot of low-hanging fruit. So I would imagine a 10-time speed up if our thing is just maximize the objective we're currently on. But I don't see how that generalizes at all to come up with the right objective in the first place. Just thinking doesn't necessarily buy you the right objective in the first place. Speaker 3 (13:33) I mean, yeah, I think this is really the key question to any kind of very rapid RSIs from current AIs. It's like, how well can AIs generalize to learning their own objectives? Because to have any kind of self-propelling automated loop, you need the AI to propose objectives, optimize them, figure that out, propose a new objective, and have this not go off the rails at any point for a long, long time. Coming back to Morvex Paradox, there might be a case of Morvex Paradox where we think there's autonomy and being self-encapsulated so we can think of what we should do ourselves and then go do it and have this loop. It's super easy because we always do this. Obviously, evolution needs to create creatures that can survive on them by themselves for a long period of time. This just might be something that for some reason is really hard for the AI in the same way that... locomotion stuff is really hard. It's like math is super easy despite being super hard for us. I don't Speaker 5 (14:19) know. Doesn't the time horizon increasing Speaker 3 (14:20) suggest that that's... Yeah, exactly. I mean, this is another possibility, but I agree. There's no obvious evidence for this. In fact, the fact that our agent is now super persistent and it's quite easy to do this is kind of evidence against this. But this would be potentially one of the reasons why we just don't get this immediate takeoff is if this is hard. Speaker 5 (14:38) If you look back from 2012 till now, or maybe... from when you started doing your research till now, what part of... Speaker 5 (14:48) Of all the innovations that have happened since that time, including purely engineering ones, including purely conceptual ones, what seems like the thing that is the thing that would be the last things humans would have to do before AIs totally automate AI R &D? Speaker 3 (15:05) Probably just like iteratively asking the right questions. Like if you can get the AI to like do any experiment, but like you need to decide what experiments to do. And like right now I think AIs are not very good at this compared to coding the experiments at all. Whenever we talk about research, they propose a bunch of miscellaneous things which are very, very tiny steps. Or Speaker 4 (15:22) even going from DeepMind's approach of we're Speaker 5 (15:24) going Speaker 4 (15:24) to solve intelligence by learning to play games at a superhuman level. That's going to be the approach to one random researcher like Radford being like, I'm going to try and just predict the next token off a very wide swath of data. And then even once Radford had discovered that, it took a while. before people decided to scale it up because we had to come up with the idea of scaling walls and the fact that you could very reliably predict these things. Yeah. Speaker 1 (15:47) I would say that the last job for humans or the role for humans that will last the longest is defining the objective and deciding what we actually want. So in that vein, something like deciding what... how the assistants should behave or what it means to be helpful or what's the objective when we're doing our all-from-human feedback is one such thing. And then later, defining constitutions and model specs is another one. And I think even if the AIs can do all the technical work, we'll have to still do a lot of that and decide what we actually want. Speaker 5 (16:31) Yeah, alignment is the final job. Speaker 1 (16:33) Yeah, alignment is sort of the answer, but it's also alignment itself can be kind of decomposed into like specification of the objective or figuring out what the right objective should be. And then like actually like achieving or optimizing the objective you've defined. And I think the first one is not going to go away anytime soon. And like... If I think about a post-training team and why you need a lot of people to be on the team, it's just because there are a lot of different areas where you have to figure out how the model should behave. It would be very hard to automate the whole thing just because someone has to think about how should the model behave in this area. Jane Street started Speaker 5 (17:24) using Antithesis to test their software in early 2025, and they were so impressed by the product that they decided to invest in the company. I recently caught up with Ron Minsky, who co-leads Jane Street's tech group, to ask about how Antithesis actually plugs in. Speaker 2 (17:38) The thing that I think is most impressive about Antithesis is we started using it in a team that was building high-assurance software and being really careful, and nonetheless, it was able to shake out bugs that were otherwise going to be really hard to find. important both because it helps make those systems more reliable, but also because it helps the teams that build it to just move faster. This matters more and more as code production is increasingly automated. I think in general, as we've been using agents more and more, the key problem that you run into is the verification bottleneck. Just the time it takes from people to look at code and figure out is that actually something you want to accept in your production software. And tools that make testing better are just incredibly helpful there. They just ease the verification bottleneck and make it possible for you to get more stuff done and move faster because you can have more confidence that the code generated by the agent is actually not introducing new problems. Speaker 5 (18:30) To see how Antithesis fits into your development process, go to antithesis.com slash thwarkesh. Speaker 5 (18:39) What is the story for why there isn't huge consolidation in model providers? There's just so many things that point to centralization here. If we step back over the course of years, is there something that is going to prevent that? Speaker 1 (18:54) Yeah, I think distillation is the main thing that fights against the centralizing force. Because basically anything that can be learned through RL can be distilled very easily. Speaker 1 (19:07) It's a small number of bits. It's something that you can learn from a small amount of data. So if you can get trajectories from the model that show a behavior, you can easily distill it. So I think distillation is one of the things that fights centralization. There's also... Speaker 1 (19:29) I mean, there is a possibility that there'll be company-specific models, that it'll be possible to learn from deployment and have a company continually improving its own model. And such a system could be provided by the current oligopoly of model providers or some other currently smaller company. But I think that'll change the game a bit. Yeah, Speaker 3 (19:54) and I also want to point out that like continual learning and it doesn't stop distillation, right? Like even if your model is improving every day, like people could be distilling it every day. So it's like the loops could just operate at the same pace. Right, that makes sense. Speaker 5 (20:05) Okay, so copying model behavior. Speaker 5 (20:09) I guess you need to know yourself what the right distribution to prompt is in order to get the relevant model behavior. Speaker 1 (20:16) Yeah. For just distilling with supervised learning, the prompt distribution is extremely important. So it's very non-trivial to distill a model even if you have full access to it and have the cot, the chain of thought and everything. Yeah, it's non-trivial to distill all of the useful capabilities from it. because you need to prompt the model with something. You need to prompt it with realistic prompts. You need to have a really wide distribution of realistic prompts. So yeah, one thing that's been coming out recently is some of the Chinese companies are probably using these router services, which are designed to allow people in China to use the US frontier models, which would otherwise be blocked in China. But there are all these... router or proxy services that allow people in China to use these models mostly for coding. And these router services are collecting and selling some of the data. So I think this is like a very useful data set for distillation because it gives you the perfect prompt distribution. Speaker 3 (21:25) I think this is one of those things where AI is helpful out here. If you actually look at the frontier pipelines, let's say the Chinese models that they actually put in their papers, it's a lot of humans or they get seed prompts from somewhere, which is some combination of humans, this kind of data, and then they synthesize a vast coverage from those seed prompts using their existing models or the other frontier models. You can automate an awful lot of this prompt distribution gathering and environment creation. It's just like humans need to provide increasingly fewer amounts of bits, it's like the models get better. Speaker 5 (21:52) Right. It still seems you're bottling. by having Speaker 3 (21:57) a service which has Speaker 5 (21:58) users or Speaker 3 (21:59) users are Speaker 5 (22:00) going through. Speaker 3 (22:00) Not necessarily. I mean, yeah, that's obviously very helpful, but theoretically you can just think about what users want or a lot of tasks. Speaker 5 (22:07) The whole point is that the user says, make me an application like this. Oh, that didn't work. I actually want you to make this new feature. But actually, let's step back and do this other thing. And capturing that whole trace is the... Or to the extent you could have done that anyways, then you just have RSI anyway. Speaker 3 (22:23) Yeah, I mean, ultimately, if you have this fully automated loop, that is basically RSI, right? The AI is deciding the data, it's deciding the training, that is the loop. But yeah, I mean, it depends how much human information you need. At some point, if you're just like, I want traces that look like this, you prompt that to the model, the model will be able to come up with a pretty good approximation. Speaker 5 (22:40) But what if you want to do, like, make me a really good politician, and then just anticipate de novo? How would a discussion in the Senate halls go or something? I just feel like there's going Speaker 3 (22:49) to Speaker 5 (22:49) be a lot of things. Speaker 3 (22:50) Ironically, this is actually, I think, easier for the distillers than the frontier labs, right? Because the distillers are just like, I want a good politician. They go to the frontier model. The frontier model already knows how to be a good politician, so it just generates those traces. Whereas if you actually want to build the first model that does this, you have to actually somehow get data on what politicians do every day and build that. So it's actually much easier to say, I want something like this, and then get the AI to produce a billion variations, than to actually create the thing like this to begin with. I Speaker 4 (23:15) think you can actually make a really concrete prediction based off this observation that the Chinese labs have this router data. So I think the thing that Jess stated this originally was I was saying, isn't it weird how Sonnet 5 and Opus 5 are almost objectively worse models than... GLM 5.3, Kimmy K3, even though they've had access to not only distillation, but logic distillation from Mythos. And so the counter here was that the prompt distribution really, really matters. You need to see what users are doing so that you can distill these behaviors and things in. I think the prediction from this is that the frontier labs don't necessarily have much of an advantage, if at all, in RL environments now. Because yes, user distribution matters for general behavior. and so on. But the best measure of a capability is the very, very hard RL environments you've made at the frontier. And so if you have access to those RL environments as Anthropic, and you have access to logic distillation, and you've still made a worse model, then maybe... Then Speaker 5 (24:13) real-world deployment matters more than the environment. That's really interesting. But they had to incentivize those capabilities in the first place in Fable or the frontier model. And so it's weird that they can't incentivize them again. Speaker 5 (24:28) Or like with a smaller model or something. Maybe Speaker 4 (24:30) we're just in this weird uncanny valley where actually trying to copy that frontier model too much, like the student-teacher gap or whatever it is, is just too large. And I think people made this point with Opus is it's like the difference between Opus 4.6 and Opus 5 is that Opus 5 really feels like it's got this AI as a judge checking every possible thing it's done. That's why it uses so many tokens. It tries to think about all these things, but it doesn't necessarily have the big model smell of Fable to know when to stop doing that or when's a good path to go down or whatever. Speaker 1 (25:00) The Speaker 5 (25:00) reach exceeds the grasp. Yeah. Speaker 1 (25:02) Yeah, Speaker 5 (25:02) I Speaker 1 (25:02) would offer a slightly different hypothesis. So I would say there are a couple of different axes for the environments you can create. And one of them is difficulty and the other is realism. It's sort of easy to create. It's comparatively easy to create a lot of difficult environments that involve doing a much more complicated task or doing something that requires a lot more cleverness. And you could say this is like the benchmarking distribution because a lot of the most prominent benchmarks just involve doing some... very hard puzzle-like task that's easy to verify.