AI researchers debate how close we are to recursive self-improvement

Dwarkesh Podcast
11 September 2026 1h 37m
0:00 --:--
Episode Description
New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.Watch on YouTube; read the transcript.Sponsors* Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithes

Summary

In this episode of the Dwarkesh Podcast, AI researchers John Schulman, Beren Millidge, and Charlie O’Neill discuss the technical challenges and timelines related to recursive self-improvement (RSI) in AI. They explore the current state of AI research, the role of reinforcement learning, data efficiency, model scaling, and the potential for AI to automate scientific research and other complex cognitive tasks.

Chapters

Introduction and Guest OverviewDwarkesh introduces guests Baron Millich, John Schulman, and Charlie O’Neill, setting up the discussion on AI progress and recursive self-improvement.
Why No Superintelligence Yet?Discussion on technical reasons why superintelligent AI might not emerge by 2036, including challenges in generalization and meta-learning.
Current AI BottlenecksExploration of current limitations in AI models such as judgment, self-checking, and productivity gains from AI-assisted research.
Scaling Laws and Paradigm ShiftsAnalysis of how scaling laws and reinforcement learning have driven AI progress, and speculation on potential future discontinuities.
Automating AI Research TrainingHow future AI models capable of automating AI R&D might be trained using human feedback and synthetic environments.
Distillation and Model CentralizationDiscussion on model distillation, its role in decentralizing AI development, and challenges in replicating frontier model capabilities.
Learning From Deployment DataInsights into how AI models can improve by learning from real-world deployment data and the economic incentives affecting this process.
Sample Efficiency and Continual LearningChallenges of sample efficiency in weight updates, catastrophic forgetting, and the limits of continual learning in deployed models.
Data vs Architecture in AI ProgressExamination of the relative importance of data improvements versus architectural innovations in driving AI capabilities.
Reinforcement Learning’s RoleWhy reinforcement learning has been more successful than expected and how it improves models despite limited bits of information per episode.
Creativity and Diversity in RL ModelsDiscussion on RL’s impact on AI creativity, diversity of outputs, and concerns about monoculture in model behaviors due to distillation.
Future Predictions and TimelinesRapid-fire predictions on when AI will function as general remote workers, automate AI research, and achieve superintelligence across fields.

Topics

recursive self-improvementreinforcement learningmodel distillationsample efficiencyscaling lawsAI research automationcontinual learningdeployment dataalignmentgeneralizationAI creativitymodel scalingdata efficiencyobjective specificationsim-to-real transferAI timelines

People

Baron Millich (guest) John Schulman (guest) Charlie O’Neill (guest) Dwarkesh (host) Brian Greenblatt (mentioned) Ron Minsky (mentioned) Dario Amodei (mentioned) Jerry Hahn (mentioned) Liam from Periodic Labs (mentioned)
Key Concepts (15)
Marvell's paradox — The phenomenon where AI systems solve many benchmark tasks but fail to achieve transformative general intelligence due to persistent bottlenecks in generalization and meta-learning.
Generalization bottlenecks — Current AI models excel on benchmarks but struggle with out-of-distribution generalization and self-correction, limiting explosive capability growth.
Scaling laws and discontinuities — AI progress follows scaling laws with occasional paradigm-shifting innovations (e.g., RLHF) that reset diminishing returns, but future discontinuities may be harder to discover.
Recursive self-improvement (RSI) — The process by which an AI improves its own design and capabilities iteratively, potentially leading to rapid capability gains if certain technical challenges are overcome.
Model distillation — Technique to transfer knowledge from a large 'teacher' model to smaller 'student' models, which can decentralize AI development but requires realistic prompt distributions.
Learning from deployment data — Incorporating real-world usage data into training to improve AI models continuously, though economic incentives and technical challenges affect adoption and speed.
Catastrophic forgetting — A challenge in continual learning where models lose previously learned knowledge when updated with new data, limiting the effectiveness of iterative fine-tuning.
Data vs architecture improvements — Analysis showing that improvements in training data quality and quantity have driven more compute efficiency gains than architectural changes at small scale.
Reinforcement learning signal-to-noise — RL training provides a high signal-to-noise ratio by focusing on success/failure bits, enabling efficient policy improvement despite limited bits learned per episode.
RL creativity and diversity tradeoff — While RL can enable novel solutions and creativity, it often reduces output diversity and can lead to behavioral monocultures due to reward hacking and judge limitations.
Sim-to-real transfer — The challenge of transferring skills learned in simulated environments to real-world tasks, especially those involving complex human interactions and long time horizons.
Objective specification challenge — Defining the right objectives for AI systems is a persistent human bottleneck, crucial for alignment and ensuring models behave as intended.
Sample efficiency gap — AI models currently require far more data than humans to learn effectively, particularly for weight updates, limiting rapid real-world adaptation.
Phase transitions in learning — AI models exhibit discrete emergent capabilities (e.g., induction heads) during training, which average out to smooth loss curves but represent significant qualitative changes.
Horizon generalization — Models trained on longer-horizon tasks generalize the ability to sustain longer reasoning sequences across different domains, improving performance on complex tasks.
References (10)
Kaplan Scaling Laws by OpenAI Researchers paper
RLHF (Reinforcement Learning with Human Feedback) by OpenAI method
Chinchilla Scaling Law by DeepMind paper
Edgebench by Research Paper paper
Switch Transformer by Google paper
Antithesis by Company/Product tool
Composer by Open Source Project project
Claude by Anthropic model
GLM 5.3, Opus 5, Sonnet 5 by Chinese AI Labs models
Delphi by Open Source Model model
Transcript (137 segments)
Speaker 1

Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk and who also happen to be at somewhat open ish, labs and companies so you guys can actually, say things on the record. I'm joined by Baron Millich, who is the CTO of Zyphra, which is developing open source models. John Schulman, who is the chief scientist at Thinking Machines, previously the cofounder of OpenAI, who led the RLHF work that led to ChechiPT.

And Charlie O'Neil, is head of model training at Base Ten. The first question I have, if we're in 2036, it's been ten years, and we don't have like crazy billions of crazy super intelligences that are running around that are like radically transformed the world, What is the most likely reason that that doesn't end up being the case? Other than sort of exogenous political shocks or, like, there's a war or they banned AI or something.

But what is the most likely technical reason that we don't like, 2036 isn't, a crazy alien superintelligence world?

Speaker 2

I mean, like, my reason would just be, like, it's got to be this sort of like, there's been a classic thing almost like Marvell's paradox. Right? Where, like, we see, like, you know we think of the AI being like, it can do this, it's going to be amazing.

Right? Like, if it can solve these hard maths problems, if it can win the chess, blah blah blah. And then it solves these things, and then it's, like, not that impactful.

Obviously, it's somewhat impactful, but, like, not everything. It's like, if somehow that continues and, like, there's never, like, the true, like, spark of generalization that occurs, I think that could lead to, like, the AIs just being, like, extremely good at kind of everything that people, like, put into a benchmark, put into an environment, but, like, there are still some persistent, like, sim2 wheel, which is somehow blocking everything. I think this is kind of unlikely.

I think we do actually see this kind of generalization even from our own practice already. But like, if it is just like ridiculously hard to like generalize meta learning, plus like we don't solve container learning, it is just like super hard and impossible. Yeah.

Like this would be my like default scenario in that case.

Speaker 3

Yeah. I agree with that. Humans have a lot of advantages over models now, and each time a new model comes out, it'll catch up in some of these areas.

But you end up getting bottlenecked by the places where the model is weaker, and where it has worse judgment or the models can't check themselves well enough. Yeah. So there's this cycle that keeps repeating where people think where a new model comes out and people are blown away and they're like, is it.

This is AGI, but then they use it a bit and then it starts to feel dumb after a month or so. So that cycle just might keep going and it's hard to predict how many times it's gonna repeat. And right now, you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you like a 100 times more productive.

Speaker 4

we would expect. For me it's like a question of how far off like this global optimum of a learner you could have on a chip is like the transformer plus like RL, basically like the current recipe. So I think people imagine that even once you have an agent which is better than all humans at AI research, even if it's like 0.

1% better than all humans, then the fact that you can run hundreds of thousands if not millions of these in parallel, you can run them much faster, like chips are gonna speed up. That's gonna outweigh every other bottleneck and you're eventually just gonna hit this very fast takeoff with recursive self improvement. I could imagine that if we continue along the trajectory that we're currently on with that paradigm where it's basically just like self retention, RL, scaling up RL environments, I guess if you think about what happened with Moore's Law, right, like we have this very like nice straight line and that held for a really really long time.

But there were so many like discrete discontinuities and innovations that had to happen to keep that scaling law going. And the same thing has kind of happened with LLMs. Like we had this pre training like scaling law and then that was kind of like hitting the diminishing returns and then we came up with RL and solved that and then we got this new diminishing returns curve to hit that made it keep looking like a straight line going up.

And so if it requires another one of those discontinuities to solve, like I'm not sure that like the current method of like training LNs with these RL environments, even like RSI targeted RL environments, would be able to discover that discontinuity.

Speaker 1

And if not, like we're probably gonna hit this like asymptotic curve where it like Sorry. But do do think the discontinuity will be harder than anything that's come since 2012?

Speaker 4

If if we had the answer to that, that that we'd kind of have the ability to implement it. But, like, maybe there's the distinguished we should distinguish between the discontinuity which adds to the current paradigm. Again, it's cumulative.

There's some thing beyond the RL that we have to discover and maybe they're capable of connecting the dots in that straight line. Or again, how far off the global optimum are we? Do we have to go back and throw out gradient descent and neural nets in general?

And I don't think if you continue to scale up the current paradigm, an LLM, no matter how many LLM's you're running, are capable of necessarily discovering that if it's too far away. Yeah.

Speaker 1

just can't get us to an AI, which is at least, can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or, like, I don't know, maybe human maybe humans would never also never have discovered the, the next learning architecture, but, to the extent humans could have discovered it eventually. But it just seems like, I don't know, if if you just look at the progress that's happened in 2012 till now and you just continue that on I mean, I know that's been powered by huge amounts of compute scaling and so forth, but, it would be weird if, like, it just didn't get to the point where it could, like, dominate humans, at least in r and d, especially over the next few years, there's gonna be Brian Greenblatt was on the podcast recently and he made this point that you could imagine as AIs get more and more capable and are capable of making progress on simulations which incentivize getting better at not only AI r and d, but generally at science.

So this is a thing that all the labs are targeting, many startups are targeting. Or another intuition pump is if you look at the ELO score of chess bots since the eighties, there's just like a very linear increase in ELO over time, but there's this huge discontinuity as they cross the human range of human experts always win against AIs to like human experts never win against AIs as this linear increase in Elo happen.

Speaker 2

so far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world. But that is because they're slowly rising in ELO relative to humans. Yeah.

Yeah. I I agree it would be very I mean, the only way for this to not happen is if, like, as you said, somehow asymptote, like, just before basically because we're already pretty close, in my opinion, to, like, where we'll start crossing, like, the human ELO score. And so we'll need to asymptote before that.

And, like, that's the only way, you know, in this scenario you post where, like, somehow we're sitting here in 2035 and, like, everything is normal for this to happen, I think. I mean, the only other way is there's some dramatic regulation on AI. It's like this is kind of what I see as the most likely way for this scenario to happen actually, rather than the technical thing.

Yeah. I think there's different kinds of research.

Speaker 4

where it's like the order research style where the objective is already specified very cleanly and you're optimizing that objective. I think everyone is picturing like if we continue along this path of like making pre training loss go down, making our own environments go up, that's gonna lead to like improvement. Like maybe what Ryan is talking about is like this much more open ended type of science which is required for paradigm shifts where we can't specify the objective and the AIs are definitely not able to specify that objective either.

Speaker 1

is that we have found In 2012, people weren't saying I'm assuming, I don't know, you guys were there, at least John, you were there, but I was in primary school. Actually, John, I'm curious for your wisdom of the ages of or wisdom of being in the trenches way back when. But presumably a big breakthrough was realizing that Next Token Prediction is the you wouldn't have thought that Nano GPT speedrun is the thing to be optimizing for in 2014.

But now that we have come to this new paradigm that's you wouldn't think to do a speedrun on that and have you guys get really good at that. But maybe there's like a next inner loop to optimize that the AIs wouldn't anticipate. And there's an outer loop of like revenue or something that eventually should be strong, but it's a very slow outer loop.

Yeah.

Speaker 3

actually just minimizing log loss wasn't gonna get you to intelligence because the important bits are accounting for such a small fraction of the loss that was gonna be overwhelmed by noise. So just training a language model on next token prediction just wasn't gonna learn the interesting things you wanted to learn. And we needed to craft better objectives that would put more emphasis on the important things.

And, like, you can make all sorts of arguments for this and you could say, oh, humans probably don't learn how to like, we don't learn how to model everything in our environment. We can't, most, like people can't create a photorealistic reproduction of some kind of scene they've looked at. So there must be, we must need a better objective, but then it turned out that it just worked anyway.

Speaker 1

training benchmarks or whatever, doesn't necessarily translate into what users like.

Speaker 3

the whole field relies a lot on generalization and it's very hard to predict when you're gonna get generalization, or when you're gonna get some kind of out of distribution generalization. So we know that if you train on the task you care about, you're gonna do better. Yeah.

But the most important advances are often types of generalization that we have no right to expect. So for example, from just pre training on this very naive next token prediction objective to various tasks of interest where some very that require understanding of the input in some deep way, or learning some skill from pre training that's very rare and not very heavily represented. And then also generalization from these verifiable tasks to less verifiable ones.

This is also a type of generalization that there's no reason a priori to expect it. Yeah.

Speaker 1

This is an interesting question because one intuition pump that you could have for why you would see some sort of singularity very rapidly without even scaling up the inputs to AI progress that are not just AI labor is that before every single experiment you run, that's like a 7 figure experiment, you spend an equivalent amount of compute on AI labor. And so you just have automated versions of you guys spending a century thinking about like what is the optimal experiment to run, doing like small scale ablations, developing literally like a century's worth of theory. So going back even before like deep learning before you decide what experiment to run, doing extremely optimal setting up of the experiment, then you do a century of thinking after the experiment is over, where you're analyzing what happened and what the next experiment to run is.

Yeah.

Speaker 3

if you think hard enough, you probably could have expected some of these things beforehand. Like there is probably some very clever way to do a small scale experiment that'll let you build the theory that then will generalize to the large scale experiments. So I would expect that like we're nowhere near the ceiling of how well you can do research.

Speaker 4

what we've seen so far. I think there's like really concrete examples of this when the objective is well specified. So again, all thinking can do is update your posterior based on the bits that you've gotten since you formed your prior.

You can't gain any new bits from just thinking. But when the objective is well specified and there is this data sitting around, I imagine there will be this big speed up in the current paradigm we're in. A good example of this is if you've an AI to think about the Kaplan scaling laws, an AI at this point would have noticed that oh they've just taken these intermediate checkpoints and didn't account for the annealing and so this is wrong.

And that would have caught that years earlier we would have made progress like would have cut off a year or two of progress just from that observation from an AI. And again once the objective is well specified which is lower pre training loss or whatever, there's many many good examples where if you just thought about it a bit more you would have been able to cut down a significant on things that you've done. So like mu p and how learning rate scales with model size and realizing the model width is important in that as well.

Like I feel like you can really back out a lot of these things and cut off like a lot of low hanging fruits. I would imagine like a 10 times speed up if our thing is just like maximize the objective we're currently on. But I don't see that how that generalizes at all to come up with the right objective in the first place.

Like just thinking doesn't necessarily buy you the right objective in the first place.

Speaker 2

very rapid RSIs like from current AIs is like how well can AIs generalize to like learning their own objectives? Because to have any kind of self propelling automated loop, need the AI to propose objectives, optimize them, figure that out, propose a new objective, and have this not go off the rails at any point for a long, long time. Coming back to products, there might be a case of Morevix products where we think this kind of autonomy and sort of being self encapsulated so we can think of what we should do ourselves and then go do it and have this loop as super easy because we always do this.

And obviously evolution needs to create creatures that can survive by themselves for long periods of time. And, like, this just might be something that, for some reason, is, like, really hard for the AI in the same way that, like, locomotion stuff is really hard, whereas, like, math is super easy despite being super hard for us. But I don't know.

Doesn't the time horizon increasing suggest that that's You would yeah. Exactly. I mean, this is another possibility, which, like but I agree.

Like, there's no obvious evidence for this. Like, in fact, the fact that our, you know, RLH is now, like, super persistent and it's crazy to do this is kind of evidence against this. Yeah.

But, like, this would be, you know, potentially, like, one of the reasons why, like, we just don't get this, like, immediate takeoff is, like, if this is hard.

Speaker 1

If if you look back from 2012 till now or maybe from when you started doing your research till now, what part of of all the innovations that have happened since that time, including purely engineering ones, including purely conceptual ones, what seems like the thing that is the

Speaker 4

thing that would be the last things humans would have to do before AI is totally automate AI r and d? Probably just like iteratively asking the right questions. Like, if you can get the AI to, like, do any experiment, but, like, you need to decide what experiments to do.

And, like, right now, I think AIs are not very good at this compared compared to coding experiments at all. Yeah. Like, whenever we talk about research, they propose, like a bunch of, like, miscellaneous things which are like very, very tiny steps.

Or even going from like, know, DeepMind's approach of like, we're gonna solve intelligence by learning to play games at a superhuman level. That's gonna be the approach to like one random researcher like Radford being like, I'm gonna try and just predict the next token off a very wide swath of data. And then even once Radford discovered that, right, it took a while before people decided to scale up because we had to come up with the idea of scaling laws and the fact that you could very reliably predict these things.

Yeah.

Speaker 3

job for humans, or the role for humans that'll last the longest is defining the objective, and deciding what we actually want. So like in that vein, something like deciding how the assistant should behave, or what it means to be helpful, or what's like the objective when we're doing our all from human feedback is one such thing. And then later defining constitutions and model specs is another one.

I think even if the AIs can do all the technical work, we'll have to still do a lot of that and decide what we actually want. Yeah. Alignment is the final job.

Yeah, alignment is sort of the answer, but it's also, alignment itself can be kind of decomposed into like specification of the objective, or figuring out what the right objective should be, and then actually achieving or optimizing the objective you've defined. And I think the first one is not gonna go away anytime soon. And if I think about a post training team and why you need a lot of people to be on the team, just because there are a lot of different areas where you have to figure out how the model should behave, and like, there's no way of, like, it would be very hard to automate the whole thing just because someone has to think about how how should the model behave in this area.

Speaker 1

Jane Street started using Antithesis to test their software in early twenty twenty five. And they were so impressed by the product that they decided to invest in the company. I recently caught up with Ron Minsky, who co leads Jane Street's tech group, to ask about how Antithesis actually plugs in.

Speaker 5

and nonetheless, it was able to shake out bugs that were otherwise gonna be really hard to find. And that's important both because it helps make those systems more reliable, but also because it helps the teams that build it to just move faster. This matters more and more as code production is increasingly automated.

I think in general, as we've been using agents more and more, the key problem that you run into is the verification bottleneck. Just the the time it takes from people to look at code and figure out, is that actually something you want to accept in your production software? And tools that make testing better are just incredibly helpful there.

They just ease the verification bottleneck and make it possible for you to get more stuff done and move faster because you can have more confidence that the code generated by the agent is actually not introducing new problems.

Speaker 1

To see how Antithesis fits into your development process, go to antithesis.comthorcache. What is the story for why there isn't huge consolidation in model providers?

There's just so many things that point to centralization here. Is there, is it, yeah, if you step back over the course of years, is there something that is gonna prevent that?

Speaker 3

is the main thing that fights against the centralizing force. Cause basically anything that can be learned through RL can be distilled very easily. Because a small number of bits, it's something that you can learn from a small amount of data.

So if you can get trajectories from the model that show a behavior, can easily distill it. So I think I think distillation is one of the things that fight centralization. There's also I mean, is a possibility that there'll be company specific models that it'll be possible to learn from deployment, and have a company continually improving its own model.

Speaker 2

change the game a bit. Yeah. I And also wanna point out that like continual learning and RSA doesn't stop distillation, right?

Like even if your model is improving every day, like people could be distilling it every day. And so it's like the loops could just operate at the same pace. Right.

That makes sense. Okay.

Speaker 1

I guess you need to know yourself what the right distribution to prompt is in order to get the relevant model behavior?

Speaker 3

distilling with supervised learning, the prompt distribution is extremely important. So it's very nontrivial to distill a model even if you have full access to it and have the cot, the chain of thought and everything. Yeah, it's nontrivial to distill all of the useful capabilities from it, because you need to prompt the model with something, you need to prompt it with realistic prompts, you need to have a really wide distribution of realistic prompts.

So, yeah, one thing that's been coming out recently is some of the of the Chinese companies are probably using these router services, which are designed to allow people in China to use The US frontier models, which would otherwise be blocked in China, but there are all these router or proxy services that allow people in China to use these models mostly for coding, and these router services are collecting and selling some of the data.

Speaker 2

distillation because it gives you the perfect prompt distribution. Yeah. I think this is one of those things where AIs help a lot here.

Like, if you actually look at, like, you know, the frontier pipelines, let's say, like the Chinese models that they've actually put in their papers, it's a lot of, like, humans or they get seed prompts from somewhere, which is some combination of humans, this kind of data, and then they synthesize a vast coverage from those seed prompts using their existing models or the other frontier models. And so it's like you can automate an awful lot of this prompt distribution gathering and environment creation. It's just, humans need to provide, like, increasingly few amounts of bits.

It's, like, the models get better. Right.

Speaker 1

like, having a service which has users or users are going through. So, Not necessarily. I mean, like, yeah, that's obviously very helpful.

But like theoretically, you can just think about like what users want or like Nobody's a lot of tasks. Like, the whole point is that we don't the user says, make me an application like this. Oh, that didn't work.

I actually want you to make this new feature.

Speaker 2

then you just have, like, RSI anyway. Yeah. I mean, like, ultimately, like, if you have this, like, fully automated loop, that is basically RSI.

Right? Like, the AI is deciding the data, it's deciding the training that that is the loop. But, yeah, I mean, like, it depends how much human information you need.

Like, at some point, you're just like, I want traces that look like this, you prompt that to the model, the model will be able to, come up with, like, a pretty good approximation. But what if you wanna do, like, make me a really good politician and then just, like, anticipate de novo? Like, how how would a discussion in, like, the senate halls go or something?

I just feel like there's gonna be a lot of things which are Ironically, this is actually, I think, easier for the distillers than the frontier labs. Right? Right.

Because the distillers just like, I want a good politician. They go to the, like, frontier model. The frontier model already knows how to be a good politician, so it just, like, generates the traces.

Whereas if you actually want to build the first model that does this, you have to actually somehow get data on what politicians do every So day and build it's actually much easier to say, I want something like this, and then get the AI to produce a billion variations than to like actually create the thing like this to begin with.

Speaker 4

that the Chinese labs have this routed data. So like I think the thing that that Jess did this originally was I was saying, isn't it weird how Sonnet five and Opus five are almost objectively worse models than GLM 5.3, QumuKate three, even though they've had access to not only distillation but logic distillation from like Mythos.

And so the counter here was that like, okay the prompt distribution really really matters. Like you need to see what users are doing so that you can distill like kind of these behaviors and things in. I think the prediction from this is that the Frontier Labs don't necessarily have much of an advantage if at all in aural environments now.

Because yes, user distribution matters for general behavior and so on but the best measure of a capability is the very, very hard aural environments you've made at the frontier. And so if you have access to those aural environments as anthropic and you have access to Logic Dislation and you've still made a worse model then maybe like That real world deployment matters more than the Yeah.

Speaker 1

That's really interesting. So but they had to incentivize those capabilities in the first place in in Fable or the Frontier model.

Speaker 4

Or like with like a smaller model or something. Maybe like maybe we're just in this weird, like, uncanny valley where, you know, like, actually trying to copy that frontier model too much, like the the strategic app or whatever it is, is just like too large. Like I think people made this point with Opus is it's like the difference between Opus 4.

6 and Opus five is that Opus five really feels like it's got this like AI as a judge checking every possible thing it's done. That's why it uses so many tokens, tries to think about all these things. But it doesn't necessarily have the big model smell of Fable to know when to stop doing that or when's a good path to go down or whatever.

The reach exceeds the grasp. Yeah. Yeah, I would offer a slightly different hypothesis.

Speaker 3

So, I would say there are a couple of different axes for the environments you can create. And one of them is difficulty and the other is realism. It's sort of easy to create, or it's comparatively easy to create a lot of difficult environments are just involve doing a much more complicated task or doing something that requires a lot more cleverness.

And you could say this is like the benchmarking distribution because a lot of the most prominent benchmarks just involve doing some very hard puzzle like task that's easy to verify. And then there's sort of like the realism axis where you want the model to be good in the realistic coding agent setting where there's like multiple back and forth as the human and there's like multiple objectives. And like, I'd say, like labs who are crafting the model behavior for the first time need to push in both directions, and to get good model behavior, you need to really push on the realism axis and have like rubrics or some kind of human feedback that's informing the reward function you use there.

But I think when if you try to do distillation naively, you end up just sort of matching the teacher on the benchmarking distribution. And but if you don't have enough of the environments that really exercise the capabilities in these, like trickier realistic settings, then you're not gonna get those into your student model. And I think maybe one thing that's happening is the big models, generalize better from the tricky, narrow tasks, to these sort of, more realistic tasks.

So if you have a really good, like realistic, prompt distribution for distillation, you can match, the big model really well. But if you only have this distribution of easily verifiable tasks, then you can match the big model on all the benchmarks, but, you do worse on, this broader distribution. So you might, that might even explain something about the smaller anthropic models, like SONNET five, though it's hard to predict exactly what they're doing to post train those models.

It could also be that they're always changing their post training stack, and they just got a few things wrong in some of these models.

Speaker 2

something too high and created some quirks that people really don't like. So it's like really easy to screw up post training in some way that doesn't show up in benchmarks. I mean, just one other sort of very basic point is just like the aicom like the Frontier AI labs buy all their data from big data companies.

And like the Chinese can also just buy the same data from data companies. And they are. Right?

And, like, they are. Exactly. There's a lot of people, like, you know, being annoyed about this.

But, like, if they have exactly the same data, like, they can buy that, they can also distill. It's like it's it means it's quite easy to, like, keep up, really. Yeah.

Yeah. Yeah. Okay.

Speaker 1

how the first models that are capable of automating AI R and D will actually be trained? Because there's a toy version which is this thing that Ryan was talking about which is you just have GPT eight try to build GPT three size models that are really good at like inner loop type challenges of beating video games that require continual learning or just getting to a certain loss with like the least amount of compute, etcetera. But John, I think you had an interesting point that maybe that's not the way it actually will happen in practice.

So I'd be curious about yeah. By the point in which you have you guys that are actually capable of automating AR and D, how are they probably trained?

Speaker 3

like the researchers taste, and, just like creating a lot of practice environments, which involve like doing multi step research projects. So I think, yeah, people will in practice do some combination of those two things and just each iteration, like patch whatever seems to be most broken in the last iteration. So like researchers will be using the AIs, a lot and, and we'll notice that they have some consistent weaknesses and then, those things will either be patched by like collecting human feedback or, like creating environments.

Speaker 4

Yeah, makes Maybe useful way to think about this is how much of the lineage we roll back and then let self play from there. I think in the limit you're picturing just giving them a GPU and maybe neural nets or something and saying like, okay, out how to train a model to do these particular tasks. The way it currently works is we go up to the very edge of the lineage and say, okay, here are the bugs Anthropic has found in their training stack in the last few months, we'll turn those into environments.

You need to train and get better on the frontier. And so you obviously lock in all the previous history of the lineage, but you could imagine a world in which you roll back to like before GRPO or something and then you have environments which like trying to get it to discover like the best form to like RL models on and then maybe roll further and further back. But I think we will be still so compute bottleneck that like people will just keep like staying at the frontier and like diffing essentially the bugs and whatever improvements they found since the last model version turning those into training environments.

Speaker 1

like new data between your model generations.

Speaker 4

like RLHF type stuff back into the model itself. And and it is distilling. Right?

And that's maybe why some of us feel like it's asymptotic.

Speaker 2

It's like you're always like just trying to get the last three months of progress. And that that progress is being contributed to by Ayers, of course, but it also still has humans in the loop and it feels like, you know, you're just constantly inching closer and closer to what the human researchers are like finding and capable of doing. Yeah.

I mean, the one thing I will say though is like, obviously, if you're just distilling on like trajectories, you can never go above it. But environments can go quite a far away above what a human can do. Like, it's very easy to design an environment that like no human can solve, but the AI can obviously still try and solve it.

And so that would be the path to like go ahead of just like what the human AI research is. Do you have like an example of like, in terms of RSI or, like like, know, working on a train set? But doing it even faster than a human speed runner.

Yeah. I mean, I feel like in AI research especially, it's very easy to define, like, goals, which, like, you know, you could say, like, the loss needs to be, like, 1.3 or something, and, no human can get there, you know, now.

But like that's a very extremely measurable verifiable task in the air. If the AI gets there, then then great. Right on building like a 100,000,000 parameter model that beats Minecraft.

Speaker 1

That's maybe too easy, but like beats like a much more complicated game or something.

Speaker 4

Isn't it crazy that a 100,000,000 parameter models to beat Minecraft, we're calling that too easy? Like imagine if you said that like five years ago.

Speaker 3

I would say a lot of research is not exactly like that though, where it's like hill climbing on a well defined goal. It's sort of more like, here's an intuition we have about some way models should be better, and then we also have some idea for an algorithm that seems to go a little bit in this direction. So let's come up with a task that is sort of designed to show signs of life on this approach and, like, see if we get some, get those signs of life.

And then if we do, we can make successively more realistic versions of the task. Right. It's like a lot more guided by intuition.

Speaker 1

And then the the inner loop is to elicit the or make test for that intuition rather than like the the the test itself leading to the insight. Right.

Speaker 3

for the eventual objective you care about or the practical, like, production objective. It's it's sort of you're you're relaxing your objective a little bit. You're saying, yeah.

Let's relax on the realism axis a little bit and find some methods that actually work, then try to get back to realism later after the method matures a little bit. Yeah. And then there's also more there's research that's more oriented towards explaining things and developing a theory or a sort of yeah.

Speaker 2

more informal theories for what's going on. Yeah. I mean, like, presumably, the models will be trained on, like, some combination of all of these tasks, like, some will be very easily verifiable.

Some will be like, oh, I misjudge, or, like, just ask the human, like, does this look reasonable? And then you will set the hope would be that like, these would all generalize to like these much sort of hard, of more vague fuzzy kind of tasks. Like, it probably will to some extent whether it generalizes enough that like, we could the loop can become like self sealing without humans being in the loop at all.

It's like unclear. Yeah. Yeah.

Yeah.

Speaker 1

seems to me that the plan for AI research going forward is, and you tell me if you think it's going to work or if you agree with this characterization. So the bet is that we will scale up our LVR training across millions of diverse environments, across hundreds of different kinds of domains. And what will emerge at the other end is an agent which has learned these basic skill or less than basic skills around being persistent, being able to triage information in context, eventually having like end to end optimization of working with other agents and things like that.

And such an agent will be very sample efficient within the context. You know, you've done research on how you actually scale up in context learning to make it like arbitrarily long, but you just keep scaling it up. And so what comes out the other end will something will be something that it basically functions like a drop in remote worker over the course of a week or a month.

First of all, do you agree that that is a bet the labs are making? And second, is it is that enough? Like, basically learning how to learn within the simulacra within a data center and then but getting deployed into the real world, but not actually like learning from real world deployment, only learning these meta skills from the simulated environments in the data center?

Speaker 4

Yeah. I think it's now hard to separate out like how much of the lab's effort is going towards like direct RSI versus like making generally intelligent models that they can continue to deploy collect revenue to fund the next big training run. I think for the latter, like yes, that's probably just the bet they're making.

And it's very clear like the pattern of like where these environments are going over the last few years. I mean like Anthropics lineage of environments is like a very clear example of this. Like, you know, first like we just focus on coding and like we're gonna get really really good at that.

And then the task horizon that we've got from coding which is probably the lowest hanging fruit in terms of like data available on the internet to create environments like their own internal stuff that they can turn into environments. Then we're gonna generalize, we're gonna go after finance next and literally just so much Excel data and all that sort of stuff in the RL training. And then it's PowerPoints, it's this long tail of the working economy and that seemed to work really well and like a lot of the other labs and thing, even the open source labs have now realized that that was the correct that And so but what is the implication from that?

Speaker 1

which will be human like in their ability to learn on the job, Why would you try to bake in all these skills of like working with PowerPoint or something? Wouldn't you just expect the model to be able to pick that up on while it's deployed? And so yeah, there's multiple different explanations.

One is just that this is we expect models to get there soon but they're not there yet. So why not amortize these skills into the model training? Another is that we're not concentrated on making it really good at widely deployed work.

We just wanted really good at RSI and this is just like a way up for us to like get revenue so that we can pour it back into a model that is actually like really good at doing RSI development and then like once the singularity happens, the thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, don't know if have a taste on like what how once you construe why there is so much task specific knowledge in these models, if the if the path is like this kind of generalization.

Speaker 3

Yeah. I mean, if the models were good enough at learning in context, then in theory you need to train them on finance. They would just be able to figure out, read all the books on the fly and figure out to do everything in the appropriate jurisdiction.

Yeah, and you could argue that you need to do a lot of this domain specific training just to make them more efficient. So even if they were smart enough to figure this out on the fly, you still might wanna do a bunch of RL and bake all these intuitions into the weights, so the model would be more efficient at runtime. Yeah.

Yeah, I'd say in practice, it does seem like model providers are going domain by domain and trying to strengthen the models in the highest value domain. And I'd say that that's one of the answers to why the models have gotten so much better.

Speaker 2

the most common types of skills. I mean, I think another thing is just that, like, it's not that expensive to do both at the same time. Right?

Because, like, the models are massive. They can easily afford in terms of that parameter to, like, learn everything. And there is likely some transfer in sort of even even if finance is not specific, the information is important for RSI, just the general meta learning of how to figure out what's important, how to have taste, how to do long horizon work is potentially generalized.

There's not that much RSI data in the world as well. It's kind of hard to generate, and that requires a lot of effort to if you can sort of amortize in this other data, get some transfer from it, you already have masses of compute and massive parameters space, so why not do that as well as obviously the direct commercial intent of selling your model. That makes sense.

Speaker 3

I mean, there's one question about whether this current paradigm of doing like sim to real will be the dominant one So basically, you look at what the real world tasks, are like, and then you try to create a bunch of environments that can be simulated, like in the data center, and, you can do RL on them. And I think, obviously, has been very successful successful, but it has a lot of weaknesses because a lot of things are just kind of hard to simulate, especially if they involve interacting with a bunch of humans in real time. Yeah.

Speaker 2

sim to real will be the dominant framework forever. I think sim to real has to be the dominant framework, while sample efficiency is kind of low. Because right now you need thousands and thousands of interactions with the humans, and no human is gonna sit there and deal with this, basically be in the loop of RL training.

Yeah. And so we kind of have to simulate that now to get But the samples you obviously if sample efficiency improves a lot, you'd expect learning from deployment to become a much bigger part of it.

Speaker 3

off policy, you can not take all the traces and even without re simulating everything, you can potentially learn something from them.

Speaker 1

Jane Street just launched a new competition and it's their most ambitious one yet. Design a Protocol Emulator ASIC. Basically, if you have a chip that you wanna test, you can connect it to this ASIC, and then this ASIC will simulate realistic traffic.

That way, you can see how the chip responds without having to plug it into a live system. Jane Street is looking for flexible, general purpose designs, not single protocol emulators. When I was chatting with them, they suggested that I start off by trying to implement what are apparently three very common protocols, UART, SPI, and I2C.

Jane Street also mentioned that they hoped that more ambitious designs will also tackle low speed USB and Ethernet and any other protocols that flex your chip's specific architecture. Importantly, your design should be reprogrammable rather than smashing a bunch of specific protocols onto a chip. If a new protocol comes out after your ASIC is taped out, your chip still needs to be able to handle it.

How exactly does it is up to you. But there is one hard constraint. Your design must target an open source 130 nanometer process node.

That's because Jane Street will pay to tape out the most novel submissions and send the physical copies to the winners. The competition is open till 01/18/2027, and working in Teams is highly encouraged. Go to janestreet.

com/thorcache to download the template code and get started. I wanna ask more about this because it's sort of weird that you have 50% of compute that's spent on inference that is not directly helping the model become better. One of the key advantages you'd expect eventually digital minds to have is unlike a human who gets to have fifty years of real world experience, a model will get to through all its instances, get to experience millions of years of deployment across all kinds of economically relevant work in the economy.

And right now that data is just not in a meaningful sense helping the model get better.

Speaker 2

across all these deployed instances. But when do you expect this kind of hive mind kind of crazy shit to be start I think broadly, like at a very basic level, this is already happening, right? Like, just in the next generation of models.

So, like, right now, can always see take your deployment data and put this in the pre train or the mid train of, like, future models, especially if you do, like, some kind of filtering or some kind of, judgment or annotation or, like, recent, you know, synthesization of that. How much do you think that explains the generation over generation improvement? I think it explains like quite a bit.

Mean especially like, I mean this is, you know, I don't know whether the labs do this because theoretically they claim not to train on people's data. But like the Chinese 100% do and like they definitely get this advantage both like obviously deploying. Is This basically what distillation is, they take out the models, they get some of their deployment data, they get some fraction of that by pinging the model, and then they train their next generation models on it.

And they can suddenly do it on their own models as well, there's no reason not to whatsoever. I completely agree with this. I think if you zoom out far enough, this is definitely happening.

Like you're picturing this like, and we're all picturing this, this is like what continual learning, like the holy grail is. It's like this very, very organic live loop of like an individual model, like getting an experience and like live updating on the spot and learning from that. And a lot of things break when you zoom into that level of granularity but the big labs are doing this, the closed models are doing this.

There's also early signs of life of people using open source models doing this at a much faster cadence. So a good example is probably Composer.

Speaker 4

You have some sort of model and are able to or Harvey's doing the same thing with legal agents. It is getting very specific environments from the data that you have for that particular task and things that users are complaining about and all the feedback that you're somehow extracting from your specific deployments and a lot of these companies have the advantage over the BigLabs in that they can use this data really really well. And then they will create environments, they will do a big post train of Kimi K3.

They'll go deploy it. They might do some online learning as well like Composer did online basically like re:Inforce for a long time. So yeah, there's still a human in the loop, there's still a human saying okay, are signals we care about.

Here's how we're gonna create environments from the data that we have. And there's still a longer cadence than maybe the one that you're thinking of. But it really is happening.

Eventually that loop will become faster and faster. I mean, composer thing is interesting because this is where the model like in cursor people like press tab or they don't press tab on the next completion that the model suggests. And based on that, every single day composer gets better at like predicting the next So that was that was the old tab model.

Like they actually did the same thing for the actual, not just like the tab model but the actual like generative model. Oh, that's interesting. And it's hard because when you do online reinforcement learning, you don't have groups, right?

You just have one user saying one thing and then you get one rollout. And so like you have a big variance reduction problem. And like Kursar's kind of fuzzy answer to this was like, oh you know we have very good heuristics which are able to estimate how much better than average this response was or how much worse than average this response was.

And then they would do this big reinforce update, and then their solution to whether it got worse or not was if it improved on cursor bench, they would deploy the new model every five hours, and if it didn't, they would throw that version out. Interesting. Yeah.

Think your biggest problem is actually just not knowing what the reward function should be for natural data.

Speaker 3

the code, the the edit,

Speaker 1

you might that might get reward hacked in some way. But but is it this seems like a bigger issue with the Sim2Real thing where the longer and longer Horizon tasks get, the harder they are to simulate within a data center. Right?

It seems to me already potentially at least even in coding, we're getting to over the point where there's like not some year long coding task that doesn't eventually require you to like talk to a client or interact with the company or interact with users. And if you think about the gamut of things we would want AI to be capable at, you want eventually, superintelligence should be able to like run a business or like start a new business and make it profitable or like have a profitable day trading in the markets or win a court case. And these are all things which are very hard to simulate in a data center.

Like an inherent part of the learning there is interacting with the real world. And so maybe they they may learn how to get better at these things from, like, the transfer between sim to real. But alternatively, maybe you do need weight updates from these kinds of interactions in order to get better at them.

And then if that is the case, if transfer isn't strong enough and you need do need weight updates, then the fact that the models are quite sample inefficient is like maybe a deeper problem. And the reason I'm curious about this, I feel like by default, I I don't see how you don't get some kind of crazy recursive stuff from prevent within the next ten years. But the one reason why that might not happen is in terms of like weight updates, the sample efficiency of weight updates, they just seem way far behind humans, right?

Like plausibly million fold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from, you know, like cold start to like finish finishing training. And so, yeah, this is all to say. First of all, is there is there gonna be a transfer between simulations and extremely long horizon, really complicated real shit that we want the AI to do in the real world?

And if not, does that really mean that like the the lack of sample efficiency in these models comes to bite us?

Speaker 4

I think maybe the way I'd break down like the two types of tasks in which models get good and models will like still continue to struggle is whether the task is cumulative or you have this non stationary distribution you have to keep learning and relitigating a bunch of stuff. So maybe an example of a cumulative task might be RSI. Like it's theoretically possible to maybe have less than a million token Python file which from scratch trains a model that is capable of recursive self improvement.

And every discovery that you make is kind of a line in the sand that you hold. If it's true that for RSI we don't need to discover a new attention variant or whatever. Like once you've discovered attention and then once you've discovered a mixture of experts and once you've discovered gRPO, you just add that to the training stack and that's there.

And a good example of this is 5.6 sole training, 5.6 terra or whichever one OpenAI told it to train.

It didn't have to go back and discover retention. Like it basically probably would have called a bunch of like scripts which is like pretraining.sh and posttraining.

sh and just did that. So like that's an example of like a cumulative task. I think the real world and the reason like people are thinking so much about like continual learning is it's not really a cumulative task.

Like imagine like in a law firm you have an agent acting as a legal associate. Like that's a very non stationary distribution. You have to be able to fit in your context like all the relationships between all the important people at that company which are also changing all the time.

You have like all these implicit ways about how things are done, where to find information, etcetera, and like that's not as clean of an example of accumulative tasks like RSI is. So I think that there will be this breakdown between tasks but you know, it's like if the labs realize that and they they do believe that RSI is cumulative in the sense that like we don't need to go back and discover some brand new like architecture or whatever Yeah. Then maybe more and more effort and compute gets focused on that versus the It's so unfortunate that RSI happened to be easier than apparently all.

Speaker 1

Yeah, don't know if you guys have thoughts on this.

Speaker 3

models are, today's models are weaker than humans in a lot of different ways, and some of them might have to do with sample efficiency in a certain regime, I mean, in some regimes, models are very sample efficient, like learning in context, but then there might be some medium length regime where they're less sample efficient because humans can do some kind of weight update more efficiently than models. So I think being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that. For example, having lower diversity of thought than humans, or being bad at certain kinds of long horizon judgments.

I mean, I think a lot of what people call taste is something about behavior that works in the long run and that people have realized works in the long run. Not everything, but like some aspect of taste, and especially for something like software engineering, like I think a lot of taste is like what are the systems that are gonna be maintainable and work well in the long run of this project.

Speaker 4

yeah, there's a variety of them and some of them are related to sample efficiency and some of them aren't. Maybe an interesting thought experiment is like, if you were able to give a model like a context window of, I don't know, a trillion tokens or whatever you would have needed to fit in like your experience prior to like let's say RLHF. And it's got all that experience in the context window and it has the same sample efficiency and in context learning ability as it does at a million tokens.

Do you think taste is then solved? Would it be able to make the same judgments that you did or is there something fundamentally missing apart from just a longer context window with the same sample efficiency? Yeah, I mean it would have to be trained to learn from that context.

Speaker 3

So I'm not sure, yeah, either it would have to be trained to learn the right update

Speaker 4

to make from that context or So you you have to take all

Speaker 2

in? Like your whole life research experience. I mean, you still need the data to train it long context, right?

Even if you could theoretically get a trillion context, you would need a trillion lengths of data to train it. Like right now, you have, like, eight context econ just missing if you had that. Like, I know.

I think, yes. I mean, this really just comes down to the question of, like, how meta learnable is taste from, like, shorter horizon episodes. And, like, I feel like there's no obvious reason it's long because, like, humans somehow develop taste with not having many long episodes.

Like, we don't live to be, 10,000. We have, like know, we we develop pretty quickly. Right?

And so, like, you know, if you think about, like, even, like, in a PhD, the difference between, like, a first year PhD student and, like, a final, like, postdoc or something, that's, like, five years maybe. And they've only done maybe ten, fifth, 30 research projects in total, but somehow they develop taste quite quickly from a relatively short succession of small things. And so theoretically, it's possible to develop it like that.

The AI obviously will have vastly more experience in which to develop tastes like meta learned. And then it's like, how well does that generalize to, like, really long horizon things? Is I think the question which I think is really unsolved at this point.

Like, we don't know.

Speaker 1

Going back to this question, eventually, should be a regime where AIs are learning a ton from each individual instance of deployment that they have. Well, currently, you could say there's a meta fuzzy process by which models do improvement for deployment, but I feel like it's a very weak very weak feedback loop. Do you see this around the horizon where there's this hive mind kind of learning that's very rapid?

If so, how exactly does it happen?

Speaker 3

Actually, would say that around will we get a hive mind that learns from all of its deployment experience? Mean, a big part of that is actually about incentives rather than being a technical question. So companies aren't gonna wanna have the model provider learn from all of their deployment because that might just reduce the advantage of their business.

Speaker 4

I think that maybe the economics of this will pressure not necessarily weight updates to one big common shared model, but modules that get subbed in. So a very obvious example, this is a LoRa, but it might be something else like there's been a lot of work to try and fit like an arbitrary context length into a fixed size. Like this is all the linear attention stuff and all that sort of stuff and like cartridges which are essentially KV cases trained to be very very compressed KV cases to fit in a lot of information.

That's another example of something that companies may be willing to sign up for if that gets subbed into the model and it's not like actually changing the base underlying model itself. So there's many different versions of learning from your data in real time and the latter ones are not really helping the big labs because they are just these modules. But I think the economic pressure will force the labs to go down that path first before they can embark on this.

Speaker 2

Which economic pressure there? Because I feel like even if you have a bunch of cartridges or laws or whatnot, can still just take all these traces and just like distill this, dump this to the pre training of like your next generation of models. Yeah.

So it may be a more indirect form of learning that the big labs are getting and that's obviously still really valuable to them. But I can't imagine a world in which we start off with like, you know, we're going to just directly train this one big model on all the exact data that we're No. Think it will definitely go through stages because I mean this is assuming there's one discontinuous event where suddenly we fix weight updates continuously and in practice I think it's much more likely to be like, the cartridges and stuff allow you to specialize in deployment.

Then you generate traces, you put that in your model. Three months later, come out with a model which is better with this stuff. You specialize it again, you consolidate it again.

And then eventually, we'll just make this leap faster and faster. Instead of every three months, release a model. Now it's like every week and then every day and then every hour at which point we basically have obviously solved it.

Yeah.

Speaker 4

you asked like kind of how far off the current paradigm we are from being able to do this. We've done a bit of research to this and people have done a lot of research. At a really large scale, when you wash out enough noise and you have large enough patches, this outer loop process of putting data into mid training, creating our own environments, it does work in some sort of continual learning regime.

But the problem is when you zoom in close enough at a micro level, it's like I've got one model and I'm trying to update it again for like a law firm or something and I'm trying to do that very continuously like with a relatively small amount of data. Like all the methods kind of break down a bit. So like if I SFT the model on just like successful traces off policy on policy, eventually in the very iterative regime, like when you're doing hundreds of these micro updates, you see both catastrophic forgetting, you see forgetting of previous information I've learned on top of the base model that was much earlier on and I see degradation of general use, general capabilities.

On policy distillation seems to push this horizon out a little bit but it still eventually succumbs to the same thing. And RL is not very good at it is good at getting capabilities in but it's not as good as getting like knowledge in. And like just this very explicit knowledge of like, okay, like this person does this at this law firm and like this is a very specific process we find.

And you have to put in a lot of compute to create the right environments to get the the knowledge in the Do you think that the fundamental issue here, why you get worse at these other skills or there's forgetting and stuff, do you think it's fundamentally an issue of capacity or it's an issue of techniques? A little bit of both. I think like SFT and even like distillate, like on policy distillation can be like way too destructive.

The reason RL is so nice is because it changes a very very small amount about the model and there's like a lot of evidence for why this is the case. And so it kind of just tweaks it in this very very very small like loss value to like get it into the right point.

Speaker 2

But that also then limits what you can do with RL, like how much you can actually change the model. Well, I'm sorry. You're saying like the reason this isn't the winner take all potentially is that it is just like very hard to distill that much information into the base model?

Without ruining something in an iterative version. Yeah. Like, it's easy to distill it into like a different base model.

Like, this is where I think it's mostly technique. It's not like it's definitely not like just like there isn't capacity. Like, if you had some model, you know, with all this data and you take like literally the same size model and pre train it from scratch with like all of this stuff in mid training, it will be better, and I think that's a lot of what's happening today.

And so it's very much like there's a bottleneck that stops us from just keeping training the same model forever versus just getting all the data from the old model and training a new model from scratch. And this is exactly as Charlie was saying, some combination like plasticity and catastrophe forgetting.

Speaker 4

train on like non stationary data because you're adding new data as you go, basically this is messing with the data distribution. So like the old stuff is just forgotten. And we don't really have good methods to like stop that from happening.

And maybe at the in the limit you're like just bottleneck by retraining the model from scratch with all this new information. Yes. Which, of course, is, like, very expensive.

Like, training model from scratch is is expensive. But you're gonna do that anyways. And so I mean, not necessarily.

I mean, like, maybe eventually, if you have continual learning, you never train a new model. You just, like, just have a model and it keeps learning and then it keeps expanding. Right?

But but there's there might be, like, some deep technical reason why that's very difficult because of these like I mean, that's a question. That's a question. Yeah.

I think we have pushed back like how much from scratch we need to do. Like, it is definitely possible now to take like the pretrained base and like do very good mid training on top of that, like, kinda continuously plus some RL from, like, different checkpoints that are later on in the training. That's looking more like a tuning learning, certainly not the case of, like, you know, take the most recent model, apply a couple of very small updates, and, like, iteratively, like, never lose it.

So but isn't this like I'm a bit confused because isn't this literally what happens during training or during post training or something? You just have like you have a model that's already gone through so much training and then you like distill some fork that's been further RL ed or something. Isn't isn't that literally what happens?

And like why is that It's it's still at a large enough scale I think that you're washing out a lot of like the noise. And you're not just focused on one distribution which as Brian said is like, that is now a very if you're just focusing on one task, right?

Speaker 1

I don't know, there's billions of deployed instances.

Speaker 2

washing out of noise and stuff from that. Right? Maybe that's scale.

Yeah. Yeah. I mean, I I think, like, definitely as as I was saying, like, can do continual mid training for, a long time and you can, like, roll back to a checkpoint given you mid training data.

But at the same time, like, can't do this like indefinitely. Like, if you just keep continuing mid training the same base forever, it just like get it does it's sort of asymptote at some point. Like, you can't just learn new stuff in that base.

And this is why people end up training new bases. Like, otherwise, you would just keep mid training the same base forever.

Speaker 1

Whenever I finish recording an interview, I immediately brain dump all my thoughts into Slack. Things like, what was the most interesting and what should get cut? This ensures that my editors have all the context they need to start editing the episode.

But it's not like these brain dumps have any clear timestamps, and my unedited recordings are many hours long. It can take a ton of editor time to even find the exact moments that I was referencing. So we decided to try adding a Grock Bot producer to our chat.

Now whenever one of my editors puts a rough cut of the episode, Grock Bot opens a transcript on its own computer and starts working, usually before I've even seen the message. It takes the notes that I dropped into Slack, and it highlights the relevant snippets in the transcript. It also uses a big case file that I've compiled with all my preferences so it can suggest potential edits.

And when it's done, it sends me its top clip candidates so that I can review everything from my phone. This has worked really well. Being able to send informal messages like I'm texting my editor and then having the transcript immediately reflect my preferences has just been so helpful.

Try Grock bot yourself at x.ai/bot. Okay.

Let's talk a bit about data now. So I'm generally interested in this question of how much of AI progress is just explained by data progress. Doesn't mean it will be necessarily hard to automate, but that's a separate question.

Speaker 4

that totally dominates human experts across every single field. Are we talking about, like, pre training plus, post training data, like, environments as well?

Speaker 2

And, like, I think the existence of this is obvious. It's just, like, whether we can create the right environments. Yeah.

And in the trivial case, we could just train it to output the Python file which like trains the actual superintelligence.

Speaker 4

Like just how to memorize in the way it's Like, yes, there's probably like a ladder of RL environments that is possible to construct such that you would get a AI researcher which is at least as good as a human researcher. But the effort to climb each successive rung grows like kind of exponentially. And that's going to be the two things that you have to trade off against as to whether how fast we're going to hit that final rung where it's better.

I think that's fairly clear. And I think there's like, we're still relatively early in our own environment creation. There's a lot of asymmetries that we exploit in order to create good environments.

So one of those asymmetries which we've talked about before is there's environments where it's easier to go backwards and forwards. What I mean by that is it's very easy to define this complex data generating process and this is this kind of latent variable you keep hidden from the model. You can generate arbitrarily complex environments and the model has to do a lot of irreducible token spend and irreducible work to figure out what that data generating process was.

There's asymmetries in terms of like you can inject information from the real world so like Anthropic finds a bug through tens of thousands of human and LML was combined and like turn that into a very very neat environment which in a single LM could theoretically find within like a few million tokens. So like there's all these asymmetries which we're cherry picking and like we're counting on like kind of this task horizon generalization. But I think, yeah, again, this is just gonna hit diminishing returns at some point.

At some point there's diminishing returns and how hard it is to create those environments in the first place coming up with them because you can't necessarily just have these really these process where it's easier to go backwards than forwards. You actually have to sit down and construct something that looks with humans a long enough time horizon, it's gonna be a really complex task to create. And then there's also gonna be like the compute and time bottlenecks for the agent to actually do those tasks.

Yep. So like, I think you're just gonna start seeing this like curve to flatten out.

Speaker 3

which is only trained on data up to 1930 on this, like, modern coding agent data, and and it did better than Claude three Opus on SweetBench. So so this model that has, like, no knowledge of code whatsoever can be fine tuned on a moderate amount of data and, like, behave better as as a coding agent than this much larger pre trained model is pretty crazy, and it kinda shows you that, like, once you have an example of, like, the right expert behavior, it's actually surprisingly easy to, like, copy that into a, like, relatively weak model.

Speaker 4

Yeah. But a counter example to, that kind of is there was a paper recently where they trained it up to like fifth grade maths. Uh-huh.

And also like primary school like English and stuff, so it was like a decent language model. And they tried to RL it to do like late high school and college maths and the gap was just too large, like they couldn't get it to climb at all. But if you did successive rungs of like, you you're seven maths and then you're eight maths and so on, you could obviously climb to year 12.

So like, again, it's just like what is the distance between the rungs on those letters and how hard is it to create? Yeah. And this just comes back to like the RL signal problem.

Like, RL is not very good at like exploring right now. So if you if the model can't like get in like, you know, a 128 rollouts, it's very unlikely to get signal to like progress. And this is why like in RL we need like curricula, whereas in pretraining we don't because that's not a problem for pretraining at all.

Yeah. And again pretraining data is different to post training data. And I imagine as we continue on, humans will be involved less and less but that doesn't change the fact that you're bottlenecked on how much signal you can extract from the real world.

So like there's a lot of signal in the world and that's true. There's people doing spreadsheet tasks, there's people doing legal tasks and all this sort of stuff but the capability frontier of where the models are at now, like how many bits in the world are actually really relevant to improving the model's capabilities? How many new math problems are being solved that are just beyond the reach or grasp of the current models?

How many new coding problems are being created or solved that are beyond the reach of the current models? And I think that's why the diminishing returns kicks in because even the world as a whole is not giving you the bits that are useful for tipping you into the next basin of capability. Mhmm.

Yep. I totally agree with this. It's really a question of where the signal is coming from.

Speaker 2

in pretraining, the signal is, like, already in common core. Right? Like, for the tasks that you care about in pretraining, the problem is it's not just, like, getting signal at all.

It's, like, filtering out all the noise that exists. Yeah. And that's quite an automatable process.

But, like, as the models get better, as we enter enter mid training and post training, the signal just doesn't exist anywhere in the original data we have. No amount of filtering will get this there's no hidden proof of the Millennium Prize problem sitting in Common Core. We can just filter until we see it.

Right? And so at that point, you have to get bits some other way, either from humans directly asking them to write out their reasoning or by creating environments where humans decide what environment should be created, what the objectives of these environments are, or some kind of training on like the human data that exists in deployment. Like, you have to get the bits from somewhere.

Yeah. Yeah. There's a question of how much of the progress in pre shredding is being driven by data.

Yep.

Speaker 1

Jerry Hahn, who's a student at Princeton, where we basically trained all the recipes from twenty nineteen till now pairwise with all the datasets from 2019 till now. So you're saying like GPT two on the newest dataset like UltraFineWeb and you train Delphi which is the newest training recipe or the open source training recipe on like the pile or some old dataset and you do like the whole grid and you see the getting to some level of capabilities, how much less compute does it take across this grid? And you see that the data seems to explain like nine x of a compute efficiency gain but the architecture improvements explains like a three x compute efficiency gain at a very small scale.

And so to the extent that that is true at large scale that most of the pre training computer efficiency gains are coming from better data, how much can that continue? Like, can you keep just filtering data more and more and building more and more synthetic data until yeah. Do you have a sense of how much this kind of pre training progress can continue?

Speaker 4

I think my my prior is that like, again, the low hanging fruit is like somewhat exhausted with like feels like we got the Internet as this big block and like, there's it's not like the Internet is necessarily like growing at the same rate. All the useful stuff on the internet is growing at the same rate. We've probably got a bunch of 0.

1% loss drops to go, definitely not as many as have currently occurred. But that's also really interesting that you find this cumulative 27 times improvement across both. I think like it was Epoch or someone who estimated like three times a year since 2019 which would imply something like three to the seven, like over 2,000 times improvement.

Where's that missing a 100 times or whatever coming from? That probably gives you a good signal of how much of this is post training.

Speaker 1

I think the explanation has to be that a lot of the compute efficiency gains are scale dependent and we're studying at extremely small scale. And that raises a question of do the data compute efficiency gains or the algorithmic compute efficiency gains have more scale dependence? I don't know if you guys are prior on that.

We just didn't have enough compute to investigate that question. I mean, like, just naively, right?

Speaker 4

the scale dependence of the architecture is, like, fairly well known. Yeah. And, like, you can fit a straight line to it.

Whereas, like, I would have no idea how to do that for, like, combining pretraining plus post training data and mid training data. Funny enough. I feel like data is actually more important with scale.

Speaker 2

saying just like an x percent efficiency again is kinda misleading because what an architecture does is let you reach a qualitatively new regime which you couldn't reach with the old architecture. And then within that regime, obviously, the data is like the primary thing determining it. But like, you know, if we say it didn't have like even like GQA, we're doing like full attention all day.

We wouldn't be able to do like a million we like we'd expect to do a million context. And then like because of that, we couldn't we could never use the data which is like actually at a million context. And so we couldn't get these capabilities, even though, like, if you just do a naive, like, how much does this do at, like, two k context where the architecture isn't unlocking anything than, the data, you know, will look much more important than in some sense it is.

Right? It's unclear to me that these things are, like, really just, like, multiplicative gains in this way. I see.

Sorry. But then but what does this take away for the scale dependence of data? So, I mean, on scale dependence, I think, like, a lot of the, like, mid training and post training data we have now is, like, actually gets better with scale.

Because, like a lot of it, like the very long context horizon environment stuff, really requires like big models to be able to like make use of them. Yeah. And like this is not know, if you try and train like your 100,000,000 parameter model on like three bench traces, it's not gonna get anywhere.

Like it's not gonna show you same kind of improvement that you would get if you train an actual sensible sized model on it. Yeah. And it's hard as well now because so many of the architecture changes.

Speaker 4

or DeepSeek. They're doing these architectural modifications with not just dropping the pre training loss in mind, but for instance how the models are going to be used in the real world. The inference efficiency, having some form of compressed attention in the deep sea models is not necessarily geared around this is fundamentally a computer It's just like, okay, we're considering how the models are going be used.

Right, right, right.

Speaker 1

parameter scaling will go as we're getting into more of a RL heavy regime. Like, I don't know I don't know how fast historically. Yeah.

You can look at sort of open source architectures and see how fast parameters have been scaling and maybe it's like roughly 2x every year for frontier open source models. And to the extent that like even frontier closed source models have like 100B or 200B active parameters. Do you think that like keeps 2x in year over year or another in an RL regime where you also want to conserve compute on rollouts?

Also maybe there is like a threshold effect where you have enough capacity and at that point does increasing parameters arbitrarily doesn't matter as much. Do you guys have a sense of in 2030 how many active parameters will a frontier model have?

Speaker 4

Yeah. I think for the next few years we're going to be like because we're so focused on doing longer and longer horizon rollouts for RL where like inference efficiency matters a lot. Yeah.

It feels like the the models aren't necessarily saturated on their ability to that where the bottleneck is still the environments. And so we might see like a little bit of plateau. Like I I I have a feeling that, you know, like mythos and and the GPT models are much smaller than like, you know, the 10,000,000,000,000 parameter range that people are talking about.

Even just naively comparing them to open source models you can probably back out that conclusion. So yeah probably for the next few years I wouldn't imagine a huge growth in the number of parameters but again there's so many different things to trade off here.

Speaker 3

there's lot of emphasis. It like depends on how quickly you know like Macore and then in house these guys can scale up the complexity of the oral environments they're training on. I would expect the models to keep getting bigger just because people are scaling up compute and the GPUs are getting bigger.

But I would say exactly how much they get bigger depends a bit on the scaling laws in non obvious ways. So one thing is that I think, like, data efficiency is gonna be a bigger driver than compute efficiency of, like, the exact architectures people use, now that we're getting to the regime where we're sort of running low on, like, high quality pre training data, so that might affect how sparse you want to make the model. And then I also think we don't understand sparsity that well, and it's like parameters are a different resource than active parameters, but it's like, sparsity has definitely increased a bit, but it's not clear that it's gonna keep increasing with outbound, there might be some kind of sweet spot.

There's an argument that sparsity should make data efficiency worse because you might have to learn the same thing on multiple experts, though that's debatable.

Speaker 1

plateau at some point at a certain level of sparsity.

Speaker 3

there should be less sparsity but what are the other implications on parameter scaling? I guess just that the scaling law you're necessarily looking for the most, you're not trying to optimize compute efficiency, so you have all your, choices you can make on the architecture, and, each of these gives you a different scaling law, then, like, traditionally you would look at some kind of envelope based on compute, so you would look at performance versus compute Yeah. And take the envelope of the best models.

But if we're making that decision based on data, so it's like we're assuming we can spend a lot of compute, and like we're sort of data is on our x axis instead of compute, then we just get a different set of Optima or a different set of models that are on that frontier. Yeah.

Speaker 4

the size of the models every year for the last few years. People have been training 1,000,000,000,000 parameter models for at least a few years. There was even an open source one called Falcon but Liam from Periodic Labs I think posted yesterday on Twitter about how an early experiment at OpenAI was training a 1,000,000,000,000 parameter model that was very, very sparse.

Oh yeah that was what they did before OpenAI. That was like, Oh before, they Google the switch transformer, yeah. So it was very, very good at knowledge but terrible at reasoning because it was so sparse.

Speaker 2

100,000,000,000 to up to 2,000,000,000,000 parameter range for at least a little bit and it certainly hasn't been. This is nice linear increase. Yeah.

I mean, I feel like there's two things, so as as Charlie was saying, like, inference efficiency is super important for our rollouts. And so, like, this will really push down active parameters quite a lot. And then I think the total parameters really depends a lot on the hardware as well.

So like you really need to get like very high memory bandwidth and like VRAM size to like actually be able to serve like multi trillion parameter models. And so like, you know, right now, you know, people still still using a lot of like h one hundreds and stuff. And so as everyone moves to GBs and then we'll get like more actual the ability to scale and actually serve and do large RL inputs at larger scales.

The data question I think is interesting because naively larger models are much more sample efficient in the actual data points. And so even if you're not saturating the model, it's still best to go bigger because the models larger models generalize better and get to a better loss for the same amount of data. And so right now, I think we kind of have a lot of data, that's not the constraint rather than computing.

So we're having small models, which are very inference efficient. But if computers are no longer the bottleneck, it might come back to larger models which are sort of undersaturated, but they have this generalization ability because they're much larger.

Speaker 1

If you just look at the basic Schinchilla scaling law and you just maximize out parameters

Speaker 2

Yeah. It actually decreases the amount of data you need to get to the same loss very little. Yes.

If you go to infinity on parameters, the amount of data you need, I think, goes on less than 10 x Mhmm. Just because the nature of, like, the the power of We're we're now on the way too much data side of the Chinchilla laws. Right?

So right now we overtrain more of Okay. The Chinchilla, so we could easily get back to a point to which as we're running out of data, move back to, like, the Chinchilla optimal point. We're even, like, a bit on the overtraining, like, you know, undertraining model side.

But surely, like, even with these new chips that come online and stuff, like, we're just gonna be so compute bottleneck for the next few years that that won't necessarily be I this could well yeah. This depends on the ratio you have training an inference compute really. It's like if you're super bottlenecked on data and not on compute, you should go bigger.

If you're super bottlenecked on compute, you should always go smaller. And then like, yeah. You can also use compute to generate synthetic data, so it's like one of these very hard things to predict.

Speaker 3

Yeah. I think part of the reason it took people so long to figure out the scaling laws in the first place was that if you don't get all these things right, then you don't get such a clean relationship.

Speaker 4

scales and where you don't have to change your hyper parameters as you change the model size. And and bugs have their own clean scaling laws as well, right? Like, you know, like West Kaplan forgetting the cosine annealing thing or like even just like not considering embedding parameters I think, And so that messed up the estimate at smaller models because embedding parameters are a decent size of the model.

Speaker 1

A bit on RL. So I feel like a year ago a lot of people were making this argument that RL will not be super successful at scaling for models. I think John you wrote a research paper where you were pointing out that models learn one bit per episode when you RL.

They basically learn did I get the answer right or did I get it wrong? And I wrote some blog posts earlier this year. Was like, it's even worse than that because when the pass rate is low and the models are very unlikely to get the answer right, it learns almost nothing at all from an RL episode.

But we I look at the models today and they seem pretty smart and it seems to be the result of scaling up RL. Baron, you had a post, I think a few weeks ago where you're trying to explain what's going on. But why has RL been more successful than one would have naively thought?

Speaker 2

I mean, so I think the success of RL comes down to a bunch of different things. So first, think what is slightly underestimated is actually the mid training. So an awful lot of, like, what we see as successes of RL actually comes from, like, very, very good mid training data, which is basically where we're essentially doing pre training but on synthetic reasoning data and the kind of environments that get the model warm started for RL.

And so this actually takes the model almost 80% of the way to the final RL checkpoint often. And then what RL does on top of that is it does a lot of essentially tweaking to the policy. And so this is one of the reasons why it doesn't need as many bits as you would naively think.

It doesn't have to learn all of these behaviors from scratch. It needs just a few bits from these episodes, which you do get. And then the other thing that I really point out in my blog is that these bits are actually extremely high signal compared to, like, regular, like, pre training, which is why you need RL at all versus just, like, SFTing on, like, successful reasoning traces.

Because it's exactly the bits about how to get the answer right. Well, there's two things. So, yes, one, it's exactly the bits about how to get the answer right, but, like, this is not exactly how you think of it because in SFT, you have a trace.

Right? You have, like, see a bunch of math reasoning and then the answer at the end. The bit is still there.

Like, you're still SFT on the answer token. So that bit is still there. What's important is that the objective ignores all the other bits.

So in SFT, you like have like, you know, to try and match like the exact reasoning token to the model produces. So you're essentially getting like too many bits about like the exact way this other model you're training on reasons. For RL, you only get the one bit.

And that means that that signal is not drowned out in the noise of all the other bits the model has. Yeah. And so that's what really it's really a super dramatic increase into the signal to noise ratio during training, which is why RL is so dramatically efficient in terms of steps.

Speaker 4

I don't know if you guys have thoughts on that. Yeah. I like, there's there's been so much debate about like what RL does to the model versus like, you know, mid training or SOT or whatever.

And like, you know, everyone talks about how, you know, parser one will go up, but parser two fifty six will go down. Like, very rare correct reasoning traces will be like down weighted and kind of like outweighed by a gradient signal from like easier kind of reasoning traces. And I I think the simple like way to view RL now is that if you have a large enough like a large enough amount of compute to sample, a large enough group size such that your probability of getting a bunch of correct answers is like pass some like not insignificant probability then like it will be up weighted.

And like to to Baron's point, basically mid training and more pre training like the the pass at one, the starting point for RL, like scales in a log number of pre training tokens. You can ask some very basic questions.

Speaker 1

I I I guess that answer makes sense and maybe there's empirical research which shows that this is what's happening. But then I just look at the models themselves and I don't know what's happened. Maybe you can give me a sense of what is the basis of the AI progress over the last year.

If it's Yeah, maybe it's just up weighting the the policies which were gonna do the correct thinking anyways. But it just seems like qualitatively, the models have gotten so much more capable.

Speaker 2

the model seem to be gaining? So, like, one thing I wanna point out here is that, like, it doesn't necessarily imply that RL has a small, like, effect. Right?

Even if you have a few bits and, like, you only change the parameters a small amount, like, the actual impact on, like, function space, the model lands. Like, the input to output mapping can still be, like, super dramatic. Know, even if it's even, like, one bit can change, like, your function space a lot and it can rule out like half the hypothesis space Yeah.

Which is huge. So like, I don't think it's necessarily the case. It's like small amounts of bits, small amounts of RL.

Once you're starting from a really good point means that like you don't have dramatic impact in behavior. At least like not necessarily. I I think it comes down to two things.

I think the first thing is that everyone was hoping that RL would like generalize this reasoning across like all these different domains. And I don't think we necessarily got this like horizontal generalization.

Speaker 4

Like just training on math doesn't necessarily make you the greatest coder. Like you do have to do RL on on code environments. I think what we did get though is like horizon generalization.

Like the models just learned how to use more tokens for longer and still make progress on some sort of task. And so like you can train on environments where they get longer and longer and longer and then put them into a completely new environment and yes, they may not have generalized the reasoning patterns which allow them to do well in that environment, but they've at least generalized the ability to continue on that task longer which is correlated with success. Think there's a paper called Edgebench which showed that the rate at which models can work for longer is like doubling every three months.

And so that's a clear evidence of generalization. And I think like the final way to think about it is like in pre training there's this idea of like quanta. So you have this very smooth like pre training loss And when you actually look at what's happening in the model, like the model is learning all these like very discreet like tasks and there's like all these like emergent points.

There's like kind of a phase transition like it didn't have induction heads, now it has induction heads. And there's like tens of thousands, millions, probably like hundreds of millions of these things. And you average them all together and you get this very smooth loss curve.

I think like to an extent, like a similar thing is happening for RL. Like there is this very slow out of loop as Baron mentioned of we will train a model and then RL it and then the next kind of model iteration of training we will dump a bunch of these synthetic reasoning traces into the mid training data. We're kind of hitting all these quanta for all these different tasks and on an individual task level it may look like a phase transition and like you're suddenly going from like a point 5% pass rate to a 90% pass rate on like a particular like finance task or Excel task or whatever.

Speaker 2

better models. Yeah. Mean I think a lot of this as well is just like I think RL does generalize a bit.

Like suddenly you get like some transfer between like math and code or like puzzles and math and this kind of stuff. Also, the sheer amount of environments I think that people are targeting is just vastly greater. So before when you try to do some task, which you do in your daily life, two years ago the labs wouldn't really care about this.

They wouldn't train the model for it. And now it's just so much broader. They have a lot of environments targeting this specific thing.

Speaker 1

probability on solutions the base model were already done and causing relatively sparse updates in the policy. But when I think, like, when I think about I think there's also another story about RL, which is going back to the Atari games and then AlphaGo coming up with Move37, the super creative move that because it was never initialized on human data, it can like think in ways that humans are not even thinking and come up with extremely creative solutions. Yeah.

Do you have a sense on when we should expect or if we should expect RL on LLMs to result in things like Move 37, just extreme creativity, even beyond human creativity because the like, there's just de novo de novo initialization of intelligence?

Speaker 2

I mean, so a couple of things here. Like, first off, I think that the AlphaGo is using MCTS, which obviously does, like, more exploration and, like, stuff than regular policy gradients. But I kind of also think that, like, RL doesn't necessarily, like, reduce the creativity.

And, like, I mean, even if we I think, you know, this is obviously qualitative, but if we look at, the, you know, the open air hugging face incident, like, these models were coming up with, multiple zero days at a time to, like, break out of their sandbox. And, like, this is clearly, like, some level of, like, Move 37 creativity, I think, already, which we just get from just, the general generalization properties of the LMs. Like, don't think it's definitely not the case of like, RL is like totally destroying their Yeah.

Like Especially on long horizons.

Speaker 3

Yeah. I mean, one thing that people call creativity is just solving hard search problems. So that's like Move 37 is obviously an example of that, or like writing some kind of poem that satisfies a ton of different constraints, so that's something AI is obviously gonna be extremely good at if trained for it.

Then, there's another way in which the models, like, the diversity of their outputs, is a lot lower after RL, and they sort of develop these ticks, and like, even though the models seem like they're good at writing, when you do some kind of like distributional analysis, you find that like they're reusing, certain themes, like all the time and they're using the same character names all the time. So there's actually It's not like you're getting the same kind of diversity that you get when you Like from human authors, you're sort of getting one really good like style. So I think that that kind of diversity has definitely been cut down by RL a lot, and in fact, since we were talking about distillation earlier, that's sort of something Yeah, one thing that's happening is that so many people are distilling mostly from Claude that like all the open weight models write the same way as Claude and use the same like, have the same ticks.

Speaker 2

kind of concerning to me that we're having this, like, this monoculture emerge. Yeah. Again, I don't think this is, like, fundamental to RL as, like, a method, though.

Mhmm. And same with distillation. Like, even with distillation, like, you're just training on the data.

It's like, just because your data is not, like, super bored, that doesn't mean, like, the training method itself is somehow wrong. It's a problem with the data. I think a lot of, for instance, the RL entropy collapse is basically due to exploitation of fairly simple verifiers when you don't have a huge diversity of environments.

Mhmm. Because for instance, the writing, think the writing is presumably graded by some judge and, like, the judge has some specific ticks and, like, the model is learning to award hack the judge, and that's why, like, it collapses. But, like, this is really a problem with the judge.

It's not a problem with, like, RL in general. Okay.

Speaker 1

Super rapid fire predictions about the future. So I want timelines on the following couple questions. By when do we have models which you can here's what the it feels like to a user.

You basically hire them as a drop in remote worker for all kinds of white collar work, not just coding, but I don't know, video editing, law, paralegal, etcetera.

Speaker 4

projects and require interacting with other people, etcetera, etcetera. It's, everything a human worker could do over a month. If you, like, mandated to use, a browser or or whatever rather than, these the again, the firm setting up the information to be like programmatically accessible like maybe a couple years.

But if it's not like browser based, like it can send Slack messages, it can do all this stuff, but I'd still probably say around a year. Yeah.

Speaker 2

Mean, I would say maybe like for the, like, full generality, maybe, like, three years. But I think, to Charlie's point, we will end up with, like, a lot of people, like, making their organizations easier for the AIs to use. And so you get, like, 90% of the way there before that.

So but the thing that's the the the diff between one year and three years, there was just literally, like Like, that I think there's gonna be, like, a long tail of, like, miscellaneous stuff, which, like, some human can do, which, like, will take the models, like, quite a while to do. Yeah. Like, you mean are you thinking of sort of computer stuff or, like, basic cognitive capabilities?

I mean, I think this is this really comes down to a question of like how quickly can we solve this kind of like online learning and like Yeah. Whether we can like get like 90% of the way there with like compaction and like writing files to yourself and stuff. And like that's my big uncertainty.

Speaker 4

Yeah. I really don't know. And and another like maybe an example of something that I wouldn't be good at is like, you know, if I have to like yell at someone to get something at work or like really push someone to get something done, like the model isn't just gonna do that.

It's just gonna be too nice. Yeah. Yeah.

I'd say there's a wide variation in quality of human remote workers.

Speaker 3

you try to hire someone like off of Upwork to do a software engineering project, there's gonna be a huge variation. It's like often quite hard to get them to do to do like a good job, or like pay attention to all the feedback you're getting, and like I would guess that in some cases it'll be worse, like the pre AI version of this was worse than what you can get now from existing AI. So I think it might end up being a little complicated, because maybe to some to some extent we already have this, like, for some, like, not so high quality of work, but then, like, then it's obviously, like, we're not, yeah, we're not matching human level in certain, like, higher quality, like, forms of work.

So but I basically agree with Charlie and Baron that maybe yeah. We'll, yeah, we'll have some version of this in a year or so that's, like, okay.

Speaker 4

will be improving from there. Like we ship the goalpost based on the very long tail all the time. Feel I like you've used this example before of doing your taxes or something.

This year I literally just told Codex to go get everything I needed to do and send it to the accountant. There was this massive list of stuff, had to use computers to click through and download some stuff and it didn't. It was like fine.

Was So like, I don't know. A lot of this stuff it can already do. Yeah.

Okay.

Speaker 1

Give you 10x total productivity uplift. Basically if it takes you a year to make a breakthrough now, you make a breakthrough every month. I think I would just refuse to give you a scaler on this.

Speaker 3

We might already be past that in some types of work. Let's say you're just trying to prove, you're trying to do certain types of math.

Speaker 1

I was just saying, but for you as AI researchers trying to make advance the state of AI research. Because how much are AI researchers sped up or uplifted?

Speaker 4

Somewhere between five and ten years.

Speaker 2

Oh, really? Okay. That's far away.

Really? You think it's longer than, like, for general remote worker? Yeah.

Interesting.

Speaker 1

I think you're probably I think I'm realizing you probably have very different definitions of fully general remote worker. I could have specified that. Yeah.

Is true. Because I mean, like yeah. Because obviously, like an AI researcher can be a remote worker.

And so like Yeah. No. I I'm picturing, like, you know, normal white collar work over the period of a month.

Yeah. I think it starts to diverge a little bit past a month. Like a very competent white collar worker, but not necessarily like a super creative researcher.

I would say like two years. Two years? Yeah.

10x? Okay.

Speaker 2

How about you, Bernd? I can kinda see that actually. Because like, it really is just like right now it's already like definitely more than 10x of like coding stuff.

And so it's like, if it can do even like one or two loops of like experimental feedback that would actually be massive already.

Speaker 1

uplift of AI researchers within two years, if you just plug it into like a very naive model of like AI progress and how much is coming from AI researchers and they're like there's like a 10 x increase in their productivity. Yeah. You you have, like, radically accelerated pace of AI progress starting two years from now.

Yeah. I mean, I think, like, this will mean that AI progress doesn't get bottlenecked on, like, AI researchers' ability to run, like, small experiments. It gets bottlenecked on other things.

Of course. Of course. But it it just, like, happens 10 x faster.

For sure. Yeah. Which is a huge deal.

And that also, like, helps the next thing, which makes gives you a 100 x speed up happen sooner, etcetera. Yeah. I'm happy to stick with longer on that one.

And what's what's the what's like the crux?

Speaker 4

Like my capacity to absorb information and make the, like, Bayesian optimal decision

Speaker 2

on the next experiment. Makes sense. Yeah.

I mean, assuming that, like, you can delegate some of this to the AI. So, like, the AI is becoming decent at, like, deciding, you know, it's run this experiment, it's got this result, it runs, the next experiment.

Speaker 1

in, like, uplift. And and okay. Final question.

An AI which is which dominates top human experts across every single field of could work that can be done over a computer.

Speaker 2

the AI will still do better than humans. So this is basically just like ASI? Yeah.

Okay. I would say like three or four years. The fuck?

Speaker 1

I mean, that doesn't seem wrong.

Speaker 3

AI is obviously being more getting more attention. So it's like one of the harder things, but it's like a lot of energy is being put into it, and it's also like not one of the hardest things for AI because it's like, it involves a lot of code and math, which models are really good at. Maybe for things that involve like three d and like spatial stuff and physical stuff, I think that that will take a little longer.

Speaker 1

especially if it's not like yeah. Yeah. If it's like mechanical engineering or something and it's not getting like the most attention right now, that might take a little longer.

But it also just include fields where there is relatively little data because of the nature of the field and it has to like learn that data on the fly. So for example, it has to become superhuman at like being an engineer at TSMC or something.

Speaker 3

Oh yeah, so you would have to assume that like the onboarding yeah. You can give the AI the same onboarding material and oh, yeah. Then there's some like, something has to be solved about, sort of longer horizon learning or Yeah.

Yeah.

Speaker 4

I'd say five to 10.

Speaker 1

It means that yeah. You think automating AI research is like ASI complete or something?

Speaker 2

Yeah I think so. Yeah. I think there's so many things in the world which like even if you have some sort of memory system external to the model and even if like context length grows a little bit, like there are just fundamentally things like even if you could research the information or write notes yourself, like you'd need more than a bit of the context field today.

Yeah. Yeah. I mean I kind of agree in like the the five year range, at least for like the stuff that like labs are focusing on.

But I think, like, there's gonna be a long tail of stuff, which, like, the AI could theoretically go out and learn about, but, like, no one has bothered to do it. And, like, the the computer hasn't been allocated to that, so that might take longer for, like, literally every single human expert. Yep.

And sorry. But by this, I also included, like, the ability to learn as fast as a human, a new domain. I mean, I think that's not necessarily necessary, actually.

Because, like, the AI will have vastly greater experience than, like, any human. Right. Thanks so much for doing this, guys.

Speaker 1

debate and discuss things together. It was very productive. Cool.

Thanks for having us.

Shared via Hopper