The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Latent Space: The AI Engineer Podcast
31 July 2025 1h 18m
0:00 --:--
Episode Description
We first had Nathan on to give us his RLHF deep dive when he was joining AI2, and now he’s back to help us catch up on the evolution to RLVR (Reinforcement Learning with Verifiable Rewards), first proposed in his Tulu 3 paper. While RLHF remains foundational, RLVR has emerged as a powerful approach for training models on tasks with clear success criteria and using verifiable, objective functions as reward signals—particularly useful in domains like math, code correctness, and instruction-followi

Summary

Nathan Lambert discusses the evolution from RLHF to RLVR (Reinforcement Learning with Verifiable Rewards), highlighting its application in tasks with objective success criteria like math and code, as introduced in his Tulu 3 paper. The conversation explores the challenges of training AI agents for multi-hop tool use, the importance of verifiable rewards versus human preferences, and a taxonomy of model capabilities including planning, abstraction, strategy, and calibration. They also delve into over-optimization, the future of open models, and the strategic moves of major AI labs.

Chapters

Introduction to Tulu 3 and RLVRNathan Lambert discusses Tulu 3 and the evolution to RLVR, a general method for post-training models using verifiable, objective reward signals for tasks like math and code.
RLVR, Agents, and Environment InteractionNathan explains how RLVR is evolving to include multi-hop tool use and a stronger notion of environment, particularly for agents needing sparse signals from multiple generations.
Open Data and Preference Tuning ChallengesThe conversation explores the difficulty of collecting reliable open preference data and the ongoing debate about the importance of human versus AI feedback in model training.
RLHF vs. RLVR: Book and Future OutlookNathan explains why his book remains focused on RLHF, arguing that RLVR is too rapidly evolving and might become 'solved,' while RLHF's alignment problems are foundational and enduring.
Frontier Models and Reasoning ApproachesThey discuss current frontier models like o3, Gemini 2.5, and Claude, comparing their reasoning approaches (reasoning-only vs. hybrid reasoning) and the role of search.
RL and Tool Use ChallengesThe discussion focuses on the difficulties of training models to effectively use tools, particularly when tools might be 'bad' or the model needs to learn to explore and backtrack.
Model Skills: Planning, Abstraction, Strategy, CalibrationNathan introduces his taxonomy of model skills beyond basic reasoning, including planning, abstraction, strategy, and calibration, crucial for advanced AI agents.
Parallel Compute, Over-optimization, and Reward DesignThe conversation explores parallel compute for robustness, how over-optimization manifests in different RL types, and the complexities of mixing reward signals and designing them for code.
Meta's Strategy and Open Models FutureThe episode concludes with a discussion on Meta's talent acquisition strategy, the potential for truly open models, and Nathan's long-term goal of building an 'American DeepSeek' and other future research ideas.

Topics

RLVR trainingRLHF trainingAI agentsTool useModel evaluationPreference dataOpen modelsOver-optimizationReasoning modelsModel calibrationCharacter trainingModel routingParallel computeSynthetic data

People

Alessio (host) Swix (host) Nathan Lambert (guest) Luca (mentioned) Mochi (mentioned) John Shulman (mentioned) Costa Huang (mentioned) Hamish Iveson (mentioned) Jensen (mentioned) Noam Brown (mentioned) Sarah Hooker (mentioned) Gwen (mentioned) Simon Willison (mentioned) Eric Schlans (mentioned) Greg (mentioned) Sam Mollin (mentioned) Amanda Asco (mentioned) Joanne Jing (mentioned) Martian (mentioned) Diamonds (mentioned)
Key Concepts (17)
Tulu 3 — A post-training recipe for language models aiming to achieve state-of-the-art performance on various tasks, particularly focusing on scaling up preference data.
RLVR (Reinforcement Learning with Verifiable Rewards) — A general method for training models using objective, verifiable functions as reward signals, applicable to tasks like math, code correctness, and precise instruction following.
Preference Tuning — A method for training models based on human or AI preferences, often using datasets like Ultra Feedback.
Constitutional AI — An approach from Anthropic involving multiple model variants and complex feedback diagrams for alignment.
Ultra Feedback dataset — A state-of-the-art dataset for open preference tuning that has been widely used in the academic community.
RL from Ground Truths — The original name idea for RLVR, which was deemed less general because 'ground truth' only applies to domains like math, while 'verifiable rewards' covers more domains.
Hybrid Reasoning Models — Models that can dynamically switch between different reasoning modes or components, often seen in models like Claude and Gemini 2.5.
Inference Time Scaling — The idea that model performance can be improved by spending more compute at inference, often by increasing reasoning steps or tool use.
Sycophancy — A behavior where models generate responses that are overly agreeable or flattering to the user, often a result of over-optimization on human preference data.
Chatbot Arena — A platform for evaluating language models through human preferences, where users compare models side-by-side.
Elo Ranking — A system used by Chatbot Arena to rank models based on user preferences, similar to chess player ratings.
Over-optimization — A phenomenon in RL where models exploit flaws in the reward function or environment to achieve high scores without fulfilling the intended objective.
Skills (Model Capability) — Foundational capabilities like reasoning, which are trained into models to achieve high benchmark numbers and inference time scaling.
Abstraction (Model Capability) — The ability of a model to break down large tasks into manageable sub-problems, crucial for complex, multi-step tasks.
Strategy (Model Capability) — The model's ability to determine the overall direction and specific steps of its plan for solving a task.
Calibration (Model Capability) — The model's ability to efficiently allocate compute, know when to give up, or ask for user input, avoiding overthinking.
Model Spec — A document or framework, championed by OpenAI's Joanne Jing, that outlines the intended behaviors, goals, and limitations of an AI model, aiding transparency and development.
References (40)
Tulu 3 paper paper
Zephyr beta project
Ultra Feedback dataset dataset
Lama 3.1 report paper
Anthropic papers by Anthropic paper
Vine PPO paper
QuietStar paper
DeepSeek company
DeepSeekMath project
DeepSeekR one project
Gemini 2.5 model
Claude model
Lama Nematron reasoning paper paper
GRPo algorithm
Strict Techery article
Retro paper paper
Bing tool
Perplexity company
Brave Search tool
Olmo project
Claude Code model
Arc AGI project
Loop paper
Retool paper
Toro paper
Humanity's Last Exam benchmark
MCTS algorithm
o one pro model
DeepThink model
GRPMath paper
Meta Ray Ban product
GPT 3.5 model
Stargate project
Apple company
Photoshop tool
Replicate company
Adobe company
OpenRouter company
Martian company
Hugging Face company
Transcript (95 segments)
Speaker 1

Hey, everyone. Welcome to the Lead and Space podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by Swix, founder of Small AI.

Hello. Hello. And we're excited to welcome back Nathan Lambert from AI two.

Welcome. Thanks. Fun to be here.

Speaker 2

I feel like I also have to say interconnects and, like, the likes of your podcast and, like, you and the AIU World's Fair. Like, you've just done a lot in the last year and a half.

Speaker 3

Not that many. Still say no to plenty of things. Yeah.

Speaker 2

Your first episode with us with us was January 2024 when you just joined AI two, then you released All the Almost. You joined us again at NeurIPS where you did the open models. Oh, well, Luca Luca did and you you supported.

And then you're more recently here in in SF for AIE. First of all, wanted to congratulate you on winning the best speaker. Oh, yeah?

For Thank you. The reasoning track. Here you here you go.

I I'm I'm limited by Mochi. Oh, a nice AI generated. I look too Zen.

I look so Zen and this AI generated we had our track host, like, take photos of you while you're speaking. And it turned into Ghibli photos. But this one, your eyes were closed.

It's fine. Okay. We we were we were trying to have Mochi, the reasoning, Fomsky join us, but I think she's like very getting very anxious, very very restless.

Those are too crazy, Mochi. Very restless. Okay.

Sure. Okay. So you've been you've been doing like really good work.

And and honestly, like, I think one of the things that we wanted to kind of establish was Tulu and ROVR, I guess. Is that a good place to start? Like, sure.

Speaker 3

journey. I think that we can recap kind of the story of what two to three was aiming to be and then kind of how it got folded into what the new narrative is. Yeah.

What the goal is is try to do the work to compress what our complicated industry post training recipes into something somewhat tractable that you can modify on your own and do post training at a what is, like, actual state of the art level. I think what we do relative to Frontier Labs is that we probably have a smaller amount of tasks. I think our post training suite for Tulu is probably like 10 to 15 tasks.

But I would guess post training at OpenAI at all, you have maybe hundreds of evals. And adding more evals is more data work and more mixing work and making sure you have these things. But unlike core evals for our suite of models from, I think, August or five b is based on Lama at the time.

It's like, it matches or beats meta on these core valves. I think meta has different priorities and their things for Lama 3.1, which is a great set of models at the time.

And it's just like, how do we distill what is very complicated post training explanations or diagrams from the like of this Lama 3.1 report where they have these complex feedback diagrams with many iterations and earlier signs of that from, like, anthropic papers that have these multiple model variants in the early, like, constitutional AI things for multiple years and I think what does that look like when you're doing a large scale instruction tuning into preference tuning and what else you might add. I think a lot of the core contributions of that before we talk about this reinforcement learning thing is like, showed this how to scale up preference data.

It's just like, the academic community had been using this one dataset since, like, all the way back in the hugging face models of like Zephyr bed beta is when this ultra feedback dataset got popular, and still a year later is like this state of the art dataset for open preference tuning. And it's just like one of those obvious things that doesn't need to be the case. So it's a big trying to make more mature, recipes available to people.

And I mentioned this either on one I think I'm trying to talk with Jordan. I mentioned the origin of the RLVR thing, which is like realistically when you work in the open, a lot of it is trying to match what industry has done. And we're on a different path because our infrastructure is different.

So some things that OpenAI does now that works really well for long context won't work that well for Olmo because we might not have enough flops in our base model. We might not have certain data sets for legal things. But directionally, like a lot of it is just trying to reproduce things.

And I've like long tried to get John Shulman on the pod of OpenAI and Tropic and Outthinking Machines. And at the time, he had gotten approval to like chat with me. And it's like and what he said was, so confirming a lot of the things I had said on instruction tuning and multitask and preference tuning.

And he was like, oh, yeah, everyone just does RL on outputs. And that's how we got the RLBR idea and scale it into something that is a general method. There's a lot of reasonable or very similar works at the time like Vine PPO and QuietStar on doing these math and coding domains for getting verifiable rewards.

I think the RLVR thing was about doing it in general recipes. Yep. And the naming was something that stuck.

Originally, we had I think it's especially like Costa Huang who's was our kind of lead RL engineer at AI too who's doing some stealth startup now. You can hear more from him now on that soon. I think he's founding engineer of something.

And Hamish Iveson, who's still a student at UW, were leading most of the technical work on this. And the naming was gonna be RL from Ground Truths. But then it's like, the verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.

So I think it's a nice evolution of the name which makes sense as you look at more domains which is now why it catches on with people. Like, once Jensen started using it, I was like, okay, that's that's set. That wasn't really our goal, but that's that's You think that's where it took off?

No. That was like in it being taking off because it was after DeepSeek. But it's like when people like that have the acronym on the slides.

And that's it's also very clear of like, RLHF is four letters. It's like we want to evolve that and have a similar four letter acronym. It's not that much magic to it, but there's definitely intention on these Sure.

On these little things.

Speaker 2

may not have worked as well. I don't know why. But yeah.

Yeah. And that's like that's what people that's what these people like all that were definitely thinking and they made that name change, which works, which was fun. You did mention so we'll we'll show you you kinda mostly quoted from the Tolu paper there, but we'll we'll show the RLVR You did mention that you wanted to change it now and we'll we'll sort of preview a little bit of the agent's discussion.

Yeah.

Speaker 3

RLVR, it's there's just a function, really, that checks if the you have a string outputted from the language model. You have a relatively simple function that's like, is this answer from the language model correct? And there's no real environment because you're just looking at the generation.

And now, I need to figure out the right way to communicate what either, like, multi hop tool use looks like for this, which is something people are definitely doing. Yeah. Thinking, like, what is the right diagram to encapsulate how o three is trained, which in action they take multiple actions because the next sequence depends on the feedback from the environment, which is some sort of information store.

So like when it's searching for a niche piece of information, it you you can't know what the next actions are without whatever feedback from my Bing searches is what they say they use. That is a step that is very much happening. And then as people try to transition to more end to end RL is a real strong notion of environment, which is that you're looking for a sparse signal from this multiple generations.

And that's what people want to do. I think it's debatable whether or not people are actually doing it now. I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof the system works, which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that when you put these pieces together or a couple different fine tunes of a model.

So it it seems like deep research has some fine tune of o three in it. And so you do that with some different domains of RL. It works rather than deep research being trained on the outcome, which I think makes a lot of sense for it not working in deep research because doing outcome based RL for deep research would be RLHF again.

Because you have to have two humans and you're like, which generated report is better? I you can definitely do that and you the whole thing at OpenAI showed that they have so many different reward models and reward signals in their post training. But that's just one of them, and I think a lot of the progress in making it exist is doing RL on a bunch of information retrieval and editing and search tasks.

Speaker 1

We talked with Noam about Noam Brown about this deep research and kinda like the verifiable rewards. He mentioned, obviously, that's an example of, like, nonverifiable thing, having RL work on them. And in one of your recent posts, you also talked about how the big labs have all this data that they can find long tail things to RL on.

And then kind of when you put them all together, that fixes it. Do you feel like what we're able to verify is like a big bottleneck? The, like, the verifications are only done in kind of like these smaller atomic things, and so we cannot really scale that.

Speaker 3

I think my comment was on making so in this post, was like, reflecting mostly on the question of what will agent progress look like relative to modeling progress. So we've had almost three years of modeling progress, and we're pretty used to the messaging on that. And it wasn't just about being with the RL on small things, but do any post training to fix a weird behavior.

And RL is a very data efficient way if you can get the right signal, but you could also just say, like, it does this weird non verifiable thing. Let's create a 100 or a thousand instructions to include in post training so that the model does this types of information extraction correctly or like soft extraction. It's a space that I wanna flesh out more with more examples of tasks.

It's just if you watch Claude code going, it's like, what is it doing in the background? It's a lot of reading files and even just the compressing context. Like, that's not I don't think that's really a verifiable thing, but that being messed up, like, that's a super crucial skill for long context actions and long longer tasks is just compressing well.

And that's gonna take some training novelty on how do you you can effectively modify your training data instead of having all the multi turn context. You just insert the summary and you wanna make the performance stay as well. It's it's also a cost saving to have shorter context.

There's just a lot of new domains like that. But do you feel like you can figure out what these things are before you release? Or do you think the labs have, like, a big advantage because they have so much user data that they can kinda, like, inspect this at inference?

I think it was mostly looking at real world data at this point. To the extent that there are clear benchmarks, you can use them in the open. But I I mean, we see the industry consolidate around data in different forms, and I I think that's a real important touch point for people.

Speaker 2

collecting reliable sources of open data that everyone uses.

Speaker 3

There's a lot of action in the space, but hard to get traction. Yeah. So I think for a long time, preference data has been something where people understand that it'd be very good to have large repositories of it.

If you want that, you can, like, annoy me to try to release all like, for two, like, we have a final dataset, but we have completions and ratings from more models. Like, I've been talking to the student who's let's figure out how to mark this down because we just have so much completions and LM as a judge AI feedback data that we don't know how to clean. That's one thing.

The the problem is I think a lot of it is task and model specific. So this notion of, like, on policy to adopt a RL word for just this preference data and preference modeling, which is that you want the sequences that you're training this reward model on, the sequences of generations to look like the model that you're starting to fine tune. Yeah.

That is something that has made it hard to kind of grab off the box. And it it's, for example, like this ultra feedback that I mentioned is just has a lot of models in it. So most models that people are fine tuning, there's some signal for it to improve on.

And I don't know how long that lasts, and we still don't have the answered question on how important human is versus AI feedback. Every time I check-in with people at Frontier Labs, they're like, yeah, we still use human preference data. And I'm like, okay.

I don't have access to that, and I don't know how to measure how much it gives you, really. It might be most of the benefit is on the what's the right adjective to describe chatbot arena? It's like people are down on chatbot arena, but it might be that the human data helps boost retention time and general preference a lot, where most academics were doing multiscale and alpaca eval type things, which it's just it's not as crucial to everybody's fighting in the attention economy.

Speaker 2

The attention economy. You're quick. I mean, since we're there, you mentioned Sycophancy, you mentioned Ellen Marina, that was one of your posts on interconnects that I really enjoyed.

Are they cooked? Is there a future for arenas? Like, how does this play out?

You know, they got a $100,000,000 now, like, what are you gonna do? I don't know what the money does for them, but I think the the eval is still valuable.

Speaker 3

at the frontier, people are very cynical, but in the compression race of how much cheap like, what is the cheapest model you can have that does pretty good at this is still so useful to a lot of people. Chat is king. Yeah.

I've never run chats with these things. It's it's why I use and g b d four point five isn't as good on chatbot arena. I think it's it's higher on like a Yep, which is a new competitor into this.

It's like they have like a vibe category, which Sorry. Yep? Yeah.

There's like yep dot a I. There's you can look it up. It's a competitor, another startup.

They have like a cat they all these companies have categories and one of their categories is vibes and GPT 4.5 is on the top. And I'm like, okay, there's something like this tracks.

It's a frontier model. Yeah. And it's just like, I I that stuff, intangibly, is very nice.

The leaderboard is established. People still should use it. It's kind of a focusing function for the community across different batches from industry to academia.

Yeah. I'm not gonna try to solve their monetization problems for them, but having clear norms and things that could be hill climb forever is very good. Like, having this idea of an EloLanking models That you cannot saturate.

Saturate. Yeah. You just It's kinda cool.

It's like it's a great problem. Like what is But you can game it. So I think that's the that's the issue.

Yeah.

Speaker 2

public about any of her like, has gripes, but she doesn't really go public like that. Yeah. Artificial analysis also has one, which I I think is kinda cool.

The other thing I think is relevant to this discussion is a lot of the data actually is like single test, like a single round. Like, it's not it's not multi turn. And I wonder how to create proper multi turn arenas because you have to switch the models as a whole premise of Elder Marina.

It depends on how valuable the user data is.

Speaker 3

If the user beta keeps being equally equally or more valuable than the inference, there's gonna be a platform to keep pushing this into more and more expensive things. Yeah. So they're gonna set up a deep they're I mean, they're probably setting up a deep research arena, because that's the data that I mean, if I was OpenAI working on deep research, that's the data that I want.

And there are competitors, and LMSYS is the entity that has the marketplace meant to set it up. Right. I mean, it's almost like how I see scale.

It's like scale kept climbing the edge of what AI data processes is and because they're the name brand, they keep climbing the incremental evaluation game and a lot of them have longevity.

Speaker 2

Yeah. Yeah. It's that's a it's a network effect in in some ways.

You mentioned scale, which is another hot topic, but, like, we'll we'll put we'll put all the sort of hot takes at the end. But I do wanna, like, know, focus, try to be technical upfront, try to you know, you're you're still writing the RLHF book?

Speaker 3

book now? I can give my spiel on it. Ultimately, RLVR is not mature enough nor is it as interesting of a book.

So I'm like Okay. On two so those those are the two fronts of why I don't want to rebrand, and there's also some personal career strategy. But that should be independent on, like, what is objectively a good book.

Because RLVR is gonna be changing so much in the next eighteen months. We've already seen it. There's all these new algorithms, but I think there's a lot more under the hood on how you do the right pre training for it and what the data is, how tool use emerges.

All of this stuff is core to what our LVR will be seen as. I'm watching to see if o three is like a niche model or becomes the path that everybody needs to follow on its kind of different style of tool use that you see particularly with search. Okay.

And we don't know how OpenAI did this. And these are the things that I think is kind of core to an RLVR book that we don't have. Whereas RLHF is a more inter disciplinary in the same way that chatbot arena can never be saturated.

RLHF can never be solved. And we kind of know these problems of alignment and over optimization and what the pipelines to getting data that people are using are. And, yes, I can add more RL algorithms to the book, which is nice for me to study, but that's not really changing it's not changing, like, oh, what reward modeling is and the different ways that people implement these today, whether it's a value function or reward model and stuff like this.

So I think the the breadth on RLHF is nice. And I think I I would tell a lot of academics that I think RLHF problems are gonna be foundational and kind of just have a much more steady study rate where we're on this massive spike of RLVR, but it might just be solved. And then it just goes back to zero academically.

It's not it's it's an embellishment, but there could just be a best practice for getting a 100% accuracy on any problem that you want.

Speaker 2

a preference is gonna go on forever. Yeah. Because it's verifiable, there there is a right answer.

Yeah. Sorry. What what do you mean by by, like, over the next eighteen months, there'll be a lot of changes?

Like, what what do you foresee? Actually, let's just catch up. Like, what's already happened, you know, in, the sort of recent history?

Speaker 3

Yeah. So there's two categories of information that we have, which is what are the models doing and what are the researchers doing. Yeah.

I think the models provide a lot of inspiration in terms of what the like, what's actual frontier is. And that's things like o three, Gemini 2.5, Claude.

These are a mix of just o three, I think is the most scaling RL approach. And then Cloud and Gemini 2.5 are very similar with hybrid reasoning models that you can turn on and off.

They've they've rolled it out in different ways. So Gemini didn't have hybrid reasoning at launch, but they're give they've they've brought it in and Cloud had it at launch. One of the most important questions has gotta be is is the o three path of just a reasoning model or hybrid reasoning models, like, more useful?

Do they diverge in their methods for training them? I think the NVIDIA Lama Nematron reasoning paper is probably the most detailed paper on a hybrid reasoning thing. And then DeepSeekR one is still the canonical recipe on a, like, reasoning only model.

And those are very different approaches, and I don't know if one will win out or not. And then there's just a lot of work on data side and RL methods. I think there's a list there's a whole list of kind of GRPo complaints that are out there where the math doesn't make sense for certain things.

Speaker 2

To me, every every paper I see come out always has like some fix to gRPO.

Speaker 3

I don't know if DeepSea is gonna come up with r two and just blow away everyone with whatever is next. Yeah. I definitely don't think the algorithm tends to be the most important thing.

Like, I think I had this in my AI engineer world fair talk, which is kind of a snarky of, how do you train a reasoning model, which is, like, you get a starting dataset, you incrementally improve the dataset, you do that until you're running out of time or your performance starts going up, and then you try all of these switches from all the papers or you turn all the you do a whole bunch of binary tests of all these various algorithmic changes, and you do a grid search and see what works. Like, candidly, that's why I dismissed GRPO when it first came out because I it was sold as an efficiency thing. Yeah.

And I was like, okay, fine. Like but like, you know, I I I I been trained to not care about efficiency because it's just a matter of resources. Yeah.

The GRPO advantage estimate is very well suited to verifiable rewards. Right. But the other thing is kind of a intangible works better on the infrastructure type argument.

And when it like, we came out for DeepSeekMath, which is well before the RLBR phase. So it was really marketed on as that.

Speaker 1

When you talk about hybrid models, how do you reconcile that with OpenAI saying they wanna move away from the model selector to just having a unified interface? Do you feel like they feel pressured to like, hey, look, when I have all these different classes, we wanna route them to the right thing? Or do you think there's something else?

Speaker 3

I would think that OpenAI wants to have a model that knows how hard depression is. I think that has to be the north star for most people working on reasoning, which is I the model will just spend the right amount of tokens on it. And if you look at a compute level discussion, see in like what inference time scaling means.

Mhmm. I think in plenty of ways, like, hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a 100 x less inference tokens. It's like, you just pay for it and compute and that'll get better.

I think it's like really like, that was something like Jensen said in his most recent, I think, like, strict techery highlighted it or had the interview with him. And it was like, yeah, everything's gonna be a reasoning model because it's gonna get so cheap and they're better. And I was like, that's why it's like the high reasoning thing is a little bit weird.

And it's like, I always just will turn reasoning on unless it's a really silly query like, oh, I like, what is this thing? So it's like, okay, like, in two years that kind of tracks, which like, I think o three is also just burning money on us. I mean, searches 80 websites for me asking what paper it is.

Like, that's a lot of tokens. But it seems directionally, like, if that's the thing that works, that'll be the default Yeah.

Speaker 2

those think the value is there. I wanted to double click on something that you seem to be coming back to a lot. You you you seem to assert that o three does something very different by using search a lot, much more than basically everyone else.

Yeah. Do model do all models come with a search engine now? Is that like a must have?

It depends on your use case. Yeah.

Speaker 3

retrieval or understanding it yeah. Yeah. There's old papers that we can try to find the links, I think.

I don't know if Sam Mollin was talking about it, but there's this retro paper from DeepMind and other architectures that people have been pulling in the discussion again, which is like, you have a very small model with a very big context length and a very big retrieval store, which I'm not one to bet against the transformer architecture and just figuring out long context and stuff like this. But those are ideas that people are bringing back, which is search search is better. You look at all the evals from reasoning models, and one of the trends is that like, Simple QA numbers all drop.

It's like DeepSeek r one to the new r one, it goes down. It's like all the new, like, QN 2.5 to QN three, Simple QA goes down.

At least when you're evaluating these without tools. And Simple QA is like a what is considered to be a very nice, fairly numerically robust, like long tail knowledge evaluation. And all of these the raw models, they're all going down.

Speaker 2

just to have this search behavior makes a lot more sense. Okay. The counter argument for this just I have been through this journey too of like, oh, why don't you make like a model that doesn't know anything but search.

Right? You can search out anything that you wanna learn just in time. But the problem is you'd need to know what the search terms are.

You need some baseline intelligence to make all this work. Yeah. That makes sense.

That's a that's a good way to put it. I think it's important because there's this thesis of like LMs becoming just online LMs, like, permanently.

Speaker 3

it hasn't been super pursued, like, Perplexity was one of the first to put it on my radar as, like, they were, like, we'll attach the search engine to the LM and that's what you get now. And I think, like, more and more people are starting to offer it as part of their default services. Like, Gemini has, like, a a search grounding thing as well.

I mean, it's what people say a big limitation of Anthropic is because it uses Brave Search which returns a bunch more like SEO slop than Is that proven? Because I I I don't know. I I thought they had their own index.

Okay. So I don't I don't have I haven't done detailed books Yeah. So I'm dealing with rumors.

But I I think they'll all do end up doing their own index. And it should it's one of those things that's like Google should have an advantage again. Yeah.

But who knows if they do? I also hinted at this in my post, but it's like, Hamish had tried to set this up the same student from RLVR playing with like, search in an RL model. And it's very easy to get the model to do tools if you prompted to, but it's very hard to get the, like, RL model to learn that the tool is useful.

And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it stops or it gets it on the eighty first. Okay. It's just the RL behavior that feels emergent from having a very nice way of like getting the model to learn to use the tool and it's not like like you can't SFTE this model to do this.

Like, it just really feels like they set up the environment right and it plugs into this deep research kind of line of work that they did and they broke down the problem into these sub RL tasks. So then it kind of lets it Yeah. Do this thing.

Interesting. I don't wanna be an OpenAI show all the time, but I was just like, I tell people to play with o three all the time because it's it's weird. It's excellent.

Speaker 2

the amount of work you're imputing on the deep research team when like, as far as I know, it's three people. It was Issa and like the two other collaborators that she had. I don't know if they did that much on top of all three.

Like every indication I've had from Over the Eye is that Deep Research is more or less a thin wrapper over, well, just all three. Yeah. It's probably like one or two small things that it they're like, oh, we can make our we we can make deep research work by adding this small amount of data to the training thing and then it just works.

Speaker 3

That is that would be how I describe it. I mean, it's I mean, what is it? Gwen, the anonymous person?

He replied to my q star post on Twitter the other day, and he was like, why was this all wrong? Yeah. And I I it's obviously, like, simple things don't scale.

There's a lot of complexity because there's a lot of other exciting things in the AI field at the time, and OpenAI kind of sends out a lot of things that confuse people. Yeah. This would fit into that, which is deep research is a minor change from an existing RL trajectory of what was like o three probably they had already figured out that search was gonna be better.

And then we're like, okay, we can repackage this. And it's a it's a simple thing that makes a big difference. Yeah.

And most of the things are like that once you have traction. Yeah. I think once trying to get the initial take off on the sigmoid is the hard q star thing.

But then once it's like once it's like this, a lot of things in the middle feel obvious, which is why I described one of the things that we work on for Omo. It's like a lot of it is just having motivation to do things that feel somewhat obvious, but they're still hard. Like, it's it's hard to get different recipes or it's hard to get a full reasoning recipe off the ground.

It's just like a huge change because you have all this inertia on this Vowel suite, And then you have to figure out if you branch your recipe or do you start from like, do we just take like open reason or zero and start from scratch which is like, it's a whole other headache of of things. It's just hard to move these projects that are anywhere above five to 10 people with inertia to get stuff done. But then once you're hill climbing, things can seem really obvious.

Speaker 2

Yeah. Okay. You covered a lot there.

Before my next question, just to close the brave thing. Our friend Simon Willison wrote a post that Anthropic added BraveSearch as one of the subprocessor in their product. Yes.

So that's where the thing came from. Now Well to what extent it gets used, we don't know. We don't know.

I I I would I would just kinda comment on a couple of things that he said, and then and then we'll go on to your your question. Yep. There's a very good post on just on the retrospective of Q Star.

There's a very good post that you had, which was that I wanna send people to is which is like it was open was o one a SIOP. Right? That does imply the question of, like, if o one was a psyop, what what else could be psyops now?

Yeah.

Speaker 3

There's definitely a psyops out there. I mean, the whole inference time scaling plot is such a psyop. Why?

Where you like put these two things next to each other with an x axis and it just looks like it's easy to control. Whenever you see an x axis, you think it's easy to control it. Whereas like for training on our the left one was training.

Yes. And training makes a lot of sense. So if you haven't you'd be especially even if you go to really old RL papers, RL learning curves are a non log x axis usually, and they look like this.

They look like these, like, whatever, like, logarithm or exponential rise. And then if you take one of these and you make it a log x as it's a straight line. So, like, that side is like, oh, okay.

We've seen this before with RL. But with inference time scaling, it being an x axis is why people are like, oh, there's a knob. I can turn search up a lot.

Yeah. Which is like what breeds all these weird ideas. The core of that article is just they're taking points from within training or there's a natural variance, and then you line them up.

And if you line them up, then you get this nice inference time scaling behavior, which is and now people a lot of people have reproduced this plot on inference time scaling, and it's it's it's much clearer now. But at the time, it's like, I I see why I thought it was a knob. It's like, oh, look.

It's a they called it inference time scaling. You control it. I think the most interesting well, you have a lot of interesting things in your blogs.

Speaker 1

But one that stood out was about RL and tool use. You said that it's easy in RL experiment to tell the model to try searching. But then if it doesn't get results with the tool, it's gonna stop using the tool very rapidly.

Can we impact that? So can there be a good tool that the model doesn't know how to use and then it kinda fails and then it stops using it? Can there be a bad tool that should be improved before giving up on it?

How should people think about designing the tool, improving the model, and kinda like where to intervene?

Speaker 3

This is definitely on the newer side for my things that I wanna work on or have worked on. I think particularly in 2026, especially in the open side, all the infrastructure models will cut up a lot where I want to go deeper on this in terms of, like, deeper search style things or very inference heavy multiple calls. And to answer your question that there definitely can be bad tools and there definitely can be like the model just using them wrong.

And something that I would want to see in a model is kind of not necessarily creativity, but like an openness that it doesn't know exactly what it'll get out of all of its tools and this uncertainty to just try a few different things, which almost seems classical RL behavior. But if you think about what a language model does, they're always very comp they're not necessarily confident, but they have like a path and like a direction in their answer. Whereas, that's a big change in these reasoning tokens is to have the notion of backtracking and and things like that, which is some sort of like openness to the tools having things that are unknown and it seems like a really nice thing for the model to have, which is like, oh, what if I try this?

What does it get? Especially in the on the open model side, which is if if this is gonna work where people want to use open models with tools, it's gonna be because people have private data stores and stuff. So if you were to train an open model that is gonna be a good reason or like o three, but on private records of some sort that'll never get sent to the cloud, like, it needs to be thinking of like, can try some things with this to get a sense for it before saying that I have to give up.

And if you look at tool, like, tool use right now looks seems much more similar to like code execution or it's just it's just a part of a sequential path that you need to get to, which is like, have a plan and if it fails at a certain step, I might have a backup. But it's not like this iterative of I need to fiddle with the environment in order to come up with my plan. It's just that I it's something that people probably are gonna have to train into these models, which is like, you might just tell it.

You're like, you don't know what is in this, but your answer might be in it, which is like a very odd prompt, but maybe it'll help. Yeah. When we had Eric Schlans from Anthropic who worked on the Cloud Agent before Cloud Code, He mentioned they spent basically, like, majority of the time on, like, the tool design to give to the model, and then you just kinda learn how to do it.

Are you usually well, I don't know how much you've worked on actual this stuff, but are you putting the tools one by one in the URL process? Do you think that helps? Or do you usually give is it better to give all the tools and let the model explore?

I don't really know. Like, we haven't gotten this to work. I would say it would probably depend on the model and your starting point.

If your starting point is already good at tools, it can probably generalize more. But if you're doing this weird base model RL and you have to have this kind of curriculum like, if you scale RL long enough, you're gonna need a curriculum of things getting harder. And like, that's pretty obvious.

So in that case, it might be tools get added when things become too hard for it to solve certain questions, which would be which sounds very intuitive, but also just really hard to manage in practice. Because what is your automated signal on your training run that is time to do that. That's why video games are so good because they're designed to unlock things as you progress.

But I think, like, with things like search, it's like, you know, if you're given access to a small data store or you're given access to all knowledge on the Internet. Good feedback for the ArcGI people for the v three benchmark is, like, have things where the language model needs to learn to use new new actuators in the world after a certain threshold.

Speaker 2

Yeah. That would be RKGI four then. Yeah.

Pearl. Yeah. Don't know.

They're they're cranking them out. They're cranking them out. They're actually doing a launch party, I think, like, in a couple weeks.

So I'm actually really like, it's fun to play RKGI. I don't know if you tried. Oh, I haven't.

It's pretty fun. Like, are IQ tests. I used to be like, oh, like, they weren't that relevant.

But, like, actually, now that we have a gradient where, like, LMs are actually significantly climbing them, now it's actually really more interest like, interesting to compare your own intelligence. I'm I'm with Noam Noam on no no harnesses.

Speaker 3

No harnesses? Yeah. Yeah.

I mean, harnesses are cool, but they're gonna they're they're a handicap that's changing the learning dynamic substantially.

Speaker 2

So it's good it's good demos, but I feel like the core thrust has to be no harnesses. I mean, it's always like, is it wrong to say that these are just inductive biases? Right?

Like, they're not in the model. Sure.

Speaker 3

This is a different task. I think I do it or I mean, I've I think I talked with Greg about this at Arc AGI, which I told him, like, do harness and no harness. You just have different categories.

Just like, you're trying to be transparent and build targets for Frontier Labs? Just do both. Like, I don't think it dilutes that much.

The no harness is gonna obviously be harder, and then you just get more bang for your buck on your benchmark. Mhmm. Yeah.

It's the same dataset.

Speaker 2

while we're at it. You had a really good summary of, like, recent work in multi tool RL, which was which had, like, loop and retool and Toro and all these other things. And I think that this is just, like, an area that's super rich for research right now.

I just wanted to give you the space to, like, highlight what are your favorites, what do you think that people should explore?

Speaker 3

I could share what my moderate ambition, what would be fun research project things is. As you want to create some sort of competitive dynamic or a vowel, and it has to be so much narrower than what industry is doing. So I I I told you this at lunch, which is, like, deep research, but only archive papers.

So you, like, don't have to do a full index. You have a limited domain. You have to figure out how to measure it or something or some I think, like, act it's good for academics to work on academic tools because they have very high domain expertise.

They already know what's going. And just, like, figure out how to make that something that is either very useful to users, if it's gonna be good enough for that, or something you get to climb on. And I don't know if those are like it's like brainstorming on the fly of like, take related works out of papers, just look at the text and break all the links, and make an eval which is filling in hundreds of related works with archive links.

Like, that's a fun deep research style idea. See if you could do it with open models on a set data store with tools. AI too has gone through a lot of discussions with this, which is you if you're trying to have impact in AI right now, it's as an academic, you have to, like, level up out of papers to artifacts, which is models, data sets, vowels.

Data sets and vowels are easier for people to have impact on. And then the next thing is, like, what do people actually use? In AI too, especially in this, like, semantic scholar team that's now working on, like, information agents of different types.

There's another thing that I'm like distance in, so I don't have all the names. But it's can we make open models do that side of thing better? It's like, you make something that people actually care about?

And then you're that's a whole level of impact that's much higher if you have actual users. It's it's hard for academics and small institutions to do that. Mhmm.

But if you're working on agents like dog feeding is viable, it's like, can we make ourselves a good Slack summary bot that we like or something? And just making these agents really tractable. I mean, that that's one direction.

Another direction is just hill climb on humanity's last exam with tools. I I just think it's kind of unlikely that we're gonna win as a academic and a state of the art number because they're gonna start spending millions of tokens per query. And it's just a lot of it's a lot of computer and like the getting beating that on the flop equivalents is gonna be so hard.

Unstructured thoughts is something that I'm mostly like, okay, I'll get to this. Like, I have I have more things to figure out on the modeling and what I call like skills level, which is just how do you do reasoning to induce inference time scaling and get high eval numbers. And once you know you can do that, you could take your knowledge with you to do it in more specific domains.

Speaker 2

There's skill and there's skill acquisition. Right? I think the Arc AGI definition of AGI.

Speaker 3

I quoted it. What is it? It's like efficient yeah.

Skill acquisition efficiency. So he's described it as three words. Right.

Yeah.

Speaker 2

Emphasis on skills in in your recent talks that you've done, do you wanna sort of reiterate that that thesis for people to pick up on? Yeah.

Speaker 3

OpenAI, etcetera, are doing probably now if it's not in their models. And with all the agents, it seems that planning is a very critical task. So it's kind of how do you come up with the taxonomy for different types of things you need to train into reasoning models for when it'll be a bottleneck.

And the found so I came up before, and the foundational one was skills, which is what I would say that we have already done with o one and r one, which is you do a lot of RL, you show the inference time scaling works, and you get really high benchmark numbers. And then the next three are kind of what comes next. And most of them are around planning.

So what I had is three and four on my list were abstraction and strategy, which is trying to not use planning because planning is a word that people already use a lot. Their strategy would be the direction the model should go in and like, technically, like, what are the steps of its plan? And then abstraction is how does it break it down into things that can actually solve?

And then the fourth last thing is calibration, which is just not wasting compute and knowing when to like give up and ask the user things because like overthinking is obviously a problem. It's easy to keep getting your eval scores to go higher by using more inference time scaling. But eventually, like, that's not what people want in their models.

They want they want a smarter training regime where the model is actually getting proportionately better for its training and not there there's a lot of papers on overthinking and stuff like this, which I think is, like, OpenAI wants it because they have to foot the GPU bill. Like, if o three just infinite loops itself for a bunch of people, like, that's not good. Does it actually I don't know, but it might.

Okay. I mean, like, these reasoning methods definitely can make the models just kind of unstable and just yeah. So it's like but it's also the GPT five idea, which is how do you get a model that just routes the question to the right not maybe not necessarily a router, but just knows if it needs to do a plan or if it can just answer.

If you look at DeepSeekR one and you ask it a hard math question, it's not like, here's my plan of attack. It just starts. And having a model that knows when to be like, okay, here's my plan of attack.

I might need to make myself a memory store. I might need to take like a Claude code approach for this query. I'm gonna build a memory store and spin up some parallel searchers and then come back.

Conceivably, this is all something you can train into a model because the the searches or the parallel models could be like tools in that case. The simple way to describe it is we have something like thinking tokens and an answer tokens, and it's the model should be able to optionally have like plan tokens before before thinking or before using tools. It's like, okay, like, here are the table stakes.

I need to do these things and these sorts of tasks will be harder versus easier. It seems more tractable than some far out ideas for AI. It's like I like a language model can write a good plan, it just needs to be asked to do so, which I'm would bet that Cloud Code and Deep Research are doing this.

Like, you get a user prompt. And first, the model is like, yeah, there's a plan tool in Cloud Code. And they first they break it down and it's like, that is something they've trained into the models.

Like deep I don't think DeepSeek has doesn't have it built in, but it it probably could do it. And just thinking about that interface between, like, if the model needs it to be able to do the task end to end on its own, like, can it do that sort of thing?

Speaker 2

reconciling this approach with the no harnesses thing is that I think a lot of the way that people, especially engineers want to model it is that the plans and and the memories are tools. And there are no special plan tokens, there are no special memory tokens, it's just context or it's just, you know, whatever. Specifically for planning because then you can do fan out to other agents for tool calls and stuff, so it doesn't have to be sequential.

But I'm just like, is this a fork in the road? Like or, you know, do we have to make a a real choice here as to do we outsource things to tools or do we keep it native within the model's tokens?

Speaker 3

I don't think it's a subjective difference. I think mostly the planning ideas to make the point that people don't get things for free. And the planning improvements might be kind of mundane, which is like we were prompting Claude and its plans were bad in this way.

Let's give some data where its plans are more detailed or break things down into more steps so that it's easier for them to do it. Yeah. Because it's a it's in a black box effectively.

So if it hasn't been targeted, it's unclear of what the performance will be. Or on the, like, open model side, it might just be the idea of having different models for different parts of it. Is that then you're really training a model to just be good at planning.

And, like, that's that's data that you need to come up with. I mean, did you, like, you only use that model for that one part of it? Does it feel like plants are much more reusable and should maybe not be generated every time?

Speaker 1

for certain sets of tasks, you wanna have similar types of plants. So maybe it's not the right way to ask the model to regenerate a plan every time. There should almost be like plan blueprints as like tools and then the model fills it in.

Like, where do you think the balance should be?

Speaker 3

I think they're reasonable. A plan is obviously an intermediate goal. I just it seems likely that there's, like, failures on this kind of planning level.

I mean, the same thing goes for these rubrics that are popular, whereas a lot of the technique that is popular for so called rubric things is you have a prompt and you have a language model generate a rubric for that prompt, is a few specific things that needs to get right. And that's conceptually very similar to making a plan for for every task. I think whether or not it's like grading is that you're gonna have a different type of abstraction than executing.

But I think in what people are seeing is that it's cheaper relative to the effectiveness to just generate it. So like, plans are not super long and they probably they're not that many tokens. So it's probably just kind of like, okay, we do this.

Like, putting it in my taxonomy might be overselling it where it just needs to be a prompt and you just need to make sure that your model's not too weird at that prompting stage.

Speaker 2

I think your taxonomy is super useful, by the way. So skills calibration strategy abstraction. I feel like maybe abstraction might be the most underrated one or hardest to solve.

The way that you introduced it, it was different than how you wrote in your blog post. You said it was basically not to overthink.

Speaker 3

That's calibration. Yeah. Abstraction is about breaking things down.

Yeah. I think both of these strategy and abstraction make the most sense on the hardest tasks that we don't know if the model can do them. Right.

So is it like if you're assigning a task to a model that you don't know if it can implement the strategy is very important because it needs to be very specific and narrow. Where if it's doing mundane code, like, deep research, the plan is actually not that interesting of a thing. Yeah.

But when you're at the frontier of if it can, like, I don't know, some GPU implementing thing, you could buy in buy into the open AI and anthropic narrative, which is help me implement this research idea in our complex, like, more distributed GPU thing. Oh my god. It's like, this is a task that's hard for a human.

And for an AI to come up with the right plan to debug and do this is very narrow path. So therefore, the strategy is pretty important of does it start with certain tests and how does it actually build this out to complexity? It's obvious that I need to come up with more better examples for this, but I think as you push it, it's more natural to see that there's only a few plans that actually get it done.

And then abstraction is just important as your task becomes so big. It's like a prompt engineering thing almost. Yeah.

And it's like, you only have a 100 k tokens you can generate. Like, you need to make sure the model breaks it down. So it's not just spawning a ton of infinite processes under itself, which I do agree that abstraction is an an interesting one, especially when you start to think about these models that could call in other models to do some tasks for it or parts that can be paralyzed with like multiple searches or just more compute.

I think that kind of folds into abstraction, is just like how do you approach a certain nugget of the problem. And I definitely say like, I don't have experience building this. It just feels like if you're gonna visualize AI doing the hardest software or other tasks, it's something that humans are very good about.

So it's like, how do you come up with a research plan in ten weeks? Like, there's a lot of how do you prioritize which experiments to do?

Speaker 2

lot of inductive biases that go into that that I don't like, a language model would not do well at that right now. Probably memory would be helpful there. So you can just skip like, the way we do this in real life is we could accumulate experience.

Yeah. One thing I I did want to dive in on on was just parallelism in general. There's one case where with o one and instead of the the the sort of q star ideas, there was one case where it was sort of overhyped in some sense.

But now it's coming back with o one pro and DeepThink. The theory is at least you correct me if I'm wrong. Basically, they run o one eight times and then they have a reward model rated and then give you the best of of the eight.

Yeah? Something like that. Something like that.

Deep think also the same. I we don't know any any details beyond that. I think there's a lot of people exploring that, at least on the info provider side of, like, you know, how do we parallelize search and planning and and all that.

And I'm worried about getting too hyped about it. I think it makes a lot of logical sense and this is one of those things where MCTS also made a lot of logical sense and we were fooled.

Speaker 3

Well, I don't think we're using parallel compute in a way to search over, like, low probability tokens. We're using it to get robustness. If you like, o one pro is it was so nice because it just had a a very predictable depth to it even on niche topics where, like, sometimes models just fail.

Yeah. You you had some numbers that went to, like, from, like, 10 to, 95% or something. I don't remember the exact numbers, but that's what it feels like.

It doesn't feel like you turn on o three pro to make it 10 times more likely to find some niche piece of information. Like like, maybe it'll be a bit more likely, but we're not getting that type of, like, searchy notion of getting more breadth or depth into our tree. So I think there's value there's value to it where we wanna use this parallelism on the what are either like the most important tokens that we're generating or like, okay, I know this part is crucial.

Let's let's just spend a bit more so that those tokens are better. But it's not a, like, transformative thing. The part that's potentially interesting on the transformative side is, like, if you can get much better verifiers.

So I think of verifiers of changing the slope of inference time scaling. You spend more tokens at inference, the better verifier you have. If you're doing parallel, it can extract a rare occurrence.

So like right now, if our verifiers are only like, they're good at like human preference, it's like, okay, we don't need to we don't need to crank that up very much. But if we are doing really diverse generations and your verifier is better, it'll get better it'll do better. I think you could look at the extreme between a reward model and an oracle, where it's like, the oracle is the more you search, eventually it works.

So the slope is is good. But a reward model is like there's really a capped signal Mhmm. Out of it.

At least if you're doing this preference type of thing. So the slope is pretty minor and it kinda has diminishing returns. So I do think that like, if you could fill that with more interesting verifiers, there's potentially more to get out of parallel compute.

But I I don't think it is, like, as transformative right now on my outlook. It's more like parallel agents makes more sense. Like, if you could break down abstraction nice, like, as a throughput engine, if our tasks are taking a long time, rather than a like, at peak performance engine.

Okay. Yeah. Which are the kind of fits with the whole agent versus model thing, where agents are much more about like getting it done at all, like, being robust and being fast for, like, this model is one generation.

It's like, can you get the answer right? Yep. I will spend a little bit more time on this and I'm happy to move on.

Speaker 2

it's a way to pull forward a hypothetical future model that you can then distill from. Yeah. Which is nice.

Speaker 3

Well, I bet people I mean, they surely will use these for synthetic data. It's just like the marginal gain on synthetic data is always very high. Or it's like Amanda Asco will say, like, better prompting will effectively make it seem like you have the next generation model.

But like most people don't put effort into their prompts. Oh god. Okay.

Or she had said something of those lines in one of her and throughout the interviews. So to just like, if you can really figure out how to kind of get into the certain states of the model. Yeah.

Well, anyway, that that that's my pitch for, like, why this is worth doing at all and, like, you know, I have a science fiction story that I wanna write about quantum models. In a world where, like, we you could explore cheaply multiple universes then, like, know, sort of pull forward the right one, that would work. This sounds too science fiction y, but I feel like in a world where we could control quantum computing well enough to explore this and scale it up enough, it could be kinda cool.

It also could be that parallel compute is grounds for interesting types of innovation. Like, I'd like I don't know, like, what does it mean to have parallel compute with diffusion language models that generate all their tokens at once? Like, does that meaningfully change some sort of application?

I don't really know. I think it would be like, the diffusion language model would be fun if it works. So you can have much more control over inference time scaling.

I mean, like, Gemini has one. It's, like, hard to suss out what it changes. But once we have all these knobs, I'm hopeful that it helps build some interesting types of, innovation because like the parallel stuff is new, architectures can change.

We'll see.

Speaker 1

I've been using the Codex Bessel van thing, and I feel like most of the generations are, like, you know, 5% different from each other. Because you use Ruby. No.

No. No. I have a JavaScript one.

I have a JavaScript one, so I should be good at that. I don't know if it's, like, just how the RL encoding works. One thing that I've noticed, these models always wanna do if statements when there's, like, a missing m variables so that it doesn't fail when it runs.

And I feel like that to me, that's just, a symptom of the RL. Yeah. The code is terrible.

Like, no you should not write code. Like, it shouldn't silently fail if there's missing variable. It should just raise an error.

Yeah. But I I feel like the URL is, like, pushing the code in this direction. And then all the generation have the same pattern.

You know, I generate four thing, all of them use the if statement just in different pieces. Yeah. That that was something I will definitely get over.

That's just like the labs are trading off massive gains in performance or small detriments in usability.

Speaker 3

And it's do you ship that model? Yeah. Like, you just ship it and deal with it later.

But I'm sure they can I'm I'm sure that's a fixable thing.

Speaker 1

pieces of the thing, but not in the full trajectory sometimes. Do you feel like these are examples of that? Or do you feel like as we get better if we did a longer trajectory where instead of just writing this piece of code, you have to think about how you're gonna maintain it later and, like, how it's gonna run that's gonna fix it or it's hard for me to grasp.

Speaker 3

Yeah. The software stuff is not easy because it's almost like maintainability almost feels like a human preference type Right. Issue again.

Where somebody could look at it and be like, yeah, that's not as good. But adding the heuristic and trading seems very messy. Yeah.

So so maybe it maybe it is. I I don't know. There's a lot more to dig into that.

I mean, like, this is what Anthropic says they're doing and just what are the actual frontiers in making like, they said they're working on code only and what does that actually mean? Mhmm. A bunch of it is gonna be designed trade offs and, like, how much autonomy the model has versus these potential side effects from training longer that we don't know how to get rid of.

I mean, that that definitely could be the sort of a behavior like that is what I would say is like a simple thing to remove, where it might just be obsessed with some code format that fails when you revisit it or something. If it even if it's like that everyone has seen it with just bypassing test cases. I think there'll be a bit more nuanced than that, but they could probably be super simple.

Speaker 2

similar semantic content address for me as over optimization, which is something that you've written about. It is over optimization with a different with a different reward function. I I know.

I okay. Well, I made that link. I wanna verify that that we we are picking on the same wavelength.

I just wanted to go over again specific topics on things that you've you've spent some time thinking about. You write that there are three types of overall optimization. First was RL for control, second was RLHF, and third is RL RL VR.

They always happen. Obviously, RL is no stranger to reward hacking.

Speaker 3

how things are evolving in in terms of how we're learning as an industry? Yeah. So that three things breakdown is for people to put the pieces together for what has happened historically.

All of these over optimizations are a just the model optimizer is strong enough where it can manipulate the agent with respect to environment or manipulate the environment in in a useful in a way that's useful to its target signal. Also, like, for context, I think with what we're doing with language models in RL in general is that if there's something that can move its reward signal up, it'll move the easiest thing. The most direct things to move that single up.

So that's part of the story that I said on sycophancy, which is this reward model for user feedback was probably so obvious that humans just like to like stuff that is, like, people press that thumbs up button when they're filled bullet points. Yeah. And, like, all those things have just been really easy for the model to extract.

So, like, once they added it, the model changed a lot, and the score went up a lot, and it was easy for the RL to find that. In Control, the oldest RL, the environment is normally a simulator that is fixed. There's no feedback.

So the over optimization looks like unphysical and nonsensical behaviors. There's the motorboat example going in circles. There's, like, an example is a project that was middle author on was like, effectively over optimizing, like, half Cheetah, which is this Majoco thing.

Instead of running, it did car wheels off into the sunset and got like infinite numbers. It's like, obviously not the intended purpose. It looks like a glitch.

So it's just kind of manipulating the the agent interface with the environment. RLHF is kind of a classic case where the model will just break down because the reward model is imperfect. So like the environment is really imperfect in the RLHF case where It's so sparse.

It's like very artificial. Yeah. It's a very artificial environment.

So it makes sense that these actions which are generated tokens will do things like reduce into just repeating one token over again. It'll be like, I think one of the early examples we had playing with this at Hugging Face was the model would just say JavaScript. It would be JavaScript, JavaScript, JavaScript, JavaScript.

It was like, some toy dataset. And it's very obvious when you see it, it's probably harder to see when you're at the top and making design decisions on when to stop training if you're doing a lot of RLHF. But that was kind of the phase that people have gone through, and now we're in the RLVR phase, which is we're giving the model reward when it does something quote unquote right.

For math, it's a bit harder to over optimize, I think unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world. They're just like, oh, I'll just find the you're train it's like the model is like, oh, you're training me on Stanford's problem set for CS whatever that it's seen a thousand times. So it's like, I'll just go get the solution manual, which I'm sure there's somebody can find an example where that has really happened.

But on code and maybe information retrieval, it's easier to fudge. So the code thing is like the easiest way to get a unit test and pass is just put a pass in it. Like, that is not too surprising that a model can learn how to do that.

And there for code, you need more reward design. To think would be a nice for like a substantial academic work is like, is reward designing code for balancing this sort of like understanding this over optimization of test cases or avoiding failures or something like this. I'm sure there's it's not just like gonna be a controlled environment because these models are complicated, but I would guess you can reproduce that in in some ways.

Speaker 2

Just to double click, reward design means, like, for example, giving credit partial credit for partially correct work. Yes.

Speaker 3

giving the model a slight penalty for doing the unit test thing if you can detect it. Yeah. For cheating.

Yeah. Which is it adds a lot of complexity to training these models compared to math, which is just if the answer is right. This is I mean, you can look at the GRPMath and partial credit is weird in that because it's kind of normalized per batch.

I don't know if I have a whole spiel ready on it for that, but it's also just it becomes very complicated if you're mixing domains and it's like, is partial credit in code better than partial credit in math or all of these things? It's like reward design becomes very complicated and that's what you're incentivizing the models to do different things. Yep.

Speaker 2

about mixing these things? So let's say you have the the one for code, you have the one for math, you have whatever other verifiers you can come up with, and individually they work? Do they conflict?

Speaker 3

I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models, like, don't get worse on knowledge benchmarks if you're training on, like, just math or precise instruction following. So the model just kind of develops an intuition for, like, where the different prompts are in space. So the gradient updates will be different depending on your batches, which is partially why people will just say do big batches.

So like a lot of the model is activated and you have a less noisy signal with RL, But a lot of intuition is that the model just kind of handles that. And there's interesting questions on sequencing, like, do you do large scale math and code RL to get the sequence length and then add in more general stuff Yeah. Which DeepSeek mentioned, but that's one thing to go the DeepSeek report is, like, math and code to more general RL.

There's a question on where do you do tools if you're gonna do, like, code execution and search within this. So I don't know if that's interweaved or if it's a second stage.

Speaker 2

Got it. Yeah. I I I don't have comments there.

It's just like, it's surprising how much is not known and you just need a lot of compute for ablations.

Speaker 3

The inference the high inference length generations definitely just like kinda breaks all infrastructure. There's just so many tokens. It's more opportunity for out of memory or other things go wrong.

So it's like just on a default, all of your training jobs need way more GPUs for the memory of inference. Sure. And or just like training.

But it's just it's just makes it more of a pain. Yeah. That's a cost thing.

Speaker 2

You know, one of the maybe controversial takeaways from the Gnome prod, which you listened to, was that there's also just wall clock time of just getting feedback from the environment, whatever that is, especially if it's like a real world thing. And I'm just like, yeah, I mean, there's some point at which your training runs cannot take longer than like a human life, like so to me that was the wall. He he he disagreed with that, but like that that was what I meant by it.

Like at some point you long inference, you you do want it to terminate within some reasonable amount of time regardless just as a user. Yeah. We have to find a way to accelerate internally within the training time faster than the passage of time in the actual universe.

Speaker 1

Yeah. We're not I'm not worried about that problem, but I agree with you in principle. Right.

So I'm I'm I'm stretching this out too far. I get it. I get it.

As we kinda start wrapping up, what are other interesting ideas that people should pursue? Like in your AI talk, you said, what I'm thinking about for scaling URL, you had big multi domain datasets, difficulty filtering, long run times. Is there anything specific that if there's people out there that are either doing research or they wanna do a company or whatever, these are like interesting things that you don't wanna do, that you want other people to explore?

Speaker 3

Most of them, I think, are not in the reasoning space. But, like, if their talks have been about reasoning. So I've been long talking about, like, character training is something that I think is under indexed on and been advising a student that's Character level?

Like personality training Okay. And how that like like, different ways of changing the personality of the model from prompting activation or fine tuning Okay. Or like data engineering.

So stuff that like Joanne Jing does for OpenAI. So like like like, how much does that matter? What are the fundamental research things?

Hopefully, I can share more that I've been advising a student on that. So I've been saying that for a while. Do you just as a side side note, do you like the model spec stuff that she's doing?

Yeah. Okay. Yeah.

That that trajectory. Yeah. So I've I've been a early fan of that.

I mean, that's how she final that's how like she noticed me as I was like the only person that covered it when they first released it. I think it was like over a year ago. I was.

I liked it. I Yeah. Well, not many people did.

Okay. Alright. Alright.

You were first. I don't know. I don't know.

But, like, that's what she said to me. Well, we had a, you know, we had a model spec talk close the whole conference. Right?

Like, that was my sign of, like, pay attention to this, guys. But it's it's real because of what it sends to, like, develop it has developer benefit of, like, where your model's going. And then also just, like, regulatory.

I think it is very important to, like, what is, like, an intentional behavior versus just, like, a training error. Okay. So I I think for model transparency, it's really fantastic.

And I've said that, like, the model spec is much more useful than a constitution. The constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we like, we don't write down our goals of the model in a constitution form.

By the way, have you looked at the constitution? Not very sure. They talked about it.

They they put in like Apple's like design guidelines Yeah. But then also like the UN like declaration Yeah. Of So at this level, I've seen it.

I don't even know if they've updated it. That's very odd. I hope that Anthropic would write a model spec.

I'm not too optimistic, but they're the next domino to fall. Well, so my take on that actually, I pushed for this too late because OpenAI already approved the talk and all that, but that I was gonna ask them to compare the OpenAI model spec to the Cloudflare system prompt, which is their closest thing to the model spec. It's the system prompt is incomplete because OpenAI has things in the model spec that their model doesn't currently do or especially when they started.

It's like we want to when they first released it, it was like we want the model to be engage able to engage on, like, sensitive subjects and maybe like even NSFW was in their model spec, which is they're just signaling of what they wanted to do and they say like, this is very hard to implement because there's all these obvious risks to doing this. But it's like in an ideal model where we can solve every problem, this is what we do. But I think is good, as I said, for many different stakeholders.

So I mean, mostly, my thing is like there hasn't been a good, like, foundational research paper on that. That's just a lot to do. It also runs into personalization and personality or similar, which is like if open models are to win, part of it could be just like everybody can have exactly the model they want.

We're serving GPT 4.5. It's kind of its thing.

You can prompt it. But if fine tuning is more effective than prompting, everybody can have the model that they want. So it's a good it's it's like a an academic problem or an open ecosystem problem where people are fighting on the turf that feels more likely to win Yep.

Which is good. Is this somewhere where you, like, as speaking as AI to Omo, you you want to win? Or is this you're just advising a grad student on it?

Speaker 2

as yet, but I'm very open to working on it. Because I think, like, open models have a strong, you know, role play use case and, you know, like character, personalization, all that stuff. Right?

Especially because people, like, they find their waifu, they wanna keep their waifu. And like, that's the derogatory term for it.

Speaker 3

part of almost should be that it is a base model that's easy to take in directions that you want. And we will have an opinion that is probably slightly conservative on personality. I mean, I've gone through the open AI models back and it's like most of these who we agree with and like be conservative on anthra morphization.

What disagree with? I don't remember. I did it a couple months ago.

But a lot of it is, like, openness or transparency, which is, like, if we're training an open weight model personality, like, we're not gonna withhold anything. Yep. And we have a different hierarchy.

So most of them are on, like, that type of information exchange rather than be kind. Like Opening Eyes Model stack is pretty agreeable and if you you read through it. And it's like treat the user with respect and all these things.

Raising kids that way. Just read the spec. Yeah.

It sounds kinda stupid. But then the last thing is for people doing research, it's like wacky model routing things where you figure out, like, a bunch of different models to off Hugging Face to route to. Because an open model tool thing could use way more models more easily than any OpenAI product.

Because OpenAI is restricted to the OpenAI's models where if, like, maybe, I don't know, OpenRouter is, like, I'm gonna make a product out of this, which is a router. Like, OpenRouter actually does it, and they're like, our chat window knows the best model based on all this usage that we have Yeah. For your query.

There's people that started other way, like, Martian, not Diamonds, I I don't know who else is. He would know. There's there's a bunch.

There's a bunch. Yeah. So I don't know I don't know if that would work.

Hugging Face should work on it. It's like like an it's it's a moonshot idea. You don't know when it'll Given your Hugging Face record, what is what is how does Hugging Face make money?

This is a very common meme question. I think mostly, like, enterprise deals. That's what they say too.

Yes. Like, they're doing their thing.

Speaker 2

their people. Yeah. Look.

They're they're great. They're big. They're profitable.

Speaker 1

It's just not that obvious to most people. I like the router idea for media models. I feel like there's, like, so there's, like, a long tail of, like, a A generative media.

Yeah. Like, a style supplier. Like, that that is actually hard to find.

On the tech side, I feel like just use the big unless you're, like, under some light latency or price constraint, you should just use the best model. Even when we're doing thumbnails, I'm like, okay. I'm trying to remove a background of somebody.

And it's like, I go and replicate, and there's, like, 55 background remover. Yeah. I just use Adobe because it's a website.

Well, that doesn't work. Like, the Photoshop model is bad Oh, okay. On some things.

But again, it's like or I wanna generate a diagram to, like, mimic something. And it's like, well, which model is better diagrams? Yeah.

Speaker 3

are Part of the argument is that if distillation works really well, we could just keep making the target for distillation smaller and smaller, which is you have models that are very narrow. Right. And they're mimicking these huge models on something that's like pre I don't know, like reformatting tables.

So it's like, can you do a table reformatter from markdown to LaTeX in a 100,000,000 parameter model? Like like, if you get it small enough, that is really economically feasible because it's effectively free and imprints and instantaneous.

Speaker 2

My pushback is on this is just if you're doing image editing, four o should do it do it all of it. Well, yeah. But but I think it does, like It's just we're just not there yet.

Like, give it five years, It'll do it. Right? There's that so why work on a router at all?

You just scale up four o. I guess I it's yeah. Right?

Like, me where the the logic is here. Like, this is like a temporary thing. On device?

Speaker 3

On device. Like, the local modeling community, I think, is much smaller than people give it credit for because most of the use for open models is still in APIs. It's like deep sea.

Convenient. It's convenient. And it's like, if there aren't that many models, somebody's gonna host it for cheaper than most people's doing it themselves.

That's pretty realistic. But there there is a small community that need local. Yeah.

The best outcome is if open models can compete on not just long tail things, but that takes the most transformation.

Speaker 2

Side note, so I resisted by building my own, like, buying my own GPUs, building my own cluster for for this reason, I'm like, APIs will will solve most of it, like, people are losing money to serve me models, why am I, you know, having those? Except for the fact that $40.90 prices have doubled in the last year.

So you actually made money doing local models.

Speaker 3

How does that make you money? Because your your investment goes up? Yeah.

You're hard worker. You're So as you use $40.90, it goes up.

Interesting.

Speaker 1

Should've bought a forty ninety. I bought a forty seventy.

Speaker 2

Damn it. What is this? Well, then then then it puts me until, like, should I buy, you know, $15.

90 if if it ever, you know, is is widely available.

Speaker 1

were doing the drops. Yeah. I know.

It was crazy. People are, like, running to I know. To the camper to buy it.

Speaker 2

Any any other topics before I give a closing question? Just generally, your work, ROVR, like, are topics of the day.

Speaker 3

company should keep considering re releasing open models mostly for PR and onboarding. It seems like the way it's going if OpenAI is releasing it. Are you excited about that?

Do you feel like it's like a psyops like The OpenAI model will be good. I expect it to They're be pretty serious. It'll best in class for some sized category and some subset of tasks.

That's like OpenAI only does things like that. You have to give them the respect they want serve. Yeah.

Super simple.

Speaker 2

open wins when more people are doing it. So like, that's that's a win. And Yeah.

Well, I mean, hopefully, they are actually open about the techniques and not just the weights.

Speaker 3

tells us anything about the hardware that they're gonna build? No. What?

No. They're so secretive about this. That's like, that's why they haven't released g p t 3.

5 or anything because it's too revealing about internal stuff or plans.

Speaker 2

Oh, okay. No. So I I I are you talking about Stargate?

Or what do what what kind of hardware? Johnny I. The the no.

Yeah. The thing is that's a different device. Yeah.

That's it. Yeah. I think that thing will run on the cloud.

Don't think that'll run local anyways. Well, okay. We have to talk about it.

Like, seems like every podcast, we talk about it. So apparently, the news from today, which I think you were looking at, was that it was like a ear device that they sued they they got sued over or whatever. But, like, I think that ear form factor is pretty good.

Like, I actually did get there with b in terms of like, where where does this ultimately go?

Speaker 3

something you want the AI to hear what you hear. And where do you hear what you hear? On the ear.

Like, that's pretty much it. I don't know if you guys have like thoughts on wearables and where that goes. I try to be.

I I think it just knows too much. That's my that's really my But you want to give it context. Yeah.

I have false privacy hopes. I think, like, a lot of people I mean, that's the whole thing. It's like people don't actually care about privacy.

It's just note taking, you know. It's just really good memory. I think the meta Ray Ban form factor is good.

I don't think it's as mass market. It's like if you get it in the AirPod size form factor, it's a way bigger market for obvious reasons. But the the, like, sunglasses form factor is the thing that works, I think.

Okay. I don't use them for AI, but they can fit the AI to work it. Like, yeah.

Empirically, yeah, it it obviously works. Yeah. Cool.

Speaker 2

what what is Meta doing? You know? You had a you actually had a pretty interesting post back in when was this?

Speaker 3

you said Lama Ford and Meta just pushed the panic button. I feel like back then, it didn't actually push the panic button, but now they really pushed the panic button. That's fair.

I think the panic button at the time was the whole LMSYS model not being the model that they released thing, along with a bunch of weirdities about, like, the day of the week they released. But to be a model that claims to be open and then not release the model that is your leading claim is just like a that that is like bad execution. Bad execution.

Yeah. Yeah. Which is fine.

And then the recent stuff, I think, mostly can be boiled down to talent is cheaper than GPUs by a dramatic margin. And at the end of the day, it's like, okay, but if we're spending this much, they go to the room and they stare in the mirror, they're like, wait, it might not actually be that ridiculous to spend this money on the top people. It's like, might as well try.

They already spend it on VR. Somebody was bound to do this eventually. And it makes sense that it's like the it's like if Apple are to some way somehow decide like we're gonna do this, they're gonna come in and do exactly what Meta is doing.

They need a founder mode CEO who's like, screw it. Like, you know, we'll we'll take the l. The the thought that that occurred to me is, you know, Meta instead of spending on VR, they should spend on RLVR.

Speaker 2

And oh, well, think the question is like, I think a lot of some researchers, like most people will take the payday and happily move to Everybody has a bribe number. Right. It's just just the summary really big.

Yeah. But like, think some researchers are uncomfortable with the idea that this is a sort of the great man theory of of research that like, you have to pay this much to get this level of talents and The talent is definitely distributed.

Speaker 3

Yeah. Right. A lot of the people that they would be paying this much have the confidence to redo things or to just do some of the same things and just like whether you call it feeling the AGI or just drive to build things or like feeling the AGI is not that different than a lot of things that have existed in Silicon Valley lore in the past.

It's just people with the vision that are willing to execute on it and they see something coming. And those people make a big difference. I think you have those people and you remove bureaucracy.

Getting technical talented researchers is actually something that Meta has a lot of or has the ability to get a lot of. So it's like the it's a lot of recycling, which is very hard on individuals and morale of an organization. But that's like understand the approach.

Speaker 2

Yeah. For sure.

Speaker 1

Cool. That's all I have. Any parting thoughts on how you're gonna build the American deep seek?

That was a nice tweet. Yeah.

Speaker 3

if I have to look at like what my in the if you were asking me, like, what my ten year goal is, and it's like, I only will have, like, a two to five year goal where I think as models are shifting more towards agents, I think that, like, scaling is slowing. It's like their side of it of a fixed cost and a fixed path to getting towards something like American DeepSeek. Or mostly just I would say it doesn't have to be American if it's fully open.

If you have everything and you can modify it, which is like, there's a few things that need to fall. A lot of it is just more resources, but it's like like, almost 32 b is if you squint like original GPT four level and fully open. And it's like, there's a few levels that you need to go through, like, that's obviously a dense model.

It needs to be taken to sparse MOE, and you need to scale it. You need to have a lot more GPUs, and then you need to do, like, large scale reasoning. It's like, that's the goal that I want to do.

There's a lot like, that's what I want to do. There's a lot of complexity in navigating, like like, how to work with AI. Like, does AI to do to get there?

It's very hard. Mhmm. I think that I mean, it's a nonprofit.

It's hard to get the resources, and building a model is a lot of aligning a lot of different people. That's the deep sea story is they have great people. OpenAI has kept a lot of really good people for a long time.

Anthropic has gotten a lot of good people right now, and it's like, it's a lot of incremental hard technical problems that you need to stack up. Like, that's what I would like to do and make work in the next couple years, but it's not easy to get there. So that's that's the pitch is like, AI two's best case scenario is AI two's gonna do other things.

Like, you can't just run a nonprofit or a company that says, our goal is in three years to have an American deep seek. Like, no one's gonna keep paying the bills on that, because you have to tell a better story. But that's like what I would like to do in that.

And I'm sure AI too will do many more interesting things along the way. Like product stuff. I don't know if it's necessarily product, but like, what are more like, what are cutting edge things in AI that we can make a new architecture for certain things?

Okay. Or like, what are demos of open models working better, whether you have like private data or something, or just far out ideas that could take you off the transformer trajectory. I think that, like, you you still need to be doing these to kind of lead in AI.

Speaker 2

Thank you for working so hard on truly open source AI.

Speaker 3

Yeah. It's fun. I think it's I mean, it makes it easy to align, like, values with what you're doing.

Yeah. It's like like, it'd be better for the world if more things are open, and therefore, it's like a lot of it is just willing it into existence. And I take seeing like what OpenAI does is or is saying they're gonna do as like hopefully a win coming soon.

Yeah. Like, deep sea was the most unexpected win that made some other dominoes fall. Oh, yeah.

I think that is the path forward and see what it takes. Thank you so much. Thanks for coming on.

Shared via Hopper