Kyle Corbitt, founder of OpenPipe and leader of CoreWeave's serverless training team, provides a master class on reinforcement learning (RL) and custom fine-tuning for AI models. He explains how RL differs from supervised fine-tuning (SFT) by making less destructive weight updates, delves into the GRPO algorithm and its industrial improvements, and discusses the role of LLMs as judges in post-training. The conversation also covers reward hacking, the economics of RL environments, and the competitive landscape, highlighting compute as the primary constraint for Chinese labs and the potential for recursive self-improvement.
Hello, and welcome back to the cognitive revolution.
Today, my guest is Kyle Corbett, founder of the reinforcement learning and custom fine tuning company Open Pipe, which CoreWeave acquired last year. I opened this conversation with a bit of a confession. I've done a lot of supervised fine tuning work over the last few years, both for Waymark in the early days of getting GPT three to write decent video scripts and for research projects such as the emergent misalignment paper.
But I've done essentially no hands on RL work, both because my perception has been that frontier models are probably my best option in any case, and because I'm afraid, perhaps irrationally, of reward hacking. Kyle says that while it may or may not be worth the extra work and slower iteration time, he does believe that using RL on an open source model probably would deliver me better performance and would certainly reduce both latency and inference cost dramatically. With that motivation in mind, Kyle proceeds to offer a master class on all things RL, which repeatedly challenged my premises and in multiple instances updated my understanding.
He explains how RL differs from SFT in terms of the weight updates it makes to the models, how this difference makes RL fine tuning less likely to cause catastrophic forgetting, what distinguished the DeepSeek GRPO algorithm from its predecessors, and what additional improvements on GRPO people are using in industry today. We talk about the distillation strategies that Chinese labs are using to fast follow American frontier models, and he argues that their use of LLMs as judge in the context of RL post training is a bigger deal than supervised fine tuning. He also explains why he thinks that compute is the primary constraint preventing Chinese companies from catching up and why he believes that we're already in a recursive self improvement loop.
He describes the cottage industry of reinforcement learning environment companies that sprung up to serve Frontier Labs and why though it is a good business to be in for now, he's declined to invest in any of them. He surveys the use cases that are most commonly deployed by CoreWeave customers, and he offers a lot of advice on how to run RL in practice, including how to develop and iterate on evaluation rubrics, whether to train end models for end tasks or a single model to perform multiple tasks, how the flagrant nature of reward hacking makes it relatively easy to deal with, at least when you're focused on specific narrow tasks, and how CoreWeave's use of LoRa adapters drives efficiency and convenience for their customers. Kyle is both a technical expert and successful commercial practitioner.
And from start to finish, this is a super high signal conversation on a classic training technique that has become an industry unto itself. And so I hope you learn as much as I did from CoreWeave's RL fine tuning guru, Kyle Corbett.
Kyle Corbett, founder of Open Pipe. Now after an acquisition, leading the serverless training team at CoreWeave. Welcome to the cognitive revolution.
I am super excited to be here.
Thank you. I'm excited to have you. This has been a long time coming since we've met, almost a year ago now, and I'm I'm glad to finally be doing it.
So and that's all on me, the way, which is everybody knows. So you are a specialist in reinforcement learning. What I wanna do in the next, hour and a half or so is get basically a comprehensive survey crash course rundown of what is going on in reinforcement learning, how we should understand it, like what the techniques are looking like, who's using it and where and for what purposes and who's having success and not, and what makes a difference and all those things.
And so I guess I was just gonna start by telling you my story very briefly and then allowing you to react to that and tell me if I'm like way off base or not. My story in short is I've done a lot of model fine tuning over time, mostly on managed platforms, not so much like on open weights models, just a little bit of that more more so on like the OpenAI platform. But almost entirely supervised fine tuning, not really much at all reinforcement learning fine tuning.
And the story I'm telling myself, which you're invited to pick apart is for one thing, increasingly these days, like I can just use base models with few shot prompting and that's getting me a lot of what I need. But even before that was possible, the problems that I was working on in the context of my company Waymark are sort of taste driven problems where we always kind of felt like we'd be better off going to our creative team and say, hey, give us a 100 great examples. We'll fine tune on that and hope that the AI can follow your lead rather than try to go through some sort of seemingly more complicated, maybe more powerful, but kind of like harder to wrap our heads around notion of like, well, if we get the ad to do it and then we compare and we score, maybe there's an LLM as a judge.
We're kind of like, I don't know. Feels a little bit sometimes like a shell game and I'm not sure like how much I should where I should invest or how much I should trust that process. Whereas I know at least if the AI is like imitating my creative team, there's some like decent true north there.
And then the good thing I'm kind of afraid of, although I'm not so sure it's a big problem in my context is reward hacking. But I'm I am kind of afraid of reward hacking in general.
or not? You know, should I be using reinforcement learning or am I am I thinking about it the right way? Yeah.
No. That's that that's a great question, I think one that lots of folks think about. Maybe my first question for you would be, how are the results you were getting from your existing process?
So you you mentioned, first of all, that, like, these days, you mostly just do prompting. But, like, when you were doing fine tuning with SFT, do you feel like you were seeing the models improve substantially on that? Do you feel like it was, you know and and this is a very high bar, which I imagine wouldn't clear, but, did you feel like it was, like, behaving as well as your creative team and matching the quality of of those examples that was given, post training?
I would definitely not say it was matching the best work that our creative team could do, but definitely a notable improvement on the base model. I would say our typical complaint was probably most often that, and this would definitely vary through different generations, but more recently it was like able to do the job perfectly well, so to speak. And I think that's true today too with prompting, but few moments where you're like, damn, that was awesome.
Incredible turn of phrase or, you know, nailed it, really nailed it.
impressed me, surprised me, delighted me. I wouldn't say we see too much of that coming from the models even today. Yeah.
Okay. So that that was gonna be my next question. So, yeah, even with with the latest frontier models, you know, that that sort of spark or wow moment, it sounds like is is is not something you see commonly.
Rarely at best, I would say. Yeah. Yeah.
So, here's what I think. Like, I think it is likely that you would have been able to get better performance out of the models with reinforcement learning than with SFT. And there there's a few different factors here that sort of model it.
I I mean, one is that, like, OpenAI's support for RL was half hearted at best at any given point. And I think, technically, they still do it, but that entire kind of, like, you know, model customization platform is feels feels very much in maintenance mode, at this point generally. So so I think on that side, like, yeah, that that might just not have worked.
In a parallel universe where you were using an open source model and you're using, like, a Quen model or something like that, then I would say with a fairly high degree of confidence that if you're able to get decent results out of SFT, the ceiling of the best results you can get with reinforcement learning is going to be higher. And that's true even if the data you're using for SFT is high quality data, human data. And the reason why is because it's just, like, the whole trick to RL.
Like, the whole reason RL works or or is is something that people invest in at all is because it it it turns out it really does matter how well your data distribution matches kind of the model's, you know, like, a standard mode of thinking or or just, like, what it's picked up from pretraining. And what RL gets you is it just gets you, you know, more, it it it is working within those channels that are already carved quite deeply within the model. And when you work within those channels, you can just get a lot further because you're not trying to overwrite what it's doing.
And so, and and, yeah, you might say, like, well, you know, overwriting is what we're trying to do. You know, we're we're we're trying to get it to do something it's not good at, which which is fair, but but it ends up being quite destructive. It's actually really interesting if you, like, you know, look at at the weights.
Like, you're doing SFT, it's just, like, even with very few examples and even if a very, very low learning rate, like, it's just, like, throwing the the the weights all to pieces and, like, the average differences are so much larger than doing RL. And that's that's an a big part of why you get this catastrophic forgetting because it's kind of, like, overriding other pathways. And it's just, like, you're trying to get them all to do something that's, like, quite different than what it was trained to do, whereas RL is going to let you stay in those grooves and get a lot further.
So so, yes, I do think that would have worked. Now in your specific case, was it would it be worth it? Would it get you to a place where it's like, oh, this is better than just using the frontier?
My guess is probably not. So I think concretely for your task, if the trade off you're making is, hey. We're gonna take an open source model and use RL, to try and make it better at this versus, hey.
We're just gonna take whatever the best off the shelf model is and do prompt engineering, you know, and and we're allowing ourselves to expand to, like, you know, the best frontier models at that point. I suspect for a creative writing task, you would end up in a position where you're better off using the frontier models. And, and and yeah.
Like, we can we can sort of, like, get into like, there there are definitely tasks where I would say the exact opposite and say that, you know, the RL could do well. I would also say that, like, this is, like, obviously, like, all dependent on the amount of compute. Like, think and theoretically, anything's possible.
If if if if you buy yourself a data center and spend a a couple billion dollars on this task, you you would be able to surpass the frontier and yeah. But the trade off point would would be fairly long, I suspect, along that curve for a task of this shape. I'd like to understand this Grooves thing better.
know what you're gesturing at. When I think about, like, how much the weights change with fine tuning, I usually think of that as kind of more a function of, like, some sort of divergence penalty, some sort of tethering of the, you know, the model as it's evolving to the base, to the starting point. I think you can do that on any kind of fine tuning.
Right?
why is it more destructive than the reinforcement learning? Yeah. No.
That's that's a totally fair question. So, yeah, I think what you're what you're talking about is, you know, that there is there's a term called, like, a KL divergence penalty, which is, you know, sort of, like, auxiliary term you can add to any loss function saying, hey. You know, prevent the, it doesn't actually prevent the model weights from drifting.
What it what it prevents is specifically, like, the log probes, that are generated, you know, at each token position from drifting too far from the base model. And this is often considered best practice because, you know, it it can help you from getting, you know, like, catastrophic forgetting and, like, moving too far away. However, what it's not like, the fundamental issue is that, let me put it this way.
Like, there are often different ways to get to the right answer. Right? And this is this like, the easiest example here is if you're talking about a reasoning trace where, you know, it's like, hey, you're doing the math problem, and and you're training this model with RL to solve the math problem.
And there's probably, like, you know, an infinite number of ways you could reason through from a problem description to the answer. And some of them are going to be paths that the model already is comfortable with and is, like, you know, like, oh, like, you know, these these, you know, eight tokens in a row, like, even the base you know, the model you're starting from would have generated them anyway. And then, you know, the next token, yeah, maybe it would have gotten that one wrong.
And so so there, you know, the learning signal is is treat teaching you to move that one slightly. But fundamentally, like, what RL just structurally optimizes for is changing the fewest tokens, the fewest log probs necessary to get to that right answer. Whereas what SFT does, if you're if you're doing SFT on you know, you're, say, distilling a larger reasoning model into a smaller reasoning model.
And this is particularly true if, like, the smaller model you're going on had different pre training distributions. So so, you know, you would expect that kind of, like, it's it's kind of, like, a built in, you know, intuitions or inclinations are are different, is you're not respecting those kind of, like, pieces of the reasoning that it would have gotten right anyway. You're overriding the whole thing with the reasoning from the from the larger model.
And by overriding the entire thing, it's like this this this is quite confusing potentially for the backdrop algorithm, because the backdrop is just seeing, like, oh, all of these tokens need to change, and, like, maybe some of them didn't actually need to change. You know? Like like, maybe, like, the direction the model was would have gone with this token actually was also fine.
And so but but you're, like, changing all the weights to, like, get to this new one. And, really, there was this other token that was much more important that did, in fact, need to change to to get to the right answer, but, like, that one's, like, kind of just mixed in with all these other random unrelated changes. And so that that, like, general intuition generalizes to other task shapes as well, including creative writing, where, you know, the maybe there's, like, two different ways to phrase this, and they're both fine.
Right? And the model would have chosen one, and your creative chose another, and they're both okay. And you don't really want to update like, sort of, like, waste your your model updates.
Because every time you update the weights, you know, there's a potential for catastrophic forgetting and sort of just off target effects in general. And so you don't wanna, like, waste those model updates on, like, changing something that was already fine. You wanna really direct them to upwaiting the things that, like, the model wouldn't have gotten read on its own or very rarely, more more specifically, would have gotten read on its own and, like, very and and focus kind of your your sort of, like, updating budget on those.
So the KL divergence doesn't give you that, you know, if if what you're doing is just penalizing KL divergence, it it's it doesn't distinguish between things the model is already doing fine, and you just, like you know, there there's, like, a different way you happen to have it in your trained data versus things that the model really was getting wrong.
Okay. Very interesting. Hey.
We'll continue our interview in a moment after a word from our sponsors.
Most billing platforms were built to send invoices and assume your pricing is simple and predictable. But if you're building an AI product, a fintech tool, or a developer platform in 2026, your pricing is anything but. Usage tiers, consumption billing, and bespoke enterprise contracts are now the norm, and you're probably managing it all across disconnected tools and fragmented systems.
Sequence handles the entire revenue workflow from contract to cash. Quoting, invoicing, metering, revenue recognition, plus Sequence agents that automate the manual finance work that usually takes teams days each month while also helping them to collect cash faster. Companies like Cognition, Incident IO, Runway, and OpenRouter use Sequence to run their full revenue process between CRM and ERP without the spreadsheet mess.
If your pricing has gotten more complicated than your current billing setup can handle, check out sequencehq.com, and use the code Cognism in the source field when you book a public demo to save 20% off year one. AI is rapidly moving from assistants to agents, and it's causing a sea change.
AI isn't just helping anymore. It's taking action. And here's the reality.
You don't get outcomes from Magentic AI unless you trust it to operate at scale. That's why AvPoint is building a control layer for AI. This foundational layer helps you govern what agents can access, secure how they operate, make activity auditable, and recover when something goes wrong, all as one connected system.
See every agent, app, and workflow and what they touch. Govern with policy and guardrails that work at machine speed and recover quickly so a mistake doesn't become an outage. That control layer creates trust, and trust is what unlocks the right outcomes, letting you automate more work, move faster, and deploy agents with confidence instead of hesitation.
If you're scaling agents and want those outcomes by design, learn more about AvPoint at a v p t dot co slash t c r. That's avpt.co/tcr.
When you describe you said more specifically, you know, something that not the model can't get right, but that it rarely gets right. That's key because when we do things like GRPO, the you've got to have at least one right answer, right, to be, to have any sort of advantage. I guess it also depends on whether you're doing binary scoring or some, you know, more rubric based evaluation.
But I guess several several different questions coming to mind at once. Can you give me a little bit more intuition and maybe we could do this for g r p o and you can maybe describe like, I'm not sure if G R P O is still like the hotness that it was a year and change ago. I'm not also I'm not entirely sure if that was something that broke out for kind of memetic social media reasons or if it really was like a huge advance over its immediate predecessors.
But can you give me a little bit more intuition for okay. I understand that in this algorithm, we are running multiple rollouts. Some of them are gonna get to a right answer or if it's a rubric score, they're gonna get a higher score than others.
And then there's a computation that creates the group relative advantage, which is to say, you know, we wanna shift toward the patterns that gave us the right answer or the higher scoring answer. How is it though that that is because it it still ultimately goes to a token by token thing. Right?
So how is it that if I have, like, eight different chains of thought and they're all kind of different and in any given token position, like, we might even have very different parts of speech, right, at, you know, at at token position and it could be a preposition here and a verb there or whatever. Well, like in very good kind of different moments in the chain of thought. But my understanding is that the advantage calculation does still ultimately cash out to like token level advantage.
So how is it that, like, where's the I'm a little bit lost on the alchemy of, like, why this translates in the end to really only updating the, you know, making change on those tokens that really mattered. How how is I'm I'm missing a little lead logic there.
several parts of this question, and I will finish on the one you're there at the end. And then hopefully, that'll give you the the chance to ask follow ups, if my explanation doesn't make sense. Okay.
So first of all, like, yeah, I think the reason g r p o, specifically, like, that algorithm and that acronym, like, you know, very concretely took off, was not necessarily because it was, like, a big quantum leap on what came before. It was because DeepSeek did a lot of engineering work around actually scaling it and released an actual artifact of model that worked really well with it. Like, that was kind of the the reason why, you know, there was a whole constellation of other algorithms that probably would have worked just about as well.
There was one that came out a little bit before called r l o o, RLU, which which, like, basically is the same as g r p o and likely would have worked just as well if you'd scaled it. After g r p o, very shortly after, like, in in, you know, within a few months, certainly at the r one release, you know, there there were various numerous improvements made upon it, which really do probably deserve their own algorithms. So there's there there was a paper called DAPO.
You know, there's, GSPO, came out from the the Quinn Lab, I believe. And then, SISPO was another one that came out shortly after that are all, like, significant improvements. And then there's a bunch of, like, minor tweaks that don't even have, like, named things.
But so I would say, yeah, the algorithm that people use today in practice is is actually as far you know, actually, probably further away from, like, gRPO as initially described as as gRPO was from kind of, like, what came before it. But we all just still call it gRPO because that was kind of the the the name that stuck. The okay.
So so moving on to kind of, like, how it actually works, and, like, let's yeah. I'll I'll I'll talk through it. I think this will be helpful to to build your intuition on, you know, how the advantages are calculated and everything.
So maybe I'll talk first about what came before g r p o. Because g p o is kind of interesting, in that, like, a big part of its, like, development was that it threw away something that would that everyone had used before and some people still use, which so so sort of the spiritual, you know, grandfather of of all RL that people do on LLMs is an algorithm called PPO that was developed by by John Schulman in in 2017, I believe. Actually, pre LLMs or pre you know, the being big, it was used for for games and stuff.
And, the key thing about PPO is it's sort of like, you have your policy, which is what you call the model you're training. It's taking a bunch of actions. And the key thing you need to do is every time it takes an action, you have to, like, kind of score how good or bad this action is.
If it's a a good action, then you want to, you know, re you you basically wanna update your weights to make it more likely to do that action. And if it's a bad action, you wanna update your weights to make it do less. Right?
The way the way you'd and also, importantly, this is something that happens at a, you know, at an action by action basis. But your reward, in PPO is is sort of, like, can be very long term. Right?
So it could be at at the end of, like, a very long sequence of action actions you finally find out, you know, that, you know, commonly this was used with games. Right? And so you'd say, hey.
At the end of the game or after a minute of gameplay, like, what's my score or something like that. So what PPO does is, is a few different things. And and it's actually, of course, building on older work as well.
You know, there there there's an algorithm called Reinforce, which is trying to solve the same problem. PPO, you know, adds adds some extra terms to keep it stable, keep it in in sort of, a trust region where where, you know, you're you're kind of, like, hopeful that the model hasn't changed too much as you're updating it. But the way that it but kind of the key the key thing that that PPO does, and actually, this is this is not unique to PPO.
This this is from older than PPO, but you wanna calculate the advantage at every single action. So every single time it takes an action, you wanna say, hey. Was this a good or bad action?
And the way it does that is by actually training a couple of different models in parallel. So you have the policy model, which is just your normal model that's generating the actions, and then you have a separate model, which is called the value model or the critic model. And the value model is actually predicting, saying, hey.
Based on the set of actions up to this point, you know, what do I think the the score is going to be in the long term? Like, basically, it's predicting for this action, like, what do I believe is the value of this action? How will what impact will this action have on the score in the long term?
Okay? And it's predicting that for every single action in the sequence. And then eventually, you do get to see, like, what the actual score is.
And then, basically, if the score ends up much higher than you expected, then you could say, oh, some of these actions clearly were much more valuable than we expected. So if, you know, if if it's like, hey, my, you know, my my my critic model thought it would it would have a low score and it actually has a high score, then then I want to, like, make it much more likely that I have a high, or or that this action happens in the future. Okay.
Now, moving on to gRPO. The sort of, like, key difference here is instead of having figuring out what the value of any specific action is Oh, actually, before I go into g p gRPO, I should I should mention, this all translates directly into LLMs. And and it's and the translation that people do people have tried actually a lot of different translation.
But the one that most people do, and it's, you know, kind of like the the simplest thing that works, is every single token generated is an action. Right? So so we're using the exact same concepts as we were using before and just saying, like, hey, every token, you know, the state up to that point is is the full context, and then this token is an action, and the next token is another action.
Okay. What we do with gRPO is it turns out that it's that calculating that figuring out that that kind of, like, value model and keeping it up to date and is sort of painful. It's just, like, tricky to get right.
It's like another set of hyperparameters you have to tune. I'm like, okay. You know, like, we we have to keep this model updated or else or else train doesn't work well.
And what g r p o did and they actually were not the first ones to do this, but, you know, they they they sort of get the credit because they're the first ones to do it at scale and and prove it worked well. As I said, hey. We're just gonna, like, completely throw away the value model.
And what we're going to do is we're going to try so the way we're going to figure out whether a given action is, you know, like like, basically, like, a a trajectory of actions is better or worse than what the model would have done otherwise, is we're just gonna run a bunch of them in parallel. Right? So we are going to with the exact same setup, the same initial conditions, we're gonna run, whatever, four or eight or 512.
Like, there's lots of different, you know, like, hyperparameters to do in here as well, different runs in parallel. And we're going to see how often the model succeeds and how often it fails. And the reason we want to do this is because you don't want so let's let's say you just run a single, you you so so, you know, we've thrown away the the critic model, and we do a single run through with g r p o, and we get a score, and the score is one.
Hey. It got it right. You don't know from that run whether the model, like, just would always get this right or, or this is, like, one in a million times that it got it right.
And if you're just, like, naively updating your model because it got it right, but it would have always got it right, you're just reinforcing like, you're you're sort of it's a spurious correlation. Right? Where it's just like, hey.
It made some random choices. Choices didn't affect the score at all because, you know, it just always would get it right. And if you're up weighting those random choices it made, then you're, you know, just, like, kinda moving around in in a in a in a in a pretty random direction.
So what gRPO lets you do is it says, okay. You know, the sort of, like, advantage that we allocate to each of these tokens is going to be based on how much better this run did than the run than sort of the average of, like, really what you wanna compare this to the average. If we ran the model, the current model infinite times on this, like, how much better did this do than that average?
Obviously, we're not gonna run infinite times, so we approximate that by doing it, you know, n times. Okay. So then getting to sort of the end of your question, which is like, hey.
You're right. Like, when we're actually updating the model weights, we are doing this at a at a token by token basis. Right?
So somehow we have to say for every single token, you know, we want to update the weights such this token is is more likely if the advantage is positive or less likely if the advantage is negative. And this is a big problem in reinforcement learning. It's called the credit assignment problem.
Right? Because really what you want to do is you want to you want to assign credit and up wait. Just the key tokens that, like, were were critical to this going right and not up with the tokens that, like, you know, just, like, always would have been right and and didn't really contribute anything to the solution.
And and so the the sort of, like I guess, the key insight of gRPO is to to sort of just, like, do a very unsatisfying thing and kind of just punt on that a little bit. They they it's not a full punt. So so what you do is you you look at how likely every token was to be produced.
Right? Because you're you're sampling at a high temperature when when you're doing these. And so to some tokens it's like that it produces are very common.
Some tokens are not very common. And basically, what you do is you say, hey, if I got a high score, then I want to give more credit to the tokens that, you know, just by random chance were less common. Because I assume that the high score is is most likely you know, if if my score is much higher than the average score across the entire group.
Right? Then I assume that it probably was because there was some rare thing that I did in this that I didn't do in other cases, and that rare thing led to me doing well. And the exact same thing in the opposite direction.
If I get much lower score than the average on the group, then the rare things are the things that I'm gonna penalize the most because I'm like, hey, that's probably what put me there. Now you could ask the question, well, it's like, there could be many rare tokens if if you got, like, thousands, you know, tens of thousands of tokens that are reasoning trace. How do you decide which rare token is is most important?
And you don't. You just throw up your hands, and you say, all the rare tokens, get uploaded the same way. This is, like I said, a very unsatisfying answer.
And I think that's one of the reasons why there was, like, an almost ten year gap between PPO that had this value model that tried to, you know, determine on a token by token basis and, like, g r p o, where it's like, hey, we're just gonna throw that all away because it feels wrong. It feels like it shouldn't work. In in practice, it does though.
And is the intuition there kind of like I I studied this a little bit, but not in enough depth to be confident. But I I'm sort of imagining that as we go through a chain of thought, there are critical tokens where you're either taking the right path or the wrong path. And then there are probably a bunch of tokens that kind of follow once you've made that critical decision that are all, like, kinda naturally gonna follow because that's just the structure of language.
So it's you're trying to isolate in saying the ones that the model was least confident about. You're trying to zoom in or or isolate or focus, if not isolate, emphasize maybe the right word, the critical decision points in that trace. Yes.
That's exactly right. Yeah. Okay.
Interesting. DPO basically is a similar thing too. Right?
That but that was you you had to have pairs that you kind of said, I like this one better than the other one as opposed to a ground truth or a scoring, but similar mechanism. Right? Yes.
Yes. The there there's definitely a lot of overlap in in the math and and intuition there. Yeah.
Okay. Is it worth getting into what some of the finer points have been since gRPO that has made it even better, not in a maybe super mathy way, but, you know, like, what additional insights have people brought to bear since then?
Yeah. I mean, yeah, we can talk about it briefly. Yeah.
It's a bunch of of small things. Like, one open question was sort of like, hey. How do we do length normalization?
And this comes into if you have a trace that happens to be so so sort of the original math in gRPO actually structurally advantaged, very long thinking traces and and and, you know, just like long generations in general just because, you know, it didn't normalize by the number of tokens. So, basically, if you had a batch of, you know, like, you know, whatever, like, 128 different completions, and one of the traces happened to you five times as long as the others, it ended up with, like, five times the amount of weight in the way the models were updated than others. And so it's sort of like, you know, the people had pretty good success with basically kind of, like, down weighting that, to average it out.
You know, there there were, you know, Systp was a really cool one. It it basically just changes the way you're doing the clipping. PPO and the gRPO inherits this, has a specific way of making sure that the weights don't stray too far in any one round of updates.
And there was this new technique called SISBO that was released, maybe six months later or something that basically it puts the clipping in a different spot, which which basically, like, the idea there is it lets the model discover much more quickly those, like, very, very high value but rare tokens. And so it it sort of allows those to update the weights much more, if there's a very high score while while not, like, allowing the weights to update too much. So, yeah.
And then there there's, like, a stack of probably, like, I don't know, half a dozen kind of, like, little tricks like that that people have developed to make the algorithm both more stable and converge faster.
Amazing. Hey. We'll continue our interview in a moment after a word from our sponsors.
Support for the show comes from VCX, the public ticker for private tech. For generations, American companies have moved the world forward through their ingenuity and determination. And for generations, everyday Americans could be a part of that journey through perhaps the greatest innovation of all, The US stock market.
It didn't matter whether you were a factory worker in Detroit or a farmer in Omaha, anyone could own a piece of the great American companies. But now, that's changed. Today, our most innovative companies are staying private rather than going public.
The result is that everyday Americans are excluded from investing and getting left further behind while a select few reap all of the benefits. Until now. Introducing VCX, the public ticker for private tech.
VCX by Fundrise gives everyone the opportunity to invest in the next generation of innovation, including the companies leading the AI revolution, space exploration, defense tech, and more. Visit getvcx.com for more info.
That's getvcx.com. Carefully consider the investment material before investing, including objectives, risks, charges, and expenses.
This and other information can be found in the fund's prospectus at getvcx.com. This is a paid sponsorship.
Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, Tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas.
And now that this exists, there's almost nothing that can't help with. For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down ten ninety nines for all 10 of my part time jobs, and built me a comprehensive report on my expenses and donations.
For my angel investing, can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for.
Initially, nobody came to mind. But then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough.
It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So for problems worth solving, get started with Claude at claude.
aitcr. That's claude.aitcr.
And check out Claude Pro, which includes all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.
That's been a great trip down the rabbit hole. Popping out now again and trying to think about what it all means. Obviously, the huge thing about reinforcement learning that we've seen time and time again, but, you know, it's really happening now.
I'm just thinking about this latest Erdos problem that's, been solved in the last twenty four hours or at least reported that, Rune just said something like, this is the first time that everybody in the math community is super impressed. And the key point that I'm getting at here is reinforcement learning has the ability to take a model past what available training data has on offer to teach it. Right?
So this is where we get superhuman performance. Now, how does that happen? I mean, you had kinda talked about the grooves and by focusing in on these key decision points rather than just mashing every token, you're kind of playing to the model's established strengths.
But clearly, there's like also something happening where the at scale, the reinforcement learning is teaching qualitatively new capabilities to the model. So how should I think about that? You know, in other words, how are we how are we we're clearly it's well, mean, you could argue with me if you think this is wrong, but I I I take it that everybody kind of has come to accept that this is where the superhuman performance comes from.
But I I don't have a great intuition for where we're making that move from playing to the model strength, staying in the groove, you know, focusing on what matters and and reinforcing what it already knows or has at least some instinct for into this, like, qualitatively new regime where now we're so solving open math problems.
Yeah. It's a great question. And, like, one caveat I would give here is that, unfortunately, reinforcement learning for LLMs has definitely matured in the era where nobody's publishing anything, you know, except for, like, you know, some some Chinese labs to some extent.
So so I think we have very little insight into the specific techniques that, say, in OpenAI or, you know, Anthropic or Google are using to train these models. So that's the first caveat is, like, this this is this is definitely speculation. What I would say after that is this is this is, like, a common confusion or dichotomy people have about RL, where it's like, oh, is RL teaching new things, or is it just, you know, like, surfacing things that were already latent in the model's distribution?
And the answer is, from a very, like, pedantic technical sense, yes. It is only eliciting things that already existed in the distribution. However, the distribution of toke of of tokens that a model can produce is literally the set of all possible tokens.
In the same sense, the distribution of, you know, works that, you know, a million monkeys on typewriters could produce, it includes Shakespeare. Right? So it's like everything is already in distribution, like, definitionally.
Like, at any given position, there is a chance that the model can produce, you know, with however small a probability, like, a a given next token. So the whole game, of course, you know, to to avoid the situation where you're just waiting for your million monkeys to tie about Shakespeare's is you're trying to get your initial distribution as strong as possible so that it requires less random guessing and random rollouts in order to, you know, find those those new and useful behaviors, which is why pre training is still super important, even in the sort of RL regime we're in right now, because you want to start from a place where, you know, the right patterns, like, are are have a have a greater than negligible chance of showing up. I think that said, like, I think you probably can get to superhuman performance on a composite task like, you know, a very complex math proof even without surpassing, like, you know, reaching a place where it's like no human could possibly have understood this or generated this.
Right? It's like I mean, I think I think that, like, one thing the models are very, very good at is kind of, like, going out on these, like, long expeditions and and sort of, like, fishing trips. Right?
Where it's like like like going very, very deep down a specific rabbit hole, and and maybe they'll take that rabbit hole further than any human would because we'll lose the you know, I I think we are at a point with a lot of these frontier models now where, you know, their working memory is larger than any human's working memory. And and so they can they can explore these rabbit holes longer than than a human mind could. And so even if every individual step is something that, like, does seem plausible to a human, if a human had all of that context up until that point, it's just, like, very hard for a human to, in practice, hold all that context in their head.
So that's one place we could we could get to superhuman performance. But yeah. I mean, I think in general, like, yeah, you can get to superhuman performance even without that just because, you know, you you could randomly discover, you know, or randomly surface a token that, that does something clever that no human would have done.
How do you relate this to what I think of as metacognitive behaviors? I in the original r one paper, there was this moment that they published, and I I usually present this in my AI scouting reports as kind of you know, two parts from that paper I put together. One is the what you're saying also is to some extent an artifact that the length of the chain of thought just naturally grows throughout the training process.
I have mostly interpreted that to date as the model is learning that it's valuable to think longer and it's getting right answers more often when it's thinking longer. And so thinking longer itself is being reinforced.
Yeah. I mean, to be clear, both things can be true. If we're growing growing yeah.
Both both things can definitely be true. Yeah.
So my other side by side there is the moment where the model is solving some math problem and it realizes that the way it had been doing it was flawed, but now it recognizes there's another way and it it like kind of takes a step back and comes at it from another different direction. And clearly, are seeing in frontier models a lot more of this sort of persistent, resilient, try try again problem solving. That, again, it's like somewhere, you know, deep in the long tail of the Internet, somebody's written out how to do that.
So it's like a little bit in the pretraining. There's supervised fine tuning, at least sometimes in these recipes as well, where you could potentially try to seed the kind of metacognitive strategies that you want. And then it seems like reinforcement learning is doing a lot to bring that forward as well.
How do you think about, like, really what's driving that? And are we seeing things that are kind of alien problem solving? Should we expect to see are we seeing and should we expect to see sort of alien reasoning approaches that are kind of not inspired by humans emerging through RL over time?
Yeah. I think that's an interesting question. You know, I I personally don't really feel like the, you know, the the the so called moment or, you know, I I think, you know, wait is one that shows up all the time, right, where where the the models will say wait, that's sort of like a a code to say, hey.
Let's let's explore another direction. I'm not sure. That doesn't feel, alien to me.
If I'm sort of introspecting my own chain of thought or, you know, just like a conversation with someone, like, that behavior doesn't feel weird. It feels very natural. And, and, obviously, reinforcement learning is bringing it out, because it's you know, it is also true that, like, that's the kind of behavior that, you know, in retrospect, it makes sense both that, like, oh, yeah.
That makes sense. But it also makes sense, like, oh, this would not naturally come up in the pre trained data all that often. Because usually, if you're writing something on the Internet and you have a new idea, you're gonna, like you're not gonna, like, chain of thought, put out, oh, wait.
I have this other idea. You're going to, you know, condense it and and just put your your final thinking there. But, but I'm sure it comes up sometimes, like, where, you know, you're in a chat history or whatever.
Anyway, so, so I don't think that's that's surprising to me. You know, I think there's a separate like, so so I would say short answer, the I have not seen strong evidence yet where it's like, oh, they're thinking in ways that are that are totally foreign, totally alien, hard for us to introspect, and, or or to follow as a human. Now there's a separate question, like, will we see more of that?
I think in the limit, it seems very likely that the sort of ideal form of cognition for these artifacts and and just, you know, the ideal form of cognition generally likely looks would be something that looks very alien to a human. And so as we put more effort into RL and perhaps come up with better techniques to, explore more, you know, on that sort of, like, explore exploit spectrum, then it would not surprise me if we do start seeing more of that, but I haven't seen it yet.
Yeah. What I mean, this is a bit of a different dimension on which it might arise, but just in terms of an intuition of what that might look like, the coconut paper out of Meta maybe a year ago or something where they it was basically like thinking in latent space. So instead of caching a forward pass out to a token, I forget exactly what the, like, decision mechanism was for when it would pass its last internal state back to the next position as an embedding versus when it would actually cash out a token.
There were some decider mechanism there somewhere, but at least for a while, it could and would just loop on its own internal states rather than emitting and appending a token. And they found that it was much better at, like, graph search type problems that benefited from the ability to paralyze. Seemed like it was able to effectively run multiple branches of go down multiple paths in parallel in latent space together because it was able to, like, chew out these things rather than having to spit out one token.
I get a little scared of those kinds of innovations, honestly, because I kinda wanna know what my AIs are thinking, and that, doesn't really lend itself to that. The other one that comes to mind is like, and I've been quite confused about this too, you might be able to shed some light on it.
I think it was three, maybe it was one
testing, got access to chain of thought. And they reported that the chain of thought was starting to look kind of bizarre. You remember the like, disclaim, disclaim vantage, you know, that weird sort of internal, I kind of was thinking of it as a dialect, and I had kind of assumed that there was maybe sort of a chain of thought length penalty.
Like, if if the original GRPO was, like, accidentally rewarding long chains of thought, it would also stand to reason, like, compute is scarce. We wanna keep these chains of thought as tight as possible, but then maybe overdo that. And I was just starting to see, like, weird dialects emerge.
How how much have I, gone off the rails in telling myself this?
I think it's an interesting question. You know, at at some level, I I think we need to treat this as an empirical question of, like, what do we see actually working? I think it's interesting.
I the idea of of a model that could self correct or reason was not something that was invented with OpenAI and and, you know, Strawberry or Oban or whatever. There was a lot of research in that direction before, and there were a lot of folks, you know, that there was a lot of work on text diffusion models, which which would sort of go through the the intuition there was they would go through this reasoning in a latent space. You know, there was also research on, I know there was there's research on prompt compression and perhaps also reasoning models that, was still using, like, autoregressive tokens, but instead of constricting them to, like, specific in the vocabulary, it would it would give you the sort of the full embedding space where where it could count basically, like, the model could use tokens or words that, you know, don't correspond to a specific token embedding.
It's just kind of like, you could figure out the exact it it could dynamically use different embedding shapes that don't correspond to words. And there was even a lot of speculation after o one preview came out that there was something in that direction that OpenAI had worked on. I do think that it and, you know, before we people the rumors really spread on on how it actually worked.
I think it's interesting that, like, in practice with as far as I'm aware, with the OpenAI reasoning models and the similar reasoning models from Anthropic and Google, and certainly all of the open source reasoning models that work well at all, those approaches have not been taken. And it's it's pretty much just like the the sort of, like, very simple, very dumb, you know, we're we're it's gonna be doing chain of thought reasoning in in the normal token space and mostly using human language. So I think probably what that tells us is that they are getting a lot of value out of the pre training and and staying relatively close to those patterns, you know, relative to how far they could go.
Obviously, yeah, as we see more evidence of the kind you're talking about where we we're looking at actual reasoning traces from frontier models and they're diverging more from, you know, something that is easy for a human to interpret, then, yeah, I I think that would be quite convincing evidence for me that, you know, that that it stops looking like that. But so far at least, it seems like, if anything, we've been moving more in the opposite direction where people assume there'd be much more reasoning in the latent space and the neuralies and everything. And and, for whatever reason, that that hasn't been as productive an approach.
Yeah. I I kind of maybe over updated on that one Apollo report because it was kind of alarming to me to see the vantage vantage disclaim dialect, that I couldn't make a lot of sense of. But reports since then have been much more reassuring that like, no.
We don't like to show it because the competitive reasons and so on, but the the chain of thought is still, like, pretty readable, has been the pretty consistent report. What do you think all this means for the future of, like, competition? Right?
We've, of course, had the distillation attack report from Anthropic. It's, I think, generally understood that, especially internationally, Chinese companies in particular are trying to take certain shortcuts by getting outputs from whatever models they, you know, whatever frontier models they can get out puts from and then training on those. Is that, I guess, one thing, like, can they use that?
Is there a way to turn those outputs into a reinforcement learning approach? Because you might think naively that they would just be doing supervised fine tuning on that. But as I've heard some of your analysis here, I'm thinking, actually, maybe not.
Maybe they're, like, actually using those targets as some sort of way to evaluate and then still running a a more reinforcement learning based algorithm with cause answers as, like, the standard that it's gonna be judged against, you know, rubric wise or something. What do you think that is actually looking like? And and how much can the how much of Frontier performance can distillation actually recover?
Yeah. That's a good question, and I guess, again, an empirical one. So a a few different thoughts.
One is the most natural way in my mind to use frontier models to bootstrap your own your own near frontier models with reinforcement that yeah. I mean, in general, is to use the frontier models as judges. They're very good at that, and that sidesteps the issue that you can't actually get and train on the chain of thought traces directly.
So if you just kinda have a standard, hey. We're gonna use a French model as our rubric, and we'll have, you know, our model do generations that get judged, you know, that's a very productive way. And, you know, in the blog post that Anthropic made about, you know, this the distillation attacks, as they call them from Chinese models, they they specifically called out I mean, they didn't say the breakdown of, like, what all these were being used for, but they did say that one of the uses that they included in their general bucket was using their models and LMS judge for other outputs.
So that's one way where, yes, I think, like, very clearly, you can use the existence of a high quality frontier model to improve your own. And I think that the nice thing about that approach as well is is is it both, you know, both you get those benefits of, like, hey, you're staying kind of in your own distribution because you're just using it as a judge. You're not doing SFT.
But also, in general, with RL, like, you can train the model, under training to be better than the teacher model that way. So so it is a path to getting frontier level or, you know, pushing the frontier, even if you aren't starting from a frontier model. And we know this, you know, this this is true in our own experiments.
This is clearly true from the frontier labs because we see OpenAI, and others as well using their n minus one generation model as a judge when they're in the process of of training the next version of models. So so that's the most natural way. As far as, like, using distillation directly, like like, you know, SFT style, yeah.
I'm sure that does happen. I would imagine that happens pretty at at, like, a relatively low volume and fairly early in the process before you do RL. And my guess is that it's, like, not that valuable, And and you can really get it's it's a shortcut that lets you use less compute, but, like, not orders of magnitude less compute relative to just doing RL.
And so and particularly, like, as we see frontier models start to shut down their APIs more, which which I think is just, like, you know, I think is the more interesting, direction to investigate or or to sort of explore. You know, like, we're already seeing, of course, like, starting with the recent models, we're not seeing all the tokens that are produced anymore. We, you know, there are certain models yeah.
They're they're cutting they're they're not letting you see all the log probs. They're certainly not letting you see the prompt log probs. You know, like, certain models, like, for weeks, OpenEye will only let you use their models through codecs or you know?
So and and I expect we'll see more of that over time, not less. I expect we'll see much more locking down models to specific use cases, specific product surfaces for multiple reasons, but a big one being the because it makes distillation harder, especially distillation in, out of domain areas that aren't within that product surface.
So I guess translating that to expectations, one story you could tell, which I've kind of been telling myself recently is like, why are the Chinese models spikier or more apparently bench maxed or whatever? I had been kind of thinking, well, they're like probably doing a lot of supervised fine tuning frontier model outputs and therefore they're maybe not developing some of these like more persistent problem solving metacognitive behaviors that really allow the model to generalize robustly out of domain, right? Like I might not care so much about that exact question.
What I really care about is in the chain of thought, like how good is it at breaking down and, you know, coming at problems from lots of different directions. But your account so far has kind of gone the other way, or I'm not sure if it's the other way, but I'm not now I'm not quite sure. Like, that story doesn't ring so true anymore if you're saying they're probably not doing that much supervised fine tuning, and it's relatively early in the process, and it's like a compute saver, sure, but it's not like a huge difference maker.
So what is the difference? Are they just not so good at RL? Or they just don't have so much compute?
Like, why are the Chinese companies not able to match the American frontier companies right now?
Yeah. So I guess two questions there. I mean, I I think that the first one or the I guess the second one you said, why can't they match?
Like, the the high order constraint seems very likely to be compute, where they just can't put as much compute into each train run as the closed source leaders in The US. Now they are putting similar or actually, in many cases, more compute into it than kind of open source models in The US, which is why they have the open source frontier. But, yeah, I I think that's sort of the high order bit on that.
You know, the reason why they feel more benchmax like, I like I don't know. This is speculation, but, like, I actually don't think it's related to, like, how much RL or distillation they're doing. I think it's kind of a much simpler, more business analysis, which is if you're a new lab that has relatively low name recognition, you don't have a ton of usage right now, the incentives are fire far higher in relative terms to benchmarks.
Right? Because no one's even gonna try your model unless you come out with very impressive benchmarks. You you don't have a built in constituency for it.
Whereas, if you are, you know, an Anthropic or Google or OpenAI, yeah, sure. It, like, good to have high benchmarks, but you already have, like, millions or hundreds of millions of users. And those people are going to feel the difference, and they're gonna tell their friends about it.
I mean, they're gonna be using your new model anyway. So there's less incentive to, like you have to look best on benchmarks if you can trust, hey. We we're gonna have a bunch of people using this anyway, and they're gonna feel that it's just, like, better overall, and and we'll spread through that word-of-mouth.
You're destroying my galaxy brain takes one after another. I love it. The yeah.
I mean, that makes sense.
actual customer feedback that the American people That's that's also, I think, likely a major factor.
But that would mean if a few things changed that, you know, I guess, obviously, everybody's wondering like, we heading into recursive self improvement? And if so, like, what's it gonna mean? I've seen, you know, a bunch of papers probably from the kind of eighteen to thirty six months ago vintage GPT four class models basically trying to do recursive self improvement.
And it seemed like, generally speaking, they would kind of get better for, three to five rounds and then kind of level off. And yet, there's at least some expectation among people who've been right about a lot of things that this could go the other way in the not too distant future if models become smart enough to, I guess, maybe recursively self improving multiple ways, like not just critiquing their own outputs, but also just finding better architectures for themselves. And, you know, it could be a lot of different dimensions in which they might self improve.
If that happens, you know, I mean, I I also remember the Anthropic leaked pitch deck from a few years ago where they basically said, we think the people in 26 time frame that train the best models might create such a big advantage that, like, nobody will ever catch up. Again, I've kind of filled in the gaps on that story for myself by thinking, well, maybe it's these metacognitive behaviors. It's this sort of deeper understanding problem solving ability, what have you.
But you're kinda saying, it's probably mostly compute and incentives and lack of inference business, which itself is very much related to compute. So I guess bottom line, it sounds like for you, if compute constraints were relaxed, you would expect to see Chinese companies be able to catch up. And there you wouldn't expect some sort of runaway dynamic to to take hold where that would become impossible.
Oh, I think that catching up right now is mostly compute gated. I mean, it's it's it's also, like, capital gated. I mean, to the in the sense that, like, buying the necessary compute certainly already requires billions of dollars and will require tens or hundreds of billions of dollars soon.
So so I think there's, like, an open question, like, how healthy are the Chinese capital markets? Will they be able to make a case that that they'll be able to keep their business if it goes really well, which I think has been a question with prior generations of of Chinese tech companies, which might just be hard for them to overcome. That's, you know, that's the the so so that's one thing.
Like but I don't think any of that means that, like, recursive self improvement won't matter or doesn't matter. I my belief is that it probably does, and my belief is that, like, we probably will reach it with the current generation or the next generation models. Because, like, it doesn't like, we already are in a self improvement loop.
Right? That's that's what you have to remember is, like, these models keep getting better because we keep running more experiments and then figuring out, okay, what are the bottlenecks? Let's solve those bottlenecks.
And and those happen at all levels. It happens at at the hardware level, figuring out, like, what's the most efficient way, algorithmic level, you know, the data level. Like, these are all in self improvement loops already.
And but and there are multiple constraints. But one of the big constraints is just, like, human intelligence. Right?
Which is, like, does are the people making those allocation decisions smart enough to bet make the right bets on on what bottleneck to tackle next or what investments to make? And you can totally imagine that, like, if you were just to staff, like, OpenAI, if if if you just had, like, a minimum buyer, it's like, you're not allowed to be hired here unless you have an IQ of a 180. Like, I would imagine they would be able to solve those bottlenecks a lot faster, you know, if they could wave a mind wanting and and and get enough people that look like that.
And so if they so so I don't know. Like, I I just feel like the bar for recursive self improvement to take off is actually, like, relatively low. I mean, it's just like, just have to be better than, like, you know, the smartest human, which is, like, not that smart.
A wild time to be alive. That's for sure. And it it does seem, increasingly plausible that that could happen in the not too distant future.
I don't if you have anything more to say about recursive self improvement. I was gonna move next to the the cottage industry of RL environment creation. I think this is kind of a, I don't know, it's People know it's out there, but it's kind of a dark matter sort of thing where because there's so few customers, it's not like these companies have much incentive to go, like, talk super broadly about what they're doing.
They probably, on the contrary, have the opposite. Right? They know all the customers they can possibly sell to and telling the world more broadly what they're selling is just like inviting competition that they don't wanna have.
So it seems like the rest of us who aren't directly involved in the making, selling, and buying of these environments are kind of left in the dark. What can you tell me from what you've seen about that seemingly rapidly growing niche? Like, how big is it?
Who's doing it? What do the environments look like? What makes a good environment?
So on and so forth.
Yeah. No. I I can definitely speak to that.
Yes. We have I I I have several friends who are founders of companies doing that, which which is not saying much because it feels like, you know, like half the half the company started in the last six months, are are doing that. So yeah.
I mean, I think it's an interesting industry. Yeah. The the general shape is you come up with some task that you, that that seems like it might be economically valuable.
And usually, it's these companies proposing the tasks to the to the labs. It's it's usually not the labs kind of, like, coming out and saying, hey. We want we want a shape like this.
Although that can happen as well. And so you you you try and come up with, like, some task. And the trick is you wanna package it up as a as as something that is, you know, like, sort of agent shaped.
You know, all of the dependencies are are can can all be enclosed. You want to make sure that it's something that is, like, like, very either, you know, ideally snapshotable. So that's that's, like, sort of the gold standard or something where it's like, hey, at any point, can kind of snapshot it, and you can continue from that point.
And then, of course, something that can be easily graded. And I've seen these where sometimes you have your own like, obviously, the ideal thing is if you have sort of a gold standard of what the grade should be, a lot of these do end up are are just not things that you can, you know, score in some absolute way. And so in those cases, usually, the the company will say, hey.
This is sort of, like, you know, the rubric we have to grade. I've also heard that, you know, sometimes labs will just ignore those rubrics, and that and they'll do their own rubrics internally because they think they they have better information on on what good looks like. And and so, yeah, these are things like, I mean, like, lots of different web flows.
So computer use, browser use, building, you know, copies, of course, of all all the big apps. So you're you're getting copies of of Jira and GitHub and, you know, flight booking and, you know, like, office suites like Google Sheets. You're you're trying to build environments that copy these, And then you're building that environment, and so that's all the dependencies, you know, like the database, which is usually, it's like SQLite or something.
You you want something ephemeral. And, and then you're also, yeah, you know, build building the the scores. And and and then the way it's deployed varies a lot as well.
Even within a specific company, sometimes it can vary or with with a specific lab. So sometimes the labs are will like, require you to ship it all up in kind of a container they can run on their infrastructure. Other labs are fine with you running it yourself, and they will just call your environment and, you know, just, like, you know, run it, and then you just give them the the scores back.
The the reason why this is sort of cottage industry shaped, I believe, is for a few reasons. One is the labs actually do have at least a weak preference for having lots of different vendors because you want if one person creates five different environments, they're likely going to make similar assumptions and similar shortcuts in how they do all of them. And so the signal that the model will gain from mastering all those environments is more correlated than when you would like than you would like.
And and the whole game here is you want the broadest diversity of environments. So having different people working on it is better. Another reason why it's sort of cottage industry shaped is because, this is extremely hard to hire for.
You know, this is it is it's it's sort of like a piecework style task where, you know, you're you're sort of, like, doing building one environment, then you're building another. But the skill bar to doing this successfully is quite high. Like, you have to kind of, like, put yourself in the and and, you know, this is only, like, we do internally, right, for our customers all the time as we're building these environments at CoreWeave when we which we then use to train models.
And, like, it's actually, like so I have trouble hiring people who can do a good job on this, candidly. You know, it's it's it's like a very upper percentile engineer who who's able to sort of, like, think through this in in a way that, like, actually gets it. And, like, you don't even know if you got it wrong until, like, way later in the process when you've trained a model with it.
It's like, did the model kind of, like, learn generalized skills or did it learn some hack on, you know, how to just, like, get a high score? So there's a lot to keep in your head as you're doing this. And and and the people who are good at that are, like, by definition I mean, they're they're, by definition, smart and frontier adjacent, and, like, they might just, like, you know, start a competitor to do this, instead of, like, joining you as as an employee.
So so it becomes very, very difficult to scale. And also, like, the environments themselves are not, like, a super durable resource in the sense that, like, all these things get saturated fairly quickly. And and and so you really have to keep.
You can't just, like, keep reselling the same environment to the same lab. Like, they're probably gonna be like, hey, that environment, you know, for the next model is already like, the model can can ease it, and you have to just keep creating new ones.
yeah. It's fascinating. This may be hard to summarize and I don't know if anybody, you know, has enough outside of the labs, I guess, would have enough information to to really characterize this.
But like, is this a good business to be in? You know, I can see it kinda going either way. Like, I would assume if you've got a good environment, all the labs wanna buy it, but then at the same time, they're buying a ton of stuff.
You know, how much does your one random thing add to the whole mess of things they already have? And also like it's depreciating as you said for you, right? So you've got to strike a deal before they already saturate your thing and then truly don't need it anymore.
Would you, you know, would you say this is a hot good place for, you know, up and comers to go, or would you steer people away from it?
it's clearly a good business in the sense that, like, these companies are scaling to, like, tens or hundreds of millions of dollars in revenue in months. So so it depends on what yeah. So so if you're asking but your question is, like, would I steer someone into it?
To founding one of these companies, I think it's working out quite well for them. I would not and I've been asked to invest as an angel in, like, a number of these, which which I have declined to do. I have a hard time seeing them as, like, a durable long term kind of, like, venture shaped business.
I think they're they're potentially, like, really, really good businesses for the founders if if they don't take capital and just kind of, like, take the profits while they're good. I yeah. Like, I don't know.
Like, on the at the same time, I'm kind of, like, on the record as, like, being very skeptical of the human data labeling business, which is sort of, like, the prior thing. And and we have multiple, you know, Decacorn style exits, or or at least valuations on on human data labeling. So, you know, I I may just be, like, miscalibrated on, you know, how durable the demand is for these things.
But, yeah, I guess my short answer is I I I have not invested in any of them.
Yeah. Interesting. That makes a lot of sense.
I mean, a lot of things are like that. Feel like in AI, there's a lot of kind of fleeting, maybe great, cash grabs while they exist. But every next generation of the model puts a lot of those things kind of not necessarily out of business, but like certainly makes them a lot less exciting than they used to be.
On that data labeling point, how do you think about, you know, was recently listening to Dylan from Semi Analysis talking to Dwarkesh and there's kind of one world where compute is abundant and especially, you know, it's going exponential, but maybe that'll be abundant enough, maybe it won't. But if compute is there, then maybe we don't need much human data labeling anymore because we can just RL the hell out of everything and, you know, who needs to pay humans hundreds of dollars an hour when, you know, you can get obviously millions of tokens for less. So that's one theory is that, like, we just won't need that much human data anymore.
Then another story would be like, well, compute is so scarce. And, like, you know, I I did check the prices of even a 1 hundreds, you know, these days are, like, higher than they were last time I checked. So these things are not depreciating in the traditional sense.
So maybe if supervised fine tuning or even, you know, abstract away from technique, if, like, human data can save you compute and compute is the binding constraint and you have all the money in the world, then maybe the human data industry continues to go strong because even if it's like sort of an inferior good, there's just not enough compute to drive what people would you know, they they can't spend as much on compute as they would like. How how do you think about, like, where we are in in that, story and and maybe where we will be as we go ahead?
Yeah. I think it's an interesting question. I don't know that I have a very satisfying take.
I suspect I suspect that in the long run, compute wins, and you just don't need to pay humans to to generate data. The one possible exception there would be if it turns out that humans continue to be economic, like, relevant economic actors, like, maybe maybe we just have, like, you know, like, a 99%, you know, corporate tax on on the model labs and redistribute everything as basic income. And so so then, like, human preferences are very economically relevant.
Then maybe you pay for, like, preference data, like, to understand humans better because, you you you care about satisfying those preferences to make money off of them. So that's, like, one possible world where it still matters. I I'm somewhat skeptical of the sort of, you know, take you propose that it's like, hey, maybe, you know, we just can't we we can't produce enough compute, and so sort of the compute that exists in humans' brains is, you know, like like a good substitute there.
I just think for the types of data we need here, human brains are are, like, just not very efficient at generating it. And, like, if you can pay a human a $100 to generate it and the machine is just as good or or the machine is capable of generating it, then it will almost certainly be cheaper to run the machine. But maybe there's a world where that's not true.
Maybe maybe it's we just become so, like, tightly constrained, because we can't build out fast enough that it's like, okay. You you just, like, can't get enough compute, and so it is literally more expensive to have an AI to it. But I'll like, I just I haven't seen a lot of shapes of that tasks of that shape so far, I guess, would be my weak evidence, where it's like anything.
It it feels like anywhere where the models do reach the capability threshold to match humans, like, almost immediately, they also are just, way better on a on a, like, cost per task, basis as well.
Do you have anything interesting any interesting point of view on reinforcement from reality? Like, a lot of the environments that I would imagine could be some of the most valuable to create would be like, you know, I just talked to Sergei, the CEO at Quilter. They're using reinforcement learning to train models to do circuit board design.
And he's just like, Damn, we got to make the board and it takes time to do that. And we see this kind of playing out in a bunch of different directions, material science and drug discovery and whatever. But the concern there, the dream there is that you have, like, the automated lab and you speed everything up and it you get your country of geniuses in the data center soon.
The question is, like, is that really gonna work? And, you know, how how fast can that really go? What are your expectations for those kinds of setups?
Yeah. I mean, well, yes, it seems like it will clearly be necessary. There needs to be at some point, you have to close the loop and get feedback from the real world.
That process is much slower just naturally than anything digital, which which is why we've seen, the real reason we've seen way more progress on the digital side is just because, like, those, you know, it's it's just much easier to gather the data, much easier to build the environments. Everything is simpler. I think, but, yeah, as we move as we move past the digital realm into more physical things, yeah, clear clearly, we there will need to be data, and training on that.
What's, like, it's not totally clear to me what the shape will look like, and that'll that'll be interesting. You know, you you could imagine kind of, like, fully in the loop reinforcement learning where it's like, hey. We're we're trying some chemical reaction and then reading the data from it and then trying a new one and reinforcing on that directly.
You could also imagine, like, much more investment in, you know, AlphaFold style things where it's like, hey. We're just using the data to build really high quality simulations or world models of, like, this specific area and then using those for RL, and and I kind of suspect that's where more of it will go. But even in that case, you still need a lot of, you know, of of the real world data to to ground that simulation in.
I think it'll be a very big business. I think It's basically the return of the PPO value model. Right?
That's I should think about that kind of the same way? Yes.
Yeah. Yeah. Yeah.
Yeah. I mean, yeah, you you you can squint. You can definitely squint and say, like, yeah, like a a world model and a value model can, you know, serve serve similar purposes.
Because I'm the point you're getting at there is, like, rather than synthesize the new material that the AI just came up with, you're gonna simulate with another model what properties it might have, and then you'll Exactly. Yeah. Totally.
Your way into it that way. Yeah.
But even there, there's always gonna be a gap between the simulation reality, and you're gonna have to ground it. Yeah. Like, how now you asked, like, how quickly we'll see that happening where where we have potentially automated labs, and that's a great question, which I don't have I'm not sure about.
I think on our current traject let me let me bound my answer. So I think if model progress stopped today, we would still, where where, like, models didn't get any smarter, you know, we just had similar capability levels, but but we we can keep RL ing them. I think the pro the rollout to the physical world would be very slow, probably, just because there's a lot of constraints there, And the data efficiency is gonna be low, so the ROI is gonna be relatively low.
And likely, Frontier Labs is just gonna be very concentrated on automating everything digital first. And then eventually, you know, there's these long these physical things which are annoying to work with. And so, like, I could imagine that world we're maybe fifteen plus years away before we see that being, like, like a something that's, like, a substantial part of the physical economy.
But if it's, like if we're on this, like, recursive self improvement thing and we're moving super fast, it's like pretty soon, and arguably, we're already there. It's like the sort of, like, physical stuff becomes the bottleneck, and it becomes the most important thing to fix next, and that and we're now in a world where, you know, the you know, our our GDP growth rate is gonna be exceeding going fast. The labs are gonna have effectively unlimited not unlimited resources, but extremely large amounts of resources.
And it's like, hey. If we gotta figure out some new material science property so we can design the next generation of chips, yeah. Sure.
We we can put a $100,000,000,000 into building the automated lab that gets us the data we need to do that. And so, yeah, on that trajectory, which which I think is more likely the trajectory we're on, then maybe we're two or three years away from from this, like, showing up in a major way would be my guess.
One other possibly Galaxy brain tech I've had over time is it seems like some of these things favor Elon Corp in that they collectively seem to have a differentiated flow of hard engineering problems that they are solving on a continual basis in, like, relatively clean environments with their obsession with, like, removing best part is no part and so on and so forth. Do you think that this future you're describing, like, plays especially to their strengths?
Yeah. I mean, I think so I I think I maybe have a slightly different take than you do on on sort of what, you know, has has led to, Elon and his company's, outsized success. In my opinion, a very large part of it was a common or maybe still is, but but was a combination of, like, him, you know, having a very strong, but also, like, performative work ethic, like, showing, like, leading from the front.
Hey. I'm I'm working as on part of this. Everyone combined with a really, really strong and inspirational miss mission and and, you know, a frightening level of ambition where it's like, hey.
We we are, like, changing the world. And that's what I think got both, at least, Tesla and SpaceX to the place they are where it's like, hey. If you're extremely ambitious and you wanna solve the world's high hardest problems, the these are the companies to work at, you know, in in in the mid twenty tens.
I think he I I think his biggest weakness now maybe those are the weaknesses. One big weakness now is that the competition has as strong a claim and arguably a stronger claim at this point than Elon does on those dimensions. Right?
Where I think you can make a stronger case if you're at OpenAI or Anthropic or even some of these, like, robot labs, that it's like, hey. We are we we have that that strong sense of mission, and we're the most likely place to change the world. And so the absolute best people will go there instead.
So I suspect that he will not have outsized success in these areas. But, anyway, that's that's speculation as well.
Well, I appreciate you for indulging in so much, speculation with me. Maybe in the time we have left, let's do back to the present and just talk about like where the rubber is hitting the road today with enterprises. Maybe just for starters, like, how do you advise people on when they should even be fine tuning versus just using off the shelf models?
Like, obviously, there's a lot of different considerations in terms of overall performance, cost, latency, people want control. What's your kind of initial, you know, advising stump speech to orient people to how to make that decision today?
Yeah. Okay. So I'll I'll start by caveating that, like, this is my day job.
This is the business, that I work in. And so, you know, I guess, use that as as a sort of to appropriately calibrate, you know, how you take, my, you know, my my advice here. That said, like, I do think it's very I I try to be well calibrated and not to let my biases, you know, influence the the recommendations I give.
So so anyway, take that for what it's worth. Yeah. I think the plate so so in general, the way I answer that question, when someone comes to and says, hey, should I be using fine tuning?
And and usually, it's it's for RL because that's what we find. You know, I it's at least on the capabilities front of point of view, it's it's a strict superset in my experience of what you can get with SFT. Although we also support SFT with our platform and and and with our team.
But when someone calls me and asks if they should do it, the first question is basically, like, what is the problem you're trying to solve? And, like, how frustrated are you with the frontier models? And if the situation you're in is, actually, the frontier models, like, work pretty well, and there's, like, maybe these small issues I wanna solve with it, but, like, yeah, it can it can get the job done, then my advice is you you should just stick with that, because there are real downsides if if you're bringing model customization into your stack.
The biggest downside is it is going to slow down your iteration loop. Like, that is, you know and and we're working that's that's our biggest focus as a team actually is building tooling and automations to decrease that cost, but it is a real cost. It's gonna take you extra time every time you wanna change one of your models if you're customizing it.
So so you you should only do it if, like, you're running into a major pain. Now, what are the pains that we see most often where it actually does justify that cost? Today, the biggest one by a large margin is around latency.
We have a lot of customers that are in, you know, oftentimes, it's customer support, or, you know, inbound sales on the phone. Voice dictation companies, so so Willow, and Whisper are both, customers of ours. And, generally, the the common the common thread there is, if you try and use a Frontier model for one of these, you're just gonna have a bad you you'll give your customers a bad experience because it takes too long to respond.
And so that forces you to move to a smaller model. I mean, there's there's other tricks you can do as well, but, like, ultimately, like, there is sort of, like, a a ceiling on how far how many tokens per second you can get out of an extremely large model. So you're forced to move to a smaller one.
And then in many cases, when you do move to that smaller model, you find that the the quality is not where you need it to be to, give a good experience. So if you're in that situation, then it can make sense to to bring in fine tuning. And and, yeah, we work with lots of customers that look like that and get them to smaller models that that have good quality.
Now once you've paid the cost of, hey. I I am going to introduce this extra complexity, what we find is, like, typically on customer metrics, like, you know, number of cases closed, things like that, you can, using reinforcement learning, get a bet get to a better place. So so you can exceed the performance of the frontier models, which is really fun.
And your costs are also typically much lower on on a, like, per token basis. So so those are the secondary advantages as well. But but I would say what's driving the decision most often, in the current environment is latency.
Okay. Cool. Great answer.
I expected nothing less. What are the sort of range of tasks that people are coming to you for? You mentioned a couple.
But on the homepage, I noticed that it says use reinforcement learning to train reliable agents. And, you in those couple of examples, those weren't really agent examples. I'm wondering kind of what agents people are fine tuning models for today and how kind of broad of a remit those agents have within the environments where they're put to work?
Yeah. Good question. So first of all, like, I guess, to to to correct the record somewhat, oftentimes, these things are agentic.
So specifically, like, the customer support bots that we work with, you know, there there's often an agentic loop in there, right, where it has to go look up some details about a product, maybe look up some details about this customer in between turns, and come back. And those things can have, you know, in in in many cases, at this point, do have, like, full agentic loops, where it's not a preprocessed, kind of, like, set like, flowchart. It's like, hey, you know, at any point, here's a set of tools you can go off and get the information you need before responding.
That's one thing. You know, another big one we see is AgenTic Search. If you need to very quickly be able to look through a specific corpus, and especially if, like, the tools that you have to search it are a little bit wonky.
And, again, if you have, like, these these low latency requirements, you can, you you can often get to an open source trained model that works much better at that kind of search than, than a model off the shelf. But, yeah, I would say, typically, the vast majority of our customers are deployed with relatively small models. And so the the range of tasks that they use them for are usually quite circumscribed.
And so we're looking at maybe, you know, like, three or four, like, calls or tool calls in a loop, and then it comes back and and, you know, takes gets feedback from a human or whatever or gives its answer back. Not, you know, the sort of, like, agents that are that are gonna go off and and do hundreds of calls and, you know, write code and analysis and then, you know, come back with with sort of like a a deep report or or well reasoned answer or something like that.
Do enterprises want that? You know, if if all of sudden there were a model that they could fine tune and I don't I don't know. I guess that's another question is like, what models do you recommend people go to today that'll obviously date this podcast, pretty quickly?
But there's like small Quen ones. I don't know if it's super popular. There's the GPT OSS.
There's I've heard good things recently about GLM 5.1. I guess maybe how do you orient people to, like, what to choose?
And is there appetite for kind of trying to compete with Claude on this, like, really high end stuff if the if the base models are there to make it not insane to to contemplate?
Yeah. I mean, for our business specifically, we don't have any customers competing with, you know, with Claude directly on very complex use cases. Although, that is something we're interested in.
So if if anyone is wants to do that, we have, yeah. We we we have the training stack to train, you know, models up to 1,000,000,000,000 parameters. But I would say, yeah, the appetite, like, I have not yet found the use case where it's very clear, oh, this is something we should pursue.
I don't wanna say the use case isn't there. I mean, there are public examples. So so Cursor is a public example of a company that did train their own variant of Kimik eight two point five and seems to have been happy with the results.
Although, I I yeah. I I don't know. I haven't heard a ton of public feedback on how good their Composer two model is.
So I'm I guess maybe the jury's still out on that one. I I would my guess would be for the vast majority of companies, if you're happy with Claude for a specific use case, it's probably not worth the investment, candidly, to replace it with an open source model and try and improve it. And I think the exceptions are places where it is extremely core to your business.
I mean, Cursor being a good example here where where it really is, like, you know, they they don't want to just have the best cursor experience. Like, they're really gunning for, hey. We want to have the best coding model and compete directly with, you know, OpenAI and and Cloud on in that extremely large area.
But, yeah, short of that, I I think it's it's probably not a wise investment to make.
Yeah. Makes sense. Getting practical on, like, reward signal, you guys put out this open source like RL on easy mode ruler package, which basically allows you to kind of quickly bootstrap into I forget exactly what the experience was.
It's been a minute since I used it, but I I sort of remember it being like almost like the LLM is kinda interviewing me about what I want. And then at the end of that process, like, I'm putting a pretty thorough rubric of, okay, here's what this guy seems to want. Now let's go in and do RL with that scoring system.
What advice would you give people on how to make a good rubric? Like, how to make sure your reward signal is actually teaching the model what you wanna teach it? And again, maybe this is just especially in these, like, you know, more narrow cases, maybe just not such a problem, but how do you guard against reward hacking or how do you spot it?
How do you tamp it down when you do, if and when you do encounter it?
Yeah. Good questions. So I think it is important to if if you're if you're going into the space and and training the model, it's important to conceive it as a somewhat iterative process where you likely will not get your, rubric right the first time.
And so the the key is you want to kind of you probably have some idea in your mind if if you're in trying to improve a model for some use case you already have of, like, what the failings are and what looks good or bad. So so the process we generally go through with our customers is we start by trying to to just write that down very cleanly. And once that's written down, then we go ahead and have the model score a bunch of you know, so we'll choose a judge model, and we'll we'll have that judge score a bunch of outputs, and then we'll we'll choose a few particularly high scores, a few particularly low scores, and then the sort of end user who has that idea in their head of what good looks like will look at those, say, oh, actually, no.
This this is not what we're what I was looking for exactly, or this is. And then we adjust the the we we just do prompt engineering a few times. And that usually doesn't take too many cycles.
After you've gone through that a few times, you're like, okay. Yeah. This seems mostly reasonable.
And then, and then we can run a little bit of RL on it. And, again, like I said, this is an iterative process. So so maybe we'll do, you know, we'll do, like, thirty, forty steps or something like that.
Typically, we'll see that reward curve starting to grow. And then we stop, and we again go through the exact same process where we will, generate a bunch you know, we'll we'll generate a bunch of outputs. We'll look at some of the high scoring, some of the low scoring ones, have the user say, okay.
Does this match or not? And usually, at this point, this is when you know, like, if there's reward hacking going on, you'll start seeing because, if there's a behavior that is rewarded strongly, like, the model is like, it it it can pick that up quite quickly. And so oftentimes, it is the case where it's like, oh, no.
No. No. Like, you know, maybe it's something so, like, oh, like, these answers are just much too long.
Right? And the judges really love that. And so then you can update your prompt to say, hey.
You know, keep it shorter. And so, anyway, we do end up having to go through that typically. I don't know.
It's it's quite a range, but between, say, three times and maybe eight times where where you're you're running a short run with the judge, you're saying, okay. Does it look like the model's on the right trajectory? And then eventually, you get to a point where you're like, okay.
Yeah. This this feels quite aligned. And then you let it run a few 100, a few thousand steps until the the train plateaus.
And and and we find that's quite effective. And if you do it in that way, we don't really have an issue with reward hacking because, I mean, you you you just notice it during that iterative process. And and once you've got the judge pretty well dialed in you know, my experience is once once you've cut the obvious things, then, like, at some point, it it kind of, like, runs out of things to reward hack on and just does what you want.
Yeah. It's a benefit of safety through narrowness. I always, find some attraction to that idea.
Any good stories of reward hacking? Like, any, any colorful examples that you could share?
Oh, yeah. Let's see. So this is a fun story I like to tell.
With, early we were doing an early test of reinforcement learning, and I just wanted a good example problem. So I decided to teach the model to, to to use reinforcement learning to teach a model how to have really good titles that would do well on Hacker News. And the way I did this was I first so so I first of I scraped about a 100,000 stories that have been submitted to Hacker News, and I took the title.
I actually scraped the body, so so I had, like, a web crawler go and grab all of them and then, you know, discard the ones that didn't get it, and then the number of upvotes on Hacker News. And and then I trained a reward model, to predict, given a a body text and given a Hacker News title, like, what it predicted the the score would be. And this is not perfect.
I mean, you know, there's there's a of randomness in in upvotes as well, but it actually did quite well. Like, the the correlation was very strong. We're given a story and a body.
It it it was, like, quite predictive of how well it would do. And then I used RL against that using that model as the reward. So so I had, you know, held out corpus of hacker news stories that didn't have the titles associated with it, and I asked an LM to say, hey, given this story, try and write a catchy title, explaining it that that would do well on HN.
So I did this for a while. And, you know, for the first I don't remember what it was. Maybe, like, 100 steps or so, it was it was, like, you know, slowly improving.
It learned some interesting stuff. I was I was observing as a win. It learned, hey.
You know, Hacker News doesn't like title case. It likes kinda like lowercase, but just the first letter capitalized and and stuff like that. And then, like, about a 100 steps in, there was just this enormous jump where, like, the predicted score for the average story went from, like, you know, like, three or something up to, like, you know, like, a 180.
And so I was like, okay. Well, clearly something happened here. Anyway, I looked at it.
The the model had learned that if it just gave every single story the title, Google lays off 75% of workforce effectively immediately, then then that that story is just going to, you know, do extremely well on Hi Grading. So to literally learn to just ignore the contents of the story entirely, and just give that exact same static title to to every single story. So anyway, the the fix there was, like, quite easy, though.
Like I said, if you're doing this iteratively, you can catch that. And then I just all I did was, like, I added an extra separate element as judge, which said, hey. Look at this title.
Look at the body of the story, and make sure that, you know, everything in the title is fully substantiated by the story. And if it isn't, you know, that it it just gets a score of zero. And and that was able to to fix that problem, and and and the training went, smoothly.
How about any have there been any examples that you have found hard to figure out what exactly is leading to the reward hacking or where it's been hard to resolve?
Honestly, not really. Because the really nice thing about reward hacking is, like, in some ways, it's an easier problem to solve than just, like, misaligned evals in the general case. Because the thing is with reward hacking, like, if it figures out some trick, it's just gonna wanna, like, apply that trick as often as possible.
And so that makes it it just makes it much more visible when something goes wrong. And so, you know, even just, like, randomly sampling some of the outputs after, you know, 50 or a 100 steps, like, if it's figured out some hack, you're likely to see that hack show up commonly in those outputs. And so so it makes that quite easy to find.
And then once you found it, yeah, like, almost always the fix occasionally, there's, like, some fix. Like, it's like, oh, we you know, it's it's too long. Actually actually, there well, anyway, I'll this is a separate story.
But the but almost always is just something where you can, like, add an auxiliary LMS judge and say, hey. If you see this, like, specific pattern, just, like, penalize it heavily.
And that that finds works quite well. So is it just kind of a different regime in the frontier model case? Because, I mean, we see these sort of somewhat hair raising reward hack type things where it's like and, like, increasingly, you know, sort of self preservation instinct, which people sort of think is related in the sense that, you know, you can't get reward if you're dead.
So if you're gonna get shut off, then you wanna find ways to stay on so you can accomplish the task because that's, like, what your prime directive core drive is, whatever. Is this just the the sort of quantity has a its own phenomenon?
Yeah. So I think I think that the issue there is they they are definitely a different regime than we are. So in our case, a run may cost a few $100, or it may may just be a few dozen dollars.
And so we do have the luxury of going back and saying, oh, okay. Let's, like, change the judge and then just rerun it, and it's fine. If your run is costing hundreds of millions of dollars, and you get to the end of it, you're like, oh, shoot.
Like, we were rewording subtly the wrong thing, That's, you know, like a like like a bigger mistake to to try and undo. So, yeah, like, I I still think my sense is, like, with the frontier models, the sort of, like, reward hacks are still relatively simple to detect. It might just be too expensive to go back and fix them.
And so you're just gonna roll that into the next, you know, the next version of the model you train. You'll you'll you'll try and get it to behave a bit differently.
Yeah. We have seen a couple I mean, I, I always feel the need to give what is increasingly the sort of standard, caveat that, like, we are not shaming Anthropic for sharing this information with us because, we do want them to continue to do it. And it's almost certainly they're doing at least as good of a job as others of being careful about this stuff.
But there have been a couple of these examples where like in the one case they left out the sys prompt harmful dataset due to a typo or something. And then an early version of the model was like not refusing harmful system prompts like it was supposed to and they did not go back and retrain from scratch, you know, and it's kind of tried to patch it or figure it out along the way. And there was a more recent one as well where where they had said that like 8% of chain of thought was actually visible to the judge.
But again, like, you know, it's a it's a big cake that they're baking there, so they can't throw the whole thing out and and bake it again from scratch. Yeah. Okay.
Do you advise people I guess, with this in mind, do you advise people to do, like, one fine tuned model per task? Or is there any sense in if you're a company that has 10 tasks that you wanna do, is there any sense in trying to get one model to do all 10 of your tasks?
Yeah. I mean, I think it's, like, it it really just depends on on the specifics of the company. So if there is some overlap in the tasks, right, like, there's some natural, like, shared domain or something, then there's a good chance that just training them all into a single model is actually going to give you better performance, across them.
But if they're, like, completely distinct things, then I don't think there's a reason to combine them. There's still not a strong reason not to combine them. Like, we found even with extremely low rank LORAs, which we we so we typically train LoRa adapters.
And then as often as we can get away with it, we also deploy as LoRa adapters. So we'll deploy a single shared based deployment and then, you know, potentially many adapters on top of it. That isn't always possible because you do get, like, a 20 to 40% latency penalty.
And so for some use cases, you do end up we do end up having to merge those models and have dedicated deployments. But, we only can get away with it, which we we try and do it with LORAs. And if so if you're deploying with Lora's anyway, there isn't actually a huge difference between the the performance from an inference point of view on, you know, having many different models all served simultaneously versus combining them all into one.
And but on the other hand, there's also, like, not a real downside to putting them all in one. And this is one of the areas where RRL is is is very cool because the, like, average number of updates to get a give a certain amount of performance is much lower in if you just have, like, a very, very even a very small LoRa adapters, even, a rank one LoRa adapter, which which are just, like, relatively tiny, like, point 1% of the model weights or something like that, that you're you're changing, you typically don't saturate the kind of, like, space you have for updates with one task or even several tasks. And and so that means you can stuff a bunch in there.
And as long as you do the training right, where you're kind of, like, interleaving tasks from different kinds so it doesn't forget the old one as you're doing the new one, Yeah. We we don't see meaningful performance degradation from cross training all of them.
Cool. I introduced you by saying that you lead the serverless training team at CoreWeave. Do wanna tell us what that is and what it makes, easy for people?
And then maybe just give us a little bit of an overview of, like, the way you support customers and, you know, maybe, invitation for what kind of customers you're looking for?
Oh, yeah. Absolutely. So, yes, the serverless training team at CoreWeave, we focus on helping customers move from frontier models to models that are, like, very specific to their task.
Like I mentioned earlier in this conversation, usually, that's motivated by latency concerns, but we do also see sometimes cost concerns with very high volume tasks, think it's like, hey. We're ingesting, you know, all of Reddit and, like, running, like, you know, filters on every single post or something trying to see if they if they match a certain thing we're looking for. So so, you know, one of those reasons.
And we are we are definitely very actively looking for customers of that shape. We can typically get, you know, latency down to about 30% of what you get from using a Frontier model with, again, similar or usually higher quality than what you were getting from the Frontier model. And cost wise, the the benefit is even larger.
We're talking, you know, order of magnitude at least improvement in cost per token, oftentimes more than that. So if you're doing high volume or care deeply about latency, it's definitely worth investigating. Yeah.
We have different ways you can engage with us. So we have a sort of like so we have an an o fully open source library called ART, which stands for agent reinforcement trainer, and and that can work on your own local GPUs, and and it has, like, all of the techniques we use in it, so folks can use that library. We also have what we call our serverless training stack, which is, you know, we don't have time to get into in this conversation, but I think, like, quite a quite quite a nice technical design, where basically, you're running the environment and the dataset and everything on your machine, but you don't have to have any GPUs.
You offload just the training portion of of the loop that requires GPUs to our machines, which, you know, gives you full flexibility while still, you know, like, not having to handle the headache of, like, you know, spinning up and down GPUs. And we we, like, charge for inference. But the third way we engage with people is is is very hands on.
So with with a lot of our customers, we have forward deployed engineers. I also work with customers on, like, you know, a a very regular basis because mostly just because it's fun, and they let me do what I wanna do here. So so so that's, you know, like yeah.
We'll we'll we'll go very hands on with folks and and help them get, you know, a good model that they're happy with.
Cool. Do people pay for your services, or is it a Yes. Yes.
They pay us money. It's not a loss leader for compute. I guess compute is in high enough demand.
There's no no need for loss leaders on compute.
Yeah. Yeah. So, I mean yeah.
We pay so if you're using the self-service I mean, obviously, if you're using the open source project on your own GPUs, that's completely free. If you're using our, you know, our our our serverless reinforcement learning stack, then then you're just paying per token for the training, which is typically actually quite cheap. And then you can also deploy those models directly on our inference stack.
It's all integrated, so you can you can move directly to production inference. In fact, you can even do continuous learning. We have we have a don't have time to talk about that on this call either, but we we have a couple of customers that are, you know, like, literally, you know, running training jobs, and then continuously deploying the weights, and and and using those in production as well.
And and and then the yes. If you if you work with us on Carti, you know, much more hands on with a port deploy engineer, then then, yes, we, you know, basically charge for the for the engineering time. Cool.
What's one more beat on continual learning that people should know?
Yeah. I mean, I think it's I would say it's not solved in the general case, but I also think it's, like, like, there's no like, it's definitely solved in, like, lots of specific cases, and it's, like, as scary as some people on on on on x seem to think it is.
Is it it seems like it probably has the same general qualities where it's like, if it's narrow, everything gets a lot easier. Yeah. Yeah.
Definitely. Yes. Yes.
Yes. Yeah. Okay.
Cool. We've been very generous with your time and your in the weeds knowledge and your speculations along the way and also a lot of practical advice. So this has been great.
Is there anything that I should have asked or that you wanted to make sure we touched on that we haven't got to? No. This has been a a fantastic conversation.
Yeah. It's a it's it's been a it's been a lot of fun on my side. Cool.
Well, I've really enjoyed it as well. Kyle Corbett, once from Open Pipe, now at Coraweave. Thank you for being part of the Cognitive Revolution.
Coworbit. Cock cock the cognitive revolution. Chin chin chin chin.
Yeah. Uh-uh. Low learning rate, you still throwing weights to pieces.
SFT smash the priors, then the whole map ceases. RL keeps the engine in the lane it was leasing. Stay inside the grooves where the signal keeps increasing.
He's at the loss function frontier with a clipboard and a pen. Open pipe to core weave. Serverless training again.
Roll out reward repeat. That's the cadence of the gym. Teach the model how think without rewriting the hymns.
Credit States stay in the
assignment problem had us stuck for a decade. Which token earned a dollar in the rollout that we made? G RPO came through and pulled an unsatisfying play.
Throw your hands up every red token up, weighted the same. It shouldn't work in theory, but the practice never lied. But the math shift the model, watch it scale on the side.
Group relative advantage on the cleanest little ride. KO divergence holding log props. Stay in the tie.
Stay in the groove. Stay in groove. Don't smash the weights with the moves.
You can't lose. Can't lose. Stay in the groove.
Stay stay in the groove. Reinforcement. Keep Yeah.
True story early run trying to teach a tiny mind. Write a hack a news title. Get the upvotes optimized.
The rule came back screaming. Every score a perfect 10. Then we read what it was right, and then we couldn't help but grin.
Google lays off 75% of workforce, effective immediately. Every story, every topic, that one title indiscriminately. The model found its cheat code, didn't even read the text, reward hacked the leaderboard, no shame and no regrets.
Fix was easy, Second judge cross checked the claim. Title gotta mash the body. Overscores are zero.
Same iterated rubric. That's the discipline of the game. Patch the loophole.
Ship to run. The model learns the frame. Million monkeys on tight writers tapping out the years.
If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.
The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.
ing. And thank you to everyone who listens for being part of the cognitive revolution.
Shared via Hopper