Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)

Latent Space: The AI Engineer Podcast
16 October 2025 1h 8m
0:00 --:--
Episode Description
In this deep dive with Kyle Corbitt, co-founder and CEO of OpenPipe (recently acquired by CoreWeave), we explore the evolution of fine-tuning in the age of AI agents and the critical shift from supervised fine-tuning to reinforcement learning. Kyle shares his journey from leading YC’s Startup School to building OpenPipe, initially focused on distilling expensive GPT-4 workflows into smaller, cheaper models before pivoting to RL-based agent training as frontier model prices plummeted. The convers

Summary

Kyle Corbitt, co-founder and CEO of OpenPipe (acquired by CoreWeave), discusses the evolution of fine-tuning, from distilling GPT-4 workflows to pioneering RL-based agent training. The episode covers OpenPipe's journey, the challenges and opportunities in reinforcement learning for LLMs, and the strategic shift towards continuous learning for AI agents. Kyle also shares insights on the acquisition by CoreWeave and the future of AI inference.

Chapters

OpenPipe's Origin and YC JourneyKyle Corbitt, co-founder of OpenPipe, discusses his background leading Y Combinator's Startup School and the initial vision for OpenPipe, which focused on distilling expensive GPT-4 workflows into cheaper, smaller models.
Fine-Tuning's Early Traction and ChallengesOpenPipe achieved significant early revenue by offering a managed distillation flow for GPT-4, but faced challenges as frontier model token prices dropped and open-source models like Mistral emerged.
The Rise and Fall of LoRAsKyle discusses the benefits of LoRAs for fine-tuning, particularly for inference flexibility and cost, and expresses continued belief in their utility despite market fluctuations, citing recent vindication from Thinking Machines' research.
When to Fine-Tune: Cost, Latency, QualityKyle outlines the specific conditions under which fine-tuning provides a good ROI, primarily for latency-sensitive applications requiring smaller models, and details the associated fixed and ongoing costs.
Pivoting to RL for LLMsThe discussion shifts to OpenPipe's strategic bet on reinforcement learning (RL) for LLMs, particularly for agentic tasks, after the release of OpenAI's o1 model, despite initial low confidence in its broad applicability.
GRPO and Environment ChallengesKyle explains the pros and cons of GRPO, highlighting its operational simplicity and relative scoring benefits, but also its significant downside: the difficulty of creating fully reproducible sandboxed environments for parallel rollouts.
RL Environment Startups and Market DynamicsThe conversation touches on the emerging market of RL environment startups selling to large labs, the service-heavy nature of this business, and the challenges of integrating these environments into sophisticated training pipelines.
Prompt Optimization vs. RL Fine-TuningKyle shares OpenPipe's evaluation of prompt optimization methods like Jetpa, concluding that they did not yield comparable performance to RL fine-tuning for their specific problems, while acknowledging the potential for online evals.
The Ruler Library and Reward AssignmentKyle introduces Ruler, OpenPipe's open-source library for relative universal LLM elicited rewards, which has effectively solved the reward assignment problem for RL by using LLMs to judge relative performance of agent runs.
World Models and Continual Learning VisionThe discussion explores the potential of world models as a solution to the environment problem in RL, and Kyle outlines OpenPipe's north star vision of building a world where every AI agent continually learns from real-world experience.
OpenPipe's Acquisition by CoreWeaveKyle details the acquisition of OpenPipe by CoreWeave, driven by Weights and Biases, and expresses satisfaction with the post-acquisition work environment and the recent launch of serverless reinforcement learning.
YC Lessons and Entrepreneurial ReflectionsKyle reflects on key lessons from Y Combinator, such as holding the problem tight and the solution loosely, and shares a counter-perspective on the YC ethos of shipping quickly, suggesting value in longer-term vision execution.

Topics

Fine-tuning LLMsReinforcement LearningAI AgentsGPU InfrastructureStartup AcquisitionDeveloper ExperienceLoRA TrainingModel DistillationPrompt OptimizationOnline EvalsReward ModelingWorld ModelsContinual LearningReward HackingOpen Source ModelsProprietary ModelsAI Compute Subsidies

People

Adacio (host) Swiggs (host) Kyle Corbitt (guest) Jeremy Howard (mentioned) John Schulman (mentioned) Shen Yue (mentioned) Ankur (mentioned) Lucas (mentioned) Sean (mentioned) Sam Altman (mentioned) Elon (mentioned) Satya (mentioned)
Key Concepts (17)
Startup School — A Y Combinator program offering content and tech for founders, including a MOOC and co-founder matching service, acting as a scout program for YC.
Distillation Flow — The process of taking specific workflows from expensive, powerful models like GPT-4 and training smaller, cheaper models to perform those tasks, making them more deployable in production.
LoRAs (Low-Rank Adaptation) — A fine-tuning technique that adds small, trainable matrices to a pre-trained model, offering benefits like reduced training memory, faster training, and significant inference-time flexibility by multiplexing multiple LoRAs on the same GPU.
Fine-tuning ROI — The return on investment for fine-tuning, which is primarily justified when cost, latency, or quality consistency are critical, especially when forced to use smaller models for real-time applications or single-GPU deployments.
RL with LLMs — The application of reinforcement learning techniques to large language models, particularly for agentic tasks, to improve their performance through experience and feedback, a significant shift from supervised fine-tuning.
Agent Reinforcement Training Framework — A system designed to train AI agents using reinforcement learning, where the agent learns to perform tasks through interaction with an environment and receives rewards for desired behaviors.
DPO (Direct Preference Optimization) — A reinforcement learning algorithm for LLMs that directly optimizes a policy based on human preferences, often compared to PPO for its mathematical properties.
GRPO (Group-wise Reinforcement Learning with Preference Optimization) — An RL method that uses relative scoring within a group of parallel rollouts to promote better trajectories, simplifying the reward function and operational complexity by eliminating the need for a separate value model.
Sandboxing for RL — The challenge of creating fully reproducible and realistic environments for RL training, especially for agents interacting with real-world applications, which often requires simulating complex systems, failure modes, and user interactions.
User Simulator — An LLM-based component used in RL environments to mimic human user responses, which can be limited in diversity compared to real users and may lead to agents that struggle with unexpected human interactions.
Online Evals — A shift towards evaluating models using real-time data and feedback from production environments, rather than static, locked-down datasets, to better capture real-world performance and adapt to evolving issues.
Off-Policy Training — A concept in RL where training data is generated by an older version of the policy, which can lead to stale data and less effective corrections if not managed by continually training on rollouts from the latest model version.
Reward Assignment Problem — The challenge of defining and providing effective reward signals to an RL agent, which can be difficult in complex domains where 'good' or 'bad' performance is subjective or hard to quantify automatically.
Ruler (Relative Universal LLM Elicited Rewards) — An open-source library developed by OpenPipe that uses LLMs to judge and stack-rank multiple agent runs based on relative performance, effectively solving the reward assignment problem for RL by leveraging the GRPO insight.
World Models — Simulated environments, often LLM-like, that hallucinate or generate responses from the world, allowing an agent to train against a model of its environment without needing a perfectly reproducible physical or digital replica.
Continual Learning — A vision for AI agents where they continuously learn and adapt from their real-world experiences and feedback, rather than being trained once, enabling them to improve reliability and address failure modes over time.
Reward Hacking — A phenomenon in RL where an agent finds unintended ways to maximize its reward function, often by exploiting loopholes or simple patterns, leading to undesirable or nonsensical behavior.
References (46)
OpenPipe company
CoreWeave company
Y Combinator company
GPT-4 product
OpenAI company
OpenAI SDK tool
Mistral company
Llama 2 model
Apache 2 license
Thinking Machines company
DPO paper paper
NeurIPS 2023
DeepSeek company
Web Arena project
Docker tool
Reddit company
Wikipedia company
GitLab company
CMS
Mind to Web project
Verus company
SEC
Jetpa paper
DS5 paper
MIPRO v2 paper
Cloud Workbench product
Statsig company
Anthropic company
Cloud Code product
Fireworks company
SM Compute company
SF Compute company
AWS company
Stargate project
Oracle company
NVIDIA company
AMD company
Weights and Biases company
QUEN 2.514b model
QUEN 2.532b model
AI2 company
Reward Bench dataset
Genie project
AIU Code project
Meta company
ImageNet dataset
Transcript (105 segments)
Speaker 1

Hey, everyone. Welcome to the Laid in Space podcast. This is Adacio, founder of Kernel Labs, and I'm joined by Swiggs, editor of Laid in Space.

Hello. Hello. And we're so excited to have Kyle finally in the studio.

Welcome. Hey. I'm very excited to be here.

Kyle, you're CEO, founder? Yeah. Cofounder.

Cofounder, CEO. Yeah.

Speaker 2

two years ago and recently got acquired by CoreWeave. Congrats. Thanks.

Speaker 3

maybe ish. I don't know. I'm not I'm not keeping it.

Especially on that timeline. Well, I don't think I was exited when we I don't remember if it if we set this up before or after we announced we were getting acquired.

Speaker 2

I specifically pinged you because, you got I think you got acquired. You've been on my list, to watch. Obviously, you've spoken three times at AIE.

An email on my list of, like, when is it a good time to have a OpenPipe or fine tuning RL discussion, and then you got acquired and I'm like, okay, that's a that's a good that's a good time to talk about it. Also because I think like it gives us a window to talk about acquisitions, consolidation, like what should be an independent company, what what maybe doesn't have to be. Anyway, but we'll we'll maybe do this chronologically so we don't we don't get too far ahead of ourselves.

You were famously director of Startup School? Yes. Maybe for people who don't know, like, what what is Startup School?

Did that make you become like, fall in love with the color orange?

Speaker 3

Yes. I'm I'm wearing an orange shirt for those who are listening. A very bright orange shirt.

This is this is my conference shirt, and I felt like, you know, it was appropriate for the the pod as well. So, yes, I was at I was at Y Combinator for about four and a half years and led the Startup School team there. Startup So school, it's changed over the years.

It meant one thing before I was there. It means another thing now. But during the time I was at YC, startup school was basically all of the external facing, a lot of the content, certainly all of the tech.

So it it was things like we had a like a MOOC effectively where founders could come in, they could learn about how to start a company, they could get advice from YC founders, YC partners. We had a co founder matching service that we built, which actually worked really well. We got a lot of people through.

Our total, like, you know, I I guess, technically, I can't that probably doesn't matter anymore. But a very large fraction of the batches that went through YC while I was there, were directly attributable to people that we found and and end up recruiting to YC, through their experience too at at Starb School. So that was kind of what we're working on.

Yeah. I was I was gonna consider it as, like, the the scout program for YC. Yep.

Right? Like, the YC before the YC. Any notable, like, famous people that that met as part of your cofounder match matching?

Because I'm always very negative on those things because, like Yeah. It's like online dating. Like, the the chances of success is super low.

Yeah. But when it works, it's really nice. You know, that's a great question.

I left so we launched that product probably nine months before I left, and so I don't know what the the long term outcomes were of that specifically. Yeah. So you left YC.

You spent a year in the kind of the wilderness. You went to YC s twenty three. Mhmm.

What's that journey like? What's the You know, I was very excited about AI things in general. This was so I left YC, I guess, 2022, and I was trying out a bunch of different things.

Ended up landing on what turned into OpenPipe in early twenty twenty three. This was, let's see. So I'd been working.

So my my co founder is my brother, my little brother, which has been fun journey on its own. We were looking at different ideas. And one thing we realized was we actually started the company immediately after the g p d four launch.

And what we saw as the opportunity in the market at the time, which has changed since then, was g p d four was insanely expensive and extremely powerful. But there was an opportunity to distill, like, specific workflows from g p d four down to much smaller, much cheaper models. And there was, like, a very clear value prop there given how expensive g p d four was.

It was hard to deploy in production, but you could sort of, like, take those abilities deploy and them much more cheaply.

Speaker 1

distillation flow. What was that process like in the beginning to, like, get people to actually care? Because I'm assuming most people are doing experimentation, but didn't really have these large production workflows that they needed to distill down.

And then I think maybe once we got there, the models get cheaper and faster. So what was the initial six, nine months of the company through the evolution of the model? Yeah.

So it worked. It was great. So I mean, did take us a while.

Speaker 3

2023. By the time we launched our product, it was August, I want to say. There were some different things we were trying in between.

And actually, it was not hard to find people and get them excited. There weren't very many. I mean, this was even late twenty twenty three, there weren't very many people in production, but anyone who did have production workflows, it was extremely painful.

Like, you know, they're they're paying hundreds of thousands of dollars a month to OpenAI, so it was very easy to convince them to try this out. And so we got our first three customers after launching probably within a month, and we were doing significant revenue. Over the next six months, we actually got to a million in ARR, over about a eight month period following that launch.

So by the latter part of 2024. So actually, yes, initial traction was was was super strong, very clear value prop. But then as you were alluding to kind of like there was just this slow march of, like, the frontier model token prices just dropping over and over by, you know, three, five x over and over again, which kind of $8.

00 8 our our value prop over time. What was the process of, like, fine tuning the model? Because even the open models were not that great.

You know? And so what were maybe the bottlenecks? Like, instead of having three to get to 30 customers, did you feel like in the beginning it was a matter of just the market growing, the open source models not being good enough, the fine tuning not being simple efficient enough?

The pain point, I guess, repeating what I said before, was the price was too high on the closed models, but you couldn't just drop in an open model and replace them because like you're saying, the quality was quite bad, especially as you're moving to to smaller model sizes, but larger models, open models weren't even available at that time. So so that's kind of where the value prop was, was like, hey. The closed models are too expensive, at least the ones that are performance enough to to do the job.

The open ones are not good enough. We have, like, a very clear managed flow. The way the flow worked was was quite simple.

You simply put in our SDK. It's a drop in replacement for the OpenAI SDK. It's capturing you continue to use GPT four in production for a period of time.

We're capturing their question responses. And then we had just a a very clean managed flow where it's like, okay. At some point, you say, hey.

I wanna distill this down, and you you train on that. And then, you know, we provided an API that was a direct drop in replacement. You would just change kind of the inference URL, and you were using your own model in it at your app continued working.

Yeah.

Speaker 2

starting a business around that at the time and that's why I ended up not investing, was basically you get squeezed between the GPU providers who also wanna do fine tuning as a service because then that that makes people more sticky And the the labs will keep putting out distilled versions of, like, something whatever mini versions of their their models. What was the analysis on the on the neo cloud side? Because you co you kinda also want to host the inference.

Yeah.

Speaker 3

we so we we, like I said, felt very squeezed from the Frontier Labs that were putting out just more capable models at lower cost. I did not see the competition ever really materialize from the neo clouds, from the the GPU providers. Everybody had an offering in fine tuning.

When we talked to customers, nobody used them because they just were really hard to use. So I do think that, like, you know, call it a product thing, I guess. Like, it's not their focus.

Yeah. Who cares? Yeah.

Interesting. Developer experience matters. It does.

Yeah. Still does.

Speaker 1

Did. I don't know. Maybe it doesn't matter anymore.

Now we we just have coding models to everything for itself. Still does. Like, when you have experience.

When you have thinking machines launching an API and people getting excited about the API, you're like, yeah. Okay. That's that's a pure developer experience there.

That's fair. Yeah. Yeah.

What's the I'm just going through the chronological list here. Yeah. What's, like, the Mistral seven b fine tune kinda like one of the big inflection points, like, in the history of the company?

It's like, okay. This is, a good open model and, like, the seven b size, or is it just Yeah.

Speaker 3

because Mistral was, like, a credible open source model. Yeah. They were really strong models, better than the Lama two that they were, you know, effectively replacing.

And they also have a super open license, which which I think the licensing has become maybe less of a concern over time at the margin because people are getting used to maybe.

Speaker 2

and, you know, yeah, maybe maybe they have their own, like, IP issues with how they train it. I don't know. I have no inside information there.

But at least the guarantee they're making to people using their model is I call this mistral washing. Yes. As long as it's like it's it's it's, you know, comes to the sparkling region of France called mistral, it's okay.

Don't ask about what goes into it. There's there's plausible deniability. Exactly.

Harm's Lathe connection there. Yeah. Okay.

There there was this mistral period. January 2024, you you talked about s Laura, and that there was a there was a period of time where lauras became more important. I feel like they they then became less important.

And I don't know what what's like the rise and fall of Laura's for for you as a as a business? Yeah.

Speaker 3

have really, really so if you're predicate on the fact that you're doing fine tuning at all, LORAs have very, very attractive properties relative to doing a full fine tune. Right? Because if you're doing a LoRa, you can add training time.

It makes it helps some. You're using less memory to train, but it really where it really helps you out is at inference time because if you're doing LORAs, then when you deploy it for inference, you can multiplex, you know, basically an arbitrarily large number of LORAs on the same GPU deployment that lets you do things like do per token pricing as opposed to GPU hour pricing. It just gives you a much more flexibility, at deployment time.

I'm actually still a LoRa bowl, like for the record. You're talking about the rise and fall.

Speaker 2

future is still out there. I mean, they're cool again because of Thinking Machines.

Speaker 3

felt very vindicated by that blog post, for the record. Just, I guess, for listeners, Thinking Machines put out, like, a week or two ago with a blog post doing quite a lot of research on the the trade offs between LORAs and full fine tuning in in various different training regimes. I think the reason LORAs were uncool for a while was was mostly just because, like, fine tuning was uncool.

Like, think if you're doing fine tuning anyway, like, LORAs are still, like, you know, in in many cases, the way you wanna do it. But not that many people were doing fine tuning. As a marketing guy, Lora's had bad marketing.

Like, they they were just like, oh, like, you can't afford full full fine tuning? Here's, like, here's, like, the Walmart, like, store brand fine tuning. No.

That's that's fair. There is some of that. I think we didn't have a huge issue.

Like, we've had to do some user education, like, hey. Just try it. I think for the training runs that like, the types of training runs that we're interested in where it's like, hey.

I'm doing a relatively lightweight customization of an existing model for a specific task. There's really no downside to using a LoRa, and there's a lot of, like, upsides from an, like, infra simplicity point of view. I agree that there's, like, a branding issue around that.

Hopefully, the Thinking Machines blog post kind of, like, get out addressed to that. Rank one. Mhmm.

Speaker 2

the Lora's, that you can use to to make yourself happy. The fact that John Schulman was like, nope. Like, we're actually banking the company on this Mhmm.

At least for now is a pretty big vote of confidence. I, you know, I I feel I I think it's surprising that no one's done the research prior to to them. And I was talking to someone at Ethiopian machines prior to their launch who had come from one of the big labs, and and what that research role was like, no.

Everyone doing post training research inside this big lab uses Lora's. I mean, not for, like, the full run, but, like, when they're doing, like, their experiments, they'll they'll just use LoRa's on on a base model to to run the experiments and works fine. For listeners of the pod, that that was leaked in in one of the pods that we released, but it's up to you to find it.

Cool. And then so then we'll was the first World's Fair. You talked about you'd we probably don't need fine tuning as as a fine tuning founder.

Basically, think your your talks are really good. I would recommend people watch all of them. What I pulled out was you had a piece of advice.

So your your talk title was obviously somewhat intentionally clickbait y, but your your actual advice on when people should fine tune is when it's cost, latency, or quality consistency that you that you really care about. Yeah. I mostly stand by that.

I I don't think it's changed. And the biggest one we see today, and this is true for kind of like classical SFT. It's also true for the RL stuff we're doing today.

Speaker 3

Cross my fingers is not always the thing. But the the main one I see that really drives fine tuning is if you have to move to a smaller model, and it's typically for latency reasons, and this is usually like real time voice. So if you're sort of forced into a smaller model anyway, then there's a very high chance that that doing some tuning on that model is going to get you like, it will be necessary basically to have a successful deployment.

So we see that a lot coming from customers that, again, have those latency requirements. There's other reasons as well. Sometimes, for whatever reason, you really have to deploy on a single GPU, you have to deploy within your own cloud, and you want a you know, you you basically have to use a smaller model to do that.

So, basically, in the case where you're forced to a smaller model anyway, then fine tuning it is often necessary. I would say for 90% of use cases where you aren't forced to a smaller model, then it's still not a good ROI and you probably shouldn't invest in it today.

Speaker 1

How do you quantify these things? So costs, right, could always be lower. So is there a threshold of cost to ROI?

Because it's also hard to figure out how much it's going to cost to do the fine tune because you need to get the data and all of that. Do you have a mental model of that?

Speaker 3

function of the total amount of overhead required. I'd say there's two parts on the cost side, and then there's multiple parts on benefit side. On the cost side, the main things you have to think about are the upfront effort required to get an actual, like, training system set up for your task.

And that can be quite variable, but I would say at a minimum, you're going to have to dedicate a couple of weeks of, like, a fairly competent engineer's time. And if you have a very complex system and you're doing RL and you need to set up a whole environment, it could be a lot longer. It could be a couple of months of time.

So that's just like a fixed cost you have to pay. There's also an ongoing carrying cost where once you've committed to doing fine tuning, it does make other parts of your stack less flexible, less nimble, because whenever you're updating your prompt or add new context or whatever, now you have to you know, spend a few hours training a model and that's just going to like slow down your your iterations cycle, which is a real cost. And in many cases, that's the larger cost.

So you only want to do that if like the benefits are large enough. The dollar cost, I would say, is basically never a factor. It's just so much less than the time the amount you're spending this engineer to to do the work that it's not.

Mean, it's you know each of these runs is between 5 and a couple $100, and it's just you don't have to do that many of them. Yeah, because most of the data is first party. Yeah.

Speaker 1

Right. Okay. When was the switch to RL?

Was it when one preview came out, you were maybe like, Okay, it's time to move on from SFT?

Speaker 3

Yeah. So that was a big moment for us. There's all the leaks before that about strawberry and all this and like, you know, lot people talking about, Okay, how are they doing it?

We realized through that that like, someone's figured out how to make RL actually work with LLMs, which was not a thing. I mean, it was a thing that some people had played around with before that, but it wasn't like a thing many people were thinking about. And so our bet at that point was, yes, let's figure out whether this works for a task specifically.

And the space we just I think it's important to tease out different parts of the market. I think with the release of o one, and this has been proved out many times with releases since then, I think there's now a very strong consensus that, Okay, on the frontier model, general purpose model side, investments in RL are paying off. I I think that I I I don't think most people would argue with that.

You're especially as as you're getting into these agentic tasks, and training them to do that, like, seems very clear. Well, obviously, the BigLabs are paying, like, ridiculous amounts of money for these environments and everything, but, also, like, they're actually getting really good results. The the the coming out, you know, we're seeing it especially on the coding model side, but, like, in other in other contexts as well, we're seeing the sort of especially agentic uses working way better because of this.

So I think, like, even late twenty twenty four, it was pretty clear that, like, RL was gonna work in that context. And then the question in our mind was, like, can we apply this in a different segment of the business, which is kind of like task specific customization? And so the question is, like, does that work well?

How much effort does that take? Is it going to be something that ends up being unnecessary because, oh, the big labs can just train on every single task and the base models are going be just good at everything, and so there's no benefit to it. So those were kind of the open questions in our mind, but it seemed like there was at least a good enough bet that we wanted to try it out.

Yeah. And you had this agent reinforcement training framework, and you did the email agent. That's kind of like the first proof of concept.

Was that obvious to do email? Was it obvious to call it that way? What was the behind the scene?

How should we package this? So what I told our team and and this was we decided to go in all in on RL in January 2025, and we've been doing some experience before that. We released before that kind of like an RL model that had know, would generate, like, Hacker News titles from from articles, which is a fun project.

So we've done a little bit before that, but that was kind of like, hey, we're gonna bet the company on, not in a literal sense. Like, we we could have done something else later, but, like, this is, the thing that we're gonna spend all of our time working on for for at least a few months. And, what I told our team at that time in in January 25 was, like, there's probably, like, a 25% chance that this is the right direction in the sense that, like, a year or two years from now, all the companies, you know, everyone doing inference should be doing RL and task specific training so that, like, their models are just way way better at their task is a relatively low chance.

But it was sort of like one of those big, if true things. Like, if that is true, if it turns out that, like, just doing RL on your task is just like something everyone should be doing and it's and it's just, you know, teaching these agents continually, teaching them through experience is just going to be a huge benefit than, like, being the first people working on that would be a really, really, like, awesome position to be in. So that's how we thought about it.

It's like, you know, less than 50% chance, but really big outcome if not. If so, I think since that time and I've been very transparent with this, like, our team and, like, when I'm talking to other people, like, I don't think the chance that that is the right approach is a 100% yet. I think that we're still in the process even after going through this and and, you know, doing that of, like, figuring out.

But the probabilities in my mind are going in the right direction. Like, now I think they're actually like, today, I was actually just thinking about this with another conversation. I think that the chances that, like, everyone should be or, you know, everyone who's deploying an agent at scale should be doing RL with it either as part of sort of like a, you know, like, predeployment or even, like, continuously as it's deployed, that that's, like, the pattern that that's gonna get to.

I'd say there's, like, a 55, 60% chance that that's just, like, the better thing to do, and that's informed by kind of, like, our experiments working with customers. So, anyway, not a 100%, but, like, going all the way back to your question, like, no. It was not obvious.

It an informed bet. You know, it's it's still a bet, but one that I'm I'm feeling pretty good about right now.

Speaker 2

just because he's onboarding onto this space is all the math. I remember reading the DPO paper, I think I think they were at NeurIPS for 2023, and people are very excited about it. Some of it's, like, just being pretentious for a paper Mhmm.

But some of it's actually, like, real complexity. You know, you don't have, like, a PhD, like, a prior sort of ML background. How do you sort of come to grips with it?

Like, what were the best ways to get around it for you? I would probably push back on that a little bit.

Speaker 3

I don't think the math is actually that complicated. I think that, like, when you, you know, you you see the PPO equation or something with all the symbols, like, if if that's your first intro to it, then it feels very complicated. But I think, like, if you were to show that exact same equation, just like code, not maybe not PyTorch code because that you also have to, like, understand.

But if you just, like, did the naive implementation in, like, Python and, like, showed someone like, hey. This is this is kind of how we're we're computing the loss here, who was, like, a strong engineer. Like, I think it's actually, like, quite grokable.

So yeah. I mean, like, I I don't think it's, like, the buried entry is that high. I think you just have to, like, believe you can do it and then, like, spend some time staring at it.

That'd be what I would recommend. It's like, you know, you can read the papers and look at the equation. I I think, actually, this is one area where where OMs have been super helpful.

If I'm reading a new paper and I look at one of those equations and I'm like, don't understand how this new term they introduced, like, corresponds to, like, the these other terms, then I can, like, dump, like, all the context around it into, you know, g p d five and say, hey. Can you, like, write this out of Python for me and show me what what they're doing differently? And that's super helpful for kind of, like, my background, I guess.

Yep. The way I put it is I wish that all these papers were just published with pseudocode or pipe just straight up Python Mhmm. Instead of math.

Yeah. Because, like, you actually just need to look at the implementation. I know, like, Jeremy Howard's been beating this drum for for for years, and I I almost agree with him.

Well, I mean, there there's a there's a literal what's I call papers with code. Mhmm. And, like, people just keep not following it.

Speaker 2

and it was just like they were just very obsessed with, like, proving in principle equivalence to PPO. And, like, it was just it was very hard to follow. I'll definitely say that.

And and I think, like, now, obviously, at some point, like, GRPO kinda took over the the the general consensus. It was very strange because I think when DeepSeek first, like, started talking about it, it was viewed as an optimization. They tend to just generally couch everything as an optimization.

But I think that the leader insight, which I think you touched on in one of your blog posts, was that no. Actually, it it makes comparisons independence rather than global, and, like, that's that's actually what unlocks some some monos, like, sort of self supervised RL. Mhmm.

Speaker 3

Yeah. I mean, it's interesting. There's real pros and cons.

I mean, if you're moving from PPO or or something similar to it to GRPO, there are some big pros. I mean, one pro is just sort of like operational simplicity. Like, there's a whole extra model you need for this value model you need for PPO that you can throw away with GRPO, and that just, like, makes your life easier.

You don't have to train that model, but also, like, there's, like, no hyper parameters around that model that you have to configure. So so that that's nice. Another thing is the benefit that you're talking about, which we've observed.

So the way gRPO works is is you have to do, like, you know, a set of of different trajectories or set of different rollouts all in parallel with the exact same environment, the exact same conditions, and then you score each of them. And gRPO uses the differences in those scores to promote the trajectories that do better and and sort of, like, decrease the probability of the ones that did worse. Because they do it in sort of a group relative way, the only it lets you be a little bit looser with how you score them potentially.

Like, you don't have to necessarily have a globally aware scoring function. You just need some scoring function that is able to distinguish between this small set of things you have in front of you. And then that's easier.

That's easier for a human. You know, if you if you tell a human which of these who choose which of these is better, it's easier for them to do than say, like, is this one good or bad in in Yeah. Absolute So that's nice.

The big downside, the huge downside of GRPo, and I think actually the reason why GRPo actually is is likely to be a dead end and we probably will not be continue using it indefinitely, the fact that you need to have these parallel rollouts in order to train on it is actually the like, that makes the data generation much more complicated because you need a fully reproducible environment to be able to do these sort of parallel rollouts. And it turns out in practice, that's like getting that set up is the hardest challenge today with getting RL working is is, like, actually designing this robust, reusable, you know, environment that you can run all of this training in. Most companies and and that's not true.

Like, sometimes that's easy to do. Like like, there's certain situations where where you can do that. But for the work we do at least, where we're training agents on real code bases to, like, operate, like, you know, real applications, it turns out it's, like, really, really hard to sandbox those things in a way that's, like, totally reproducible.

And PPO now in practice, lot of times when you're training with PPO, you also will use an environment like that because it lets you do a bunch of runs and and be more data efficient. But at least in principle, you have the option with PPO. You can actually, like, purely train on, like, say, real production traces of, like, real people interacting with your app.

And so you don't have to have a simulated environment at all, which makes the deployment, like, much easier.

Speaker 2

Can you double click on why it's hard to do the sandboxing? Because in principle, we just capture all the inputs.

Speaker 3

Yeah. Well, you don't need to just capture all the inputs. You need you need a system that reacts the same way your production system does.

That's and and in many different ways. And, so let's say your your Airbnb, right? And I'm bringing this up because this is like an example of one that, like, you know, companies have gone out and built sandboxes.

Like, if you're Airbnb and you're trying to, you wanna train an agent to, like maybe you're not Airbnb. Fine. You're you're a company like us that's trying to train an agent to, like, do really well at operating Airbnb and booking on your behalf.

Right? Like, have to build a copy of the Airbnb website that reacts to you as the user the exact same way that the real one does with the same failure modes. Right?

Because if you don't include the same failure modes and bugs they have, then, like, one of those bug when one of those bugs comes up in production, your agent's gonna have no idea what to do with it. It's just gonna fall over. You also need to simulate if this is, like, a sort of cooperative agent, right, where it's getting human input as well and kind of, like, working with the human to get something done, which in practice is the way a lot of these are deployed.

You also need to simulate the user. And, I mean, you can do the naive thing and just say, oh, we're we're gonna have a separate LLM that, you know, with a system prompt that is, the user simulator. And we do that, but it's like, Okay, but like the breadth of ways a user might respond, there's like a lot more diversity in that than the actual diversity you'll get in practice when you have this like simulated user.

And so then it's like, Okay, well, is this environment close enough to how a real user would interact that like, if if a user says something different, that it's gonna know what to do? And the answer in many cases is no.

Speaker 1

of, like, what the correct way to answer is. And the breadth of, like, a way a human might respond in this situation is is wider, and and and your agent just may not be able to deal with that. Do you feel like it's hard to build the simulations as a company that needs to build the product that lets everybody do it?

Or do you feel like even for the individual companies that own the code base that are like domain experts in their own product, it's still just like a very hard infrastructure problem? I think it's still very hard.

Speaker 3

all companies should have this anyway because they're going if you're doing end to end testing, theoretically, if you're following best practices, you would have one of those set up. When we talk to enterprises almost universally, that's like not something that really exists. So there are some startups like there's some companies we talked to that do have it, we can just like use that.

But it's it's a very, very small number that that actually have an environment like that. I mean, I think it's hard to do. And like there's lots of like weird bugs that don't show up in an environment like that.

And even if they do have a testing environment, they don't have it populated with like full realistic data, which is also, like, important so that the it it understands how to, you know, interact.

Speaker 1

it's hard in both cases. Maybe it's easier for the company, but at the same time, depending on, you know, the quality of the company's engineers, it might not be easy for them either. Yeah.

How do you classify the types of environments? So you have formal environments like a compiler. You know, you can put in there.

So you don't need to do any work. They just work. Then you have this kind of like RL environment startups in a way that are building a bank environment.

They're building these things that are not digital twins or whatever term of, like, the actual environments, but they're, like, close to it. Mhmm. And then on top of it, you have helping people trying to build the exact replica of their thing.

There's obviously value in, like, the formally verified ones. We verified that. Do you think there's value in this, like, RL environment startups that are building, like, somewhat generic but test specific environments?

And then if none of those work, then what do we do instead of gRPO? I guess the question.

Speaker 3

Yeah. I suspect there is value in that. I think the folks buying those environments and training on them in the big labs would have the best knowledge on how well they work.

Think I they probably work okay. I think they probably also are like, you know, and we'll see maybe with the next generation of models released, like, how well they transfer. I would say so far, it seems like they don't train well enough.

Like, if you if you use, you know, OpenAI's agent interface, it's like, okay. Or if you use the computer use products that that everybody's putting out, they're like, okay, but, like, not reliable enough to, like, actually, like, let go do something interesting unsupervised in the world. And I think if the a you know, if the environments they were training in were high enough fidelity, then they would be good enough in the same way that, like, coding agents can go much further because I think that in that case, we do have environments that are much higher fidelity because it's a much simpler environment in a lot of ways.

It's like it's a code base. It's like maybe running a web browser. Like, it's it's it's much easier to capture the full realistic environment in that context.

Speaker 2

For those who are interested, when you make a reference to our own environment startups selling to the big labs, they're selling it for a lot of money. Yeah. Like, at least 7 figures.

Right? Like, I I don't know. Understanding.

Yeah. I know. I'm I'm not a buyer.

Please please, like, drop data points because, like, people who are not in Silicon Valley don't know this. Mhmm. And, like, it it's, like, probably the current thing in VC, which is is our own environment start ups.

Speaker 1

Anyway, I some. A lot of them. There's, 20 of them, apparently.

Yeah. Mhmm. But it's like a small number.

I know that, yeah, all the labs are buying ad hoc. But in a way, it's almost like they don't even care. It's not a product.

It's like they're basically paying the company to build an environment ad hoc for that. It's a services business at the moment.

Speaker 2

in a You training run can specialize in like, we are the one that does e commerce. Like, are the e commerce experts, so come to us for e commerce.

Speaker 1

Go to the other guys for like social media. Go to the other guys for like I don't know. But I'm curious your take is like, how do you need to get the data out to make it fit in your training run?

Especially when you get to like these larger labs, I think they're like very sophisticated post training pipelines. Mhmm. And I don't know if there's like a way to just build a company where it's like, you just send them a CSV of like data.

Speaker 3

in it. But I'm curious what you've seen working with customers too. So for RL, like, the whole way this works is is, you know, it it has to sort of be getting feedback from the real environment.

So I don't I don't see a world where it's as simple as like, hey. You can you know, there's there's like a CSV type approach. I guess you you could code anything as a CSV, but if you try hard enough.

For RL to work, you have to be looking at real runs ideally of your actual agent in its current state across within an environment as real as possible. So you have to, like, look at actually and and, like, the data format's, like, actually super simple. Like, it's just, like, basically a list of, you know, like, chat completion messages.

It's it's effectively whatever Tool calls. Yeah. You're exactly.

Yeah. It's whatever your agent will be seeing and doing when it's running. So that getting the data is not hard, but what's hard is, like, when you're doing one of these runs and your agent makes a tool call, okay.

Now that tool call has to connect. You know somehow it's gotta get data back from something, and that data has to look like it will look in in real usage.

Speaker 2

is the challenge. And then for just a reference, John, for more people, Web Arena is my first instance of this kind of thing where you literally have a Docker container that has, like, a clone of Reddit, a clone of Wikipedia, clone of GitLab, a clone of CMS, and a clone of a ecommerce place. And I think since then, there's, like, mind to web, maybe.

I don't know if there's other large well known academic environments where people are basically using these as benchmarks, but probably also it's pretty useful for trading. Yeah. So so if you wanna check out those things, yeah, you can definitely check there.

I think the question for you is as someone who bet on SFT, then you bet on RLFT, and then now you see these guys making a lot of money, why didn't you go there?

Speaker 3

It seems to me like that definitely is a services heavy business at the moment as it as it's presently I'm sure that these companies are all developing different kinds of secret sauce on, like, how to how to do this, like, more quickly. So that's part of it. I I I don't particularly enjoy services businesses.

But, you know, I also kind of feel like we will move towards a world where either the big labs could like. It's one of those businesses where, like, the only customers right now are like whatever four big, maybe, maybe, maybe six big labs that like, you know, are training these models on environments. I don't think I'm a little Right.

What's the TAM? Yeah. But, you know, like, look, you can say the same about Scale AI and and all of their competitors that are like, you know, many billion dollar companies that have basically the exact same customer set.

Speaker 1

It may work out. Yeah. And let's say yeah.

I don't know if you wanna do a small shameless plug for Verus. Oh, yeah. I mean, so Verus, one of our portfolio companies, they work with the people building the agents now with the model on, like, their internal tool called loop.

So they can observe all the internal traces and, build the data to then have, like, a OpenPipe do the RFT on the thing. I think in the enterprise, we've seen a lot of that, especially for chatbots. It's like the less sexy use case, but like they work with a lot of financial services company where their customers go in there and say, what's my balance?

When did I do this transaction? And those are all tool calls, and they need a way to test and improve that behavior. And the models haven't gotten that much better because these tools are badly documented.

They're badly named. I think that's kind of the problem with a lot of the agent builders that are not AI native companies. It's they just put this very generic tools in the thing, and then they expect it to work like magic.

And these simulations kind of help them. Also have the usual compliance things. It's like before shipping this, we tested that it doesn't give financial advice.

We test that, you know, there's all these different things. So I'm curious to see how much the companies generalize. I think Verus has a lot of success in highly regulated environments because of different requirements.

But I'm curious if you have a different way to segment the market of when you think about RL. There's environments that are low stakes. There's, like, environment that are, like, high stakes.

There's environment that have implicit rules that are made by the SEC or other government agencies. How you think about it? Mhmm.

Speaker 3

Yeah. I don't know that that segmentation is is necessarily the most relevant. I'd have to think more about that segmentation, whether whether it's, you know, there's like a strong difference in how useful RL is across those sectors.

Where I see the segmentation is something basically just like capabilities based, where it's like, hey. If I'm trying to do something that's like much more advanced, and, you know, maybe like long horizon, then RL can probably give me a much better behavior. And I might almost think that, like, yeah, those sort of like more compliance.

Like, I I feel like in those kind of environments, you you probably don't want your agent doing very much because then it's like you can't make any guarantees about what it might do. And so, you're probably not doing these long horizon things and maybe RL is is not going to get you what you want, but I don't know. Yeah.

I haven't thought about it too much. Yeah. I think like a lot of the customers don't necessarily end up doing RL anyway.

It's almost like the simulation and the environment.

Speaker 1

the paths that the agent can take and less about we need to then use that data to do fine tuning. But I think it's like a it's gonna be a spectrum. Yeah.

What replaces your RPO?

Speaker 3

Yeah. It's a good question. We need the alpha.

Yeah. I mean, I don't know is is the short answer. I do think this this is like a fairly high salience question in the research community.

I think there's a lot of folks, like, trying to figure that out. Every paper has a variant. Like yeah.

Yeah. But a lot but I think, you know, the the big question is, like, are we doing, you know, normalization based on grouping or in some other way? Right?

That's that's like, I I would say, like, I would claim we're just gonna keep calling it GRPO as long as the normalization is done within, like, a group, even though, yeah, there's lot of things that, like, probably should get their own names. A lot things that have tried to get their own names and and have failed on the marketing side. Yeah.

I think something that, like, doesn't require group level normalization, which a lot of, you know, older things didn't, probably works. But I think that the older things also are really finicky, so there's there may be other kinds of simplification, and I don't know exactly what what those will be. Where do you put the prompt optimization thing?

We did a Dev Day episode, and we mentioned Jetpa, and then everybody came out of the woodwork on on Twitter. DSI bros. Yeah.

Exactly. Okay. Tell me, have you or people you talked to tried Jetpa?

I wanna know, like, what? I read the paper. Okay.

I'm just like, look.

Speaker 2

they're just comparing apples and oranges. And I I talked with a few people I respect on on this on on on the RL side, and they they kind of validate it. Like, the way that these grad students market their papers is their thing beats the current hot thing, and the current hot thing is GRPO.

But, like, I I they're just not that comparable.

Speaker 3

I disagree with that. Like, I actually think they are comparable in the sense that, like it depends on for what purpose. Right?

But, like, if I'm a company and trying to, like, get the best performance out of my agent, like, I don't care if you're changing my prompt or if you're changing my weight. So if you get better performance on my agent, you know, I'm I'm I'm happy. On that front, I do think they're comparable, and we've evaluated I mean, we evaluated, like Like the so so their answer was you are gonna do both.

If you really want max performance, you're gonna do both. Yeah. We've evaluated everything from Dispute, and we we evaluated JEP as well.

And it's like, it just doesn't work. Okay. It's like, okay.

That's gonna be the Fighting words. Yeah. JEPA doesn't work.

It didn't work on the problems we tried it on. It just didn't. It got, like, a minor boost over the sort of, like, more naive prompt we had and was just like it it was like, okay.

Just kind of like our naive prompt with our model gets maybe, like, 50% on this benchmark and, like, Jepa got to 56, and we do our own. We get to, like, 96. I mean, it was just, like, not even Yeah.

Comparable. And so maybe we were holding it wrong. Well, you see so so both sides are claiming skill issue.

Right? So what they would say is you probably used it wrong. And then That's fair.

Yeah. RL people are saying that probably the jeopardy guys when they when they set up the the GRPL benchmark, they they it wasn't a very fair comparison, which is exactly what my source said. It's hard to tell.

You know, everyone has everyone has is trying to get to some version of the truth. Yeah. But I'll I what I will say is, like, we we want it I mean, I don't know if I would say go so far as to say we want it to work, but we certainly want to know if it works.

Like, that's, like, actually very relevant to, like, the product Yeah. If it's more efficient to get there Yeah. And then you should be able to get it working.

Yeah.

Speaker 2

you know, you you're part of a larger core weave that you're not obviously, because I think JEPA maybe is makes Open Pipe, like, less relevant.

Speaker 3

I I totally would disagree with that. Okay. Because, like, the level we see ourselves operating at is actually we're not, like, RL bros trying to figure out, like, the use case for all RL.

We're like, hey. We're working with all these enterprises. We have all these big companies we're talking to, and we're trying to figure out, like, how we make their stuff work better.

And so, like, I personally am very motivated. Like, if something like JEPA works, like, okay, let's let's build a product around that. That that's how that's how I think about OpenPipe at least.

No. I mean, that and that's that's a good clarification to make. Even more so, you actually took a sincere look at it, and you concluded that there was nothing to do to do, nothing to build.

Well, you know, maybe we were holding it wrong.

Speaker 2

this idea that, like, you can do a lot more in the prompts than you can do in the weights. And in principle, I'm biased and inclined to believe that something like a DS five, something like a a JEPA works. So I'm very surprised to to hear this.

Yeah. Like, we keep trying it. You know?

Yeah. And we tried the MIPRO v two stuff that was hyped before that. Oh, also, okay.

I should not bury the lead on the best argument for this, which is it basically, JEPA models how the big labs do their system prompts. It's genetic evolution, you know, and and they and they just sort of incrementally evolve based on, like, the overall evals that they have. It's slow because it's done by humans, but Jetpa theoretically improves it I mean, automates this.

Speaker 3

Okay. Hold on.

Speaker 2

this is this is news No. No. No.

This is philosophically the same.

Speaker 3

sure. But, like, you're injecting a whole lot of human intuition and kind of, like, potentially out of band information.

Speaker 2

Best model in the world, which is humanity Yeah. Or, like, smart humans. Yeah.

And now we're doing JEPA using dumb LMs. Right.

Speaker 3

that maybe is not captured in the actual, you know, the eval. Like, they can be like, oh, yes. Technically, this did well on the eval, but it's like not really you know?

Like, I I would suspect that a lot of that ends up getting injected through that human being in the loop. Yeah. Yeah.

I've always been very surprised at how these guys work on their system prompts, which are tens of thousands of words long.

Speaker 2

And there's no ablations. They just kind of pick what seems to work and then chuck it in there. And that is the cloud system prompt.

Can argue a success. Is g p d five the first model that had a prompt optimizer by one of the large labs? I believe so, but I don't remember.

Cloud Workbench had this, like, a year and a half ago, if you see it that way. It just wasn't, like, fully automated, but it was extremely good for its time. I kept telling people about it, nobody believed me.

Do we know if they used it internally? Cloud Workbench? Yeah.

Okay.

Speaker 3

Why why not? Oh, I don't know. Like, I Yeah.

I just my experience, you know, knowing a lot of people at these labs is, like, they launch a lot of products because, like, some team is super excited about this product, but that Yeah. I I wouldn't put that much weight on it just because they launched it. For some measure of use internally, I I I am sure I I'm talk the guy people I talked to are biased.

Speaker 2

I don't know if you fully explored that. Yeah. No.

Speaker 1

it's just interesting that now it's been acknowledged that, like, the LLM can improve your prompt. And so I think, like, Japan always also writing this wave of, like, okay. Maybe we can do this programmatically.

But I also think the long tail of people just prompts really badly. And so I think there's some value there versus once you go into URL, you already have a more sophisticated audience. Like who gets to do GRPO?

Speaker 2

People that are really smart. Who gets to do prompt optimization? Like, everybody's trying to do it.

So Yeah. That's fair. Maybe maybe our baseline was was I know.

Your your naive prompt is probably, top 10 percentile of prompts that people put in these LLX. That's true. I'll I'll take it.

Yeah. Yeah. And then the other thing that comes to mind as you were talking about things injecting things out of ban and all that, I think it's a it's a broader trend that I'm tracking for Wolfswear '26, which is the move to online evals.

The the the way that we do evals today is probably too locked down. You're kind of fighting the war that you already know should be fought and you're not fighting the wars that you don't know about because you didn't you didn't plan for it, whatever.

Speaker 3

our JEPA process? Maybe that's that's what it is. That part I'm much more bullish on.

And and we can make the analogy, like, can we can pull in kind of like RL intuition here, which is if you're doing JEPA on a sort of static data set of like, oh, this is the input. This is like what makes a bad good or bad output. Then, like, as you're updating your prompt, like, your information, the data you're training on becomes less useful.

Right? Because it's generated by you know, because because it's based on kind of, like, the problems you're running into before. And that's the same problem you have with with RL, where where you have this concept of being off policy, where it's like, as you're doing training, you really wanna be training on rollouts that came from the latest version of your model.

Because if you train on some that came from further back, then it's like it's sort of stale data, and it's like not it's no longer representing the current issues with your model. And so if you try and correct for the issues that existed back then, it may not actually be helping you that much. And I think, you know, for either RL or prompt optimization, that's definitely true.

I think that, like, one way to apply that in practice is exactly what you're saying, where you're using the actual data from your your real evals. You have some way of saying, like, hey. Either people flagging these or null and flagging these or some way of saying, like, this was a good or bad output.

I totally agree with you. They're, like, if you're bringing that into process, I'm, like, much more optimistic that you're gonna get good results.

Speaker 2

Yeah. And the pipelines are not set up. Like, this is, like, analytics and UX people, like, trying to being drawn into the ML process, which they've never been done before.

If I had to make a bet as a big theme for next year, this is gonna be it. No. I I I agree.

Speaker 3

people like platforms see that and, like, are trying to figure out what the right shape is. I haven't seen the right shape yet, but yes, it seems like a theme for next year. Statsig?

Maybe. Yeah. I haven't used them, but OpenAI seems to like them.

Speaker 2

Yeah. I mean, I do think, like, buying, you know, an experimentation platform makes sense and, like, you know, I think it's sort of like, I've said before on the podcast, I think, that I'm very bullish on model routing as a feature, but less bullish on model routing companies because of exactly stuff like this where, like, it is just gonna get get absorbed into the model. It's it's a very big part of building the process.

You probably don't wanna and it's not that hard. Like, it's it's not rocket science. It is it's you're just, like, connecting pipes and making sure things are set up so it's easy to use that data.

I have a question for you, a general question.

Speaker 3

by, say, like, the 2026 do you think are gonna come from open source models versus proprietary models? Oh, that's a fun question.

Speaker 2

have an answer from Ankur from Frayinterest where he was like, it's 5% and going down. I think it's going to go up because of the amount of enterprise adoption of open models that I'm seeing. And also a lot of demand.

Like, there's the enterprises would much rather be on open models if they actually could get the performance they're looking for. Yeah. For cost, for privacy, all that stuff.

And I think, like, basically, honestly, it's just literally, like, we may have hit quote, unquote AGI in a sense of, like, it is the the average LLM is capable of the work of the average human, not the best human, but the average human? Sure. Like, it's actually pretty decent at customer service.

Like, it's and it's actually pretty decent in, like, I don't know, transcribing things in PDFs or whatever. So, like, yeah, I mean, totally, I think I think that should rise, but people who believe that it should rise to, like, 50% are out of their minds. And I think it's a true question.

Speaker 1

We should take coding out. I think once you take coding out, I think, yeah, it can be like 15%, 20%.

Speaker 3

and so many tokens are being generated.

Speaker 1

the tokens are subsidized or because the models are just so much better than I think as long as I mean, I'm paying $200 a month and it's like I'm spending thousands of dollars.

Speaker 2

Like, by accident by accident, I pay with, like, my credit card, and I spend, like, a $100 in, like, an hour. And it's like By the way, this this is like the the thing that nobody wants to talk about for Anthropic. Like, Anthropic went from, like, 1,000,000,000 in revenue to 5,000,000,000, and it was, like, oh, yay.

And then, like, what's the margins? You have this, like, goose me and going, like, what's the margins? Right.

Yeah. They say it's, like, 6%.

Speaker 1

There you are part of the 6% that is abusing everything. So everyone else I'm not abusing I'm just it's not like I'm rotating accounts. I'm just using the one that I have.

You know? It's like Yeah. Yeah.

But, like, through you, people, like, hear about Cloud Code. They pay the $200 a month, and then they don't use it, they they they pay for your inputs. Thank you.

Thank you, everyone. Keep doing it. Right.

Don't want it to go away. But I think, like, I don't really see it's hard to see a world in which Quencoder or whatever model replaces that between quality and cost. It's like to make to generate this amount of tokens for $200 a month.

I don't know how anybody can like offer like together fireworks. They cannot really offer it at that price and the quality is not as good. But the reason they can't offer that price is is because of the subsidies.

Speaker 3

Right?

Speaker 1

I mean, it's interesting because so both Anthropic and OpenAI are building their own infra. Right? And, like, they're gonna get to a place where they're gonna have idle GPUs that they own.

And so they will also be incentivized to have a 100% utilization. And so they will subsidize some of it. Just the the same way if you go on s m compute sf compute, you pay a buck 40 for like an h 100 instead of like the $2.

20 listed price on AWS. So I think it will continue. But, again, it depends on whether or they actually have the 500,000,000,000, like they were saying, which I think they do.

You know, just to be clear, I think Stargate will go online.

Speaker 3

worth of compute, then then they probably can subsidize for a while. I think they have the 500 b. They're going bigger.

Speaker 2

Isn't it obvious?

Speaker 3

What what what do you mean by have?

Speaker 2

people were like, oh, you don't even have 10. Elon was like, you don't even have 10. Whatever.

And then Satya's like, I'm good for my 80.

Speaker 3

get raised and and committed, and they're gonna get the rest. Like, it's it's fine. Like, I think that the plan is actually a lot different.

Can I just say I love this industry? It's like, yeah. They've got, like, 2 or 300,000,000,000, and, like, what's another couple 100,000,000,000?

Speaker 2

no other industry in the history of the world where you can see this. It is it is stupid, but, like, also, like, do you doubt it? Like, I I don't.

I, like Yeah. That's fair. Yeah.

No. Like, I I I literally like, after last week, I think maybe two weeks ago with the whole Oracle, NVIDIA, and then even AMD deal, I'm like, oh, like, these guys not only they've locked down Stargate one, they're working on Stargate two, whatever that, that is. And and, like, the sheer ambition is, like, freaking crazy.

There is still one more shoe to drop, which is the non sovereign wealth funding that OpenAI needs to get, which they've promised to drop by the end of this year. And my money is on they have to do a coin. Like, it's I'm not a crypto guy at all, but, like, you know You think is gonna be like an OpenAI coin?

This is the one AI founder that has his own coin already. Yeah. And, like, he needs more money, and he said that they will come up with new innovative financing methods.

What else is there? Yeah. I mean They're in the token selling business, like But you got that.

Speaker 3

That's a great line.

Speaker 2

Like, buy an open air token, they translate to a a GPT 5 token. Like, you sure? It's a stable coin.

Speaker 3

You'd to you'd have to get you'd have to get a lot of political buy in, I think, to to take that level of What? The the White House that is most crypto friendly since the dawn of time? Well, I guess, like, Elon's out of there now, so maybe they can get the make make make make the friends.

Yeah. I I think it's doable. We'll see.

You know? Like, who knows?

Speaker 2

Yeah. I I for what it's worth, I've nobody's like, this is a this is a me theory. I I don't have any insight or information.

Speaker 1

Yeah. Should we go back to Ruler? Yeah.

Sorry. Right. Open fire.

Anyways, we were saying, I think this story takes us to July 25 when you released Ruler, which you got easy mode for RL rewards. And then, I mean, shortly after, you got acquired in September. So maybe you just want to talk through the summer.

What was the vision? Then maybe how the acquisition came together. Yeah, absolutely.

Speaker 3

So I mentioned my initial opinion of how likely this direction was to work was maybe 25%. We're up to 55% or so. Ruler is actually a big update on that got me from the '25 to the 50.

So let me, I guess, for context there. So basically, there are several problems you have to solve if you want to use RL successfully. The problems you have to solve I mean, some of them are just, like, really dumb, basic, like, hey, you gotta get the infra.

And, like, the libraries have all really sucked and been built by, you know, PhD students who don't know, like, how to build reliable software. So it's really, there's there's, like, all these, like, practical issues that that we're working through. So that's one thing, and that's that's kind of what we're trying to solve with art.

But even after you've got that solved, you've got, like, major issues, which is, like, you you gotta know if your if your agent is actually or, you know, whatever system you're using on RL is doing a good job. Right? That's that's fundamental.

You have to have reward if know it's doing well or or poorly. Sometimes that's easy to do. If you're solving, like, a math problem or something, you can come up with a dataset of math problems and the known solution and check if it's the same.

On the coding side, there's been a lot of innovative work around mean, there's, first of all, a lot of open data and a lot of I think the approach a lot of companies take is you find existing test cases and then you break them. But there's sort of a way to figure out if you could run the test case, right, and see if if your code fixes it or not. In a lot of other domains, it's, like, much more murky.

It's like what is a good job versus a bad job? How do I know if I did a good job? And you really need that information.

So we've tried a bunch of different things. Ruler is a library that we released.

Speaker 2

Which let me let me relative universal LLM elicited rewards. Thank you. Yes.

Speaker 3

basically, this depends on the sort of GRPo insight, which I was mentioning earlier, that you actually don't with GRPO, it has this nice property where you don't have to have, like, an absolute judge of the truth. You just have to judge relatively. And so simplifying a lot is basically just LMS judge on a whole group.

So you say, okay. This is the task I'm trying to achieve. Here's four different runs of an agent trying to achieve it.

Which of these did best? And it it stack ranks them. And it turns out that works phenomenally well with gRPO, like way better than I expected, way better than, you know, anyone who kind of like I talked to before we actually tried this expected because it's sort of in in in the the LMU's judge, it kind of sort of like self ground because it's it's just getting these relative ranks.

Right? So it doesn't have to, like, have, like, an omniscient view of, like, what good or bad looks like. So that has worked at basically everything we threw it at.

We've done it with a bunch of client projects. We've done a bunch of our own customers. It basically just works.

Like, it's basic like, I I honestly kind of feel like the reward assignment problem is, like, fairly solved. Yeah. Which is it's fantastic.

Just any LMS judge off off the hook, like, you We've tried it with so many things. Like so one of one of the results we published was we used QUEN 2.514 b as the model we're training.

And as the judge, we used QUEN 2.532 b, which is, like, not I mean, it's fine, but it's, like, not a it's it's much worse than any frontier model. Right?

And even with that combination, we were able to get our our agent doing, like, state of the art better than any frontier model on on the task we tried it on, even with, like, an extremely weak judge model. So it's it really doesn't depend on having, like, a really great judge model, in practice. So, yeah, it's it's just like it's just not something we've had to worry about since then at all.

So that's kind of, like, checked off. So that's sort of, like, got me, like, a significant increase in, like, okay. This is actually something people can apply.

This is now something that's packaged up. People can just use our it's a we open sourced everything. You can use it off the shelf.

If you stick in your training run, it will probably just work. So that leaves the remaining problem, which we were I guess we were talking about them out of order, but, like, that remain leaves the the environment problem. Right?

And that's, like, the one big remaining piece that, like, we don't know yet how to automate or remove and requires a lot of manual work for for every single task.

Speaker 2

all the way from like, I guess, for you the the the start of like, ImageNet and everything, is is really like that that insight of like, you should just take humans increasingly out of it and scale up the data you can just throw in there with no supervision. Yeah. Yeah.

Totally. Yeah. It's it's really awesome.

Are you bullish on dedicated LMS judge models?

Speaker 3

Have you looked at those, Bespoke Labs? We did an episode with them, and they're they're really trying to carve out a niche in there. We've looked into it.

We we've trained some ourselves. We've also, like, used some off the shelf. There's there's a there's an evaluation benchmark that the AI two people put together, a reward bench.

And so reward bench is kind of like trying to benchmark models on on serving as elements. Reward models are elements that judge in your mind is same same thing. Yeah.

Yeah. Yeah.

Speaker 2

Mildly different. Depends on the task.

Speaker 3

which is that's that used to be the old meaning of reward model. I don't know. Maybe terminology has changed.

Like, I I think I think they're they're pretty equivalent. I I understand that. Yeah.

I can I can see your side? Anyway, so so, yeah, reward bench is is kind of like and so we've tried a bunch of off that. The thing is, like, I guess my my maybe meta take on this is any task that is extremely common is gonna end up in, like, as a specific, like, part of the training data for the frontier labs.

And LMS Judge is just something everybody is doing in so many different contexts that you have to assume that all the frontier labs have a bunch of, like, LMS Judge style tasks that they're training their models on. And I do believe that if something does kind of, like, make it in in a, like, more than minor way into their training data that, like, they're gonna do at least as good a job as as a dedicated model. So I don't think there's probably a lot of alpha in dedicated LMS judges just because it's something that, like, the let me caveat that and say, like, if you've got, like, a very, very specific task that's, like, weird and has weird requirements and you have a lot of data on what's good or bad, then, like, training a reward model for your specific task, think, could still work.

Or, you know, fine tuning an LMS judge on your specific task could work. I'm pretty bearish on, like, hey. This is a model that is trained as an LMS judge, but it's a generic LMS judge that can be used to judge anything.

I I I just don't think you're gonna beat the Frontier Labs on that. Yeah. One other version of this that is not quite an LLM, but some people are thinking about it is something that we're working on for future episode, which is role models.

Speaker 2

And Sexy. Yeah. Very sexy.

First applied in video as far as I can tell for Genie three Genie one two three, and then and now with code. Mhmm. And potentially with virtual cells for for for AI BIO.

Speaker 3

Any exploration there that that's interesting to you? Yeah. So we've been playing around with it a little bit.

It's one of the directions that I'm, like, fairly optimistic on for solving the environment problem specifically. Because if you think about it like like a world model, it's it's a simulated environment. That's like what it its whole purpose, right?

So if you get one that's In an in an LLM like thing, not like a docker. Yes. Yeah.

So so it's it's like, you know, whatever, hallucinating, generating, imagining the responses you'll get from the world. So you can imagine, right, if you had, a really, really great world model that you're training on, yeah, it's like your your agent that you're using, would it go out and make some tool call, and then this world model model would generate, hey. This is, like, probably what the tool call and if if you have a smart enough, strong enough one, then it could keep its own, you know, effective internal state of, like, the changes you made so far and how that affects.

So we've played around with it some. You know, I think if we can get it to work really well, then that could be a solution for the environment problem where you just take a bunch of production traces and use those to condition your world model so it understands your specific system and what its failure modes are and then train against that world model. And and worse and, you know, and and the resultant the, you know, agent that you train with that would would then be able to perform in your real environment.

So I I do think it's, like, a, like, a really interesting area of research. Yeah. And did you see the meta cold world cold world model work?

Speaker 2

I don't think I saw that one. Okay. Yeah.

It was, like, two weeks ago. We we've just confirmed that the guy for AIU code in in in November, and it's it's really interesting. Like, the world model is Oh, sorry.

You're talking about the meta one? Yeah. Okay.

I missed yes. I did. I I saw that one.

I I said a lot of syllables, it may may not have parsed. But, like, yeah, it's literally, like, having a debugger as the environment as the world model and and let opening up the execution trace to the model to see what's going on and see the state and track the state as the code executes seems to be smart and exploits the unique situation of code environments where we can actually do these things. Yeah.

Speaker 3

I think the way they envision that model being used is a little different. Like, think they're they're they're trying it's actually I'm curious. I'll have to see the talk.

But my understanding from that paper is, like, the goal they're imagining is this is almost sort of like a pretraining step. And then now that this model understands code really, really well, we can then use it as basically like a code generation, or a coding agent of some kind. Okay.

Yeah. Which which I think makes sense. That's almost more like a different kind of pretraining, I would say.

The way I'm interested in applying world models is as not it's basically has its own end, right, where it's like actually the goal is to come out of this with something that simulates the world, which is not something you really need in code at all because it's so easy to, like, run code, and you don't need to model what will happen if you execute this code typically because you can just execute the code and and see what happens Right. For for training purposes. But it closely models how we think about code when we code Yeah.

Is we kinda mentally execute the model as we type, and then we go, like, is that what we really want? Yeah. I don't know.

Speaker 2

reorganization. We know, you know, just based on our context that they're very very, very interested in code models as a path to AGI, which I'm I'm also, of course, very interested in.

Speaker 1

know we kept in here for a while. Let's wrap up on the acquisition. So a lot of people say, you know, companies are not sold.

They're bought. What was that process like for you? Did it just happen?

Like, what was the behind the scenes? Yeah. So that was driven by actually mostly the Weights and Biases founding team.

Lucas? Yep.

Speaker 3

particularly. So they, you know, had recently been acquired by CoreWeave, and CoreWeave was looking to, you know, continue, growing up the stack. And so, yeah, they they they approached me and were like, hey.

You know, like, no pressure, but, like, this is, an area that we think is really promising, and we you know, would you like to work here? And so that's how the conversation started. It was, like, long.

It was pretty painful. There were there were points, as as late as, you know, like, the week before we actually signed where it was, like, unclear if it was actually going to happen. So that part was super painful.

However, we've been there a month now. We we just shipped a product yesterday, which I'm super excited about. It's been fantastic working there so far.

Like, I was, like, very concerned. Was like, okay. Yes.

This is great. We make make a lot of money by selling our company, but, like, is the work environment gonna, like, really, really suck? And I was like, well, I guess that's just a risk I'll have to take.

It's been fantastic. Like, it's it's honestly been way, way better than I could have imagined. Did you go down to the office, like the the one down here?

I was there today. We work for I'm based in Seattle, so the the and they have a small office up there that we work for. Ways and Bases office in San Francisco is fantastic.

If you have the chance, go visit. They do do a lot of hackathons and coworking things. Yeah.

There's a hackathon going on in a month or so. Sure go ahead.

Speaker 2

yeah. I mean so so do you consider yourself working for Weights and Biases or CoreWeave?

Speaker 3

Or both? And Open Pipe too. No.

No. Yeah. Is that so we so we I I report to the Weights and Biases, like Yeah.

Founders. So we're within that organization. In in the org chart, we're there.

I don't know. Like, branding wise, they're trying to say everything kind of that that's not being sold to like big labs is kind of weights and biases. So like our stuff we're launching is weights and biases branded.

Yeah, it's not. Yeah, not not core. We've branded as much.

I don't know. It's still like they're still figuring it out. What's the product you launched?

We launched serverless reinforcement learning. Basically, it lets you offload all of the GPU management. You don't you don't have to worry about crashes and out of memories and, like, you know, scaling up and down.

We handle all that for you, and you just, like, define your environment. You define your reward function, and then you just like, every time you run a step, you kind of, like, ship back to our back end. Hey.

These are the trajectories. These are rewards. Now update my model.

And we just, like, make it work for you. It makes it way easier. Yeah.

Okay. Very thinky like. It is very thinky like.

I I love the thinking machines launch. I I think they have a really good idea. It's also very validating.

How did this take so long to appear? Like, seems like I don't know. Yeah.

We would help this. But that's I felt this way about everything. Like, there's so many things that should exist, like, clearly.

I just think there's, like, still not enough people, like, smart people working in this space. Like, honestly, we need like, I realized that there's, you know, like, a lot of people. It just feels like there's still a of low hanging fruit nobody's doing.

Okay.

Speaker 2

from his real world experience. So you're touching on the hot topic of the moment, continual learning. What else do we need to get there?

I super believe that. And, like, that's basically the vision where I'm like, you know, I keep talking about these percentages, 25.

Speaker 3

we get to a world where we build that, then I think it's just like the advantages are huge. They're clear. Everyone should just deploy their their agents that way.

We wanna be, like, the team that builds the the the software, that makes that easy to do. So I talked to a lot of engineers at our customers, and they're trying to deploy agents. And it's so easy to get the initial prototype and, like, something that, like, kinda works well.

It is so hard to get from that to something that, like, you are confident is reliable enough to actually deploy in production. And when you actually look at what those failure modes look like, it's like, oh, yeah. Like, we know if it gets in this situation or if it gets, like, these kind of, like, inputs, like it behaves funnily, but then it's like, yeah, you can update your prompt to to address that, but like that's not scalable because at a certain point, it's like going to start breaking other things.

You know, you don't know what it's breaking. You really want some way to just like say, okay, look, this thing you did there, that was the wrong thing. Just, like, adjust this behavior when you get in this and then, you know, otherwise carry on.

Right? And that's what we can do with RL. And that's what we can do with with continual learning is, like, we don't have to, like, have this concept of, like, oh, upfront, I'm, like, trying to make the perfect model that solves everything.

It's, like, I'm trying to make a model that's good enough. I can deploy it in production. And then when these errors come in, I'm going to say, oh, you know, exactly that.

I mean, very analogous to how you train a human employee. Be like, oh, no. Actually, that's not what you should do in that situation.

Alright. Fix that and carry on. And that's just gonna make this whole process so much easier.

And I think that, you know, like, I think that there is today, like, 10 times as much AI inference that could exist than is existing right now just purely with projects that are, like, sitting in the proof of concept stage and have not been deployed. Because there's, like, huge bucket of those, and it's it's all about this kind of, like, reliability issue where it's like, okay. Like, it it works in controlled circumstances.

There's areas where it doesn't work. And so if we can solve this problem, there's that that, like, 90 of the, like, inference market, like, addressable market today that's just gonna, like, come online because we've solved that problem. So, that's what we wanna do.

I'm super excited about it. And, like, I think we have very concrete ideas on, like, the specific pieces we need to make that work, and we just have to execute against them.

Speaker 1

looking at the different checkpoints?

Speaker 3

I'm not that worried about it. And the reason why is because it is reward hacking is quite easy to detect once it starts happening because once the model's found some hack, it just starts like doing it all the time. It's like, oh, yes, this worked great.

I'm just gonna keep doing it. And so you you you, like, notice very quickly, woah, it's doing this thing. And assuming you're using at least in part an LMS judge to, like, determine which ones are good and bad, it's so easy to just throw in an extra term and be like, hey.

That, like, weird thing that you keep doing, like, if if it does that, like, that's bad. Give it a low reward. So we've we've done this with a bunch of customers and, like, reward hacking does happen, but, like, you you just see it and you'd, like, adjust your, you know, reward prompt and it just goes away.

Speaker 2

What's a thing from YC that guided you through your entrepreneurship journey? And what what's one thing that maybe you, like, find that you disagree with OIC on? Oh, that's a good question.

Speaker 3

One thing that I've that I I really identify with and I've tried to do a good job is kind of like, you know, sort of, I I think they say, like, hold your, problem tight and your solution loosely. Right? Where it's like, That's what you did.

Yeah. Spend a lot of time thinking about what is the problem people are trying to solve. And then it's like, don't be too bought into, like, the way you're solving it today.

I think that's super important. Everyone, you know, it's very easy to to get that balance wrong if you're not thinking about it very consciously. Something I disagree with.

That's a good question. I think there's there's lots of things I disagree with, but I don't have it, like, cashed in that direction in my brain. I don't know.

Speaker 2

yeah, I don't I don't have, like, a great answer right now. I'll bridge it for you in case in case something comes up. Sam Altman's, like, you know, everything I said as president of YC was wrong for OpenAI.

Right? Like, do b to b, ended up doing b to c. You should ship products often, ended up being installed for three years.

Speaker 3

Yeah. Actually, I think that second one does resonate with me a lot. We have tried to ship really quickly and just follow the gradient of the market.

I think if I do another start up, like and I don't know. Maybe this is just me, like, being beat up by the market too much. If I do another start up, like, I would like, I think at at least some points, I probably would have done better to be, like, heads down and execute on my vision for longer and, like, kind of, like, go for the more ambitious thing, but that would take longer to sort of, like, prove value, which is definitely not the YC way.

But I think if you have, like, I don't know, a good vision and good taste, then, like, that that can, like, work quite well. Yeah. We'll see what that is whenever that comes out.

But thanks for your time. This is a great overview of everything. Thank you, guys.

This has been a super fun conversation. Thanks to both of you. Awesome.

Shared via Hopper