Ashvin Nair, Head of Model Development at Cursor and former OpenAI researcher, discusses his transition from robotics to LLMs, including his work on OpenAI's o1/o3 reasoning models. He reflects on the surprising lack of real-world impact from achieving benchmarks like IOI Gold, the limitations of RL generalization, and Cursor's strategy of co-designing products and models to automate software engineering.
Okay. We're here at NeurIPS. We're recording a special Lanespace coverage of the the folks at NeurIPS, and we're here with Ashwin from Cursa.
Welcome. Hi. Yeah.
Thanks for having me. So I guess the like, Ashwin from Cursa is, like, a new identity. I didn't even know if I should say that cause you only joined Cursor for three months.
Mhmm. Before that, your opening eye work in 01/2003. Mhmm.
Before that, Berkeley Mhmm. PhD in RL. Just but focus on robotics.
Robotics. Yeah. Is it weird searching for robotics to language models?
Okay. This is kind of interesting because a lot of people have been kind of doing this Are you opening up? Yeah.
Got a robot.
actually was, at OpenAI in 2017 also, working on robotics. Yeah. I was interning, like, right before my PhD where I worked on robotics there.
2017. Is that, like, Javan and that's Japan after that career? He he was all famously opening as first intern.
Oh, really? Okay. Then he might have been before, but, yeah, there was, like, 15 interns.
It was a very different company. It was just, like, robotics, DOTA, and, like, 15 interns that summer all having, like, pretty exciting individual projects. Like yeah.
That set of interns, if you look at where they are now, it's kinda cool. Yeah. But yeah.
Anyone from that class that, like, you would shout out? Like, there's just, like, a lot of cool papers that came out, like, Laro Pinto, now is at, NYU. Yeah.
In Fateman. The, the person who leads, reasoning at XAI, I forgot his name. Eric.
Woah, Hila. No. Eric.
But yeah. I forgot his name, but he he worked on, like, KFAQ and stuff. Think yeah.
The vision dude? Greg? Not Greg.
But yeah. But, yeah, was it was, like, an exciting time to be there. But, yeah.
I think I think robotics is a pretty good fit for LMs because, like, the switch ends up being pretty like, you know, you can do similar things. Like, you wanna look at a lot of data. It's, like, kind of hard to get, like, stuff get hard to get stuff working in robotics world.
I think, you know, it kind of builds, like, very gritty people who, like, look at data a lot, that kind of thing. So, yeah, for whatever reason, I think, like, that transfer is, like yeah. It's happening a lot, and I think it makes a lot of sense.
One of my Neribs highlights so far, I had dinner with it's like a small group dinner with Lex Freeman yes yesterday.
And Lex used to be in robotics. Mhmm. And he was like, my assessment of robotics people robotics people are the best to talk to at a NeurIPS Mhmm.
Because they're most grounded, he says. Because they don't have a choice. They work with the real world.
So Yeah. Looking data, and then they do they're hard. The most unhinged, the most detached from reality are, like, the the simulation people.
I see. Yeah. Yeah.
Yeah. Yeah. I think I agree.
Yeah.
Yeah. And I kind of I actually did a little bit bit of both during my PhD. Like, I work in kinda, like, you know, just, like, prototype ideas in sim and then get them working on, real world robotics.
And, yeah, I mean, probably, like, robotics is where you kinda feel AGI the least. Mhmm. Right?
Because it's just so far away from working. Now I think, like, I think over the last year, maybe, there's been demos that have been super interesting from, like, physical intelligence and, like, Sunday and stuff that yeah. I'm I'm starting to be like, okay.
Like, this kind of You're saying about Sunday robots themselves? I haven't seen them. Apparently, they've been doing demos, and then I'm pretty keen on seeing them.
Yeah. I've I've seen the physical intelligence ones live. And, yeah, it's pretty impressive, like, just on, like, in, like, someone's living room Folding.
Folding laundry and stuff. Like, you know, can just, like, toss it in there. Hard work.
It's like, you must be materialized.
Yeah. Yeah. Yeah.
Yeah. Okay. And and and, Nathan, last thing on robotics, and we can kind of pivot to o one zero three.
Just Omei has, like is is, like, restarting a robotics team. Mhmm.
serious? Is that I actually know very little about it, because, yeah, I was in, like, a pretty different part of the org. So, yeah.
I mean, I I think it's serious. I think there's a ton of excitement around robotics right now. I'm actually kind of curious what drives it, because I don't think I fully understand.
You know, like, like, there's been, like, crazy raises and stuff recently, right, for robotics companies. I guess my own view on it and so when I, left robotics in 2022, I thought I would actually come back to robotics. But I think my view on it now is that it feels like LLM agents are gonna be, like, a trillion dollar market before robotics is maybe even, like, a $10,000,000,000 market.
And this is just because so, I mean, you know, LM agents already create value out in the world. Yeah. Robotics, it's, like, kinda hard to make the case that, like, you know, kind of AI robotics, like, does anything that useful yet.
And then once it does something useful, then you have to make the unit economics work out.
like, these kind of things. So I think it's kind of hard. I would say the market is kind of efficient in that the software LM companies are raising tens of billions Yeah.
And then the robotics companies are raising hundreds of millions.
So I think this very recently, it's been, like, single digit billions. Oh, really? Okay.
So I think I think that's I think that's, like, the maybe surprising thing to me is that, like, it feels like It's ahead of where it's actually at. Yeah. Like, I would say that robotics is in kind of, like, the GPT one to GPT two area right now.
Okay. And I haven't worked on robotics closely. What task would qualify as, like, oh, that's that's the inflection?
It's a little bit like you know when you see it. I thought the, like, Sunday demo as well kinda cool. Like, maybe it's, like, starting to get there where and the details matter a lot where it it's kind of, like, it it can't be it has to be in, like, a new scenario, like, in one that you haven't seen before and maybe on, like, other So it's general I OOD.
Yeah. Exactly. And I think that was kind of what GPT two was too.
Right? It's kind of like you start to see hints of, like, cool generalization. But Yeah.
Yeah. But, like, and I think that's fine. Like, you know, it doesn't have to, like, work out of the box.
But yeah. I think at this point, especially, it still feels like in robotics, you're not exactly investing in a technology probably. You're just investing in a team.
Yeah. Yeah. I'm not in this space whatsoever, but, like, that's kind of my impression.
It's actually nice when you're not in there because you're like, know as much as, like, most basically everyone else. Yeah. Yeah.
So we're just gonna speculate. Exactly. In my people, there's a robotics team at OpenAI.
Yeah. So coming back to language models, did you join 04/2001 or were you like So I joined right before Tatcha BT in, like, I think September '22. Yeah.
So yeah. Actually, yeah. I was, like, pretty burnt out from my PhD and I was like, okay.
I'm gonna go to this, like, chill research lab and and then, like yeah. Yeah. Like, ChatGePity happens and, like, you know, everything kind of blew up and, like, a lot of stuff got kind of, like, refocused.
But what I guess, what so OpenAI obviously, ChatGePity surprised OpenAI. Mhmm. What did they tell you they were looking for you to do?
And then obviously, it changed. But Yeah. I mean, so I joined on the Codegen team.
The Codegen? Yeah. By you Codegs.
Exactly. Yeah. It was like the team that shipped Codegs, but by the time we were working on, like by the time I joined, we were kinda more so working on the model doing tool use, these kind of things.
Yeah. And so, like, very related to the chat like, we're kinda like a sister team to the team that made It's like Shinichi. Yeah.
Yeah. Exactly. So yeah.
So we're just kinda working on making the models, like, smarter, like, kind of programming competitions, yeah, how to do like SAT for that, that kind of stuff. The word and IOI gold have felt reachable in that title? Oh, yeah.
No. Like crazy. Like I think I think and this is something I've, like, like, repeat to people again and again.
And these days, like, if you told me that we could have gotten I OI Gold then, I would have just assumed that we could all just go on vacation, like, you know, it's all over. Like, AI is solved. Like, no point in working anymore.
We got it. Yeah. It feels like nothing's Nothing that much has changed.
Right? Like, life is still the same. Yeah.
Yeah. So I think that's, like, super interesting. Yeah.
I I I don't have a great way to explain it, but I think that's that's actually, like, what I spend a lot of time thinking about is, like, you know, why is that the case? Yeah. Because yeah.
Mean, you kinda see this again and again in AI, right, with, solving chess and then, like, it doesn't really matter and solving Go and yeah. So you keep seeing it, but, I think it, like, surprises you every single time. Yeah.
I think maybe I think one is we keep moving the the goalposts.
Yeah. We're very good at that. And then two is I think actually our just our definitions of what constitutes AGI is bad.
Mhmm. And we don't actually mean what we say when we say, oh, when we have achieved this, then we have AGI. Yeah.
So, like, clearly, when we have achieved our goal with a language model, we have AGI. It's wrong. Yeah.
And I think I think shifting the goalpost to some extent is correct. Like, we keep good hearting whatever goalpost we have. Yeah.
And I think it's kinda hard to, like To be good hearted is, like, too negative. It's like, I will cheat to do what I Mhmm. What you asked me to do.
Mhmm. But I don't think it was cheating. It was just scaling test and compute.
At a meta level, I think the community not cheating, but, like, makes a lot of, like, implicit decisions to go after the, you know, evals and benchmarks that matter the most. And so So Subha is very fine for sure. Yeah.
Exactly. But yeah. I like Yeah.
Hopefully, not not that good hearted. Well well, but but I like, it it kind of clearly is to some extent. Right?
Because, like, you know, most programmers in the world cannot do I OI at any decent level, but, like, we're still struggling to, like, automate most programming jobs or, like, you know, there's a lot left to do. So it's like it's like like, we're, like, language models are here, like junior, senior dev, and then suddenly for IOI, you're, like, expecting Exactly. And, like, there's something suspicious about that.
Yeah. Okay. I kinda saw this at a meta level also with RO research.
So, yeah, I I did my PhD with Sergey Levin at Berkeley from, like, 2017 to '22. And that era of RL research was, like, super interesting because RL was, like, super hyped. Right?
Like, starting from about DQN in, like, 2015. And a lot of the methods that people were really excited about is, like, you know, off policy learning, like, value functions, like, these kind of things. And somehow, that that stuff hasn't really panned out, I would say.
And it's not exactly clear why, but in the academic literature, we thought we were making a ton of progress. And I think in retrospect, I had to say that we probably kind of overfit to the benchmarks pretty heavily. And, you know, how I see this in retrospect is that we gave ourselves a lot of, like, new knobs to tune and then implicitly kind of tuned those to fit the benchmarks.
Everyone knew that we're doing that at some level, but I think it's hard to appreciate, like, that it's not just happening for a single paper at kind of, like, a meta level for the whole community that's happening too. Yeah. And I think the result is that, like I I don't know.
Like, a lot of the RL research that came out of that era, I don't think is, like, that used, you know? And I think it's kind of for a similar reason that basically you were kind of like benchmark maxing. I will full out say there was RL Winter.
Right?
premise at the time basically gave up. Mm-mm. Some of them died, some of them pivoted, whatever.
Yeah. Yeah. Yeah.
I I think, in so because I was, like, in in academia, there was still quite a lot of excitement over it. But, yeah, it's it still felt quite academic and yeah. I I, in that era, was a little bit frustrated because, I felt like, you know, one of the pitfalls of academia is that, it doesn't really reward, like, simple ideas that work and instead kind of tends to reward, like, kind of math theory ideas.
Those math theory ideas also give you these, like, kind of implicit knobs to tune that allow you to, like, overfit. Well, you know, the the things that actually work tend to be kind of simple ones that have less knobs and just generalized to, like, many things with a lot of There's just, like, less secret sauce to it Exactly. Apart from just throw a lot of compute in.
Exactly. Exactly. But that those are the things that tend to like It's like not intellectually interesting.
Yeah. Exactly. And yeah.
From academic point of view, it's like, oh, like, why am I sitting in school? Like like yeah. Think I think for a lot of people who do PhDs, they're kind of wired in a way they they want to, like, think about interesting new stuff.
Yeah. And, yeah, like, you know, the the scaling era kind of like, you know, probably stuck to that. Scaling era.
Is scaling era over since we're Oh. Fired up. Well, I I think I've just been, like, patient to that from, Ace Huskyver's interview.
I don't think it's over, but there's definitely something interesting happening. I like the the thing I was saying about, like, IOI and IMO. I think we'll still continue more or less on the same track.
Like, clearly, you know, like, these labs are, like, releasing their new pretrained models, and they're, like, still doing, like like, much better than before. So I think I think scaling is still happening, but I think it's happening in a different way or, you know, it's, like, worth, like, seriously interrogating why is it that we're not just, like, automating all jobs right now. I think my view of it is something like RL, the way it's applied to LLMs right now, is kind of a weird funny tool where it doesn't really generalize beyond the training distribution that much.
It generalizes to some extent and generalizes in interesting ways, but it's like very picky. Right? Like, it it kinda it could it can kill the training distribution, like, completely.
It can be, like, best in the world at it, with, like, not that much effort really. But, yeah, it doesn't really generalize. So I think what we had to do is bring the world of economically useful tasks in distribution for RL Mhmm.
If if we commit to using RL as a tool. And, you know, it might be the case that maybe there's some, like, cool continual learning thing or something that, like, shifts a paradigm next year or something like that. Mhmm.
But, it really feels like if RL is a tool, then, yeah, a big thing that needs to happen is, like it's not it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, see it, and then you use the RL on top of that. Yeah.
Have you seen GDP Vowel? Yeah. I've seen it.
Yeah. Yeah. Is that basically what you're envisioning?
Yeah. I haven't I haven't looked at GDP Vowel closely.
Actually, I haven't seen exactly, Like like So roughly yeah. To recap, it's a 128 tasks across, like, any white collar job that takes, like, more than 5% of GDP. Mhmm.
Right? And they they basically created all the context Mhmm. To eval on it and evaluated every model.
Yeah. Famously, OpenAI OpenAI's eval's to you. Whoever runs that one always finds that Anthropix are the best for the Yeah.
Yeah. It's yeah. Props to them for We have published it.
Yeah. Yeah. It's an actual science.
I think it's good. But but like I think like Yeah. In a in a sense of like generalizing beyond coding competitions to economically useful tax Yeah.
That is it. I I I think that is the what is more important for g g six? Yeah.
What I'd like to do is kinda, like, I haven't I just haven't read the, like, GDP eval traces close closely.
It's not clear to me that, you know, like, what, like, what is the job of an accountant entail, and, what kind of context needs to be in the product Yeah. So so you can actually do it. PDFs.
I see. I see. Like, so they they try to go as close to source documents as possible.
I see. I see. Yeah.
Yeah. So, yeah, I think, like, roughly operating in this kind of thing is what I envision. Yeah.
oh, let me clean up this data for you Exactly. To make it easy for the LLM to process. No.
Yeah. PDF in an agent, go. Yeah.
I think that's what, like, roughly the right shape of the thing.
you know, operationalized is that you'd want to, like, co design the product and the model so that, like, the product for whatever it is, like I mean, coding is kind of maybe the easiest first step because most of the context that you care about is just your code base and, like, being able to run stuff in the terminal and that kind of stuff. And still, like, we're not that close really to automating it necessarily. But, know, for for, like, all the other jobs, the context is, like, insane.
Right? It's, like, all the conversations you've had with your coworkers, like, Slack messages, you know, like, for my, so at OpenAI, I was working on kinda, hyperparameter scaling research, and I actually wrote not like code. Grid search or, neural architecture search?
No. More like, understanding how diff like, science of deep learning in, 2020 where it's like, oh, you had to, like, initialize the layers in a particular way to get get good scaling of. Kind of the analog for that for Aural.
K. The thing is, like, I didn't write a ton of code. So the LM, you know, it's like writing code is not the bottleneck, but it's more like, you know, over the course of a year, I, like, run sweeps, look at, like, the interaction between different hyperparameters, and kind of build up that knowledge for, like, a year of, like, just different graphs.
Mhmm. And to do my job, the model would also need all those things in context, you know, to, like, successfully, like, you know, kind of automate my job. And you'd kind of want a product that allows you to, like, bring all that context in.
Mhmm. Did you have to build it for yourself or is there an existing one? Oh, no.
I mean, like, as no. I mean, I like, I you know, those graphs are just sitting in my head. Yeah.
Right? So I think it's it would be pretty hard to, like, go automate that job. But I think what you need to do is build a product that kind of yeah.
like, to teach the models to, like, use that that context. Yeah. Another conversation that I think has really come to a phase this year is kind of the the depth of one model fits all.
Mhmm. Mhmm. I feel like the point of the g in the AGI is, like, one model fits all.
Mhmm. Think OpenAI has, like, clearly abandoned that this year. Oh, what did you say?
Fiji Simo writing a blog post of the title that we are no longer doing one model fits all. Oh, okay. Interesting.
Okay. And I think Mark Chen or one of the other senior people that are not in SAM also saying this in the in the podcast. And so so basically, like, the idea was you you started with Codex, someone else was doing Insta of GBT, then we launched GBT $4.
04 o, I guess o one. Mhmm. And o one was kind of a supposed to be like a reasoning one model fits all.
Mhmm. And then we merged the four o and o one zero three line into five. Mhmm.
And now we're splitting it out into five and five codex again.
Well, OPI is very guilty. I mean, you know, I I don't think you should interpret those as like scientific facts about the universe. It's just more like OpenEye has a tendency to ship the orb chart, basically.
Yeah. Right? The world has the tendency.
Yeah. Exactly. So I think a lot of it is related to that.
But, yeah, I see what you mean by, like yeah. Actually, I do I do wonder if yeah. Like, the current reasoning part the current reasoning paradigm is just kind of fitting itself to this kind of peaky in certain areas thing.
Right? I don't think it's so much a matter of, like, model capacity though. It's just more of a another kind of organizational thing that, like, if you care really, like, a lot about coding, you probably don't have the data to do all the other stuff.
I don't think it's so much a matter of, like if if you had all the data, probably, you would benefit from just, like, training on all of it, and you'll get some generalization between these. But it's hard to find, like, one organization that cares about all these at once. Yeah.
Yeah.
Before I double click on just, like, the o series in in OpenAI Mhmm. I do like to ask OpenAI people who are there, do you have a favorite blip story? Yeah.
The blip was crazy for me.
yeah.
it was, like, Thanksgiving. Like, everyone remembers where they were, what they were wearing. Yeah.
Yeah. Exactly.
I was at Thanksgiving with two OpenAI friends, actually. And then one of them on, like, Friday afternoon is like, oh, like, Sam Altman just got fired. We were just, like, co working together.
It's like I'm like, what? Oh, good joke. And then, yeah, no, it was crazy.
And then, yeah, it was it was just, like, a crazy weekend of just, like, ups and downs. Like, you know, we thought yeah. Like, like You signed you signed a letter?
Yeah. I did. It's like 95% of people signed it.
Yeah. Yeah. Yeah.
Yeah.
Well, I think maybe I had a slightly more conflict like, I actually do think that governance feels really important to me Yeah. Because it does feel like no matter if we hit AGI in, like, two years or ten or whatever, it's not clear that we have a good structure for the governance of it. K.
And so it it is a question that I think we, like, probably should spend more time on. And I was, like, during that period just pretty willing to be like, you know what? Like, let's forget about the, like, equity and stuff.
Like, you know, I think it's, like, good and healthy to have a conversation about, like, how exactly the government should work. Okay. You care about this.
Yeah? Uh-huh.
So now the open end nonprofit has this, like, secret shadow board of members that determine when where we've reached AGI.
Yeah. Yeah. Is that better?
Yeah. Don't I don't have a, like like, maybe I I would say I don't have an answer. Like, you know, like, it's just it's not Above my pay grade, but like Yeah.
And even even even back then, was kind of like, well I don't care. I I do care quite a lot. When the blip happened, one of my reactions was like, well, you know, this nonprofit board stuff, like, actually, if it takes such somewhat, like, surprising, maybe erratic actions, like, maybe you'd rather just have, like, you know, a thing like the Microsoft board, which is kind of like, you know, like, probably, like, all the pensions of the world.
But why it's like true, like, serious people, but also, like, you know, the stakeholders are kind of like the whole world because everyone's kind of, you know, through their pensions or something invested in it. Like, maybe that is a bit more of a democratic way to run things than having, like, seven people, run it. But, yeah, I don't I don't really know.
It feels like we haven't solved governance, like, at all, though. Right?
kind of good outcomes for society, maybe. Yep. So about, like, the the transitioning into reasoning.
Right? Mhmm. You shocked me by when by mentioning that the reasoning team is 300 people?
I it's If they were, you drove by It's kind of like, you know, now that like, you know, when o three was, like, of, shipped as a product, like, I think it just, like, kind of gets, like, larger and larger how many people worked on it. So Yeah. Yeah.
I think I, like, like, lost track of the numbers, but, yeah, like, a lot of people contribute to the different aspects of, safety and whatever. Evaluate. So, like Original o one, like, I saw the video.
Was, a dozen people. You know, like Yeah. Well, even then, like, if you look at all the contributors, it was probably more like 50 to a 100 people.
Okay.
Yeah. So so, I mean, like, let's let's tell that story from your point of view, figuring out what does RL mean there. Mhmm.
I guess, was this a branch of any other prior work that you wanna credit? Yeah.
yeah. Like, setting the scene, guess, you know, in, like, 2023, people are kinda talking about, oh, like, is UDS scaling scaling laws dead, this kind of stuff. Every year.
Every year. It's Yeah. Yeah.
But especially I think especially that year, it it felt pretty, like, serious, you know? Yeah. I think in general, OpenAI is really good about, like, having conviction in something and just, like, really, like, from first principles, like, going after it.
And I think, like, the people who are kind of most responsible for that is probably, like like, Ilya Setskiver and Jakob Pachacchi. I think even, like, like, Dota was kind of more or less the same template in some ways. Right?
And that was, 2017. And so a lot of the people there have kind of this, like, AGI of their bones kind of point of view, and they've basically been convinced that, like, RL would be the way to get there. So I think for a long long time, people have been convinced that something like that should work.
And it's just that it started to work once the Like, human training got good enough. Okay. Yeah.
I think human feedback is kind of like a a bit of like a side branch because Yeah. You can't really pour that much compute into it. Right?
It's like you you take the model and you, like, elicit it to be a little bit better in terms of personality. But, like, the people there are really convinced that, like, at some point, you know, it's not about copying the Internet. Like, you can go, yeah, do RL and, you like, know, that's, like, the path to, like, getting much better intelligence.
So I think it it there's kinda, like, a long line of kind of, like, returning to RL in for in, like, different ways. And then it's just that around, like, yeah, 2023 is when it started, like, really clicking. And it's kind of interesting because even, you know, it's it's not like those initial models performed, like, way better than the existing models because they're, like, smaller scale.
But people were very good at being like, oh, like, this is kind of interesting. Like, you know, the the reasoning trace that you see here is kind of not something that you've really seen be so accurate, in other models like this like this one. Kinda similar to how think a lot of people didn't really think of GPT or GPT two as something that was, like, super compelling probably.
I know that I personally didn't like GPT two that much of GPT two. Was like, okay. Whatever.
And then I think and then GPT three happened, I'm like, woah. Like, I feel a lot of FOMO sitting in my PhD. It's kinda that where I think it takes a bit of, like, first principles conviction to, like, yeah, decide that, like, oh, this this thing, like, there's something here, and we should really go scale it up.
And OpenEye is really good about once you decide that something is good, then you just, like, scale it up all way. Yeah. Is was there an internal prototype pre o one Mhmm.
That was like, okay, this is the thing. We'll fund it to scale it up. Right?
Like, there there usually is. Yeah. Yeah.
Exactly. Yeah. What was the thing?
What was the demo that, like, really sort of sold Just this, like, you know, like, like, running RL on even, like, a pretty small model, producing, like, very interesting reasoning traces and, like, getting, like, surprisingly good scores on math Yeah. In a way that we couldn't have done without, like, a bunch more pre training. And then, you know, once once that looks good, then, you know, more and more resources went to just, like, scaling up that new, like, law.
Yeah.
tool use and this kind stuff. Yeah. Yeah.
Did I think a lot of people make a lot of headlines on the the large models, but I think a lot it's very underappreciated, the minis. Mhmm.
How well this solution works. Mhmm. Any comments on just, like, discoveries on Yeah.
Nothing nothing much to say there. I was also, like, not super involved in Yeah. The mini stuff.
I think maybe one thing not not exactly related to that, but, like, it seems like externally, people are kind of very, like, oh, like, research seems to come in these, like, big leaps. Okay. But I think internally at OpenAI, it feels very smooth.
Like, you have a bunch of experiments. Yeah. Some of them have inconclusive results, but maybe you stack them.
Yeah. Exactly. You stack them and just, like, you keep scaling, you keep, like, having, like, different runs that, you know, get a little better each time.
Okay. So I think that's maybe one other aspect that's, like, a little underappreciated is that, like, I don't know. Like, in the media, there's just these wild swings between, like, oh It's so, like, Google is wearing it.
Yeah. Exactly. And I think, like, internally at BigLabs, it's just kinda like, oh, we're just, like, chugging along.
Like, maybe this month is a little better than last month or something, but it's, like, not as crazy up and down. Think the question is, there used to be more of this, and now I know there's less Mhmm. Which is, well, the stuff we've released, we're, like, you know, internally, we're, like, six months ahead.
people tried opening up wasn't that excited about ChatGPT's launch was because they already had GPT-four. They were like, oh, we just, like, put this out. Like, we're already way ahead.
I think now people are just releasing things as they have them.
think, yeah, especially because there's some, like, competitive pressure. Right? Yeah.
Think people are probably pretty worried that, like, if you if you let a lead linger for too long, that'll, like, grab a lot of market share. Like, I don't know. Nano banana pro right now is probably, like, you know, it's like Being pretty good.
A month. Yeah. So I would say, like, now the lead internal to external lead time is about one to two months.
Yeah. Yeah. Which is Exactly.
Tiny. Pretty pretty sure. Yeah.
Tiny. Yeah. Anything else on on reasoning side?
I guess you can talk about on on, so say, the work on coding. Anything surprised you or, like, is an external misconception on o one zero three side? Before I go to cursor.
not really. Like, yeah, I mean, it's, yeah, pretty cool. Like, I I think, you know, it felt already by, like, maybe early twenty twenty four.
Like, oh, wow. Like, this recipe, like, really works, and we can see how far we take it. And so I think, you know, it was, very steady progress, And, you know, by that point, it was probably pretty pretty predictable that we could, like, you know, really, like, smash, like, you know, things like IMO or IOI.
Yeah. One one funny thing that kinda happened is while this was happening, I went to this conference called The Curve Yeah. Which is about, like, kind of AI progress.
And Joseph Gordon Levitt, like, right? Yes. I went last year.
Last year. Was before the Owen stuff was released. Yeah.
And I, like, went to this thing where people were kinda making bets on where we'd be on epoch AIs, like, the the math the epoch a math exam and, like, humanities last exam and stuff like that. And, their estimates were, like, oh, we'll be at, like, 20% in, like, 2027. And I think at the time, there was, like, you know, models internally that were, like, already better than their estimates.
So they're, like it's, like, off by, like, know, two years or something. And the interesting thing is, like, those are also people who are kind of, like, you know, predicting that there'd be, like, Dyson spheres by, like, 2035 or something. Okay.
So So it seems like they're the current estimate is is way under Yeah. They're too pessimistic in the short term, too optimistic in the long term? Yeah.
Well, I don't know if, like, I I didn't there might be days since he was like twenty thirty five. Like, I don't I don't really know. Yeah.
But I think that that is like one interesting aspect is that, yeah, I think people still seem pretty miscalibrated in different ways. I I do really appreciate how that community makes predictions though. Like Yeah.
says like, oh, I I saw this the whole time. Like Yeah. So I do appreciate that.
Is this EA adjacent?
Yeah. I think having Yeah.
Exactly. It's like That's exactly correct. Group.
Yeah. Yeah. Yeah.
I like that they they like to sort of register their opinions ahead of time. Yeah.
the people who've been, you know the capabilities predictions in that group have been broadly correct if you look, you know, from, like, 2015 to 2020 or something, like, where I think a lot of people kinda thought that AI was, a sham or, like, you know, not really gonna be that useful for a long time.
that, like, it will probably reach, like, human level intelligence. Yeah. It's weird.
So, like, I I I I feel like a skeptic when I keep saying, like, everyone always predicts that AJ happens in their lifetime. Mhmm. That's very convenient for whoever.
And, like, we have a consistent view of history where you make see, like, people in the eighteen hundreds and nineteen hundreds making predictions. Mhmm. It somehow always lands in their lifetime, whatever the the thing is.
But, like, this time it might happen. Almost surely. Right?
Like, I don't know. I'm I'm pretty sure. Yeah.
Yeah. Yeah. So so it's it's an interesting observation, like, how different are we from our predecessors in in terms of developing a technology.
Yeah. Did the deepseq moment this year, also this year Mhmm. Crazy, change anything internally?
Not really. Yeah. I think that was I think more more so just, like, surprised that it created such a moment.
Like, it was kinda confusing. Right? It was, like, deep seq shows that NVIDIA chips are actually more useful than previously thought, and, like, NVIDIA's stock, like, goes down a bunch.
Like, it was kind of like a I think it's more more like, okay. Well, I'll I'll I'll do the steel man Yeah. Yeah.
That side, which is, well, you don't need the top of the line Nvidias.
You can just use the the sort of previous generation or the shackled ones they sell to China
to do an equivalent amount of work for for a PC model. I see. Yeah.
But then it was also I guess the feeling in OpenEye is that, like, well, I think we we had a better model already at the time. Right?
and it was quite valuable. Like, like, smarter models were clearly quite valuable, so you kind of wanted to be at the frontier. Okay.
I I so I wasn't quite framing this as, a race Yeah. Dynamics thing between labs. It was just also more like, well, were they right?
Or were their approaches right? Mhmm. They had r one zero, which is kind of like a really cool branch.
So more like commentary on what we learned about RL this year in 2020. Yeah.
a lot of the labs have kind of, like, converged onto some similar ish way of doing RL, and they're all kind of back at the same level of, like, Frontier again. Like, even the Anthropic, models, like the, Opus two four point five, it has this kind of, like, there's this, like, RKGI two plot that looks exactly like the AI ones. Right?
Like so I think everyone seems to be converging on a pretty similar form of RL. Yeah. It's kind of interesting.
I think people basically figured out in one way or another to, like, achieve more or less the same thing. Yeah. Yeah.
Let's talk about the move to Cursor. Yep. Why is Cursor accumulating and drawing so many cool RL people?
Yeah. Yeah. So I'd actually kind of, like, already, like, talked about this so far, I guess.
Like so, yeah, I think from the perspective of Cursor, it's, like, you know, nice not to be so, like, dependent on, like, external labs for everything. And, like, I think there's also, like, unique opportunities to co design the product with the model in ways that I think we couldn't do unless we actually, you know, built the model ourselves and, like, had access to, yeah, making it good. Yeah.
So, yeah, that's kind of like, broadly why Cursor's so excited. Yeah. I'll push back a little bit.
Alright?
OpenAI is has infinity resources. Mhmm. Infinity data has codecs.
You could've just stayed.
Yeah. Yeah. Well, actually, right around when I was leaving is when, like, I think people started actually, like, using codecs a lot.
So that was kind of like a it like, happened right after I left. So that's kinda funny. Like So so mostly people are using Cursor internally.
Mhmm. Maybe a bit of Windsurf because it's left over from the previous thing. Sure.
Yeah. Yeah. Exactly.
So it it wasn't that obvious. But, actually, I think more to the point, this thing I was saying about, like, RL is kind of a tool that doesn't really generalize that well. So what you wanna do is bring the entire, like, kind of test distribution inside your training distribution.
I saw the opportunity to do that at Cursor kind of, like, directly, and I think the Cursor folks are also just, like, really excited about that kind of vision. And it's just, like, a small place where, you know, like, the product people sit, like, right next to the ML people, and I think there's a lot of potential there. We can kind of see that.
Recently, Jacob Jackson had this blog post about, like, online tab where, like, you know, we're doing policy every two hours. Exactly. Like like a policy update every two two hours or something.
And I think that's the type of thing that you know, I think it's, like, a little hard to do it's, like, very hard to imagine that at OpenAI, for example, just because, like, you know, the product is this, like, kind of complicated thing and also, like, the product people and RL people are pretty, like, you know, on, like, different sides of the org. I think if you put your mind to it, you would. It's like, you know, TAB is an autoconfeet.
It's a smaller model. It's, you know, it's not Yeah. As complex, I guess, as below them.
Yeah. But I don't think that's really this, like, you know, I think we I don't think that's why Curcur was able to do it. It's actually more about, like, just the org itself being kind of, like, smaller and a bit more, like, focused.
Yeah. Yeah. Okay.
it's always been a big theme, it's bigger this year, is, well, don't you need to curate your data? You can't just, like, chuck whatever your users are doing in straight in because that tends to get you towards the middle of the distribution that actually you wanna spike it.
guess it depends how you're thinking about continuing. I mean, I don't know. Like like, humans are quite good about dealing with bad data too.
Right? Like, you can see something like, you can see someone doing something dumb and decide, like, you're not gonna do it. Filter it out.
Yeah. Yeah. But but, like, it's not even actually filtered out.
Like, you have, you know, presumably some kind of value function that, like, says that if you see someone touch a hot stove, like, you're not gonna go you don't need to like, it's not just filtering it out. You're actually not gonna do it. Right?
You could rediscover hot stoves on Christmas. Yeah. But, like, you don't you don't need to.
So I think there's something pretty deep there. Yeah. Like, it seems like we're kind of, like, a few orders of magnitude of, like, kind of data efficiency, basically, away from, like, that kind of, like, you know, you you you do something once or, like, you you make a mistake, like, you yeah.
You you introduce, a bug in your code. You're not gonna do it again, but the models will happily just like keep doing it, even within the same context, but definitely, you know, of course across context. Yeah.
So I think there's something like interesting and deep there is like maybe yeah. I suspect that it will be kind of like paradigm shifting in the next like year or something, but I have no idea, like, you know, what it might be. Yeah.
Yeah.
Is primarily, you worked on Composer Mhmm.
and maybe Search? So I I have I've I've actually just worked on Composer. Yeah.
And that's kind of like the main focus of the company, basically, or like the ML group is Which is shipping a better Yeah. Shipping a better composer.
I guess, the impressive
brag a bit about the ML group? Yeah. Yeah.
I mean, I I think the ML group is great. It's like you know, it's just like twenty, twenty five people. And, you know, I was, like, honestly, like, pleasantly, like, very, surprised at, like, how good Composer is, like, given the size of the group.
And, you know, it's not like a big research lab yet. And, yeah, it's I think it's like a really good model. You can kinda see that in the reception.
And I think it's kind of the start of hints of, like, co design with the product in some ways, because I think one of the reasons that people really like it is it's smart enough that you actually wanna use it. And it's also fast, so you kind of, like, stay in the loop with the model while you use it. Because I think all the other smart models had this kind of they're they're slow that you wanna go kind of context switch away and come back.
And that sucks. You know? Like, just as, like, a, like, programmer, it just sucks to kinda context switch.
It kind of, like, gives you ADHD. Like, it it's, like, really terrible. Yeah.
I agree. And I think, yeah, it's, like, one step in the direction of, like, being able to be more in sync. And I think that's, like, basically, the whole company is just really, you know, full of people who want to, you know, code, even, like, the cofounders, you know like, actually, the cofounders are often some of the best, like like, high taste testers, which also kinda gives you a lot of, like, reassurance that you're gonna ship good stuff.
So Any example test that, like, maybe Composer doesn't solve yet, but you're really motivated to solve? Yeah. Ironically, I feel like I'm actually, like, a low taste tester in some ways.
Because I don't know. Like, you know, I just, like, write, like, slow, like, machine learning code and just, like, think about algorithms and stuff all day. Yeah.
I think more broadly, I'm super excited about co designing the product so that you can actually, you know, not just right now, we're getting better and better at, like, answering user prompts, and I think that's why ComposerOne is, like, quite good. But, you know, what we're really aiming for is, like, more, like, you know, automate software engineering as a process where you, like, write code, you go look at Datadog, look at what's, like, happening, then come back and, like, you know, maybe have some hypotheses about what's better, like, re rerun stuff. I think that's the type of thing that we actually want to make the model do.
And I do think that Cursor is kind of, like, uniquely positioned to do that the sense of, like, you know, if we can kind of if if a lot of what a software engineer does kinda ends up in the product, I think we can use that to, like, get better and better at, you know, not just writing code, but kinda, like, the whole job. Yeah. I think that's very inspiring.
Just to double click on just any sort of RL insights.
and Lee have talked a lot about, like, the internal tooling that you've had for all the, like, the cluster visualizations. Is that helpful? Is that what every lab has?
Yeah.
the tooling at Cursor is actually really good. I think because, you know, it's just kind of like a people are just down to, like, vibe code stuff. They, like, do Of course.
Test their own stuff. Like so we just have, like, a lot of good tooling where you can, you know, like, have, like, a SSH session into, like, our own, like, user environment or something and, like, you know, see if, like, code runs the way that, like, users got it to run, like, this kind of thing. I think that's actually, yeah, quite nice.
I think, basically, one of the big lessons in ML in general is that you wanna be, like, really close to your data and understand your data well. And, yeah, I think because there's, like, kind of yeah. Again, kinda, like, uniquely positioned to do that well.
Especially because all internal tooling, you're not buying anything. Yeah. It's just, like, internal.
And part of it is just that we're also working on a product where you can understand it really well because it's a code product. Well, like, you know, if I don't know. In in in in OpenAI, if I was, like, to look at, like, a biology question, I have no idea, like, you know, what what this is about.
Yeah. Yeah. Yeah.
Interesting. Okay.
So I think that's a good overview of of everything. I guess other than the we covered OpenAI and Cursor, just interesting RO work that other people are doing that you're that you're, like, still mulling over its influential to your thinking, or good papers, anything like that? Yeah.
unfortunately, like, kind of gotten the habit, especially at OpenAI, of, like, not reading that much external work and just, like, reading people's, like, Slack post internally Nice. As, like, the main, like, way to, like, you know, like, learn new stuff. No super inspiring recent things have popped up to me.
I do think that this, like, kind of vibe of, like, yeah, continual learning, it's like does feel like and there's something super interesting there and, like, it feels like, maybe even in academia, people could make, like, a big crack at it. And continual learning specifically meaning kind of what Tab is doing? Yeah.
Maybe what Tab is doing, but also just, like, kind of, like, in context learning, but with, like, infinite memory or something so that you don't once you experience something in context, it should just, like, be in your weights, and you shouldn't have to, like, make that same mistake again, that kind thing. Why do you think there's okay.
capacity for the weights to remember things. Yeah. Yeah.
if you do that too much. Right? Not I mean, you know, you you start out by memorizing or, you know, like, learning from trillions of tokens.
Yeah. Now, you're gonna experience, like, thousands or maybe millions of tokens, and somehow, like, you know, we can and those the million tokens are kinda And it's you only need one epoch. Yeah.
Exactly. So Crazy. Yeah.
So it feels like if you if you could learn enough about those million tokens that you're actually in deployment on, I don't think you should need like, I don't think there's a risk of overloading the capacity of your model. Right? Because you you can train on a trillion tokens, and it's, like, fine.
Right. Right. So this proportion is a drop in the water.
Yeah. Exactly. The water bucket.
Yeah. Unless you run it for years and, you know, at some point, it's sort Maybe. Yeah.
Yeah. So, basically, I I I find it very curious. I've only had one podcast on information theory of language learners.
Mhmm. Mhmm. Like, what is the theoretical capacity?
How much are we using? And you should probably track that. Yeah.
Yeah. That's a good idea. Yeah.
Like, treat the weights. If you want to store things in weights, okay, treat it as a hard drive. What's the capacity of the hard drive?
How much can be stored in there? We know we know the capacity. It is the number of bits that, you know, that this Okay.
Right? But the language but but the parameters Yeah. Yeah.
Physically cannot store one in that. Yeah. Yeah.
Yeah. And it's yeah. I've I've heard that there's this kind of, someone recently at Cursor Jacob kind of brought up this view.
you know, neural networks and kind of, like, a CPU view of neural networks where, you know, is is what's getting the weights, like, yeah, memorizing stuff? Or is it like you're, like, having some, like, few circuits that do a lot of work? Yeah.
This kind thing? And I don't know. Yeah.
You know, I would love to yeah. There's, like, actually so many of these kind of more science y questions that I would, like, love to explore sometime, but then it really kinda conflicts with, like, empirical stuff, you know? Like, unfortunately, at any given moment in time, it doesn't seem like the most fruit for, like, you know, improving something in this, especially in the short run, but even in the next, like, couple years is, like, understanding some of these questions.
Yeah. I mean, I guess this is technically supposed to be the role of academia, but it's, like, also hard to explore those ideas there without enough compute.
But, yeah, actually, I would love to, like, go at some point, you know, like, return to exploring these kind of, like, fundamental science ideas. Okay. This is a I was kinda springing this on you, so you can take some time.
What is a good RL interview question that if somebody can answer, they should join Cursor immediately?
Oh. It's a hard, question. Assuming you do interviews.
Yeah. Well, actually, at Cursor, we do, like, work trials. Yeah.
And it's, like, two day work trials that I actually think that that's, like, more representative. Because you plug in and Yeah. See how they behave.
Exactly. So I actually think it's, like, more valuable. This is honestly less of a thing about how you understand RL and bit more, like, were you around in the, like, 2017 to '22 era.
But, it's, like, why is off policy RL unstable?
question to, like, yeah, dive into. I don't actually know, so I need to dig into it. Yeah.
Cool. Thank you. That's that was great conversation.
Do you have any sort of call section?
Yeah. I mean, you know, we are definitely hiring a cursor. So, yep.
If you're interested in working on especially, like, kind of, data and rewards for code, I think that that's, like, a huge need. Yeah. Please, like, get in touch.
Yeah. That's it. Yeah.
Thank you. Suits.
Shared via Hopper