This episode features a crossover discussion between Swix and Jacob Efron, debriefing AIE Europe and exploring the 'Agent Labs Thesis' in 2026. They cover the latest in AI infrastructure, the intense AI coding wars, the evolving market dynamics between foundation models and startups, and future frontiers like world models and personalized AI experiences. The conversation highlights the rapid pace of change, the shift towards open models, and the increasing role of agents in software development.
Isn't that crazy? That number is just mind boggling. What is the state of the AI coding wars today?
We're in a phase of sort of, like, capability exploration. The general thesis that I have been pursuing now is that the same way that 2025 was a year of coding agents, 2026 is coding agents breaking containment, do everything else. Do you worry about the foundation models just eating into a bunch of these startup categories?
Mid sized startups, yes. What do you think the end state of this market is?
there would be Today on unsupervised learning, we had a a fun episode in what's really become an annual tradition, a crossover episode with our friends at Latent Space. Swix and I sat down, and we talked about everything happening in the AI ecosystem today, what we thought of the various changes at the model layer, what's happening in the infra world, the coding wars, and a bunch of other things. It's a ton of fun to do this with someone I really respect and another great podcaster in the game.
But without further ado, here's our episode. Well, Switch, this is super fun to be with another unsupervised learning latent space crossover episode. Yeah.
I feel like a lot of places we could start, but, you know, one thing I always find fascinating about the way you spend your time is you obviously are like at the epicenter of this engineering movement and community, and you run these events and conferences and put on these awesome talks, and and I think just have a great pulse on the zeitgeist of what's going on. Yeah. Maybe to to start just, what are the biggest topics people are thinking about right now?
Yeah. So I just came back from London where we did AIE Europe, and we're doing roughly one per quarter now, which Yeah. You're really up the up the pace.
It's try we're trying to match AI speed. Yeah. Exactly.
The topics would be completely different, I imagine. Yeah. You know?
I definitely curate the tracks. Like, you can see what I think when you see the track list and the speakers that I invite. Obviously, Open Claw is, like, the story of the last four or five months.
And then be be just below that, I would consider harness engineering and context engineering to be two related topics in agents and rag. And then there's a long tail of evergreen stuff like evals, observability, GPUs, LM infra in just general just in general. We also have other updates on, like, multimodality and generative media, let's call it, but definitely the first three that I mentioned are top of mind Yeah.
People. I think Harness is particularly so interesting.
There was this tweet from Harrison Chase, the the Lane Chain CEO that caught my eye recently where he said, it finally feels like we have stability around the infrastructure around AI, and I think what he basically was implying is like, look over the past two, three years as a company at the epicenter of AI infrastructure, it was a bit like playing whack a mole, right? You were constantly moving around with however the building patterns were evolving. For Harrison, for sure, right?
right?
sharp people about this. Yeah. But yeah.
Basically Now now is finally time for stability. Yeah. Yeah.
Do you buy that, or what have you kind of make of that take?
I think that it's very expensive to say this time is different sometimes. But when you're just writing code, like, it's actually okay to just, like, try to make a call. And I think it may not even matter if this call is right or not.
Like, I just don't even care that much because you can be right on a thesis, but if you don't you know, you don't figure out how to monetize the thesis, then who cares if you said something first? That said, it does feel like, for example, we went through a lot of different ways of patch packaging integrations up with, with agents. And it feels like we've landed at skills, which is like the minimal viable format, which is just a markdown file, with some scripts attached to it.
And I don't see how it can be more simple than that. And so there is some justification for the stability around harnesses. I feel like there may be more adaptation with regards to maybe, like, the real time elements or sub agents or memory or any of those, like, agent disciplines, let's call it, in in agent engineering.
But if if the thesis is that, okay. You just want agents are LLMs with tools in a loop with a file system where they can do retrieval with with skills and all these, like, standard tooling that now seems to be relatively consensus, then probably that makes sense.
this thesis that we're there because if it changes again, just change with it. It's fine. Yeah.
That's always know, I've always been struck by how that is much more challenging for infrastructure companies and application companies. Like, obviously, I think Yeah. You know, on the application side, you've seen, you know, Brett Taylor from Sierra, Maxime Strung from Legora, like, they're like, look.
We build what's ahead of the models, and we're willing to throw everything out every three months as the models get But better and the thing you at least have there is you have an end customer, right, that's decently sticky. They will mostly stick they'll give you a shot at least of building these things. What I've always found more challenging at at the kind of, like, you know, reinvent yourself every three months at the infrastructure layer, it's like, you know, developers are definitely a pickier audience maybe than an accounting firm or, you know, a bank, and so it's definitely a a a more challenging position to be in, to have to constantly reinvent yourself.
Yeah. Yeah. And like when they churn, it's like very complete.
Like they'll leave to like the hot new thing, because there's like no defensibility, I guess. Like even if you are a database, like people can migrate workloads off databases. Like, it's it's an it's a known thing.
So I think, like, basically, what we're talking about is the vertical versus horizontal, debate in in AI startups. And, the way I think about it also is just that, like, when you are, Legora, when you are a bridge, like, you are the outsourced AI team. Right?
You your your job is to apply whatever state of the art AI methods. Yeah. Like, this translation layer between model capabilities and your end customers.
To to the end customers. And, like, well, if they didn't have you, they would have to hire in house, and they're not gonna hire in house. So they have you.
And, like, I think that's, like, a reasonable, like, very robust to any whatever trends and and discoveries that people make in in the engineering layer. I do think, like, there is, like, sort of useful horizontal companies being built, but they're all very much, like, sort of, like, the reinventions of classic cloud in the AI era, and the the primary one being sandboxes. Yeah.
Which, like, it's another form of compute, guys. Like, let's not get too excited about it.
I mean, like, the the workloads are enormous. Right. Yeah.
It's interesting, and I feel like as as part of this, you know, the questions that folks are asking around infrastructure, there's a lot around, you know, the extent to which companies should have their own AI teams and what they should be doing in house, and, you know, I think there's questions around should people be training their own models? Should people be doing, you know, RL, in house based on the data they have? I feel like, you know, one has to evolve their takes on this every every three months with paces, but where where are you at on this today?
I think mean, actually, all models have gone up.
And, obviously, I'm involved in cognition and also cursor is doing doing a lot of own model training. And I think that that is some part of the what I've been calling the agent lab playbook, where you start off with the state of the art models from from the big labs, and you specialize for your domain. But once you have enough workload and enough high quality data from your users, then you can obviously train your own models and, like, save a lot on cost and latency and all that all that good stuff.
You also get, like, a marketing bonus of, like, calling it some fancy name and putting out some research. From my seat, I can't tell how much of it is, like, actual, you know, value that's provided to the end user and how much of it is that marketing bonus. Right?
It seems some combination of the I think it's both. Yeah. No.
No. There there actually is real value. And you you know that for a number of reasons.
Like, one, even when it's not subsidized, people do choose it as, like, one of the top four or five. This is both Composer two and, Suite 1.6.
I want the top five models, like, in a in a fair market, in a free market Yeah. In a in a in a model switch where people do choose it. And, like, it's not subsidized.
So they could that's as good as it gets. But beyond that, like, domain specific models, for example, for search with with both which both companies have absolutely makes make a ton of sense. Everyone says, like, yeah.
We should always always do this. And, honestly, like, I think the infrastructure for that is becoming easier with, like, thinking machines, Tinker thing, well as as primary intellects, lab stuff. Yeah.
I mean, like, it this is one of those, like, reversal of the the bitter lesson where you first bootstrap on the large models and the general purpose models to get big. And as you get very well defined workloads that are just high quantity but not high variance, then you just distill down to a smaller model and run that on your own Right. Which, like, totally makes sense.
RL use case, which I think is really mostly around, you know, improved quality for for different things. Obviously, there's probably, like, more efficient ways to, you know, get a smaller model that's that's faster and cheaper, and it'll be interesting to see whether you know, obviously, had, you know, two, three years ago, this whole case of companies that were, you know, pre training and claiming better outcomes in in their domains, then getting kind of cooked as each model iteration improved. Know, I wonder whether that's similar story plays out in the in the RL space.
Yeah. For the focus on on on pure outcomes and quality, not the cost side, which clearly your own models for cost at scale makes a ton of sense. I think there are this there are two sides of the same coin.
quality constant or trade off a little bit of quality for a drastic decrease in cost, and that's true for everyone. One element I wanted to bring out, which is very much in favor of open models, is custom chips. So this would be Cerebras, but also Thales.
And then there's a huge range of stuff in between. This has been a huge story this past year on just, like, everything non NVIDIA is getting bit up, including, like, freaking MatX is working for which is very which is very rewarding for me. But I think one of those things where, like, oh, like, the suddenly because the number of alternative hard, hardware is increasing and the inference that you can get is insanely high.
Like, we're talking thousands of tokens per second instead of less than a 100.
doesn't hold as much anymore because the speed is so high. Have you seen a lot of companies go all in on the alternative chips?
Cognition has on Cerebras, and and so has OpenAI. And so, no, I don't think so beyond that. That's that's mostly because that's Foreshadowing of Clearly yeah.
I used to be kind of a skeptic in terms of, like, okay. So what if I get my inference at a 100 to 100 tokens per second sped up to 200 tokens per second? It's only two x faster.
It's not that big a deal. But when you I think every 10 x does unlock a different usage pattern, and you we have proof in and and some of the others that you can actually drastically improve improve inference speed. And what happens from there?
I don't even really know. Like, it's it's so hard to predict when entire applications just appear at once. Yeah.
And it also isn't that expensive. Right?
would caution people to not dismiss it too too quickly. Yeah. I mean, one other, like, infra question I was curious to get your thoughts on is, obviously, seems increasingly a lot of the cutting edge infra companies are building for agents as the buyers of their product or users of their product, right?
Oh. They're another huge theme. Yeah.
And I'm trying to figure out, like, what what do you have to do differently about selling into agents?
or is there, you know Absolutely not. I think they are easily prompt injected and, very tuned towards, like, basically comp compounding existing winners. Yeah.
So, like, if like, congrats if you won the lottery for getting into the training data Right. Before 2023 because now you're, like, installed in there for the foreseeable future. But, yeah, you know, one stat that Vercel CTO Malte Ubel dropped at my conference was that there are now 60% of traffic to Vercel's, like, app admin app architecture for, like, configuring Vercel applications is bots.
It it's not it's not human. So, like, your primary customer is agents now, and it's mostly code like, mostly coding agents, mostly people using a CLI, OMZP, whatever. But, yeah, I mean, I think I I think step one, if it doesn't exist as an API that agents can use, it doesn't exist.
Right. Right? Which I think is, like, it's a good hygiene thing anyway to to make everything API available, but now it's, like, an extra push on, like, products people to not only work on the UI.
You should probably work on the on the CLI stuff. Beyond that, I think, honestly, there is, like so I I come from the sensibility of I think everything that you are trying to do for agents experience now, which is the term that Matt Bilman at Netlify is trying to coin, is the same thing that you should have been doing for developer experience. That you should have had good docs.
You should have had a consistent API, that is mostly stateless. You should have, I guess, discoverable or progressive disclosure or, like, search or, like, whatever. And so now that people have energy in, like, finding these customers to do that, that's great.
Do I believe in extending beyond that into something like AEO, for gaming the chatbots?
Not necessarily, but obviously there's gonna be huge advantages from people who figure out the short term wins. Yes. And short term wins can compound.
Do you think these compounding advantages to, like, the the pre training data cutoff companies, like, you know, obviously over some period of time, I imagine that doesn't persist. And so I as you think about, like, I don't know, three, four years from now, what the, you know, selection criteria end up being, do you think it still mirrors exactly what you were saying before? Like, it's exactly what you should have been doing all along to sell a good product to developers?
It could be, except that I think in three, four years, we'll probably have much better memory and personalization. So then general AEO or GEO doesn't really matter as much. So I think whatever memory or personalization system we end up with will probably determine what you end up choosing much more than than what is currently the case, which is just frequency of mentions, let's call it.
Yeah. Yeah. So you just spam quantity.
And I think that's I mean, that's something I'm looking forward to. I do think, like like, you know, I I think that the fundamental exercise to work through for yourself is if you start a new, sort of disruptor company now, there's a there's a big incumbent that everyone knows. Like, like, Superbase.
Superbase is, like, kinda like the Postgres, like, database, incumbents. If you wanna start, like, new Superbase, how would you compete with them? And I don't necessarily have the answer, but I I I do think, like, people like resend, like, relatively new.
I think they were starting, like, 2023, and it's still there was there was a recent survey where, like, people checked what Claude recommends by default. If you just don't prompt it with anything, just say, like, give me an email provider and says resend as in, like, seventy seventy percent of these cases. Like, the fact that you can get in there with, like, such a relatively short existence, I think is is encouraging.
Yeah.
no. That that very short mentions this because, it's not gonna be 20 of them. It's gonna be, like, three.
No. Definitely. It feels like, you know, probably more more consolidation than ever, or or kind of, like, you know, a winner take most market than maybe the the the physics of go to market in the past might have enabled.
The other thing also is, like, semantic association is gonna be very important, in the sense that, like, you want to do, like, the combo articles where you're like, use my thing with Vercel, with blah blah blah. And, like, that all gets picked up in a in a corpus. And so that's probably one thing that you you wanna do well.
I don't know what else. It's it's it's it's one of those things where, like, I think I feel I feel I'm behind.
Listen. Yeah. With I wanna meet the person that doesn't feel behind.
But, with with AIX, right? Like, so so, like, my my stance was exactly what I said before. Like, everything that you that you should do for agents is something that you should have done for humans anyway.
Yeah. And so to the extent that you're just getting it more energy to to do things for agents, great. But, like, it's hard to articulate what new thing apart from just, like, more spam, that you should be doing anyway.
That will be my take right now. I I I do think, like, there there will be more turns at this. I think the personalization turn that is coming, will be big, and I don't know what that looks like because, like, basically, we're kinda we we feel kinda tapped out on the memory side of things.
Yeah.
and you've obviously have a have a front row seat to the AI coding space today. I feel like coding in many ways, know, people view it as this like I mean, besides being like the the mother of all markets and this massive opportunity, I think it's kind of a preview of, like, what's to come for many other spaces, both Yeah. You know, I feel like agents are most advanced in coding.
I also feel like the competition between foundation models and application companies, you know, and mirrors what we may see in other spaces. And so maybe for our listeners, can you just lay out, like, what is the state of the AI coding wars today?
It is massive. Right? Like and I don't think necessarily the last time we talked about this, we appreciated the size of No.
I wish we did. The state of the AI coding wars today. Both OpenAI and Anthropic have made it their p zeros to competing coding.
And Thropic is at, like, 2,500,000,000 in ARR just from Cloud Code. The way they recognize ARR is up for debate. OpenAI, I don't think the public number is known, but let's call it 2,000,000,000 as well.
And then Cursor is like rumored to be 2,000,000,000, you know? And, and those, those are like the public numbers that are so like huge markets that have just been created in the past one year, like, like, Anthropic just, like ClawCode just recently celebrated their one year anniversary, which is pretty amazing. So, and then I think like the other thing that I see is there's, there's some other people who are like, oh, here's, like, the the sort of relative penetration of, cloud use cases.
Right? Like and it's, like, coding 50% and then legal whatever health. It's, like, the the remaining ones.
And there was a very popular tweet that was like, okay. Look at the the empty space and all these other use cases. If you are a new founder today, should be betting on the other stuff because on on a sort of catch up Yeah.
Theory. And my consider my my pushback is the same pushback that I had on Apple versus Google, which is like, well, why is this type of difference? Like, why if it went from, let's say, 10 to 50% in the past year, why can't it keep going?
And, like, getting that wrong is actually a very painful one because you could have just didn't did the momentum bet instead of the mean reversion bet. So I I I think that that is the the the state of things now that people are very much into psychosis. They are getting rewarded for spending more rather than spending less.
And I think we're not in that phase of efficiency. We're in a phase of sort of, like, capability exploration. So I think people who are more crazy, who are more creative, get rewarded comparatively.
What's an interesting I mean, it feels like behind these, like, token maxing leaderboards and whatnot is this it's like the first phase of this transition from a workforce perspective is you just gotta show your employer, hey. I I use these tools. Which is my number of tokens I cost.
That's it. They don't care about the quality right now. It is, maybe distasteful to someone who cares about the craft and and all that.
But, directionally, everyone just wants you to go up regardless. And so, there it's not very discerning. It's and it's probably very sloppy, but I think it's net fine because we're still probably underusing AI just in generally.
Yeah. And so I think that's, like, very interesting. Like, we had on the podcast, Ryan Napopolo from OpenAI who spends a billion tokens a day.
Yeah. And that's for those counting home is, like, something like 10,000 worth $10,000 worth a day of API tokens if they did market rates. And, like, most of us can't afford that.
Yeah. But, like and and probably a lot of what he does is slop. Right.
But, like, he's going to this he's like, if there were a new capability, he would discover it first before you because he was he was trying and you were not trying. Right. And, like, you only do things that work.
Like, well, good for you, but, like, the the people who are going to discover the next hot thing are living at the edge.
Right. And increasingly living at the edge is just having that compute budget to, like, run these experiments. Mean, kind of similar to what living at the edge on the research side has always been.
You know, it was constrained in many ways by the amount of compute you had to run these experiments. It feels similarly on the almost on the builder or, like, actualizing these tools now.
that where, you know, restricting limits or restricting model releases even is, like, the name of the game, whereas Codex is like, come on in, guys. Use our SDK. Use our login.
We don't care. We're gonna reset limits, whatever. You do want to try to exploit the subsidies where you can get it, and definitely Codex is super subsidized right now.
Gemini also very subsidized. And comparatively, like, I think you should make hay, I guess, while while that's going on, it's not that bad to be a capabilities explorer on just the $200 a month plan from Cloud Code or from OpenAI.
yeah, I've I'd my sense is that people aren't even there yet. How do you think this, like, market ultimately plays? Mean, it's obviously such a big market that, you know, any slice of that market is interesting for for anyone going after it.
But I think what what makes people so interesting in the coding market particularly is it feels like it's kind of this foreshadowing of what will happen in other you know, any other kind of application market that the foundation models eventually turn to and are all their models against and gather data around. And so how do you think you know, like, does there end up being room for lots of different kinds of players? Or, like, what do you think the end state of this market is, and is that do you think that's applicable to other markets?
there will be I mean, status quo is probably the most likely outcome, which is there are two big players, and there's a small range of longer tail people that, fit other use cases that the the two big players don't. That feels right to me. I think that for it to for the market structure to to significantly change, there would be there needs to be significant change in, like, the economics or, like, the the brand building or, like, the the the value propositions of the of the companies involved.
And I haven't seen any in the last six months that that have really changed the stories materially. So I feel like they would just keep going until something something else happens. Something else happens, meaning, like, Microsoft wakes up and, like, goes, like, guys, we have GitHub.
We have, you know, we'll we'll we'll do something much bigger here than other than just Copilot. And, that will be a big change. MSL has put out a model now, and I was in a breakfast with, Alex Wang where they were like, yeah, like, we we really, really want to go after the coding use case.
They haven't done anything yet, but, like, don't underestimate them. Right? And and and similarly for the Chinese labs, like, I think they're trying to go after it.
Like, ZAI is doing stuff, GLMs, ZI, GLM, same thing. And and so it's not like everyone's trying to get a piece of that pie. I I feel like the the status quo has been pretty stable for the past, like, almost a year, I will say.
Yeah.
what service area do the model companies leave for application companies? Yeah. That's a good one.
It's very much evolving. I will say because OpenAI did not have this level of attention on coding a year ago, we just don't have that much history, right? And it seems like, for example, so the big push at opening on now is the super app.
Is that a consumer thing? Is that like a products, like portfolio rationalization thing? How much is that gonna take away attention from coding at the time when they actually do want to put more coding?
I think it's it's very unclear. So I do think, like, there's there's all these, like in both BigLabs, there's, sorry. Both of them are Anthropic and and DeepMinus and XAI as separate cases.
They are trying to see the other TAM expansion areas. So cloud code for finance Yeah. Cloud co work, all those all those things.
Whereas I think cursor and cognition are, like, comparatively just focused on coding. And so I I do think they leave space. And I do think for the other verticals that also means the same thing, right, that that they're not gonna be that intensely focused on on on that domain.
Except for, I I think I will mark out finance and health care as like the next ones that they're clearly going after. I I would say comparatively, healthcare seems more thorny. There there there have been some announcements about it, but like, I would respect the the finance work a lot more just because like the the path to money is a lot clearer.
Yeah.
I mean, like, think, you know, maybe similar to the space that's being left in these other domains, you know, there's obviously a lot that's required to actually implement these tools in enterprises
versus, you know, maybe just giving them, giving model access to folks out of the box. Yeah. Yeah.
Yeah. So the the agent lab thing is, like, we'll do the last mile for you, whereas I think the model labs tend to just trust the model and and be minimalist about it. Both of them work.
Yeah. I I don't I don't necessarily think one, beats the other, for every for every use case.
which is kind of interesting. We've we've been in this phase of of pure capability exploration, and so I think nothing has been, you know, better for for large labs. Right?
I mean, they were always gonna be at at the frontier of of capability exploration, and so I think have a very good relationship with lot of these enterprises. But ultimately, over time, like, the the incentive structure of these labs is always gonna be maximal, you know, token consumption for for the end customers they work with, and there's just, I think, so few companies that have actually gotten to massive scale. Maybe coding again is the most interesting.
It's the space that really is just completely gone. You know? Yeah.
You must live it up every day. Like, absolutely insane.
And I think you get Even okay. I mean, like, I think we we say good things about cursive cognition, but the sheer lift off of, like, both endopic and open AI, because they, they have independent valuations. I mean, let's throw an XAI in there.
Yeah. It's now IPO ing at 1,200,000,000,000. That number is just mind boggling.
market cap or valuation that that, like, you you reach and you're going, alright. It's it's gonna be chiller from now on. Like and these guys are not slowing down.
No. Well, I also think the dynamic that's fascinating about some of these later stage companies is is, you know, in the past, I feel like in in venture world, if you got to a certain level of scale, the question around you was really more a valuation question, and this is why there was different types of venture people didn't. The late stage growth people were just incredible at a little bit of what's the ultimate market opportunity of this company, but also what's the right way to evaluate it.
We know it's in some bands of an outcome that is like, sure, there's some variance to it, but it's like relatively understood what that band is, and then maybe you get over time surprised to the upside. Whereas any kind of like even the labs themselves, any later stage company, the bands of which that company might be worth right now, even in a year or two years, are so massive because of how fast the ecosystem changes, that it's like even for later stage companies, every three months could be an existential level event to the upside, to the downside, and I think that you're obviously seeing it in positive with Code, which if you think about a company like Anthropic, for a while it was unclear if they were going to have access to enough capital to really stay in the race, right? And then coding hit at the exact right time, they had the perfect model for it, they executed brilliantly, and now we're one of most valuable companies in the world.
have zero sympathy for OpenAI because they're crushing it, and they're all rich, You know, this is like a high class champagne problem to have, to to be number two at coding or whatever. Like, who cares? Like, you're you're doing great.
Yeah.
you know, even though you're in the AI coding space, but it's a lot of people I talk to think Codex is just as good if not better than ClaudeCode. Right? And and I think one thing that I've been really surprised by, and maybe maybe ClaudeCode is a better product in some ways, I'm curious of your thoughts, is just in consumer AI, with ChatGPT, you saw this big first mover advantage, right, where admittedly today, like, I don't know, Claw Gemini, great products, not sure, not abundantly clear ChatGPT is any better, but like, people stick with ChatGPT.
It's the first thing to introduce them. They stay, but they're not growing anymore. Don't know if you've seen.
Right, but that to me is more of like a product problem than it is they're not it's not like they've lost share to someone else. My understanding is the overall problem with consumer AI today is much more of a how do you take this tool, and for folks like us, like knowledge workers, it's like this incredible magic tool, but it's not necessarily a daily active use tool for a lot of people around the world today, and whatever the product it's kind of a category wide problem, like in coding, for example, the entire space has gone parabolic. There may be some relative growth in other consumer AI players, but it's not like consumer AI as a category is going parabolic, and they're not capturing most of that thing.
I think it's actually the larger problem is much more, hey, the category has hit a bit of a plateau, people haven't figured out how to bring tons more users on board, or increase the frequency of those users, and so it seems more of a category wide problem than it is a massive market share change. Was gonna draw the comparison to the coding space where ClaudeCode was the first product obviously to introduce people to this magical experience. By all accounts Codex is pretty damn close to as good, if not better, but, like, still that first product, you you would have thought that would not be a super sticky, you know, product surface area, and it actually has it turns out, it feels like the first lab to introduce you to an experience really does keep a lot of the a lot of the focus.
maybe it's, like, still still early days. You know? Chachi BT is, like, three plus years old, and Yeah.
Cloud Code is only just turned a year. Yeah. So just give it time.
You know? Like, yeah. I mean, definitely, some a lot people have switched from to Codex.
Maybe that will keep going. It's, like, really hard to tell. Yeah.
I don't think is as high as it might be in some other, areas in our careers that we've looked at. Yeah. Though, I mean, I've been surprised by the Claude code thing.
I I would have thought that, like, in many ways, I always worried about the You think you would been gone by now? Not gone, but I would have I I always worried that the that the consumer business of these companies would be quite sticky, and then the enterprise API business was actually, you know, in some ways, like, your least loyal buyers, like, would they would move to Right. But but they worked out that it wasn't the enterprise API, was enterprise product.
Totally.
with two products that by all accounts are pretty damn similar. Yeah. No fight there.
I will say I do think that Codex is still in like a catch up, in terms of personal experience. The only thing I like out of Codex is like Spark and like the, I feel like the skills integration is a little bit better. I feel like the speed is a bit better maybe because it's written in Rust or whatever.
Very minor things that you like, almost like telling yourself rather than like objectively assessing between two of them. I do think like vibes wise, I think that's going on. The, you know, I feel like the missing questions in this whole debate is like, why is this so concentrated in only two names?
Right? Like, where is the Gemini you know, presence? Where is the XAI presence?
And, like, they are trying. It's just they haven't made that much progress yet.
Well, I think what the what the Cloud Code moment does show and it actually, in some ways, makes you a little more bullish on the potential for someone else to catch up because it does feel like if you're the first person to introduce some magical net new product experience that that actually might be stickier than one might have imagined.
Right. Right. Right.
Okay. Yeah. And so is that I've heard you What do you think that new product experience might be?
I it's it's like and this is a failure of imagination on my part. Like, I always wonder, like, people always say this, like, well, the the thing that will save us is, being first to the next new thing. Like, what is it?
Yeah.
don't know. Something around, like, a consumer agent, computer use, like, hybrid. I think the obvious I I think we're, like, scratching the surface on the consumer side.
is, like, a vision of things to come. Totally. And it's kinda good that OpenAI has, like, the association with OpenCLaw, but by no means do they have the rights to win it.
The general thesis that I have been pursuing now is that the year the same way that 2025 was the year of coding agents, 2026 is coding agents breaking containment to do everything else.
And so coding agents continue to still win, but because they generate software and software eats the world. So, like, it's kinda like the trans associated property of, like, software eats the world, coding agents eat software, therefore, agents eat the world, which is, like, an interesting Yeah. And breaking containment, always an easier phase phrase in the consumer context than the enterprise one.
You've seen people run these really cool experiments in their own personal lives. I think, like Yes. Figuring out, you know, how you I I obviously, everyone's focused, you know, on the enterprise side now around how you create these experiences.
I feel like the vibes know, people love to have these narratives of like everything is completely shifted. It's like I actually you know, OpenAI organizationally, you know, volatility aside is, you know, great products, great team, great models, like everyone else in the world is incentivized for there to be two, three more, everyone would love more like great model companies, and so I feel like the natural forces of the world revolt when any one company is too much the star of the show, right?
a reversion of vibes, not maybe completely the other way, but at least a little bit more equal at some point over the next six, twelve months. I I think there's just a kind of different stages. When when you talk about the world wants wanting more model companies, I talk think about, like, the Neolabs.
Yeah. And I mean, don't know. Is it fair to say none of them have really broken through in the past year?
I think that's totally fair. Which is rough. And well, how are we gonna how are we gonna grow that diversity in in in choice?
Like, that's this is it.
Yeah. It'll be really interesting to see what what what ends up happening with that. And you've seen, you know, folks like NVIDIA, you know, very incentivized to make sure there's there's a broader platform of of other model providers.
I think I don't know. People say this, but I I I don't think they try that hard.
NVIDIA tries harder to build Neo clouds than Neo Labs. Well, they try pretty damn hard to build Neo Clouds, so that's yeah.
Like, you know, let's call it like the, the core weaves of the world, Much happier place in the you know, than any NeoLab built on top of them.
Yeah. Though one might argue, it's it's easier to to enable a NeoCloud to be successful than it is. You can't will a NeoLab into existence the same way.
Everybody has more direct control over it, for sure. What else is kinda catching your eye today on the startup side? I mean, you worry there's obviously this whole narrative of like, you know, the foundation models, you know, they announce a product and every stock goes down 15%, like Yeah.
Do you do you worry about the foundation models just kind of eating into to a bunch of these startup categories?
Not really. I think actually like as, there's there's okay. There's there's there's the there's the point of view of like being an investor in startups and there's a point of view of like, do you want to start something?
And I think honestly, like the, the downside for all of these is so minimal in, in the sense of like the worst you do is you just get hired into one of these labs anyway. So I think the, the market for people who just do things and try things and try to execute in like a competent way, even if you're like, it doesn't work out commercially, even if it just wasn't that great anyway. Like, but like that's your job interview to go into one of these things anyway.
So, I don't feel that from a, from a very, very small startups perspective. Mid sized startups, yes. I would say there's been a lot of dead LM infra consolidation, like the Lang fuses of the world getting Azure to the click house.
And I think like people have maybe worked out the domain specific playbook and like, I think that's okay. I, I'm yeah, I'm not that, not that worried about, okay. So, I, I would say I'd be more worried about traditional SaaS, like low NPS SaaS.
This is the whole AI versus SaaS debate that's been going on. And like, literally I'm going through that exact thing in my company where, so I think kind of thinking through this on a very visceral level, right? On one hand, you have the people who say you vibe coders don't appreciate the amount of work that goes into a a And, like, yeah, you think you can rip out Salesforce, so did the 30 entrepreneurs before you.
Right? Like like, you know, you classically underestimate the things that you don't deeply know and, and, and, you know, you keep talking to the audience is not you. At the same time, like we have never been able to build software so easily and customize software so easily.
And like, yeah, you're not gonna use 90% of the things that Salesforce. So like, what's the typical, what have you done internally? So we have the main SaaS that we do for event management and sponsor management.
And we pay 200 ks a year for that. Not, not huge, but like chunky for, for, for my, my scale. And like, yeah, I could probably spend 2,000 and build like a custom version of the, the, the trick has been dealing with my, the rest of my team and getting them on board because I'm the most cynical person on my team, but like, I can't make that decision myself.
Right? Like, I think in the same way I've been telling with other CEOs, team leaders as well, it's like, well, you can be super quad pilled. You can be super LM psychosis, and that you think that's okay, but you, like, you have to bring your team with you.
And I think like there the sort of widening disparity in LM psychosis in companies is causing real rifts because on one on one hand, the people who are less AI native are not getting with the picture. They're not they're actually, like, behind. They're actually not waking up to the fact that, like, you'd everything you think is necessary is not actually that necessary.
And in fact, it's that it would be better of you if you just, like, held your nose and went in and went came out the other side only talking to agents in natural language. And, like, your life would actually be better, and you just you're just, like, close minded. There's that perspective.
The other perspective is, oh, you vibequoter, you you did this in a weekend, and you got the 80% solution, and now the rest of your employees have to pick up the rest of your shit. Right. That you you thought you you're you're such hot, amazing at, but, like, actually, you didn't figure it out.
And, like, actually, LLMs are still useless at this and blah blah blah. So, like, I think there's this huge debate going on in every company right now. And, like, you know, I have a small microcosm of it, but, like, yeah, it it's making me hesitate to to pull the trigger, but like, I will at some point.
It's like, maybe I put it off for one year, but not like five. But like so so like SaaS is definitely getting squeezed. It does make me wonder, like, I I do think that there's an opportunity for a more AI native, system of record thing that is not just Postgres, or not just MongoDB, although both are very good.
Maybe it's like a convex or, like, people Yeah. Bring a convex a lot. I don't know.
Like, like, I I just feel like the sort of quote unquote Firebase of of AI apps isn't really a thing yet, beyond what we have, which which is fine. It's it's it's just we could probably start in a more sort of rapid iteration cycle first before scaling up to, like, a Postgres or MongoDB, which are more sort of old tech. I was at a dinner with, Mike Krieger, the CPO of Anthropic, and and he we're just kinda going around the room going, like, what are people most worried about?
Yeah. And, for me, I instead of security, I brought up biosafety.
Yeah. Classic.
Actually, like, I said it was cliche and classic, and the rest of the table were were like, what do you mean someone sitting at home can manufacture a virus that wipes out half of humanity? So it's like the OG Jeffrey Hinton, like, this is why you should be scared. I'm like, yeah.
Like, read the, you know, risk reports. Like, this is, like, the thing. I think and Mike was just sitting there knowing he was sitting on Mythos and going, like, actually, it's security.
And I think, like, I think the there's there's part of it is very good marketing, like, good. Like, I would actually advise and topic to tune down the marketing because, also, it's it's just a very good model, you don't have to make so many marketing claims around it. At the same time, it is not really a private model if you give it to 40 companies, each of whom have, like, 10,000 employees or whatever.
Right? It's not it's not private. It's it's like there's bad actors in there.
Yeah. Hopefully hopefully not as as bad as releasing it widely, but, no. I mean, it's an interesting, you know, it's an interesting case study for how mean, many model releases might I mean, you know, this might be the first model release that looks like the rest of them from from now on.
Right?
you know, restrict access, bundle, product with model maybe, whereas, OpenAI has definitely been a lot more sort of philosophically aligned on, like, we will just enable access everywhere and we don't know what you what will come out of it. Right? Right.
Though, I mean, this current moment, obviously, the cynical take is also it just ties to the amount of compute that both companies have. Right. Right.
Yeah. I think I think that's true. I I do think, like, the the this is the the the scale and the dawn of, like, larger than 10,000,000,000,000 parameter models is very interesting.
I don't think it I think it's a temporary phenomenon because we have much larger compute clusters coming online for everyone over the next, like, three, five years. And this is, like, already written in the cards. Yeah.
So to the extent that, like, you know, will we have rationing of models, above 10,000,000,000,000, in, like, two years? I don't think so. I think everyone will have that.
Rationing of the next phase. Right. Right.
But like that's as it should be almost like, my, my classic example, which I, this is just me theorizing, not anything confirmed by Google. When Google announced Gemini, they actually announced three sizes, which was flash pro ultra. They never released ultra.
They only have pro and flash.
which like, yeah, I mean, I actually think that's as it should be for any lab that they do that. Yeah. Just because those are the models that people actually wanna end up using, and it's just like cost per head.
Yeah.
It's cost. It's not the want, it's just the cost. I do think like it is interesting that for a while, I was I was considering the theory that models capped out at $22,000,000,000,000, and I think that's proving to be wrong.
And, well, then if I'm wrong, how wrong how wrong am I? Do we do 200,000,000,000,000? Do we do 2 quadrillion or whatever?
And I don't think we have the straight answer to that, but, like, it's interesting that we are continuing to scale number per ams when everyone kind of, like, can see that we're not going to get, like, the next thousand or 1,000,000 x from this paradigm. So, like, the others, like, the alias of the world are working on other model architecture improvements. We need a different scaling law, I guess, because, like, we're I I feel like people already feel like we're tats out on this.
Like, the the end the end state of this is we turn most of the world into data centers. And, like, I don't know. I don't know if we want that.
Yeah. I mean, if the if if if the return of intelligence are there, maybe, maybe not so bad.
that, like, is wrangling people's sensibilities right now, especially in terms of, like, context lengths. My classic quote is that context length is like the slowest scaling factor in in LMs. Yeah.
We, like, we took maybe three years to go from, like, 4,000 context length to a million, and that's about it. Like, Gemini has had a million token context length for two years now, and no one's using it. Like so, like, yeah, it's memory memory is probably gonna be the the biggest limiting constraint on all these things.
Yep. Certainly seems that way. Guess I'm curious over the last year since you recorded last, like, what's one thing you've changed your mind on?
I feel like I was kind of bearish on open models, like, last year, in a sense of, like, I I had just done the podcast with Ankur Goyal Yeah. Of BrainTrust, where he and he I mean, you know, he has a good cross section of all the top AI companies, and he says market share of open source is 5% and going down. I think that's changed.
I think it's going up.
Even if Even though the capability gap does seem to be increasing,
depending on the time. It's hard to tell. It's really hard to tell because like, okay, for listeners, capability gap increasing is like on public benchmarks.
And let's say you're comparing Methos versus like, I don't know, GPTOSS or like GLN 5.1. And, it's, it was really hard to tell because even if they were closing, you will also not believe that they were closing that much because it's very easy to gain the benchmarks.
So you just don't really, really know. All you know is like there's somewhat objective open router stats on like what people choose in a free market and people do choose some of these open models in significant volume, except that a lot of them are heavily discounted. So you need to kind of like price adjust, these things.
So even if even if that were true, which I I'm not sure. Like, I I feel like the number is just up now instead of down. I think the separation between what the top tier agent labs are doing versus the average startup in AI or the average GPT wrapper is significant enough that you should not worry about the the the sort of mean industry number, and you should cohort things into like, here's the median, here's like the bottom 80% and here's the top 20%.
And top 20% acts very differently than the bottom 80 And so top 20% is, which is what I all I care about is definitely going towards more open models. The fireworks and the together is a crushing. And so all the fine tuners.
Right? So like, I think maybe last time we even said things like fine tuning as a service doesn't work.
Well, now it's gonna work. It's a derivative of the open market open models market. Well, also in the workload scaling to the point where people care about cost and speed, you know, more and more.
And that, like, you know, moving from just pure use case discovery of, what can these models do to, okay, we know what they can do at scale. Now let's do them cheaper and faster. Yeah.
Yeah.
change, I I think is probably the most significant in in my mind. And like, I I always like to do the mental math of like, this is what I think about, scheduling a learning rate. Like, when you've been wrong once, what else were you wrong on?
And I I'm kind of working through it. I I to me, the the the other thing was the coding one, which obviously I I have now come full three sixty on. But I think, like, people are not appreciating dark factories enough, which I don't know if you've discussed in the pod yet.
No. Ever. And so this is a kind of a strong DM slash Simon Willisen term.
The the general idea is, okay. There's different levels of AI coding psychosis you can have. The the very first level, which I I by way, I encountered first incognition five months ago was zero, human written code.
Yep. Right? Which, like, seems like a reasonable thing now was less reasonable five months ago.
The next frontier that sounds as crazy today as it as as zero coding was in in the past is zero human review. Yeah. Like, you just just check it in without even reviewing it.
And very few people are doing that, but Opening Eyes is is exploring this. And I I feel like it's it's definitely the only scalable way to do this, which it just means that you have to just kinda like flip the SDLC or change large amounts of what what you normally do, which is probably things you should have done anyway, more testing, know, more automated verification or whatever. But like, that is a frontier at which, like, when you have unlocked that in your companies, you are just gonna produce much more quantity of software than than you've ever had.
It's gonna be like so much so disposable, so cheap that you can probably innovate in quality a lot as well. Like that that quantity helps you get to quality. Yeah.
people associate more quantity with slop. Right. No.
the best sign of of of productivity inefficiency, but going forward. Yeah. You but you still get rewarded for it.
So you're like, fuck it. Whatever. But, like, I I I think, like, the the the people who are who are doing well who do well who do most well in 2026 are not the cynics who go like, oh, that's just slop.
I'm not gonna participate in that. They're like, okay. Like, this is happening with with or without me.
Let's bend this the right way. Yeah. No.
I love that.
kind of overall quality, certainly for latency and cost, it always made sense to me, but for overall quality, God, you just get that for free in the models three, six months later. I think what I'm starting to change my tune on a little bit is hearing all these app companies talk about, we build stuff and then we throw it out three months later as the models improve. You're like, okay, well then what you're doing for capability improvement is just another version of that.
I still don't think that your RL or post train is gonna make you have a better model for years and years to come, but maybe I I think you still have to be pretty rigorous on, that the single best thing you can do to solve a customer problem? And oftentimes, it's literally just like, now add more data and feed more data even via connectors to these models, or I don't know, do some clever engineering on the back end or whatever it is, but if the single best thing you can do for that three month time period to improve your customers' outcomes is, you know, post training in some way that, like, really improves the output of a model, even if you throw it out three months later because the general models get up there, it still might have been worth doing. And so I think I'm like more open to You throw out the results, but you don't throw out the raw data.
Totally. Like so Right. Then you just run it again.
And so basically, there's some obviously, at the level of cost of like $10,000,000, maybe that's too much, but there's some level of cost where No. It's actually 10,000,000. No.
Of course it's not. You know? Yeah.
at which it's the equivalent of just, like, staffing four engineers to go build something for three months. Yeah. So the other thing I really as for for listeners, I'm just gonna leave some some droplets of info.
Look into like, the the long trajectory, the synthetic rubrics work that people are doing is very important, including, something that's called doctor gRPO. I'll just I'll just leave those key search terms in there. I I think it what it means is that RL is going much more multi turn than people think, And that means that you can customize the models in way more specific dimensions than traditional, let's call it SFT or, you know, like, sort of shallow RL that was done in a year ago.
So like hundreds of turns. Yeah. And I think that leads you down a path of like complete domain specificity.
What else like are you, you know, of these like unanswered questions in AI today, you like looking for, you know, in the next year, are you
paying close attention to? I have a few thesis for like, is the sort of next frontier? One is memory, which memory and personalization we talked about.
The other is really world models, which we've done a small little series on from Fei Fei Li to even Moon Lake and Geno Intuition. And there's a lot of debate as to like the relative importance of this. I think a lot of it manifests as like three d static worlds that you kind of inhabit for a little bit and you walk around and they're like, cool, but like, how does this help me with my B2B That's like all the hype now is robotics, right?
Yeah. And there's obviously a correlation between world models and embodied vision and experiences, which leads to robotics. But I think world models is very interesting in just in improving intelligence itself from the next from the next token prediction paradigm.
And so I think people are kind of testing their edges around that. One of our top articles this year so far has been on adversarial world models. I I do think, like, if you don't do anything else, just read Fei Fei Li's essay on spatial intelligence, on why LLMs don't need don't have it.
And she is she may be she may not have the solution yet, but she has the right problems And so everyone else is trying to solve that problem statement in their own way, and let's see who wins. But, like, I I don't think it does you any favor to equate world models to robotics or world models to gaming or some kind of, like or, like, the current manifestations because what is at stake is a much more important conception of intelligence than just answering questions. It is does does does does the AI understand what a table is?
Like, what what matter is, what physics is? It's almost like for for those who are movie fans, it's like Google hunting where Matt Damon, like, knows everything because he read it in a book, but he's never Great great scene with Robin Williams. With Robin Williams.
And I I I look at that scene and I go, like, that's exactly the the the difference between, like, a very intelligent LLM who knows everything but hasn't experienced anything. Wow.
That's an awesome note to end on. That's a great have you used that antidote? That was great.
Yeah. So so one thing I've done with Lanespace is I moved to, like, adding daily write ups. Yeah.
And so the one one of the times I was doing this daily write up, That's great one. I love Also, it's been a ton of fun. Thanks so much for for coming for coming I'm Jacob Efron, and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world.
As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends.
It's really what ultimately makes this whole thing work. And so please consider doing that, and thank you so much for your support and listening. We'll see you next episode.
Shared via Hopper