Baseten CEO Tuhin Srivastava discusses the explosive growth in AI inference demand and Baseten's 30x growth, arguing that inference is becoming the strategic "last market." He explains why the application layer will persist due to unique user signals and the ability to post-train specialized models, despite severe GPU capacity constraints. The conversation also covers the geopolitical implications of open-source models, Baseten's operational challenges at scale, and the impact of Jevons Paradox on AI consumption.
Hi, listeners. Today, Elad and I are here with Tuhin Srivastava, the founder and CEO of Base10, inference cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source and perhaps multi chip future, and what 30x scale in a year looks like.
Tuhin, welcome back. Hi. Good to see you.
Thanks for having me. All right. You are in one of the craziest markets, AI inference.
It's very important. There's a lot going on. You guys have grown 30x over the last year.
And I think I can say you're expecting to do more than a billion dollars in revenue this year. Mhmm. What's going on?
Tell us about scale.
Yeah. No. It's been it's been nuts.
I I think what's happened over the last, honestly, twenty four months, but just kinda keeps getting bigger and bigger, is that I think everyone is realizing that you can put AI everywhere. You have all these great options available from closed source to open source models. The open source models have crossed some sort of chasm in terms of their baseline capability.
And then I think RRL techniques and post training for specialized models has become mainstream enough and, you know, there's enough examples of it work of it working. The customer's realizing they can, you know, kind of own their inference more and more. And what that's meant for us is more, you know, the long tail of models coming true, customers in housing a lot of that intelligence themselves.
As the application layer just gets, you know, bigger and bigger and bigger and that's growing, we just someone index on that and we've been around to be able to collect the demand.
There's an existential question in here that I think everybody is continually asking of does the independent application layer get to exist at all versus the labs?
have to believe this. Why do you believe it? Yeah.
Look. I I I think it'd be it'd be a sad thing if it didn't exist in general, and I think that's, like, my but, you know, sadness is fine. The I'm sad all the time.
Oh, yeah. Sadness is fine. But, like, that that that's not the reason why I think the application layer will exist.
I think the application layer will exist for a number of reasons. One is because, you know, I think this idea that what is valuable to a company is, you know, the user signal that they can gather that only they can gather. And to the extent that that is encoded in a model, I think a lot of their business will be at risk, but to the to the extent that it is encoded in workflows, that is where they will be able to develop mode.
So a good I think a good example of that is, say, a company like a Bridge where the clinicians edits off the notes and what they do with those notes after the fact. And the thing that happens in inside the EMR three steps down, and that becomes a workflow that only Can you explain what a bridge does? Sorry.
A bridge a bridge is a ambient scribe that is used by physicians in, like, you know, almost all hospitals in The US. I think Lad's an investor. Great Shiv's amazing.
Great great company. Great team. Great product.
And, you know, they they've basically, you know, got this very, very deep integration into into hospitals and to clinician workflows. And my argument would be here is that actually, you know, it's very, very hard for a frontier model company to be able to eat up where that could they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post train models on that reward signal and and start to get long long horizonogenic models running that.
And I think to the extent that that is possible and that signal is differentiated and unique and and is somewhat rare to to get access to, there will be an application layer. And I think, you know, it's like support company is another example of that where, you know, a support a support task isn't one shotted.
Usually, at a company like BS10, when a ticket comes in, there's like, what, like, one, two, ten, twenty actions that get taken, and that is where, you know, someone can develop a specialized model. So there's almost two versions of this then. There's new companies like Abridge or Decagon or some of these other things that you mentioned that are doing these new types of applications that are using AI and they sell it to customers.
The other is enterprises building things in house or building their own models.
these new application companies versus enterprises adopting AI? And how do you think that looks in a couple years? Yeah.
I I think that's a that's a you we did I think you asked me the same question two years ago Oh, yeah. On on the file of the Internet. I have to be repetitive.
It it is crazy At inconsistent. The answer is just that it's crazy that the answer is still, I think I think if you look by inference count, it'd be 99% the fall. Yeah.
And that is that kind of represents the scope of the opportunity here is that the majority of the market hasn't come online and and added AI into the system. Enterprise adoption is well ahead of us, and I think that's one of the very exciting things about AI. Yeah.
There's just so much still to come, and people are underestimating that, I think. 100. And I but but what's cool is that we're seeing the transition happen.
Right? Before, it was like, hey. Are they are they using AI tools?
Mhmm. I don't think that was immediately obvious two years ago, and think it's obvious now that, yes, they are. Are they using closed source model APIs?
I think they're starting to get there. I think once you do that and then you kind of see what is possible, then comes the whole custom model adoption. I think that is all that is ahead of us today.
So if the majority of your customer base today is, as you described, the the former, like, application companies, AI natives, the fast growing mean, of them are at considerable scale now, like the abridge, cursor Yeah. Open evidences. Open evidences of the world.
What do they teach you? What does that push the company to do? How do you think about serving them versus evolving for the enterprise?
Yeah.
I I think, firstly, like, you just learn a lot by building with the company's greatest scale, doing the most interesting things. We we think of it two ways. Like, I think there's, like, the the the most obvious way, which is just build for the highest scale, you know, most the the customers that will push you the most from technologically and everything kinda will fall into play.
Think the Stripe evolution as a company showed that was like Stripe now, like, serves like so many enterprises. But twelve years ago, that wasn't the case. But they just built for the frontier and kind of went with them.
I think the second way we think about that this is to just think about building for companies that are serving enterprises. So, yes, we don't serve the enterprise, but our customers serve enterprises. Ibrutinib, Pfizer, OpenEvenus, Decagon, all these Ryta, Gamma, all these Calc Clay, all all these companies serve enterprises en masse.
And what we actually get is, like, a translation of the requirements from them, which is, you know, they're like, hey. We need this sort of data retention. We need this type these web models need to be deployed.
This is the types of GPUs or the latencies they're okay with. This is the model requirements from, like, a transparency perspective that they care about. And so I think that is actually the more nuanced answer is that if you listen to what their needs are, we actually get a full translation of what the enterprise will require.
and Open Evidence, we're probably pretty well suited to go serve the health care system given that they are selling and latent health given that they are selling to them. How much of a shift are you seeing in terms of the types of open source models that are being used? And so I think we've seen an evolution where two, three years ago, I think the main thing was kind of Mistral and then a few other things, and then Meta kind of came along with LAMA, and then it kind of really shifted in terms of the misperformant models or of Chinese origin in different ways.
Do you see that sort of mix reflected in terms of what's being used by your customers? Yeah. I think customers, at least the customers we are serving, are very and these are, like, the fastest growing AI companies in the world that are very forward thinking.
They they wanna use the best model, and they they are optimizing.
I think there is there there are a there's a subset of tasks, which I think is small today, where people really start to start with cost. Mhmm. But everyone comes from capability first because that's really where the economic growth is being unlocked, where the value is being delivered, and then they optimize.
And I think that's, like, actually been, you know and so with that in mind, you name everything from GPTOSS all the way to Moonshot Mall to DeepSeeks to Canopy or Orpheus, which is, like, really good text to speech models, customers generally wanna use whatever's at the frontier. And and I think the the difference has just been I think we have a lot more visibility into how to run these and how to run these really well. And secondly, that they're good now.
or is there something embedded in the models or Trojan horses or other things? A, do you think there's any real concern there? And B, people often talk about how there should be US counterweights to this.
From a geopolitical perspective, do think that's something that's legitimate or something we should be worried about?
Yeah. Sort of origins of these models versus their uses? Yeah.
Look. I I I think these these models, firstly, are fantastic. They're amazing.
We work with these teams. They're truly awesome. I'd say, look.
I I don't it it is hard for me it is hard for me to see and I I could be wrong, but, like, you know, if if I if I network bound these models that they're not magically, you know, gonna be able to cross those network boundaries Mhmm. And to data center. And, you know, I I don't and we I've never seen any real evidence except from some very early models that I think people picked up on very quickly that there is some agenda or bias built into these.
I do think that to some some extent is I I think there is importance to The US that we develop our own models. I think that that would be a massive loss if that there are five companies, you know, five different labs in China that are creating open source models, and we're struggling to get one set up. So it's necessary.
I also think it's inevitable. And, you know, like the sea the deep seak moment a year ago, I remember someone saying to me, I thought it was very well said, which is like, and the world's changed a lot, but they said, hey. You know, we should kinda just forget Mhmm.
That this is a Chinese model. We should just act like this came from Mhmm. From Meta and and build and build with that in mind.
Mhmm. It's like, know, I I think you're kinda missing the forest from the trees. Like, there's two there's two scenarios.
Right? Either America does not ever come up with good open source models Mhmm. And there's probably a fundamental problem there, or we will get there, and we need to be ready for that world.
Yeah. That makes sense.
like you, I think it's very important for The US to have a strong open source footprint here. At least for now, it looks like effectively the Chinese government is subsidizing at least a large subset of these models, and that subsidy or surplus is effectively just being passed on to US enterprises who are adopting these models. In other words, it's a way for the Chinese government to effectively subsidize US enterprise in an indirect manner, and I think that's a little bit lost right now.
But it's always interesting to weigh that against some of the other concerns that are raised. I appreciate your comments on this.
like, if it is fun like, I I think if you think of the economics here, which is deep deep sea by most, deep seq's a very good model. Mhmm. You know?
And, like, you can argue whether it's at the absolute frontier or not, but, like, let's let's go back three months and it's there. And so think about every and we're doing a whole lot of things three months ago. Yeah.
Yeah. And so let's just think about it that well. You know, if it you could run deepseq, probably 20% of the cost of running open andropic models in production with comparable better latency, probably better reliability.
If we don't have access to that intelligence Mhmm. In that form, I think it's just a massive loss. And as a country, we won't be able to innovate as fast because, like, the cost of intelligence going down in control of intelligence, what we have seen just means more intelligence.
Yeah. Intelligence being embedded in more places. Yeah.
Google, etcetera. Yeah.
Has been- Actually, maybe you can just characterize like workload a little bit. Like how of tokens being served on base 10, like how many of them are from custom models of some kind versus like vanilla open source today?
custom.
yeah. Like, it is So, like, 95% plus? 9095%.
Like and I think that's really cool, to be honest. I mean, look, we have we have two businesses. We have we have three business.
We have we have three businesses. Yeah. Should we help you count?
No. No. So we have, like, dedicated dedicated inference, which is basically custom model inference.
Your SLA is your SLA. Then we have shared inference with a shared inference endpoint, shared SLAs, and then we have a training business. I'd say 95% of the tokens today are on the the first business.
And almost all of them, there's probably a yeah. For almost all of them, the customer is making some modifications to the model with with their own data specialized for the use case. And I think what's even more important is they might be compiling it in different ways.
No one is just running the vanilla open source weights. Like, you you might be customizing it for quality, but you also might be customizing it for performance.
You made an acquisition of a research team a few months ago. You've mentioned post training customization. What was the rationale behind the acquisition?
What is that team doing today? Yeah.
So the the rationale around the acquisition was, you know, we we are infrastructure and product people. We are product people and now are really good infrastructure people. And the and we didn't have much of a research capability ourselves.
And and what we saw was the market moving heavily and heavily, like, that we could accelerate the market itself with post training resources of either product types or onto even just as resources for that market. So Parzed was a company that was a base 10 customer. So they were post training models and running them on base on base 10.
And I think what they realized was that they would eventually need to become an inference company. And what we realized was, like, hey. We we really needed that expertise because it represents a way for us to get closer to the customer earlier and be able to support them all.
And it just made sense pairing them together. And just as I said in the opening statement here, which is, you know, as more and more post training models have come up, we've realized that the demand for people to either for software loops to do post training or for post training expertise is very high, and we're really, really investing in that. There are also a bunch of Australians.
You know, I like to think that we had a bit of alpha there. But, yeah, that's been fantastic. They're working with all sorts of customers.
And it's also very interesting when you start you know, we were doing a lot of research on the performance side and less so on the post training side. It's interesting as we've started to do a lot more research on the post training side, you start to see how linked inference and post training are. And, like, you know, even even when you think about stuff like quantization and when you should do that and, like, you know, how how training how how you train the model affects how you need to quantize for inference and how paired these problems are Mhmm.
Has become, like, very apparent. And more and more, rely on the post training inference are kind of both sides of the same problem. So because inference will ideally will beget more post training, where inference creates data, you do eval, you can now post train post train on the on that reward function that you that you found with those eval and and hopefully just set up an entire look.
Ant and OpenAI, Sam, Greg, etcetera, have said in recent months that, like, inference is super strategic, inference talent is strategic, capacity is strategic. So between that and post training, these are very difficult to gather, like, capabilities. Yep.
I imagine that lots of your customers go to you guys for advice on like how to do this progression of moving to custom models. Like, what do you tell people about the life cycle and when they should invest in that? Yeah.
I I think it's, hey.
with the best in class model that you have something worth optimizing. And and I think, you know, a lot of you know, if a customer comes to us, you you there was was that meme which was like it was like two years ago. It feels like no GPUs, pre product market fit.
It's like no post training pre product market fit is what I know. Yeah. Yeah.
Yeah. It is what I'd say. So people that you're working with here are very very at scale first.
Yeah. They they they have a user signal that they know how to optimize, and they've shown that they can, you know, they can serve customer value and that value and that they have something special around that value. And once you have that value, it's like, okay.
Now how can I do that better, faster, and cheaper? With the idea being that, hey. If you need to be very good at customer support, you can you maybe don't need to be that good at coding and that a specialized model might be a better fit for that problem, and you can do it better, faster, cheaper.
What about the capacity side?
started with unifying capacity across all the clouds and neo clouds. How do you think about this when everybody keeps talking about a supply crunch and a multiyear supply crunch?
I think, you know, there's so much narrative around the supply crunch. And no matter like, as much as we hear about it, I don't think people realize how bad it really is. Like, there is, you know, there is very, very little slack compute available.
You know, we we run pretty large clusters ourselves, and we run them at, like, uncomfortably high utilization. You know, we when I'm saying we're, mid nineties utilization most of the time. There is we have made we we have we sit in 18 different clouds now.
We have 90 clusters around the world across 18 different clouds. And, like, you know, initially, we started we, like, built this technology to be able to, like, kinda create one runtime fabric that spans all these different clouds and try to abstract that away from our customers as a way to think about reliability, latency, failover, all these things that we think are gonna be very important for very mission critical use cases. That same technology, just our ability to get compute wherever humanly possible has been really, really helpful in our ability to get supply.
And and what what I mean by that is we can be introduced to a new provider in a different country and have it up and running with the whole base 10 inference stack. As part of the fabric. Yeah.
Part fabric in half a day, half maybe less. Even for and that gives us enormous flexibility. Even for us, it is hard for us to grow.
We have a we have a I think it's yeah. I'll say it. We have a a 4PM standing meeting for the company where we basically, like like, how do we, like, how do we how do we manage capacity for the demand right now?
I think the second part which people don't really the two the the second part that people don't really understand is that there are also a lot of suppliers right now that it's kinda grifty. You know, like, I I I think, you know, they haven't run they haven't run data centers before. You know, they don't understand SLAs for especially for inference.
And so, like, you know, even when there is capacity available, there's a lot of like, there's probably we we run a lot more than this, and we have redundancy, and so it's fine. But if you you know, there's probably like a dozen good clouds, and I'd probably put three or four of them in the gold tier.
run these data centers as well. How far ahead can you actually buy capacity right now?
contract length or actually like, hey. I want this in January 28.
Yeah. Either one. Yeah.
Yeah. I mean, it's more the I want this in January 28, at least I have some visibility into my future supply. Yeah.
You could buy that, but you gotta also remember how quickly the market how quickly the market is moving. And, like, you know, that gets balanced somewhat off, like, the fact that the h 100 is such a great chip. Yeah.
And, like and then set you know, it's crazy. If it's four years, four and a years old, the price is going up still. Yeah.
Maybe it has a useful life of nine years. Yeah. So, you know, that's good.
But at the same time at the same time, you know, yes, you can do that, but, you know, you're making a lot of bets Yeah. As part of that. And then in terms of I think that's the big thing that's changed over the last six months is that the term length that people want has just gone up.
So if you if you wanted a thou a thousand thousand 24 b 2 hundreds Mhmm. Which is, you know, from a good cloud, right now, you're not getting that less than a three to five year contract Mhmm. Right now with a probably a 20 to 30% t c TCV prepay.
pretty significantly. Does that does that impact how you think about going public as a company?
Yeah. I think you'd go sooner. Yeah.
Exactly. Yeah. I I think you need like, I I think the and I think there is demand for that.
But I think, you know, the pull the you know, and also, you know, one of our one of the one of the realizations that we had recently and we're we're software people. And so we don't we don't think like this all the time is that, you know, our business has, like, very interesting working capital requirements. Like, we don't and and I and I think, you know, even and that as a result of that, it has very interesting financing Yeah.
Yeah. Requirement.
debt. There's also other things we could do in terms of debt or other structures. Yeah.
Yeah, I've learned a lot about debt recently.
Given the supply crunch, in France being one of, you know, the top couple markets you've been going after, you have plenty of people who understand this problem and therefore, you know, some competition. How do you think about, like, what are the factors that create a dominant player here or a winning player? Is it, as you mentioned, cost of capital?
Is it access to supply? Is it software? Is it demand?
Yeah. Is it just being excellent at everything?
Yeah. Look. I I I think what's so interesting about inference is g Is it operations?
I guess, such as cloud. Think so. Yeah.
I think, like, GPUs as a service is not sticky. I think that's been seen. Like, customers generally just see that as as commodity.
Imprint with the software layer included is incredibly sticky. You know, like, just just like, you know, none of our top 30 customers have ever churned. You know, we're talking like 400% annual NDR Mhmm.
Around our business. And so it's, like, very it's it's very, very sticky. So I think that software layer is very important.
The optimist in me is like, oh, there's so much value in the software, and I we will build the best software layer for inference that exists. Think, you know, as I think it's becoming clear now, access to inference computers is a strategic advantage.
running inference. Yeah. Yeah.
And in a world of constrained compute, the number one thing to own is compute. Yeah. And so, you know, just owning it in and of itself is an asset, and I think people underappreciate that.
Yeah. You make a good hot chocolate without milk.
you're vegan. Unless you're a vegan. No one wants a vegan inference.
Yeah.
Well, I got to ask you, people might want alternative milk, right? So, okay, when you, the H100 is a great chip, people, you know, want a B200, they want GB200, they want, of course, tons and tons of Nvidia. When you think about making a bet, you know, several years in the future, do you believe that there's a, like, multi chip world?
Like, what do you what do you think happens from a compute perspective on the chip side? Yeah.
I think, you know, like, diversification everywhere is a Mhmm. Same way I wanna water many models. I think, you know, we wanna water many most things.
And I think You'd be sad if it didn't happen. Yeah. And I and I think everyone would be sad.
I I will say to some extent, which is yeah. And I think there will be inference specific chips. I think you have, like, decode specific chips.
Think and we're we're looking at that. And NVIDIA said this. Yeah.
Yeah. I mean, that was that was a whole Croc LP thing. It's like, you know, I I think I think that is very straightforward and and makes sense.
I think people really, really, really underestimate supply chain stuff with NVIDIA, like how good they are at that CUDA, how good CUDA is, the developer ecosystem around it. And, you know, we it the ability like, to me, like, one of the most important things as an infrastructure company in this moment is how fast you can move. And you can move fastest with NVIDIA today.
And I think that is the reality. Like, it just, like, given the scale that they operate at given the scale that they operate at, it's it's hard to it's hard to see it's hard to see the the and I'm not saying it won't happen, like, the short term like, in the next couple years, how anyone's gonna be able to compete with that. Especially with, you know, so much of the other the other players.
Like, what you need to be able to compete here is the ecosystem to form around you. And if you tie up all your supply with one buyer, which, you know, a bunch of the other chip providers have done, it's actually hard for that ecosystem to form. You know?
what do you think is like happening with the actual workloads that you have to go invest in? Right? Like obviously code agents and long horizon agents over time have become a big deal.
People talk a lot more about CPU compute, video inference is different. Yep.
I don't know if it's that sandbox. Just like what what's important for you guys to invest in now? Yeah.
Look, I I I think there's for for us, all the runtime stuff is obviously very important. And what that means is like what chips we run on, how we run, what kind of workloads we support. Like, do we get very good at diffusion transformers?
Yes. Coding agents need sandboxes. We should go build sandboxes.
There's all sorts of new speculation techniques to get faster inference. We need to do that. Even stuff like KBCatch aware routing and, you know, that stuff's a bit old now, but, like, getting continuing to be very good at that and somewhat disentangling pre fill and decode and starting to treat them as separate problems.
Think that's, you know, something we are very focused on, and we're seeing massive gains there. That's at the runtime level. I'd say, you know, beyond that, you know, everything we think about is how to create more of that loop between inference, post training, because we think that just begets more inference.
And so, like, we we will build a partner in almost everything there. So, you know, we're gonna work with, you know, the best Evalis company in the world to make sure that's very well integrated, like Braintrust into and around Base 10, you know, we will partner with all on the sandbox society, build build the best sandbox experience that will exist. And then we'll create the the best training APIs to make it so continual learning becomes somewhat of a solved problem.
It's not just like a discrete thing. That's, I think, the core base 10 product thesis. It's like how do we build that loop?
Then everything else out around that becomes how do we make sure that we can do everything we can to ensure that gets as big as possible. That's access to compute. That's on infrastructure.
Make sure we can get compute anywhere. Make sure we have access to our own compute. And then I think it's all the primitives that come off of that just that just become incredibly, like, margin accretive both for us and our customers, which is, you know, stuff like, you know, sandboxes and, like, the async batch inference, like, how do we drive utilization by having a first class batch inference experience.
To me, this is, what an inference cloud looks like. It's that you are very good at inference, and then you you start to do all the things tangential or that loop into inference and partner and where necessary and build where necessary. But we really do wanna own like, start with that core inference story and then go down to unblock supply, accrete margin, and go up the stack to unlock value.
What would surprise people about some of the issues you discover only at scale? I'll give you an example. I was surprised when you guys ran into scale limitations, like fundamental limitations with some of the hyperscaler products that you were consuming.
Yeah.
supporting infinite scale. Yeah. I mean, I I think you just and, like, again, like, I think very, very large companies that run services of big scale is probably the same stuff.
Mhmm. Is that all the edge cases just become actually experience them. You experience them.
And, like, you know and you I'll give you a few examples here. Like, you you see you know, you start seeing you know, yesterday, we had for the first time ever, we saw some kernel panic. And that only happened because some fluent bit worker was creating too many logs and and the scale was too big and it was all into one node.
And it was happening two two two terms at the same time by two different workers. So you see all, like, the systems level and current level problems. But then you start to see I think the the craziest stuff is that you start to see with with LLMs that these runtimes are pretty immature.
Even how we use KV cache is, you know you know, probably a little less sophisticated than most people see than most people see. We we we are starting to see the the limitations of the current and the next set of primitives that need to be built from a scale security performance perspective. But I I think it's really at the runtime level and the systems level and then but the edge cases are, I'd say, a lot more systems level than they are LLM specific.
What are the things that keep you up at night? Capacity.
Think I I think, you know Quick answer. Yeah. Yeah.
I think capacity. I I think the other one is probably just this market's so big, and it's so like, it it represents a moment when you should be as aggressive as possible. And, you know, really, you know, we've we've grown a ton obviously over the last twelve months, the last few months, but the answer is always just go, you know, go bigger, go faster.
And I think that's really, really fun. It's also a little exhausting, and it's also like we are we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get. But I but I think the big one is compute.
I think, like, there's no world in which there's enough compute to, you know, get the amount of the amount of value that we wanna get out of LLMs in the next five to ten years. We have to invent a lot of new stuff. Yeah.
Maybe if we just talk a little bit about what you're learning scaling. You know, 30x is like an aggressive thing to go through as a company. You've brought in a lot of really amazing talent like Danny and Samir and Stephen Day, folks on both the technical and the go to market side.
Like, what do you what do you think is working about how you are recruiting and scaling, or or what's your philosophy on that?
like, until, I don't know, twelve to eighteen months ago. I remember I went on a walk with a lot actually, and a lot of it was like, you just need leaders. And and, like, it's actually, like, so contrary to everything.
You know? As engineers, you're like, oh. It's all overhead.
Everything overhead. Is overhead.
I and I You once told me, I think, that you you didn't you're like, hey, Sarah. Sarah, what about we just have engineers instead of salespeople? Yeah.
Yeah. Yeah. That.
It'd a Everybody learns it. Everyone learns the same. And we're we're all done that.
But I remember, like, you know, you you said it so clearly at the time a lot, and I and I think that's what we've noticed. We're just, like, actually having a leadership team that you can trust that you can trust is is is so important. I I think the the two zero three thing that I'll say is, like, you want people where you can give them whole problems.
And to like, you know, if you feel like you are micromanaging, if you feel like you need if you feel like, you know, you have to be involved in everything, I think that's a bit of a cop out as a founder because you're just like, I just need be involved in everything. It's like, no. You you probably don't have the right people.
I think the second thing is be very, very clear what you're optimizing for. Because I think when you're very, very clear what you're optimizing for, the people on and, like, if it's something generic, like, we want the smartest, hardworking people, like, you can't do much with that. Like, with us, what we cared about was, hey.
Actually, we don't care about a lot of people who have done this before. We care about first print people who think from first principled print first principles. Work has to be a high priority, but they also have to be very kind and nice and, you know, care about the collaborative environment.
We don't have a hero culture. You know, very low ego. And, you know, if you need if you need a manager, like, it's probably not it's probably not the right place to be.
But I think once you have when you have that clear rubric, the the people become very apparent that will fit into it, the people that don't fit into it also become very apparent. And I think what's more like, we've hired amazing people like you mentioned, but I think what's a lot more interesting is, like, I think we've we haven't had a ton of, like, turnover there unnecessarily. Like, peep people tend to work because we because we have a very we are very clear on what we want.
It took us a while to get there though.
culture? You know, we were talking to Alyssa and Henry about this, and she's like, well, the hard thing about cloud is actually just operations. I slept with a pager under my pillow for a decade.
I don't think I've seen you detached from your Slack channel. Yeah. My phone is buzzing right now.
Yeah. I'm kidding. That's just not a step one.
Yeah. I'm anxious. So And you've been concerned before.
Like, do people get it? Like, know, what is distinctive about that?
like, one, I think if you've worked at an infrastructure company like, we were once in a meeting with a bunch of AWS execs, and this was, you know, like, very senior AWS folks. All their pages went off multiple times during our forty five minute meeting. You know?
Like, it's a I I I think, like, it's it's it's very much just a cultural thing. But, yeah, like, I I don't you know, our, like, inference can go down. And, like, you know, we you know, you learn to like, you know, what's this?
Like, I think, Amir, my cofounder, when his pager goes off, his seven year old said, is that a p zero? Oh, that is that is that a p zero? And so, you know, I I think that is you just have to get used to it.
That's the culture you live in. It just changes the speed, but it also it's you know, it becomes like a, you know, a cultural thing. I I think it's very, very it rejects people that don't fit into it very, very quickly.
Like engineers who avoid patriotism. Yeah. You know, when we when we have p zeroes, we're like, everyone on call.
Like, you know, like, there's been a joke that there may as well be a siren that goes off in the office when when they're interested. So So people have been talking ad nauseam in the AI community about Jevan's paradox Yeah.
Where if you decrease the cost of it's really a question around price elasticity and availability. If you decrease the cost of a good, say intelligence as a good, people actually consume more of it. Yeah.
Like the personal or business ROI of it, the demand for it goes up, Do not you see this? And are you are you working against yourself trying to make these models more efficient? Do people just use them more or less?
Yeah. I I think you think about this from a developer's perspective and a consumer perspective. I think, like, I think consumers just want the best answers and the and the and the best experience that's somewhat governed by, yeah, more intelligence to some extent.
I think when you go to the developers from the developer's perspective, they would insert more intelligence if you make it cheaper. Like, that's yeah. And they will they will they will insert more intelligence anyway.
Mhmm. But if you make it in more cheaper, they'll they'll insert a hell of a lot more intelligence. You see this with agents.
Agents are just longer running now. And I think that's what we have seen with the cost of inference going down, which is, you know, folks are just like, okay. We can we can run this for longer or we can make it do a bit more work, and we'll get to a larger end.
I think, like, if like, compute scales from an inference perspective as well. And, you know, I think we are seeing that with almost all our customers, which is, you know, they either they either start with, like, this is the quality of answer I I need to get to, and this is the amount of inference I need to do to get there, or this is the base level model that I can start with that I can work with to get there. And I think the more we drive down the costs, what they realize is more intelligence just means better using I just want a better answer.
Answers, better experiences, more dollars, more dollars, more revenue. Yeah, I think inference going down just begets more. But it is truly I think we're kind of in a world that it is is the last market.
Right? Even if there's AGI, all that's left is inference. Yeah.
So you do not see in your customers a, this answer is enough and this action is enough dynamic? No. Yeah.
It's gonna keep going for a long time.
do you view all this evolving towards the future? So basically, seems like it's gonna be one of the biggest markets of all times. We have this massive shift where we're moving from software and seats and digitization into actual intelligence, selling units of cognition, selling agentic workflows.
What does this all look like in a couple years? What is your view of this future world?
I think for the for consumers, it's the best possible thing. Right? Like, every everything is somewhat smarter.
You get better care because your doctors have access to better tools. There's all this stuff about there being less software engineers, and I think we just build more software. I think we just build a ton more software.
Like, you know, I see know, we're not slowing down hiring software engineers. We're just building more things.
all those all those good things. It's almost like everybody has their own team for everything. Right?
You have an agent which helps with your doctor. You have an agent that helps you learn stuff. You have an agent that helps you organize your life.
It's concierge. It's a concierge for everything. Yeah.
Concierge is everything for everyone. Yeah.
well, that that's amazing. I think that's great. And I think the the the and education, same thing.
Have concierge education. Like, you get personalized access to everything. I think then you go one step back and how it affects developers.
I think, you know, and and companies, I think if you don't embrace this Mhmm. I think it's the extension moment for a bunch of folks, which is like, you know, everything needs and I don't think that means that, you know, forward design needs figma. I don't think that's a thing.
that user value for those end consumers that we talked Yeah. Very exciting. Thank you so much for joining us today.
Yeah. Thanks, guys.
Find us on Twitter nopriorspod. Subscribe to our YouTube channel if you wanna see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen.
That way you get a new episode every week. And sign up for emails or find transcripts for every episode at nopriors.com.
Shared via Hopper