In this episode, Akshat Bubna, CTO of Modal, discusses the evolution of AI infrastructure tailored for agent experience, highlighting Modal's journey from a runtime platform to a specialized cloud for AI workloads. They explore challenges in scaling inference, sandboxing, speculative decoding, and the unique demands of agent-native primitives in modern AI applications.
We're here with Akshat of Moto, CTO of Moto, together with Vivo. Congrats on your series c. Thank you.
Your party yesterday was amazing. All the photos and all the swag.
We we had a bunch of art installations, which is kind of fun seeing, like, our products on pedestals next to, like, Rodin.
Very nice. Very nice. When you started, it was not the GPU inference company.
I mean, maybe it was in your mind. Take us back to the origin story.
I actually first met Eric, who's the CEO, through an investor. And back then, Eric was already thinking about building a new kind of runtime. And he he got there thinking through why are workflow orchestration products so hard to use?
It's because you have to run them on Kubernetes. Kubernetes is hard to manage.
custom images, and it has a terrible developer experience. And I'll I'll interject for Yep. Peep for listeners who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.
And I I actually came across Eric through a data council because he he did that talk on the the sort of serverless container stack that that you guys did, which is like that was my my first, like, okay. I need to take models very seriously moment. But it was still very unclear.
Like, do I actually need all this for just my data pipelines?
Yeah. I mean, initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. There's a lot of things that get solved by serverless functions.
Like, can do ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. But then we also were thinking about this as, like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would have been the first one, but we were thinking about inference.
Back then, was more classical inference, like computer vision stuff and running XGBoost and whatnot. But we we added GPUs to the product a year before ChatGPT came out. Nice.
We just didn't think it would be that big of a deal. Yeah. Just like add a 100.
Was there any like early key problem that really sparked off why you built it? Yeah. Primarily, it's just none of the tooling that was out there was built for, one, a really great developer experience.
And also there's a general trend of a lot of the workloads that we were seeing were very this is I wish there's a better word for it, but compute heavy. Like, they need one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for like slow scaling and more for like web server use cases. And also, there's just a lot more specialization in, kinds of environments these workloads run-in.
Like, we kind of sometimes need accelerators, sometimes need different kinds of images.
And this is just like a consistent thing that we saw across a lot of companies. That would be the the next step. Be be nice.
I don't know how much is factored into the early story, but I wrote a post when I was at Temporal about infrastructure software defined infrastructure or something like that. The self provisioning Self provisioning. Yeah.
I couldn't even remember my own post. And then you put me on the landing page. Yeah.
We we really like the term, and and so we we stole it. Because you had the insight that everything can just be in decorators next co located with the code. Right?
Yep. Was that a big part of the original story? It's just like a DX layer?
really didn't want people to spend so much time writing YAML. And it seemed like you could really condense the surface area of what you're doing, put it in code so you can actually operate on it, just so you can operate in other code and, like, build stuff that's more expressive and dynamic.
And so, yeah, that that was always a very important part. Then the pushback is, this is a DSL. Yeah.
It's your closed source. I am locked in to Modo.
Yeah. We we we never really got pushback for that, because the nice thing about Modo is you can bring whatever code you have. And, sure, the DSL is at the sort of configuration there for what hardware you're using, how you're scaling things up, but you still own the code.
part of our story, even as we do inference now. Yeah. How much of it do you think still stays the same today?
Like, if you were to build something today, DevX obviously very important, but I feel like, you know, a lot of this has kind of been changed with just hook it up to an agent, have Cloud Code, have Cloud X implement a tool.
primitives that are kind of different than if I'm doing this myself. Right? We've actually changed our SDK team to think about agent experience, sort of developer experience.
And and we think that the same benefits that apply for DX also actually apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can basically make a couple of changes in a decorator and it gets this sort of self provisioning runtime of being able to see its changes live in action. Yeah.
actually find the model is way faster to use for agents versus operating on a different substrate. Yeah. Because, like, you you again, you co locate the infrastructure requirements to the code that runs it.
While the the the negative thesis now is that nobody's looking at their code anymore, so there's no point.
Yeah. I mean, people are looking at code. One thing we actually still see is is really important is observability.
Like, how good is your dashboard? And of course, like, we have we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, you know, make judgment calls and whatnot.
maybe more important now than looking at the code itself. Yes. Because, like, you know, it's you you can try to treat the code as a black box and then use sort of see the observable action that comes out of it and then just prompt a change.
Yep. So I I think actually, I think it it takes a bit of restraint to not specialize, to say, I want to ship a new primitive and then just just be general purpose. People ask you, what are you for?
You're like, I don't know. We can do this. We can do that.
Yeah. Mean, I'd be curious to say, you know, like, okay. If if we were to ask you, like, what is model for even at a high level?
There's a lot you guys do, sandboxes, GPUs, everything. How how do you answer?
Model is a cloud platform that's built for where we build the primitives from scratch for AI applications. And right now, it basically covers inference training, batch crossing, and sandbox workloads.
But for building a lot more and I noticed you didn't say web server. So there there is still a role for, like, the always on large scale Kubernetes type of things. Yeah.
Absolutely.
of of the world because, yeah, we we think the differentiator for us is the our other workflows that need specialized compute, need to scale up and down a lot.
Yeah. And they're they're they're just shaped differently. I think you're building a lot of it alongside the startups.
Right? They're innovating quite a bit.
the Cognitions, Deckergons, ramps, whatnot, they're they're innovating with you. Right? And that's not something AWS is doing directly with.
Yeah. Absolutely. I I think I mean, this is, again, classic.
We're a small team. We can move really fast. Our engineers are working with our customers and figuring it out.
Yeah.
I walked in. There was someone wearing a Moto shirt. I was like, what are you doing here?
They're like, yeah. I just I I am embedded inside of Cog.
Yeah. I think that was Peyton. We we sent them over because the latency of communication was too high otherwise.
Yeah. You know, distributed node, you have to put you have to place one and collocate. So, actually and so I had a direct personal experience.
Right? So I worked on small developer three years ago. It was inspired by Cloud One.
I think you onboarded me at some point, like, just before. And I was like, oh, like, I need some bursty compute. Like, I was just gonna try using Modo.
And it was a pretty pleasant experience.
like, the analytics. Yeah. You blew up on Hacker News, and and we got a big traffic spike.
like, that that was a good use case. Yeah. That was so to me, that was proto cognition.
Right. If only I had, like, stuck to it. Like, that that was, if you just, like, draw the tech tree out, you're like, yeah, like, probably this will happen.
Yeah. Like, he was so close. You're just dealing with this.
but the the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing. This is like 2023. Yeah.
Yeah. Yeah. So we built You should have some new API right after that.
Yeah. Yes. Like, we built sandboxes in May 2023 before anyone was even knew this was gonna be a thing.
The And first example we published was we took small developer and put in a loop so the agent can iterate on itself.
Loops are hard these days. Yeah.
Loops in when was this? 2023?
Yeah. A small check. Yeah.
It's, like, mid twenty twenty three. I mean, so the I mean, those sort of obviously, for listeners, like, the the problem was the models are not built for any of this. Right?
Like, you you are you're trying to, like they're not post trained to understand, like, you know, looping and, like, cell correction and and tool calling there, but, like, also not that great. Yeah. I I don't remember if you used tool tool calling in this one, but, yeah, the models are just diverse after, like, 10 iterations and not produce anything meaningful.
Yeah. But, like, then so I mean, okay. Like, now talking to myself three years ago, the answer is Parts to look at.
Collect all the failures, build benchmark, and then collect all the, you know, examples, build the RO environment Right. Sell it for, like, $10,000,000,000 to Meta, and then and then also train a model and then sell that for $60,000,000,000 to to Elon. And this is Yeah.
Money machine. Like, it's like, it's actually about the heart thing.
I mean, it's hard to have that kind of inherent conviction that this stuff will get that much better.
it's so fucking obvious. Like, fair enough. Like, what else were we doing back then?
I don't know. Anyway yeah. So so so, I mean, this that that was the start of your sandboxing Mhmm.
Journey. Right? I feel like it didn't blow up blow up until, like, last year.
Yep. So there was, a couple of years of quietness. Exactly.
Yeah.
product value. Like, my experience with Model, Charles, before he had joined Model, met this guy at a hackathon, and he really insisted we wanted to run some small model, you know, not hosted anywhere. And he's like, ah, there's this cool company modal.
They'll, like, spin up a GPU sandbox so we can throw it on there. They'll take a Hugging Face link. And, you know, like, there's so much value just right there.
Right? Like, instant hosting, spin it up, spin it down. It'll stay cold, but, you know, we run the demo a few days later, it'll come back up.
And, like, all this stuff in retrospect, like, it's still what we needed, like, today. Yeah. I mean, it's still needed today.
workload shapes have changed a lot as we we run stuff for people with really massive production scale.
from, like, thousand to 1,500 GPUs very quickly in in a given region. It's the same shape of the problem. Okay.
So you look at, say, Cursor Composer. Right? They had a we'll do RL on a model every couple hours.
You guys have a whole version of RL inference gym and whatnot. Mhmm. When you look at workloads like that, you're basically doing train runs where you need to scale up, scale down every hour thousands of GPUs.
Right?
That's the example for we do need it. Right? You will also actually, I'll I'll take a step back and maybe talk about, like, how how people use model today.
Because our our biggest use case actually is Elastic Inference. And the thing we first found product market fit with was inference for custom models. So we kind of stayed away from the LM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere.
But Model is the best black box that, for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them actually, like, have a very unpredictable traffic pattern. It's like diurnal.
It's and some days, like, the company will be launched and, you know, they'll need, like, way more. And it's not just one model that they deploy. They all these companies deploy lots of different models in different regions, and so the auto scaling problem becomes even harder because then you have to scale within a certain region.
And those cycles sort of are offset. So different times, need to scale up in different regions. So that's like our sort of And that, you know, that that in and of itself is a huge category.
which, you know, provide this fireworks, does this as a service together, whatnot, base 10.
That's kind of carved into its own niche for language models, at least right now. Yeah. I mean, the thing that we we have actually specialized in is the auto scaling aspect.
Yeah. Because we we found that it's not universally true that everyone else can auto scale. And we've gone deeper into it on the tech side, but we've incorporated GPU snapshotting into the product so we can actually take the GPU state, like your TARS compiler model, snapshot it, and next call starts way faster.
And so going back to your question, it's that's why you need a lot of burstiness for inference. But then people also do a lot of on demand training. Like, for RL stuff, your rollouts are bursting, as you said.
People also do a lot of batch jobs. So we'll see a lot of companies, before they have a training run, they'll need thousands GPUs to run encoding or something like that. And I think those things are much more bursty than I agree that agents are are not that bursty.
Sandboxes are except when you're doing RL. RL is this insanely bursty. Yeah.
Yeah. Like, when when you're doing rollouts, you you sometimes need a 100,000 sandboxes and and your cycles. Yeah.
I'm curious if you've seen early sparks of continual learning. There's some people like our friends, Engram recently announced this, they're they're trying to do training. That also seems like a different workload.
Right?
twenty four seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot. But seems like something you guys would work for. As you said, we're we're fortunate to work with a number of customers at the frontier and and grab some of our customers.
And and they are taking the primitives we have and trying to use them in very interesting ways like continual learning. It's possible as the stuff gets better, some of that will be part of, you know, our offering as well if, you know, more people need it. But we're we're just waiting to see how how it shakes out.
primitive that you added after sandboxing that was the next step in the story?
I guess we've been going much deeper into LLM inference Yeah. Because we've realized that some of the advantages we have with, like, auto scaling, again, especially in different regions and whatnot, are not present elsewhere. And the place where we had a gap was we weren't working on the model there itself.
Like, we were a black box. And we realized that we actually can get to frontier level model performance, you know, by having great people who who work on all this. And we've actually been open sourcing a lot of our work in terms of recently, we shared our work on D Flash, which is a block based speculator, and we've open sourced all of it.
So you can get by using open sourced D Flash, you can get the same performance as you would with one of the proprietary providers.
And the next thing we're thinking about here I thought this was actually an interesting blog post as well. Right? Like, I think in here you make a claim that, you know, not a claim, just that how how effective speculative decoding really just gets you.
Yeah.
you know, what people should know? Yeah. Absolutely.
I mean, the high level summaries, it would help to describe what speculative decoding is. Yes. High level.
I think And, like, we so we've covered, like, EGO and all this Yeah. Like, Hydra and all those things, but it was, like, two years ago. I think it doesn't hurt.
Right? Speculative decoding is you have a smaller model called a draft model, predict tokens ahead of the bigger model. And then you have the bigger model verify all of this, all the tokens that are predicted.
And the reason it's faster is if you're predicting one token at once, you're kind of bound by memory bandwidth. But if you can batch the verification of the draft model, then you're much more efficient using compute and it's faster. And as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's multiple times of, you know, the original model speed.
And that's what we highlight here. It's like, people talk a lot about we made these kernel faster and whatnot, but improving kernel will only give you, like, a few percentage points of improvement. And increasing its sub length literally is a multiplicative decrease in Like, two to four x.
Yeah. Exactly. Without much head on performance.
Yeah. Yeah. I think it may be yeah.
Mean, you you are running a second model. Right? So it may be something more expensive in in the compute, but I meant quality performance.
But, yeah, I mean, I think So I there's no drop in quality performance because you're always you're never accepting a token that does a model. Better or the same. Exactly.
Right? Yeah. And so we've been working a bunch on Dflash, which is a block based speculator.
So it's instead of predicting one token at a time, it's predicting a block. And we've been open sourcing our work with it. The the next thing for us here is for for helping people train speculators and custom models.
It's it's something that traditionally is very FDE driven, support deployed engineer driven. Like, you work with customers and help them do that. And our vision for this is why we launched Auto Endpoints is we want to make frontier level performance available to everyone.
And so we've mentioned this in the announcement. We kind of teased it. The next thing we're relaunching is, basically, as you run an auto endpoint, we shadow traffic and Do do you want to explain what auto endpoints are?
Yeah. How lovely. Yeah.
Yeah. So this is, I guess, going back to your model is is you touch the code, but sometimes people actually don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and scalability that modal has. So we've made that easier with, basically, a way to create an endpoint from our UI, from the CLI that has all of our optimizations that we talked about, like the deflash stuff already baked in, and there's full transparency.
So we give you the code. You can go run it yourself. And if you want, you can eject out into the full model experience, which we see as people get sophisticated.
They they do wanna tweak the models. They wanna fine tune stuff. You you can still do all of that.
It's it's not a black box. And, yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to person.
yeah. I guess just to kind of understand it directly. I mean, you know, obviously, you you have the GPUs.
You have an endpoint that's compatible. You serve open model. If someone was to do this themselves, what's the delta that you guys provide?
So you do a lot of open source great work on effective inference. How does it compare to say, I take the same model, GLM 5.2 FB eight, take off the shelf inference engine, VLM, SGLANG, you know, get compute of similar capacity, similar cost.
What's the kind of delta that plugging into something this like this offers outside of the benefit of, you know, scaling?
It's kind of interesting because we we've taken the approach of open sourcing our contributions and upstreaming them. We work closely with the SGLAN team. We we actually want the improvements that our team comes up with to be there in open source for others to use, even outside of modal.
The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance first. The other thing is, with these endpoints, we are way more elastic, as you said, than anyone else. And you have true scaling to zero, You have true burstiness.
And in practice, that matters a lot more to people than just finding the GPU and running model code on the Yeah. And I I will say it's actually not that straightforward to just like, what I said is easier said than done. Right?
Yeah.
there's there's quite a bit of combinations you can make there. The trade offs aren't really known at face value. Yeah.
I mean, it's it's not just that. I think it's it's that running production grade inference is a hard infer problem. Even if you subtract out the auto scaling is controlling things like tail latency and making sure every request is delivered at least once and whatnot.
There's a lot of innovation that you can do here. Think it's very interesting that you're starting to encroach on, like, as you become a full cloud, you're starting to encroach on other people's turf. Mhmm.
What will you not do?
Well, we wanna follow our users and make sure they they get, like, a platform that has everything that works well together. So right now, we're kind of focused on the model life cycle and the agent life cycle. So both, like, going from data prep to training to inference, and then also, if I wanna deploy a background agent, let's say, you know, sandbox, super persistent storage, a whole bunch of other stuff.
who did OpenInspect. Yeah. And obviously, RealInspect also is on model.
Yeah.
great example of a background agent that was really successful because they were able to use some of parameters like snapshotting and fast scaling to just have something that feels really reactive and and works well. Yeah.
That's the new CTO of Ramp right there. Yeah. Rahul.
It was really, really fun. Yeah. I mean, okay.
You know, I I think all very bullish. Like, know, one of my reflections was also I did not originally so, obviously, when I met you guys Mhmm. You weren't that much in the GPU game, and now you're all about inference.
And one of the points that I hinged on for Jensen's keynote at GTC this year was what we're calling, like, the inference inflection. Right? That let's say in AI workloads or machine learning workloads, it used to be, like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting like, because of how much agents basically are blocked or call out to to CPU heavy stuff, the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.
Yep. GPU, CPU. And now it's like just constantly.
And you just have to co locate everything.
Yeah. And that's one of the things that actually, again, we see as something appealing about Modal, which is we've built this capacity pool that spans 17 cloud providers. So we're very good at running on various kinds of cloud capacity across the world.
You don't have your own data centers. We don't have our own data centers. We we just run across a lot of neo clouds and Yeah.
Are you metal providers. Yeah. Question mark.
Yeah. Yeah. You're you're running the math, and you're like, what what's the cut over point where you're like?
Yeah, that's a good question. I mean, part of it is we we see our differentiator in in the software layer, and being capital light and focusing on the software helps us move really fast. So far, it's worked out well because there are so many other people building data centers that we're able to work effectively with them and again, focus on what makes us special.
Seventeen gets you into like the local providers sometimes.
what was the most interesting one? There are actually a lot more Neo Clouds than you expect, and they all have various degrees of various levels of reliability. And that's why it's something we've invested a lot of time in, is actually building our own reliability layer on top.
So if the GPU falls off the bus or something happens, we user workloads are not affected.
Yeah. You know, you as a user would be able to. It's a useful thing to have because I now everyone knows, like, what layer you are and, like, you you sort of optimize for being the super cloud of Yeah.
That's that's the idea.
like they want. Oh, they pin it in, like, EU? Exactly.
Or EU, EU Like, data retraining the locality thing or performance or what? It's either data locality or latency.
Want your they're running sandboxes in model. They want them to be right next to your Yeah, it's easy sense it.
is important in all those things. So you've kind of accidentally I don't know if it's accident, but, like, you've built the perfect primitive for agents to to express themselves in it. You know, like, it's almost very funny how every extra development just involves more file system, just involves more CPU Yeah.
Just like the things that you already have. I don't know much about if there's any, like, networking usages that are interesting, but you've also done some good work on networking.
Yeah. I mean, that's exactly right. Like, we we're sort of just taking compute, storage and networking, and building stuff on that layer for, again, the stuff people need.
Yeah. We we see a few interesting networking things coming up. One is people actually want networked sandboxes.
We have a It's for like a Docker cluster type thing? Sorry. Yeah.
Docker Swarm.
What is it called? Compose. Compose type thing.
Yeah. So actually, if you want Docker Compose, our sandboxes now support this thing called sidecars. So you can a sandbox is actually a pod of containers, and you can run multiple containers in the sandbox.
Also useful because, going back to networking, people want a lot of control over outbound networking from a sandbox. Like, they might wanna run man in the middle proxy for, like, maybe logging stuff for RL or controlling how egress can happen to a domain, injecting credentials. And, yeah, so we've kind of had to build a lot of that stuff ourselves.
Yeah. But then also sometimes people actually want sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for for a different reason.
And yeah, we'll see if that becomes the Like just an open socket.
like MTLS.
We do support that, which is you can expose a tunnel inside the sandbox. Yeah. And then you can either expose it to public Internet or it can be and you can add like a HTTP Odds layer above it.
But we have this thing called I six p n, which we haven't talked about, which is this like overlay network using IPv6 addresses. So if modal containers within the same workspace, when this enabled, can actually address each other using this private IP v six address, and no one else can. Mhmm.
This is like sort of private networking for containers. We actually built it because we needed it as primitive for our disputed training product. So we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs, and they have RDMA networking.
So you can run a distributed training job that's truly serverless, and we need the overlay network for that.
yeah, what would people do with it. Build primitives and let people figure it out. Right?
Yeah. Exactly. You put that up pretty interesting.
They're like they read the docs we think. Let me use that first.
is literally not even in our docs page. People somehow found it, and they're using it.
mean, the way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, yeah, you have it built in. I'm sure someone found it, found it to be a lot more efficient before you actually made a thing out of it. Right?
Yeah. And and not to split hairs, I I guess the the overlay network actually is the TCP overlay network. The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that.
But then people found the the TCP part. Can I tell you this is like a big moment for me? Because Yeah.
2,200 submissions for for the World's Fair. Yep. And then I got this from John Osterhalt, who I don't know if do you know John Osterhalt?
The My name sounds familiar. He published a he's a well known professor, published a lot of interesting software design books, and this is the talk he chose to submit. It's on t it's on RTMA at inference.
And I'm like, you wouldn't think that this guy who is, like, kind of operating systems guy would care about RDMA.
I I mean, it it makes sense to me because This is the cloud. Right? That yeah.
Like like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is basically a systems problem of moving memory around and scheduling.
This shows you how primitive my understanding of networking stuff is. Is this like the domain of WireGuard as well? Not quite.
So It's adjacent? So explain everything. Sure.
Sure.
How do we move memory around GPUs?
Well, sorry. Yeah. That is memory.
Sorry. I I was talking more and maybe I was talking, like, five minutes back about the the private IP v six addressing that you've set up. It's basically a VPN.
Yeah. It's sort of like a VPN.
yeah. You're right. It is it is Right.
You already moved on to the topic. Similar Okay. In the same space, WireGuard is encrypted, and and this is Oh, you don't need encrypted.
Yes. It's not encrypted. That's the main difference.
TCP connection based on whether you're allowed to do it. Used to involve a full sidecar, but now you have in the Linux kernel? Yep.
Yeah. I don't know if this is a natural follow on to the topic of, like, my skepticism on distributed training Mhmm. Is that, well, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck.
Is your networking fast enough? Yeah. So I I guess you're talking about sort of fully distributed training, like a Dylocore or something, which is, like, cross data Yes.
That's the extreme. Yeah. You're kind of in the middle, and then other people would have, like, the Mellanox cables up in in, like, their actual data center.
RDMA, I think Mellanox is or InfiniBand is like a is you also use RDMA. But basically, it's a way to bypass the TCP networking stack and transfer stuff much faster between one node to the other.
internal networking, which is the standard that's needed. Okay. So I misunderstood what part of the CPU TCP 3.
2 of Okay. Yeah. I I mean, very impressive work.
So is it effectively you're you're extending sort of, like, the model philosophy to the the trading cluster? Like, yeah. Yeah.
And we're we're not going for, obviously, like, large scale pretraining runs.
The thing that we've built multi node training for is we see a lot of smaller scale post training. Like, people are post training, like, medium sized quant models so they can get higher quality on inference. This is a perfect fit for something like that.
Yeah.
labs explore branches in post training and then eventually merge whatever they find in.
Yeah. The the other use case we've seen for multi node training is even if you have a big cluster, your researchers are still doing small runs. And Yes.
You're having elasticity there better Yeah.
Do, like, this is actually, like, the current limiting factor for auto research, which is, like, you you basically need to give your model some GPUs.
We have a blog that you run on auto research, and model is Yeah. Yeah. Like, turns out to be a pretty good substrate for that.
So my impression is auto research means many things. Like, if anything, the Andre coins. Right now, it's still science fair.
Right? Like, not not not actually, like Men. I don't know how we can Yeah.
Actually do taught the same thing. Yeah. You would know.
We like, our internal both training and inference teams actually use this sort of the general shape of this quite a bit. Like, we have this one internal repo called AutoInference, which essentially we've automated our own FDE efforts using this harness, which is the agent will just spin up a sweep of different things. It'll even run, like, NVIDIA Insight Profiler, and it'll, like, tweak configs and it'll arrive at the right thing.
It'll change your GPUs.
and it actually works really well. Nice. So By the way, I enjoy that your FTE is so technical that you have to do these things.
It's very different for FTE from other people. Yeah. Yeah.
Yeah.
essentially, they're, like, applied inference researchers or applied training researchers.
Someone told me, like, they have to be able to build, but they also have to be able to sell.
Do they have to sell or are they, like, they're good? This is, post sale type of thing. It does being able to talk to a customer and engage effectively with them matters a lot.
They're all on the same thing. You know? But it's it's not really a sort of sales thing.
We we pair them with We have solution architects as well that are more on the presales side. Okay, let's spend a bit more time on auto research.
this year. Where does this go? You know, like have people explored enough?
Like, you know, there's all these beautiful charts of, like, improve it, improve, then they sort of level off a bit, and then you find the next thing. Is this basically sort of one abstraction up from normal training? Is that how we think about it?
Or do you think about it differently? Like, model level training versus basically a high like, AI driven hyperparameter search.
Some some people call it, like, neural architecture search or whatever. Right? Like Yeah.
I mean, I so the the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's basically a hyperparameter sweep that's guided by some sort of model intuition. So it's, like, much more efficient than whatever other sweeper you would have.
Yeah. I mean, you just know, it's just a question of where you wanna spend your computer. You know?
Right. Because, yeah, you can just throw infinite amounts of money on this, and somehow you'll bang on Shakespeare. You know?
Yeah. Infinite monkey. Yeah.
I mean, so, like, very good for a model. And I I think it's also very important that agents can spin up other agents. They can spin up their own infrastructure.
Like, very, like, very good for you. How good are LLMs at generating modal code? Like, you know, like, the benefit of existing pre LLMs is that you are in the data?
Yeah.
surprisingly good. I think, like, pre Cloud four, they were not, and then now they're able to one shot stuff out of the box. We were playing around with releasing, like, a modal bench for, like, the harder things that the LMs cannot do yet, and and maybe What what's an example of that?
I think the the things that sometimes agents struggle with without right guidance and a skill is how to use the rest of our observability. Like, how to something is failing, like, how do you look at the the logs and then update the right thing? It's sort of reasoning about that.
But they're able to one shot, like Yeah. Can can just add a skill to it? Yeah.
So so we have a modal skill now that which is kind of actually why we built this model bench. It's to find things like that so we can we we can address them in our Yeah. In skill.
Yeah. Yeah. No.
No. I mean, it's it's good.
Are you facing any shortages? You know, like, we talk a lot about GPU shortages, but also CPU, also memory. Mhmm.
Yeah.
We have had a lot of growth, which means that there's a we've had to be much better about
proactive Yeah. So we we have Which, by the way, like, it's like an MBA's, like, dream job. It's like just planning this stuff.
I think last time you and I talked with somebody about this. Yeah.
competent team of people that we call in the roles called compute strategy. So, yeah, anyone listening here who wants to work on Compute strategy? Yeah.
I think the norm the normies call it FP and A or something. Well, it's more it's it's not FP and A. It's it's there's a lot of interesting financial questions of, like like, what is the blend between one year and three year reservations?
How do we forecast our own capacity? How do we basically, especially since our capacity is very fungible across different GPU types and different regions, like, basically have to model a lot of it, and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to, like, take bets based on that. Tokatomics.
This is, like, probably not a real point, but at the end, I was trying to think about, like, what other industries I was trying to think about, you know, we cannot be first to, like, these kinds of problems. Yeah. And what other industries have had this?
And I was like, airlines with with fuel. And, like, they have to hedge their fuel. And, like, I think for a long time, Southwest, because they made them like a hero fuel bet, they like were like super low cost because compared to everyone else.
Yeah. I hadn't thought about that. And then we're in a fun time too, you know?
Yeah.
It's a lot of the compute business in general for us is also about being very good about capacity management. That is how you have great unit economics, but also over time, it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get if they don't care about latency, like, get much cheaper pricing and they'll get results back in, like, next twenty four hours or something, like a batch tier, essentially.
Yeah.
give people sufficient Yeah. I feel like they're not as popular. Like, like, the the Frontier Labs have all those APIs.
They're not as popular as they should be. The demand that we see for something like that is actually not for LLMs.
Although sometimes people wanna run evals do and synthetic data prep, and and there it makes sense. K.
when they get it back. Yeah. And, like, they have a reasonable it's it's also, like, a cousin to the stopping problem of, like, will this finish in time?
Yeah.
You you can you can bound it. Yeah.
on it. Yeah. Yeah.
I think what's what's interesting is, like, the the next phase of modal. Mhmm. Like, what, you know, do people expect from you now that you're sort of established and you're, like, a well known compute player among all these leading companies?
You had an inference launch week, and we talked a little bit about the launches. Like, else? Like, what else should people know?
We are building primitives that make our users' lives much easier. So I think, for example, LM inference, thousands more companies are going to post training their own models and deploy open source models for inference. So we're thinking a lot about what is the best product shape for that.
And that involves everything from our training gym to then endpoints that get frontier level performance. You know, again, but I haven't talked to anyone. It looks somewhat different on other verticals.
Like, we're also seeing a lot of real time audio video stuff in there, which is why, like, we're working on things like regional routing with with fallbacks. So you can you can get sort of GPUs that are as close to users as possible. So you you get like low latency for video streaming and and whatnot.
And then on the agent side, it's we're still sort of working very closely with our customers because stuff is changing so fast in terms of what they need.
production agents. So, yeah, we're thinking about those other things that fit in there. I wanted to ask what what the other things are.
Yeah. I probably should share it. I think okay.
So I do think a lot about the principal components of cloud, and you do talk about compute storage networking. Mhmm. Yeah.
Because so far for me, it's fine. So far for the I mean, the first couple generations of cloud, it's fine. What's different qualitatively different about agents that you you need some new permission level?
Like, lot of people you know, obviously okay. And I'll just kinda spew tokens at you until until it, like, hopefully sparks something. Yep.
Like, the the new level now is whatever Cloud Code does, which is dangerously skip permissions or, like, allow this by command or, like, whatever. Right? And and sometimes they're like, okay.
We we have, like, this adaptive thinking mode where, like, just just trust me, bro. I I will make calls for you. Is that it?
You know? Like, basically, like, of LM mediated permissions. Looping it with a goal and flooding, bro.
Yeah.
stuff that is at the sandbox level because you do want hard boundaries. Yeah.
exfiltrate stuff. But, like like like, maybe maybe that's old school thinking. Maybe we're the dinosaurs.
Maybe the AIOS or the LLMOS is really the kernel as a goddamn LLM. Like, it makes you feel uncomfortable. Yeah.
Know. But that's what trusting the LLM is. Like, imagine a spherical cow perfect LLM.
Right. Let it.
Maybe. Maybe.
I wanna test the boundaries. Right? Like, obviously, I I don't believe that, but I wanna see where I'm wrong because that's that's the non consensus.
Yeah. I mean, I think you always need hard guardrails when you want and you can pair those with softer guardrails. Right?
And the ask a deal on mediated.
There I'll get you end with a couple of your commentary on, like, the ecosystem outside of Moto. Managed agents, everyone has one. Gemini, OpenAI, Cloud, very useful for you.
But also, like, it is their way of starting to edge into your space. Yeah.
What's going on? Yeah. I mean, we're very excited to partner with Anthropic and some of the other foundation labs, well, I'll name who we're also working with.
The way we see it is the managed agent thing is a great place to start if you're starting out building an agent. And but but then when you get to building something more production grade, like you're a company that's like Ramp, that's building their own Ramp also runs their accounting agent on us, their external facing agent. You need a lot more control over your compute primitive on things like what sort of how do you persist different files that the agent has access to and how do you snapshot and restore?
How do you control the networking? Maybe you want GPUs. When you get to that point, you kind of want a specialized sandbox provider that gives you those things.
And that's the role that we are trying to play. Yeah. We don't really have an opinion on the harness, whether it runs in it's a Cloud Managed Agent and you hook it up to Modal Sandbox, or you run the harness in Modal Sandbox.
We'll see where people converge with that. Yeah. Do you have any opinions on, like, the meta harnesses?
It's just another layer on top of these things. You mean, like, the OpenPie and OpenPie is one. I think Vercel had one, which I can't remember the name of right now.
Fred Schott had one. And then to me most recently was data Databricks that had Omnigen. All these are sort of meta hardest.
Like, it's kind of pseudo agent cloud type things.
I personally have not played around with them.
mean, is bullish modal as long it consumes more infra. That's why we're focusing on the infra layer.
somewhere where our relative competence is and also
it's a hard problem to solve. Yeah. I mean, I I will say, like, just generally reflecting on don't know if you if there's if there's other topics of modal, but, like, just generally reflecting as an infra person, not as intense as you, but in that field.
This has, like, been the most exciting time in infra. Like, it was boring actually for for for a while, and you couldn't really get people excited about data infrastructure. Like, Eric would get on data console.
Everyone just watched the video. And, like, it's like, look at how many sandboxes I can spin up, and no one gave a crap. Yeah.
And, like, now everyone gives a crap.
That's true. It is a very exciting time.
the amount of scale all of this stuff needs.
make sense in retrospect, which is, the best kind, but I wouldn't necessarily have thought about it myself. It's it's just We need we need the predictions. You know?
Mean, I I think there's a lot that you just don't even see, right?
but what else, you know? What else is coming up for us? Where do you see things going?
Yeah. I mean, in general, it's clear that there's a obviously, there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much, is also we work a lot of companies that are doing things like drug drug discovery and computational bio, like the China discoveries world.
Big things are probably going to happen there.
getting good results out of them. Is there air gap model? Is there is there a version that is, like, on prem air gapped whatever?
No. We we We use it cloud only. Yeah.
Yeah. Okay. But yeah.
I mean, so what you're saying is, like, because you're focused on primitives and they're good primitives, you find use cases and all these kinds of things. I should probably diversify you a little bit away from LMs all the time. Yeah.
Absolutely. We're we're our our goal isn't to only serve the LM infamous market. No.
Yeah. We've we've had both on the Let's say, bio images. Yeah.
Mean, there's a lot here. There's QTiTTS, customizing oh, Chatterbox. You know, there was a customizing whisper.
Yeah.
reminds me of a fallen competitor, which replicate. Mhmm.
happened? This is one thing we've kind of stayed away from is providing an API for models. Because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.
Yeah. And we've always done a build for companies that are building sort of products and need sort of more flexibility that's not just an API. Which you you can build an API for a model, and this is clearly what it is.
But you can but you're saying you can wrap it into a more fully functioning back end that you that you run? Yeah. So actually, all of our examples, it's not that spin up this model, here's an API token, use it.
They're actually all code. Okay. And so the point is that this is just an example.
Starter code. Yeah. But you can you can tweak it however you want.
And if you're like a company building a product, like a computational bio, whatnot. Yeah. I guess I'm trying to tease out for listeners.
Yeah. When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product? Product.
Right? Like, what is that layer?
that people add that qualifies it to be something more? I think there's a little bit of, like, a selection effect of, a lot of companies who do want to get deeper into that level are probably building something that's more differentiated. And I think an example is like, we with LLM Inference, originally, we worked with companies that were building their own post training frameworks or they were Ramp actually early in the day was training their own tokenizer and swapping out the tokenizer in Lava and whatnot.
I'm not saying that's successful in that case. A better example is like like, let's say, Suno, because Suno does not use Modal Mikey for on the pod. Yeah.
you know, it's an API. It's interesting as well. Like, we had Ethan, most recently on the x AI Grok team Mhmm.
Make a prediction that actually, like, the next tier in video gen is not a better video model.
or agent that orchestrates video models.
write code. Like, yes, I can make my six second video or my ten second video from Groc, but actually, I want my six minute video. Mhmm.
And I'm not going there through normal video gen.
Yeah. That's interesting.
Yeah. Give it FFmpeg and just do whatever they really do. Like that, that's not enough.
Yeah. You need to give it Adobe.
Yeah. I hadn't put it together with, like, there would actually be a video production thing.
Yeah. Yeah. Well, I I think about this a lot.
Obviously.
Yeah. Sorry.
Luma Luma agent is a version of this for video production, but, you know, it's a one off. I was gonna get your quick takes on on some other stuff that happens in recent news, and just to see if you have anything interesting. Gitpod, very, like, somewhat, like, you know, different market.
They were they're in, like, sort of, the CICD market, but, actually, technically, very impressive. I don't know if you've, like, taken a real look at them. Yeah.
people on our team have talked to the GitHub team, and they've they're technically very strong. Yeah. I I actually am we're we're very bullish in modal and the CI market as well because there's there's more agents, coding agents.
Yeah. They're gonna run a lot more CI, and the primitives there can be much better. I think there's a lot of wasted CI.
Yeah. So is it just like, let's filter?
Like, what what is the highest order bit here in improving CI for agents?
preparing your artifacts and, like, you know, getting you to the basically preparing your dependencies and whatnot. And obviously, like build systems help with that. But like, if you have primitives that are like memory snapshot and restore, can you just run CI more efficiently.
Oh, okay. Okay. Okay.
Interesting. Yes. I mean, another form of, like, you know, on demand compute.
Yeah. Exactly. Yeah.
Yeah. It it needs the same, again, platform.
Yeah. So so for those who don't know, Gitpod rebranded to Ona. Mhmm.
It was like there's this whole thing. I I actually, I I, like, sort of semi sounded the alarm recognition. I was like, you should take these guys seriously because they're infrared very good.
Yeah. And but, you know, and then then they join OpenAI, and presumably we'll we'll see Codex Cloud from the owner team. Mhmm.
Like, which which I think would be very, very strong. To me, like, teams like that that can set up the networking and, like, the the secure boundaries for, like, in your like, agents to have their own cloud each effectively is what you're doing kind of. And I'm just trying to draw the analogy or or the differences.
If you have studied them, like, what is the philosophical difference? You know?
they didn't go off to the right market at the right time because we I guess, also got lucky with, like, agentic use cases really taking off and dealing, like, more of, like, a sandbox shaped thing than, like,
my understanding is yeah. I mean, like, sandboxes work Yeah. My mind.
Yeah.
Is sandboxes. Yeah. It's just, like, build time sandboxes versus runtime sandboxes.
And, actually, it turned out runtime was better. Right. And and the difference there is runtime sandboxes have a different configuration surface of, like, how you configure images, how you, like, attach, like, storage.
Yeah. It's it's fascinating. Other people, Astral, also OpenAI, also, like, Python tooling ecosystem people.
Are you still sort of bullish build building on top of Python? Also recently, Modular also Mhmm. That got bought by by Qualcomm.
Just any any of your takes there.
Yeah. I mean, we we had Python as our first SDK language because that was the language that people did data and ML in. I actually now have Go and TypeScript SDKs as well.
And our runtime is completely language it isn't in Rust, but it's it's not tied to Python by any means. We haven't seen I think with, like, inference and training stuff, people are still very Python. And the interesting thing with, like, the agent stuff is people use our TypeScript SDK a lot more because they're not actually doing anything that needs ML.
dominant. The last two languages in the world. Yeah.
That's it. Well, English and prompting is the first prompting. I occasionally talk to people who try to build new languages.
They're like even, what's his name, Brett Taylor, who's chairman of OpenAI, was like, we we need we need a new language for for LLM, so no one has come across one. And I keep looking. You know, Python and TypeScript are you have a lot of data plus, but then also they are very imperfect as just as languages themselves.
But then my close is, I think, modal used to be a big bet on developer experience Mhmm. And you've pivoted the team to agent experience. Is it, like, the way now like, do do do can entire companies and unicorns, multi unicorns be built on just having better agent experience?
Do you need something else? It's a big part of our identity.
It's not just, you know, like the very tactical, how does an agent use the CLI, but it's also how easy is it to spin something up. Like, what is your iteration time when you want to spin up a new service and you want to get something going in prod? In practice, that matters a lot to people.
And I think it would continue to matter. Like, people are building stuff even faster, if you give them ways to do it quickly and not have overhead, then I think the debate for me has been, do you do anything differently that is like very fundamentally different for developer experience versus agent experience?
this, they're like cosign We actually also have a blog post on that.
on like 0.9 or whatever.
Yeah. I mean, pretty much it's the main shift for us has been, as I said, like, built this benchmark, modal bench, to see where agents are lacking and Yeah. Actually literally add surface areas to a product if if they're reaching for something.
Like, maybe this should just be a CLI. They They hallucinate their own features. Yeah.
And sometimes it makes sense. Like, if they're reaching for this thing, it's product feedback. Like, give it to them.
And then, yeah, actually moving we used to only have, like, logs and metrics in our UI, just moving all those things to CLI as well, so they're accessible in in in that form. Simple as that.
Cool. Thank you so much. Yeah.
This is great. This is a great update, and I can see why you guys have succeeded so much. It it is really focused, but also really good execution.
Thanks. I mean, we have a long way to go. Alright.
Thank you. Cool.
Shared via Hopper