This episode features Nader Khalil and Kyle Kranen from NVIDIA, discussing Brev's acquisition and its role in enhancing developer experience for GPUs, alongside Dynamo's innovations in planetary-scale AI inference. They delve into NVIDIA's unique engineering culture, the 'Speed of Light' philosophy, and advanced concepts like disaggregated inference and hardware-model co-design for agents, highlighting the future of AI development and deployment.
Agents can do three things. They can access your files. They can access the Internet, and then now they can write custom code and execute it.
You should really only let an agent do two of those three things. If you can access your files and you can write custom code, you don't want Internet access because that's one vulnerability. Right?
If you have access to Internet and your file system, you should know the full scope of what that agent's capable of doing. Otherwise, malware can get injected or something that can happen. And so that's a lot of what we've been thinking about is, like, you know, how do we both enable this because it's clearly the future, but then also, you know, what what are these enforcement points that we can start to, like, protect?
Alright. Welcome to the Leading Space Podcast in the Chroma studio. Welcome to all the guests here.
We're back with our guest host, Vibhu. Welcome. Good to have you back.
And our friends, Netter and Kyle from NVIDIA, welcome. Yeah. Thanks for having us.
Yeah. Thank you. Actually, I don't even know your titles.
I know you're, like, architect something of Dynamo.
Yeah. I I'm one of the engineering leaders and arch architects of Dynamo. And you're director of something developers.
Yeah. You're the developers director, developers guy at NVIDIA. Source, agent marketing, Brev, and, like, dev tools and stuff.
It's kinda the focus. And we're we're kind of recording this ahead of NVIDIA GTC, which is coming to town again or taking over town, which which we'll all be at. And we'll talk a little bit about your sessions and stuff.
Yeah. Yeah. We're super excited for it.
One of my favorite memories for Nadir like, always do, like, marketing stunts. And, like, while you were Rev, you, like, had this surfboard that you, like, went down to GTC with. And, like, Nada NVIDIA apparently liked it so much that they bought you.
What was that like? What was Yeah. Yeah.
We we our logo was a shocker. We we we were always just kind of, like, trying to keep true to who we were. I think, you know, so much of startups, you're, like, trying to pretend that you're a bigger, more mature company than you are.
And it was actually Evan Conrad, SF Compute, who was just like, guys are the greatest guest. Yeah. Oh, really?
Amazing. Yeah. He was just like, guys, you're two dudes in a room.
Why are you pretending that you're not? And so then we were like, okay, let's make the logo a Shaka. We brought surfboards to our booth to GTC.
And the energy was great. Some palm trees too. They actually poked out over the walls so you could see the bread booth and no one else just from very far away.
Oh, so you remember it back then? I remember it. Pre acquisition, I was like, oh, those guys are cool.
Dude, that makes sense because we so we signed up really last minute, and so we had the last booth. It was all the way in the corner. And so I was I was worried that no one was gonna come, so that's why we had, like, the palm trees.
We really came in with the surfboards. We even had one of our investors bring her dog, and then she was just, like, walking the dog around to try to, like, bring energy towards our booth. Yeah.
Stesh. Yeah. Yeah.
She's the best. You know, as a conference organizer, I love that. Right?
Like, it's like everyone who sponsors a conference comes, does their booth. They're like, we are changing the future of AI or something, some generic bullshit. And, like, no.
Like, actually try to stand out, make it fun. Right? They and people will still remember it after three years.
Yeah. Yeah. Yeah.
You know what's so funny? I'll I'll send I'll give you this clip if you wanna if you wanna add it in. But my wife was, at the time, fiancee, she was in medical school, and she came to help us because it was, like, a big moment for us.
And so we we bought this Cricut. It's like a vine like a vinyl printer because, like, how else are we gonna label the surfboard? So we got a surfboard.
Luckily, I was able to purchase that on the company card. We got a Cricut, and it was just, like, fine tuning for enterprises or something like that that we put on the on the surfboard. And it's 1AM, the day before we go to GTC.
She's helping me put these, like, vinyl stickers on, and she goes, you son of a and she's like, if you pull this off, you son of a bitch. And so right pretty much after the acquisition, I stitched that with the Magnussia acquisition. I sent it to our family group chat.
Oh. Hey. I know.
Well, she she made a good choice there. Was that, like, basically the origin story for Launchables? Is that We And maybe we should explain what Breve is and Yeah.
I mean, Breva is just it's a developer tool that makes it really easy to get a GPU. So we connect a bunch of different GPU sources. So the basics of it is, like, how quickly can we SSH you into a into a GPU?
And whenever we would talk to users, they wanted a GPU. They wanted an a 100. And if you go to, like, any cloud provisioning page, usually it's, like, three pages of forms or in the forms somewhere there's a drop down.
And in the drop down, there's some weird code that you know to translate to an a 100. And I remember just thinking, like, every time someone says they want an a 100, like, the piece of text that they're telling me that they want is, like, stuffed away in the corner. Yeah.
And so we're like, what if the biggest piece of text was what the user's asking for? And so when you go to Revit, it's just big GPU chips with the type that want. Animations that you work on.
Pre like, pre you can do like, now you can just prompt it. But back in the day Yeah.
artisanal code. Yeah. I was actually really proud of that because it was an I I made it in Figma.
Yeah. And then I found I was, like, really struggling to figure out how to turn it from, like, Figma to React. So what it actually is is just an SVG, And I have all the styles.
And so when you change the chip, whether it's, like, active or not, it changes the SVG code. And that somehow, like, renders like, looks like it's animating. But it would we just had the transition slow.
But it's just, like, the a JavaScript function to change the, like, underlying SVG. Yeah. That was how I ended up, like, figuring out how to move it from from Figma.
But, yeah, that's art artisan.
Speaking of marketing stunts, though, he actually used those SVGs or kind of used those SVGs to make these cards.
Oh, yeah. Like, a GPU gift card Yes. That he handed out everywhere.
That was actually my first impression of that. Yeah. Yeah.
Yeah. Yeah. I think I still have one of them.
Yeah. They look great. Yeah.
I have a ton of them still actually in our garage, but just they don't have labels. We should honestly, like, bring bring them back. But I found this old printing press here, actually, just around the corner on Venice, and it's a third generation San Francisco shop.
And so I come in, an excited startup founder trying to, like and they just have this crazy old machinery, and I'm in awe because the the whole building is so physical. Like, you're seeing these machines. They have, like, pedals to, like, move these saws and whatever.
I don't know what this machinery is, but I saw all three generations. Like, there's, like, the grandpa, the father, and the son. The son was, like, around my age.
Oh, it's like a holy holy trinity. Yeah. Yeah.
Like It's funny because we we so I I just took the same SVG, we just, like, printed it, and it's foil printing. So they make a a mold that's, like, an inverse of, like, the a 100, and then they put the foil on it, and then they press it into the paper. And I remember once we got them, he was like, hey.
Don't forget about us. You know, I guess, like, early Apple and Cisco's first business cards were all made there. And so he was like, yeah.
We we get, like, the start up businesses, but then as they mature, they kind of go somewhere else. And so I actually I think we were talking with marketing about, like, using them. You should go back and make some cards.
Yeah. Yeah. Yeah.
Yeah. Know, I I remember, you know, as a very, very small breadth investor, I was like, what why are we spending time, like, doing these, like, stunts for GPUs? Like, you know, I think, like, as a, you know, typical, like, cloud hard hardware person, you go into AWS, you pick, like, t five x x l, whatever, and then it's just, like, a from a list and you look at the specs.
Like, why animate this, Judy? And and I I do think, like, it just shows the level of care that goes throughout Perf and Yeah. And now and also Dynamo.
And NVIDIA, I think that's what the the thing that struck me most when we first came in was, like, the amount of passion that everyone has. Like, I think, you know, you talk to you talk to Kyle. You talk to like, every VP that I've met at NVIDIA goes so close to the metal.
Like, I remember it was almost a year ago, and, like, my VP asked me. He's like, hey. What's Cursor?
And, like, are you using it? And if so, why? And I'm just, like, surprised at this.
And he downloaded Cursor, he was asking me to help him, like, use it. And I thought that was or, like, just show him what he you know, why we were using it. And so the amount of care that I think everyone has and the Passion.
Appreciate passion and appreciation for the moment. Right? This is a very unique time.
So it's really cool to see everyone really, like, appreciate that. Yeah. One thing I wanted to do before we move over to sort of, like, research topics and the the stuff that Kyle's working on is just tell the story of the acquisition.
Right? Like, not many people have been been through an acquisition with NVIDIA.
What's it like? What yeah. Just anything you'd like to say.
It's a crazy experience.
you know, we were the the thing that was the most exciting for us was our goal was just to make it easier for developers. We wanted to find access to GPUs, make it easier to do that. And then all oh, actually, question about Launchables.
So Launchables was just make one click like, one click deploys for any software on top of the GPU. And so what we really liked about NVIDIA was that it felt like we just got a lot more resources to do all of that. I think, you know, NVIDIA's goal is to make things as easy for developers as possible.
So there was a really nice, like, synergy there. I think that, you know, when it comes to, like, an acquisition, I think the amount that the soul of the products align, I think, is gonna be is gonna speak to the success of the acquisition. Yeah.
So it in many ways feels like we're home. This is a really great outcome for us. Like, we you know, I love brev.
nvidia.com. Like, you should you should use it.
It's a front page for GPUs. Yeah. If you want GPUs, you go there and get Internally, it's growing very quickly.
I I do remember. You said some stats there. Yeah.
Yeah. Yeah. It's I I wish I had the exact numbers, but, like, internally, externally, it's been growing really quickly.
We've been working with a bunch of partners, with a bunch of different customers and ISVs. If you have a solution that you want someone that runs on the GPU and you want people to use it quickly, we can bundle it up in a launchable and make it a one click run. If you're doing things and you want just like a sandbox or something to run on, right, like OpenClaw, huge moment, super exciting.
And we'll talk into it more, but, you know, internally, people wanna run this. And, you you know, we have to be really careful from security implications. Do we let this run on the corporate network?
Security's guidance was, hey, run this on breadth. It's a VM, it's sitting in the cloud, it's off the corporate network, it's isolated. And so that's been our stance internally and externally about how to even run something like OpenCLO while we figure out how to run these things securely.
But yeah. I think there's also, like, you almost like, we're the right team at the right time when NVIDIA is starting to invest a lot more in developer experience or whatever you call it.
or I don't know what you call it, like software. Like obviously, the end of it is always invested in software, like, this is like a different audience. It's a wider developer base.
Yeah. Right? Yeah.
It's funny. It's like, it's not So what was it called internally? What is this that people should be aware that is going on there?
What like developer experience? Yeah, it's called developer experience?
always wants to make a good developer experience. The thing is, a lot of the technology is just really complicated. I think the thing that's been really growing or the AI is growing is having a huge moment, not because like, let's say, data scientists in 2018 were quiet then and are much louder now.
The pie is complete. There's a whole bunch of new audiences. My mom's wondering what she's doing.
My sister's learned, like taught herself how to code. Like the, I actually think just generally, AI is a big equalizer, and you're seeing a more technologically literate society, I guess. Everyone's learning how to code.
There isn't really an excuse for that. So building a good UX means that you really understand who your end user is. And when your end user becomes such a wide variety of people, then you have to almost reinvent the practice.
And actually build more developer UX, right? Because there are tiers of developer base that were added, you know, the hackers that are building on top of OpenClaw, right, for example, have never used GPU. They don't know what CUDA is.
They just want to run something. You need new UX that is not just, hey, how do you program something in CUDA and run it? And then we built, like when deep learning was getting big, built Torch.
But recently, the amount of layers that are added to that developer stack has just exploded because AI has become ubiquitous. Everyone's using it in different ways. It's moving fast in every direction, vertical, horizontal.
That? You you even take it down to hardware, like the DGX Spark. It's basically the same system as just throwing it up on big GPU clusters.
Yeah. Yeah. It's a grace of Blackwell.
Yeah.
We saw the preview at the last year's GTC, and that was one of better performing videos of our NVIDIA coverage so far. Awesome. This will beat it.
That was actually We have. Fingers crossed. Yeah.
DGX Spark was first coming out, getting to be involved in that from the beginning of the developer experience, and it just comes back to involved? Yeah. Yeah.
Very direct. Go ahead. Say more say more.
Yeah. Yeah. I mean, from it was just like I I got an email.
We just got thrown into the loop, and suddenly, yeah, I I was actually really funny because I'm still pretty fresh from the acquisition. And I'm I'm getting an email from a bunch of the engineering VPs about, like, the new hardware GPU chip, like, where or not chip, but just GPU system that we're putting out. And I'm like, okay.
Cool. Natter's now involved with this for the UX. I'm like, what am I gonna do here?
So I remember the first meeting, I was just, like, kinda quiet as I was hearing engineering VPs talk about what this box could be, what it could do, how we should use it. And I remember one of the first ideas that people were ID ing was like, oh, the first thing that it was like, I think a quote was, the first thing someone's gonna wanna do with this is get two of them and run a Kubernetes cluster on top of them. And I was like, oh, I think I know why I'm here.
I was like, the first thing we're doing is easy SSH into the machine. And then and, you know, just kind of, like, scoping it down of, like, once you can do that, the person who wants to run a Kubernetes cluster onto Sparks has a higher propensity for pain than someone who buys it and wants to run OpenClaw right now. If you can make sure that that's as effortless as possible, then the rest becomes easy.
So there's a tool called NVIDIA Sync that just makes the SSH connection really simple. So if you think about it, like if you have a Mac or a PC or whatever, if you have a laptop and you buy this GPU and you wanna use it, you should be able to use it like it's a GPU in the cloud. But there's all this friction of how do you actually get into that?
That part of Rev's value proposition is just there's a CLI that wraps SSH and makes it simple. And so our goal is just get you into that machine really easily. And one thing we just launched at CES, it's it's still in, like, early access.
We're ironing out some kinks, it should be ready by GTC. You can register your Spark on Brev. And so now if you Like remote managed local glass.
Yeah. Because Brev can already manage other clouds anyway. Right?
And you can use the Spark on Brev as well. Right? Yeah.
But yeah. Exactly. So so you so you you set it up at home, you can run a command on it, and then it gets it's essentially it'll appear in your Brev account.
And then you can take your laptop to a Starbucks or to a cafe, and you will continue to use your you can continue to use Spark just like any other cloud node on Brev. Yeah. Yeah.
It's just like a pre provisioned Yeah.
Exactly. Yeah. Tiny little data center.
Tiny little data center.
One more thing before we move on to Kyle. Just have so many Jensen stories, and I just love mining Jensen stories. My favorite so far is SOL.
What is what is SOL? SOL is actually I I think of all the lessons I've learned, that one's definitely my favorite. It'll always stick with you.
Yeah. Yeah. I you know, when you're a start up, everything's existential.
Right? Like, we've we've run out of money. We were, like, on the risk of of losing payroll.
We've had to contract our team because we ran out of money. And so, like, because of that, you're really always forcing yourself to, I to, like, understand the root cause of everything. If you get a date, if you get a timeline, you know exactly why that date or timeline is there.
You're you're pushing every boundary, and, like, you're not just say you're not just accepting, like, a a no just because. And so as you start to introduce more layers, as you start to become a much larger organization, SOL is is essentially, like, what is the physics? Right?
The speed of light moves at a certain speed. So if light's moving some slower, then you know something's in the way. So before trying to, like, layer reality back in of, like, why can't this be delivered at some date?
Let's just understand the physics. What is the theoretical limit to, like, how fast this can go? And then start to tell me why.
Because otherwise, people will start telling you why something can't be done. But actually, I think any great leader's goal is just to create urgency. There's Yeah.
And create compelling events. Yeah. Right?
Yeah. SOL is a term in NVIDIA is used to instigate a compelling event. You say, this is done.
How do we get there? What is the minimum as much as necessary as little as possible thing that it takes for us to get exactly here?
helps you just break through a bunch of noise Yeah. Instantly. One thing I'm unclear about is can only Jensen use the SOL card?
Like Oh, no. No. No.
Everyone. Get the bullshit out. Because it obviously, it's Jensen.
But, like, can someone else be like, no. Like Frontline engineers use it. Okay.
Yeah. Every I think it's not so much about, like, get the bullshit out. It's like it's like, give me the root understanding.
Right? Like, if you tell me something takes three weeks, it's like what what principles. Yeah.
The first principles. It's like, what's the what like, why is it three weeks? What is the actual yeah.
What's the actual limit of why this is gonna take three weeks? If you're gonna if you if let's say you wanted to buy a new computer and someone told you it's gonna be here in five days, what's the SOL? Well, like, the SOL is like, could walk into a Best Buy and pick it up for you.
Right? So then anything that's, like, beyond that is and is that practical? Is that how we're gonna, you know, let's say, give everyone in the company a laptop?
Like, obviously not. So then, like, that's the SOL. And then it's like, okay.
Well, if we have to get more than 10, suddenly, there might be some right? And so now we can kind of piece the reality back. So so this is the Paul Graham, do things that don't scale.
Yeah. And this is also the what people would now call be high agency. Yeah.
Yeah. Yeah. It's actually really interesting because there's a there's a second hardware angle to SOL that, like, doesn't come up for all of the orgs.
So SOL is used, like, culturally for everything. I'm also mining for, like I think that can be annoying sometimes. I mean, like, someone keeps going SOL, SOS at you.
And you're like, guys, like, have to be stable. We have to what's your fucking plan? Yeah.
It's an industry balance. Yeah. I encountered that with, like, actually, just with with Alec.
Right? Because we we have a new conference, so we need to launch we have we have goals of what we wanna launch by the conference. And, like, yeah, at the end of the day Wait.
Is this GTC or something? Well, this is like so we I mean, we did it for CES. Did for GTCDC before that.
We're doing it for GTC San Jose. So, mean, like, every you know, we have a new moment, and we want to launch something. Yeah.
And we want to do so at SOL. And that does mean that some there's some level of prioritization that needs to happen. And so it it is difficult.
Right? I think you have to be careful with what you're pushing. You know, stability is important, and that should be factored into SOL.
SOL isn't just, like, build everything and let it break. You know, that that's part of the conversation. So as you're laying layering in all the details, one of them might be, hey.
We could build this, but then it's not gonna be stable for x y z reasons. And so was like one of our conversations for CES was like, hey, we can get this into early access, registering your Spark with Brev, but there are a lot of things that we need to do in order to feel really comfortable from a security perspective. There's a lot of networking involved before we deliver that to users.
So it's like, okay, let's get this to a point where we can at least let people experiment with it. We had it in a booth, we had it in Jensen's keynote, and then let's go iron out all the networking kinks. And that's not easy.
And so that can come later. And so that was the way that we layered that back end. It's not really about saying, like, you don't have to do the maintenance or operational work.
It's more about saying, you know, it's kind of like highlights how progress is incremental, Right? Like, what is the minimum thing that we can get to? And then there's SOL for, like, every component after that, but there's the SOL to get you get you to the the starting line.
And that that's usually how it's asked. On the other side, you know, like, SOL came out of, hardware at NVIDIA. Right?
So SOL is, literally, if we ran the accelerator or the GPU at basically full speed with no other constraints, how fast we'd be able to make a program go.
In training, then you work back to, like, some percentage of, like, MFU, for example. Yeah. That's a that's a great example.
So, there's an there's an SOL MFU, then there's, like, you know, what's practically achievable. Cool. Should we move on to a sort of Kyle side?
Kyle, you're coming more from the data science world. And I I mean, I always whenever whenever I meet someone who's done work in tabular stuff, graph neural networks, time series, these are basically when I go to new reps, I go to ICML, I walk the back halls. There's always, like, a small group of graph people, small group of tabular people, and, like, there's no one there.
And, like, it's very like, you know what I mean? Like Yeah. No.
solving the problems that they solve. Yeah. But everyone else is just LMs all the time.
Yeah. I mean, it's like it's like the black hole. Right?
Yeah. Has the Event Horizon reached this yet in NeurIPS? But, like, you know, those are those are transformers too.
Yeah. And and those are also, like, interesting things. Anyway, I just wanted to spend a little bit of time on on those that background before we go into Dynamo proper.
Yeah. Sure. I took a different path to NVIDIA than that or I joined six years ago, seven if you count when I was an intern.
So, I joined NVIDIA right out of college. And the first thing I jumped into was not what I'd done during internship, which was some stuff for autonomous vehicles, heavyweight object detection. I jumped into something like recommenders.
This is popular. Yeah, did Rexxis. Yeah, Rexxis.
That was the tabular data at the time. You have tables of audience qualities and item qualities, and you're trying to figure out which member of the audience matches which item or which item matches which member of the audience. At the time, really, it was like we were trying to enable recommenders, which had historically been a little bit of a CPU based workflow, into something that ran really well on GPUs.
It's since been done. There are bunch of libraries for Exis that run on GPUs. The common models like deep learning recommendation model, which came out of Meta, and the wide and deep model, which was released by Google, were very accelerated by GPUs using the fast HBM on the chips, especially to do vector lookups.
But it was very interesting at the time and super, super relevant because we were starting to get this explosion of feeds and things that required recommenders to just actively be on all the time And sort of transitioned that a little bit towards graph neural networks when I discovered them because I was like, okay, you can actually use graph neural networks to represent relationships between people, items, concepts. And that interested me, so I jumped into that at NVIDIA and and got really involved for, like, two ish years. Yeah.
Yeah. Is that you can just kinda choose your own path in NVIDIA. Oh my god.
Yeah. Which is not a normal big corp thing. Yeah.
Like, you have a lane. You stay in your lane. I think probably the reason why I enjoy being in a a big company The mission is coming from a start up guy.
Yeah. The mission is the boss. Yeah.
It feels like a big game of pickup basketball. Like, you know, if you play one if wanna play basketball, you just go up to the court, and you're like, hey. Look.
We're gonna play this game, and we need three. Yeah. And you just, like, find your three.
That's honestly for every new initiative. That's what it feels like. Yeah.
Yeah. I know. It also, like, shows.
Right? Like, Nvidia is just releasing state of the art stuff in every domain. Yeah.
Like, okay. You expect foundation models with Nemotron.
Voice just randomly or call up to your parakeet just comes out. Another one The NVIDIA voice team has always been producing. There's always just every other domain of paper that comes out, dataset that comes out.
It's like I mean, it also stems back to what NVIDIA has to do. Right? You have to make chips years before they're actually produced.
Right? So you need to know. You need to really forecast The design process starts like Exactly.
Three to five years before the chip gets to the market. Yeah. I'm curious more about what that's like.
Right? So, like, you have specialist teams. Is it just like, you know, people find an interest, you go in, you go deep on whatever, and that kind of feeds back into, you know, okay.
We we expect predictions. Like, the internals at NVIDIA must be crazy. Right?
You know? Yeah. You know, you you must not even without selling to people.
You have your own predictions of where things are going, and they're very based, very grounded. Right? Yeah.
It's really interesting. So there's two things that I think that NVIDIA does which are quite interesting.
One is we really index into passion. There's a big organizational top sound push to ensure that people are working on the things that they're passionate about.
many times they can just email someone way up the chain that they would find this relevant and say, Can I go work on this? It's actually, like, I worked at a a big company for a couple years before starting on my startup journey, and, like, it felt very weird if you were to, like, email out of chain, if that makes sense. Yeah.
The emails at NVIDIA are, like, mosh pits. Yeah. Shoot.
And it It's just, like, 60 people just whatever. And, like, they're they're Does this get messy? Like, reply all Oh, it gets in it's insane.
It's insane. Agents help, you know, manage the context. But but but that's actually, like I've actually so this is a weird thing where I used to be, like, why would we send emails?
We have Slack. I am the entire I'm the exact opposite. I feel so bad for anyone who's, like, messaging me on Slack because I'm so unresponsive.
And your email maxing. I'm I'm email maxing out. Email is a different place.
Perfect. Because because We can't work together on Slack. Email is great because important threads get bumped back up.
Right? Yeah. Yeah.
And so Slack doesn't do that. So I just have, like, this casino going off on the right or on the left. And, like, I don't know which thread was from where or what, but, like, the threads get and then also just, the subject.
So you can have, like, working threads. I think what's difficult is, like, when you're small, if it's not 40,000 people, I think Slack will work fine. But there's I don't know what the inflection point is.
There is gonna be a point where that becomes really messy, and you'll actually prefer having email because you can have working threads. You can cc more than nine people in a thread. You can force stuff.
You can fork stuff, which is super nice and just like yeah. And so but that is part of where you can propose a plan. You can also just, like, start honestly, momentum is the only authority.
Right? So, like, if you can just start start to make a little bit of progress and show someone something, and then they can try it, that's, I think, what's been, you know, I think, the most effective way to push anything for forward. And that's both at NVIDIA and I think just generally.
Yeah. There's the other concept that, like, is explored a lot at NVIDIA, which is this idea of a $0 business.
Like, market creation is a big thing at NVIDIA.
Like Oh, you want to go and start a $0 business?
just says we're completely happy investing in $0 markets. We don't care if this creates revenue. It's important for us to know about this market.
We think it will be important in the future. It can be $0 for a while. I'm probably mangling his words here.
Mercedes
with NVIDIA logos coming in.
Clara, it's it's actually Yeah.
Yeah. $0 markets are are a thing. Like, you know, Jensen.
I mean, okay, look, cars are not a $0 market. Yeah, that's a bad example.
Think he's messaging Zero today, but Or even like internally, right? Like, it's like, an org doesn't have to ruthlessly find revenue very quickly to justify their existence.
ideologically free at NVIDIA. They can pursue things that they Will you research officially? I was never in research officially.
I was always in engineering.
I'm in an org called deep learning algorithms, which is basically just how do we make things that are relevant to deep learning go fast. That sounds freaking cool. And I think a lot of that is underappreciated.
Right? Like, time series. This week, Google put out time effort.
Yeah. A new time series paper. Rexus.
Semantic ID started applying transformers, LLMs to Rexus. And when you think the scale of companies deploying these, right, Amazon recommendations, Google web search, like, it's huge scale, Yeah. You want fast.
Yeah. Yeah. Yeah.
Actually, there's a fun moment that brought me, like, full circle. Like, Amazon Ads recently gave a talk where they talked about using Dynamo for generative recommendation, which was, like, super, like, weirdly cathartic for me. I'm like, oh my god.
I've I've supplanted what I was working on. Like, I you're using LMs now to do what I was doing five years ago. Yeah.
Amazing. And let's go right into Dynamo. Maybe introduce Yeah.
Sure. Sort of top down and yeah. I think at this point, a lot of people are familiar with the term of inference.
Funnily enough, I went from inference being a really niche topic to being something that's discussed on normal people's Twitter feeds. It's on billboards here now. Very strange.
Seeing just an inference ad on one zero one. Inference at scale is becoming a lot more important. We have these moments like OpenCLaw where you have these agents that take lots and lots of tokens but produce incredible results.
There are many different aspects of test time scaling so that you can use more inference to generate a better result than if you were to use a short amount of inference. There's reasoning, there's re querying, there's adding agency to the model, allowing it to call tools and use skills. Dino came about at NVIDIA because myself and a couple others were talking about these concepts that you have inference engines like VLM, SGLANG, TensorFlow, TLM, and they have one single copy.
They think about things as one single copy, like one replica, right? Like one version of the model. But when you're actually serving things at scale, you can't just scale up that replica because you end up with performance problems.
There's a scaling limit to scaling up replicas. So you actually have to scale out to use maybe some Kubernetes terminology. We kind of realized that there was a lot of potential optimization that we could do in scaling out and building systems for data center scale inference.
So Dynamo is this data center scale inference engine that sits on top of frameworks like the VLM, SGLANG, and TensorFlowGLM, and just makes things go faster. Because you can leverage the economy of scale, the fact that you have KV cache, which we can define a little bit later, in all these machines that is unique and you want to figure out the ways to maximize your cache hits, or you want to employ new techniques in inference like disaggregation, which Dynamo introduced to the world in March. Not introduced, it was an academic talk beforehand, but we're one of the first frameworks to start supporting it.
your inference at scale.
new things.
By the way, this is why I wanted to put two of you together. I was like, yeah, this is this is gonna be good. It's very different.
You know? Like, we we we've we've talked to each other a bunch. Actually, you asked, like, why why can't we scale up?
Yeah. Model you said model replicas. Yeah.
So so scale up means assigning more Heavier. Yeah. Heavier.
Like making things heavier, adding more GPUs, adding more CPUs. Scale out is just like having a barrier saying I'm gonna duplicate my representation of the model or representation of this microservice or something. And I'm gonna, like, replicate it many times to handle load.
And the reason that you can't scale scale up past some points is, like, you know, there there are sort of hardware bounds and algorithmic bounds on on that type of scaling. So I'll give you a good example that's, like, very trivial. Let's say you're on an h 100.
The maximum NVLink domain for h 100 for most DGX h one hundreds is eight GPUs. Right? So if you scaled up past that, you're going to have to figure out ways to handle the fact that now for the GPUs to communicate, you have to do it over InfiniBand, which is still very fast, but is not as fast as NVLink.
Is it like one order of magnitude, like hundreds? It's about an order of magnitude. Not terrible.
Yeah. I to remember the datasheet here. I think it's about 500 gigabytes a second unidirectional for NVLink and about 50 gigabytes a second unidirectional for InfiniBand.
It depends on the generation.
and the transfer speeds. Also, maybe even just going a few steps back before that. Most people are very familiar with you see you know, you can use on your laptop, whatever these SD LAN, VLLM, you can just run inference.
There's all Lama. There's all run it on that laptop. You can run on laptop, then you get to okay.
Moto's got pretty big, right? GLM five, they doubled the size. So, what do you do when you have to go from, okay, I can get a 128 gigs of memory.
I can run it on a Spark. Then, you have to go multi GPU. Yeah.
Okay, multi GPU, there's some support there. Now if I'm a company and I don't have, like I'm not hiring the best researchers for this. Right?
But I need to go multi node. Right? I have a lot of servers.
Okay. Now there's efficiency problems. Right?
You can have multiple eight h 100 nodes, but, you know, is that as a like, how do you do that efficiently? Yeah. How do you, like, represent them?
How do you choose how to represent the model? Yeah. Exactly.
That's like that's like a hard question everyone asks. Like, how do you size?
Oh, I wanna run GLM five, which just came out. New model. There have been, like, four of them in the past week, by the way.
Like, a bunch of new models. You know why, right? DeepSeq.
No comment. Yeah.
But GLM five, right? We we have this new model. It's it's of like a large size.
And you have to figure out how to both scale up and scale out. Right? Because you have to find the right representation that you care about.
Everyone does this differently. Let's be very clear. Everyone figures this out in their own path.
I feel like a lot of AI or ML even is like is like this. I think people think, you know, I I was there was some tweet a few months ago that was like, why hasn't fine tuning as a service taken off? And, you know, and, That might be me.
It might have been you. Yeah. But people want it to be such an easy recipe to follow.
But even, like, if you look at an MOE model specific to you. Yeah. Yeah.
And and the model and the situation. There's so much tinkering. Like like, when you see a model that has however many experts in the MOE model, it's like, why that many experts?
I don't know. They, you know, they tried a bunch of things, that one seemed to do better. And I think when it comes to how you're serving inference, you know, you have a bunch of decisions to make.
And there you can always argue that you can take something and make it more optimal, but I think this this internal calibration and appetite for continued calibration.
Yeah. And that doesn't mean, like, you know, people aren't taking a shot at this. Like, tinker from thinking machines, you know?
Yeah. RL as a service. Totally.
It's it also gets even harder when you try to do big model training. Right? We're not the best at training MOEs when they're pre trained.
Like, we saw this with Llama three. Right? They're trained in such a sparse way that meta knows there's gonna be a bunch of inference done on these.
Right? They'll open source it, but it's very trained for what meta infrastructure wants. Right?
They wanna they wanna inference it Now, a the question to basically think about is, okay, say you wanna serve a chat application, a coding Copilot. Right? You're doing a layer of RL.
You're serving a model for x amount of people. Is it a chat model, a coding model, Dynamo, you know, back to that. Yeah.
Sorry. So we sort of jumped off of jumped on that topic.
Everyone has their own journey. And I like to think of it as defined by what is the model you need? What is the accuracy you need?
Actually, I talked to Nadeira about this earlier. There's three axes you care about. What is the quality that you're able to produce?
Are you accurate enough? Or can you complete the task with enough Performance. Yeah.
There's cost. Can you serve the model or serve your workflow? Because it's not just the model anymore.
It's the workflow. It's the multi turn with an agent cheaply enough? And then can you serve it fast enough?
And we're seeing all three of these play out. We saw new models from OpenAI that are faster. You have these new fast versions of models.
You can change the amount of thinking to change the amount of quality, right? Produce more tokens, but at a higher cost and a higher latency. And really, when you start this journey of trying to figure out how you want to host a model, think about three things.
What is the model I need to serve? How many times do I need to call it? What is the input sequence think was that what does the workflow look like on top of it?
What is the SLA? What is the latency SLA that I need to achieve? Because there's usually some this is usually, like, a constant.
You you know the SLA that you need to hit. And then, like, you try and find the lowest cost version that hits all of these constraints. Usually, you know, you you start with those things and you say you you kinda do, like, a bit of experimentation across some common configurations.
You change the tensor parallel size, which is a form parallelism. I'd say it goes even deeper. First, you gotta think about model.
Yes. It's like a multistep design process. Because as you said, you can choose a smaller model and then do more test time scaling, and it'll equate quality of a larger model because you're doing the test time scaling or you're adding a harness or something.
So, yes, it it goes way deeper than that. But from the performance perspective, like, once you get to the model you need you need to host, you look at that and you say, hey. I have this model.
I need to start it at this speed. What is the right configuration for that?
if you run the same prompt twice, you're getting, like, double digit. Yeah. Exactly.
And you get a lot yeah. But the the key thing there is you give the context of the failed try. Right?
Yeah. It takes a shot. And this has been, like, you know, basic guidance for quite a while.
Just try again. Because, you know, you're trying. Just try again.
Did you try again? All advice in life. Just try again.
It's a paper from Google if I'm not mistaken. Right? Yeah.
I think it's it's like a seven page, little short paper. Yeah. Yeah.
The title is very cute, it's just like, yeah. Just try again. Give it it has context.
Shot. You just, like, say, like, hey. Like, you know, like, take take a little bit more take a little bit more information.
Try and fail. And that basic concept has gone pretty deep. There's, like, self distillation RL where you you do self distillation.
You do RL and you have past failure and, know, that gives some signal. People take try it again. Not strong enough.
For for listeners who listen to here, Vivo actually and I and we run a second YouTube channel for our paper club where Oh, that's awesome. Just covered this. Oh, cool.
Self desolation and all that. That's that's why he's so up to speed on it. I'll have check it out.
Yeah. It's it's just a good practice. Like, everyone needs, like, a paper club where, like, you just read papers together and the social pressure just kind of forces you to just go.
We we there's, like, a big inference reading group. Feel so bad every time I I he put it on, like, on our he shared it Yeah. One of your guys is is big in that.
I forget. Yeah. Ishaan.
Ishaan? Yeah. Yeah.
Ishaan. Ishaan's on my team. Actually, it's funny.
There's a there's a there's a employee transfer between us. Ishaan worked for Nadir at Brev, and now he He was he was our head of AI, and then yeah. Once we got in So that because I I'm always looking for, like, okay.
Can can I start another podcast that only does that thing? And Ishaan was I was trying to, like, nudge you, Sani, to, like, is there something here? I mean, I don't think there's there's new Infosys every day.
So it's like it's like You would you would actually be surprised.
The amount of blog posts you see. And if you There was a period where it was, like, Medusa, Hydra, what, Eagle. Like, you know, now we have new forms of decode we have new forms of speculative decoding or new What what are you excited about?
the amount of, like, post training, the amount of tokens that the GPU rich can just train on, and it it was a hybrid state space model. Right? Yeah.
It's co designed for the hardware. Yeah. Co designed for the hardware.
And one of the things was always, you know, the state space models don't scale as well when you do a conversion or whatever, the perform and you guys are like, no. Just keep training, and Nematron chose a lot of that. Yeah.
Also something cool about Nematron, it was released in layers, if you will, very similar to Dynamo. It's it's it's essentially it was released as aggregate. You can the pretraining, post training datasets are released.
The recipes on how to do it are released. The model itself is released Just full of the benefit from us turning on the GPUs. But there are companies like ServiceNow took the dataset, and they trained their own model.
And we were super excited and set, like, you know, celebrated our work. Zoom, the Frontier Model Labs. Zoom is Zoom is CGI.
I think, you know, also just to add, like, a lot of models don't put out base models. And if there's that, why is fine tuning not taken off? You know, you can do your own bus training, but Yeah.
That's true. You guys put out base model? I think you put out everything.
so. Base can be base can be cancellable.
Base can be cancellable? Yeah. Safety training.
Do we get a full picture of Dynamo? I I don't know if we What is I'd you mentioned the three axes. Like, break it down of, like, you know, what's prefilled decode and, like, what are the optimizations that we can get with Dynamo?
Yeah. That that that's that's that's a great point. So to summarize on that three axis problem, right, there are three things that determine whether or not something can be done with inference.
Cost, quality, latency. Right? Dynamo is supposed to be there to provide you, like, the runtime that allows you to pull levers to, you know, mix it up and move around the Pareto frontier or the Pareto surface that determines, is this actually possible with inference in AI today?
Gives you the knobs. Yeah, exactly. Gives you the knobs.
And one thing that we use a lot in contemporary inference and is starting to pick up from, in general knowledge, this concept of disaggregation. So historically, models would be hosted with a single inference engine. And that inference engine would ping pong between two phases.
There's prefill, where you're reading the sequence, generating KV cache, which is basically just a set of vectors that represent the sequence, and then using that KV cache to generate new tokens, which is called decode. And some brilliant researchers across multiple different papers essentially made the realization that if you separate these two phases, you actually gain some benefits. Those benefits are basically, a, you don't have to worry about step synchronous scheduling.
So the way that an inference engine works is you do one step, and then you finish it, and then you start scheduling the next step. It's not fully asynchronous. And the problem with that is you would have essentially prefill and decode are actually very different in terms of both their resource requirements and sometimes their runtime.
So you would have, like, prefill that would, like, block decode steps because you'd still be prefilling, and you couldn't schedule because, you know, the step has to end. So you remove that scheduling issue. And then you also allow you or you yourself to, like, split the work into two different types of pools.
So prefill typically and and this changes as as model architecture changes. Prefill is right now compute bound most of the time.
it's usually memory bound because you're retrieving a linear amount of memory and you're doing a linear amount of compute as opposed to prefill where you retrieve a linear amount of memory and then use a quadratic memory. Funny. Someone XO Labs did a really cool demo where for the DGX Spark, which has a lot more compute, you can do the compute hungry prefill on a DGX Spark and then do the decode on a Mac.
That's faster. You can machine stratification. With our future generations of hardware, we actually announced with Rubin this new accelerator that is prefill specific.
It's called Rubin CPX. I have a question. When you do the scale out, is scaling out easier with Dynamo because when you need a new node, you can dedicate it to either the prefill or decode?
Yeah. So Dynamo actually has a Kubernetes component in it called Grove that allows you to do this crazy scaling specialization. It's a representation that I don't want to go too deep into Kubernetes here.
But there was a previous way that you would like launch multi node work. It's called leader worker set. It's in the Kubernetes standard.
Leader worker set is great. It served a lot of people super well for a long period of time. But one of the things that it struggles with is representing a set of cases where you have a multi node replica that has a pair, right, know, prefill and decode, or it's not paired, but it has like a second stage, that has a ratio that changes over time.
And prefill and decode are, like, two different things. As your workload changes, right, the amount of prefill you'll need to do may change. The amount of decode that you you'll need to do might change.
Right? Like, let's say you start getting, like, insanely long queries. Right?
That probably means that your prefill scales, like, harder because you're hitting these this quadratic scaling growth. Yep. And then for listeners, like, prefill will be long input, decode will be long output, for example.
Right? Yeah. So, like, decode decode scale I mean, decode is funny because the amount of tokens that you produce scales with the output length, but the amount of work that you do per step scales with the amount of tokens in the context.
So both scales with the input and the output. That's true. But on the prefold decode side, if suddenly the amount of work you're doing on the decode side stays about the same a little bit, and then the prefill side jumps up a lot, you actually don't want that ratio to be the same.
You want it to change over time. So Dynamo has a set of components that a, tell you how to scale. It tells you how many prefill workers and decoded workers it thinks you should have.
It also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduling on your actual hardware, on your computer infrastructure.
Not gonna lie. I feel a little embarrassed for being proud of my SVG function earlier.
No. It is. It's really cute.
I I like It's all it's all engineering. It's all technical. One thing I'm I'm kinda just curious about with all with you see at a systems level everything going on here Mhmm.
And we're, you know, we're scaling it up in in multi in the distributed systems. I think one thing that's, like, kind of of the moment right now is people are asking, is there any SOL sort of upper bounds in terms of like, let's call it just call it context length for one sort of better word, but you can break it down however you like. Yeah.
I just think like well, yeah. I mean, like, clearly, you can engage in hybrid architectures and throw in some state space models in there all you want, but it looks still looks very attention heavy. Yes.
Yeah. Long context is attention heavy. I mean, we have these hybrid models.
And most most models, like, cap out at a million context, and that's it. Like, for the last two years, it's been it. Yeah.
context co design thing that we're seeing these days is actually super interesting. It's my secret side passion. We see models like Kimi or GPTOSS.
I'm going use these because I know specific things about these models. So Kimi two comes out. Right?
And it's an interesting model. It's like, like a deep sea style architecture. It is MLA.
It's basically deep sea scaled, like a little bit differently. And obviously trained differently as well. But they talked about why they made the design choices.
For context, Kimi has more experts, but fewer attention heads. And I believe a slightly smaller attention dimension, but I need to check that. It doesn't matter.
pretty cool. Yeah.
Chinese Reddit. Yeah. Is.
Yeah. So it's actually an incredible blog post. Like, all the MLSYS people in in in that I've seen that on Jipu are, like, very brilliant.
But they they they talk about like, the creators of Kimi K two actually, like, talked about it on on on there in a blog post. And they say, we we actually did an experiment. Right?
Attention scales with a number of heads. Obviously, like, if you have 64 heads versus 32 heads, you do half the work of attention. You still scale quadratically, but you do half the work.
And they made a very specific, like, sort of barter in their system and their architecture. They basically said, hey, what if we give it more experts? So we're gonna use more memory capacity, but we keep the amount of activated experts the same.
and we decrease the number of attention heads. And kind of for context, what the what we had been seeing was you make models sparser instead. So no one was really touching heads.
You're just having Well, they they did. They implicitly made it sparser. Yeah.
For for Kimiya, they did. Yes. They also made it sparser.
But basically, what we were seeing was people were at the level of, okay, there's a sparsity ratio. You want more total parameters, less active, and that's sparsity. But what you see from papers like the labs like Moonshot, DeepSeek, they go to the level of, okay, outside of just number of experts, you can also change how many attention heads and less attention layers, more attention layers.
Yeah. Yes. Yes.
So and that's all basically coming back to just tie together is like hardware model co design, which is Harder model model context code design. Yeah. Right?
or like, really what is good at super short context tasks, you may like design it in a way such that, like, you don't care about attention scaling, because it hasn't hit that, like,
the turning point where, like, the quadratic curve takes over. How do you consider attention or context as a separate part of the co design? Like, I would imagine hardware or just how I would have thought of it is, like, hardware model co design would be hardware model context co design.
and the context that is produced by the harness is
a part of the model once it's trained in. Like, even though towards the end, you'll do long context, you're not changing architecture through I see. Training.
I mean, you can try.
You're saying everyone's training the harness into the model? I would say to some degree Or there's co design I know there's a small amount, but I feel like not everyone has, like, gone full send on this.
the harness that you think the model will be running into the model. Interesting. Okay.
Like, Vash is like the universal harness. I'll give an example here. Right?
I mean, or just like a easy proof, right?
Well, can provide a counter argument, which is what you want to provide a generally useful model for other people to plug into their harnesses. If you harnesses can be open source. Right?
Yes. I mean, that's that's effectively what's happening with Codex. Yeah.
But, like, you may want, like, a different search tool, and then you may have to name it differently.
a model? Would it be have you have people compared training a model for the for the harness versus, like, post training for I think it's the same thing.
post training. I see. And so, I mean, Cognition does this, CarSha does this, where you you just have to, like if your tool is slightly different, either force your tool to be like the tool that they train for or undo their training for their tool and then retrain.
Yeah.
I would hope that eventually we hit like a certain level of generality with respect to how to training. This not a new tool. This is not AGI.
stupid, like, learn my tool, bitch. Like, I don't know if I don't know if I can say that, but, like, you know, I think what my point kind of is is that there's like, I look at slopes of the scaling laws and, like, this slope is not working, man. We we're at a million token context.
Okay. Maybe next year, 2,000,000. We're not going to a 100,000,000,000,000.
You know, like, this this is Oh, there's so many interesting ways to get it work. I see. Work.
What's kinda funny is whenever there I I feel like we always want to see a trend that we can predict, but every time something's come, it's been like a leapfrog.
I I don't know how we go from one to two, but I imagine what what's likely to happen is we break through that from some new
Yeah. There's actually there's an interesting formalization of this. There there's an essay it's a pretty interesting essay by Leopold Aschenbrenner called situational awareness.
Okay. Yes. He introduces a concept of awareness called an unhobbler.
Right? So, you know, Leopold in this essay details, hey, I want to get, you know, like, I want to get to this point in intelligence. And I think that it is four orders of magnitude worth of, like, compute and data and training away.
And, you know, he says, oh, yeah, I think data centers can scale up by about this much. I think that you can do scale up the data and some other things by this much. But one of the things that makes the rest of that order of magnitude growth possible is these unhobblers, like these scientific discoveries that are discovered during model architecture search or training that really, really, really impact how you are able to scale.
A good example of this might be that we see a lot of models that are and this is probably very tiny on Hodler, but it's important for the performance perspective. We see a lot of models that are, like, trained with multi token prediction natively during pretraining. And per DeepSeek in their paper, they say, hey.
They decide this actually helped us in ensure state more stable convergence. But there are unhobblers that are like that, and then there are rather large unhobblers. Architecturally, a lot of our models, we have different types of attention, and one of the problems with attention is you have a lot of KV.
But people have found different forms of attention, like GrooveQuery attention and MLA in DeepSeek, multi head latent attention, that decrease the burden that KV has on the model, which allows you to grow longer in context. Yeah, and that was very drastic for DeepSeek. Yeah, context, I context length of DeepSeek is 128,000 tokens, or might be 256,000 with rope extension.
That entire context, I think it's 128,000, fits into eight gigabytes. And previously, think the LAMA 405B context of a similar size was, 40 or 80 gigabytes in the same precision. Yep.
So, like, those unhaulers, like, really decrease the stuff of that size.
an unhobbler showing up. And it's just science. More deep learning algorithms is what it is.
Could actually
give you an example of like a theory, not a theory here, but something theoretical. And a hopper that you're excited about it? Well, and a hopper that I mean, I haven't seen.
So it could be a tar pit and it could just not work. But I would be really excited to see a model that does prefill and decode differently. So a model that does prefill locally, document wise prefill, like it does it in chunks, and then you do decode globally across the entire sequence.
Because logically to me, it doesn't seem like you would necessarily need to have KV be associative between documents that have no mutual association. But that places a lot of burden on decode and pure attention within the decode phase to make those connections since the KV is, like, static at that point. And you see other techniques that are interesting like this too.
But if if you're able to do that, like, if prefill becomes local and decode is is still global, you solve that prefill quadratic scaling problem because you have a bunch of, like, small chunks that you prefill independently.
Okay. Alright. Well, let's wait and see, but I I think it'll be pretty exciting.
Fingers crossed. Yeah. Fingers crossed.
Yeah. I'm excited for prefill decode on separate hardware. So, like Yeah.
Grok acquisition. Right? Can we decode on the Grok?
Can we get super fast? I don't think I'm allowed to comment on this.
Mark is gonna shoot arrows at us. He's got a blow dark. Yeah.
He's in the side of room just like like Go to sleep. Yeah. Yeah.
I'm I'm super excited to see the team come in and, like, you know, I've gotten the the pleasure of working with some of the the grok people coming in. So, you know Yeah. I know Sunny, we've had him at the same conference that you were at.
Yeah.
And I I think you're you guys are gonna be doing some sessions at GTC. I don't know if you wanna this is a good place to plug them. Yeah.
Yeah. Yeah. So I can't speak to any LPU related sessions at GTC.
I have no idea about that. No. No.
No. You you.
the on the Grok side. Yeah. I use the associative NVIDIA U.
On the on the NVIDIA Dynamo side, we're we're giving it there are a large number of sessions. For those that aren't aware, you can actually search all of these sessions for GTC online. Just go to the GTC website.
I don't know what the URL is, but go there. Google it. Yeah.
And you can just look up Dynamo, you'll get all the sessions. Are about 20. There are a couple that are hosted by the Dynamo team.
There are a couple that are hosted by people that use Dynamo that wanna show off the results they've been able to get. But there are two that I'm really excited about. One is just the general Dynamo tutorial.
And this is the I'm going out with Harry, who's our lead product manager for Dynamo. And we're sort of talking about, like, how to use Dynamo to get better performance and also, like, where we see Dynamo going in the future. And then there's another session that I'm doing with one of our agents teams at NVIDIA to talk about sort of the future of agents in production inference.
So we're talking about there's this new horizon with respect to agents because we have these harnesses that actually impart structure upon calls. If you compare the past and the present with respect to how LM calls work, in the early days when there were chatbots, like, every call was, like, very different. There was basically no structure.
You could assume that, like, people you if it was conversational, there might be, like, some implicit structure because you have, you know, a multi turn conversation. But agents, you have this this harness that, like, abides by rules. Right?
So it imparts direct structure onto the context. And you see this there's an interesting Twitter post about how Claude code, like, structures its context so that you get as many caches as possible. I think it was by one of the PMs for Clog code, and he wrote about it.
Type of structure that the harness can impart actually goes hand in hand with the inference code design. I'm doing a talk. I don't know the session name or the session number, but I'm doing a talk.
for agents going in Dynamo and in inference in general. Yeah. I think there's only one PM for Cloud Code and SkyWoo.
The rest are there's there's DevRel, there's Boris. Maybe it was maybe DevRel. Yeah.
Exactly. I mean, let's go into agents. I think this was, like, the last part of the the the this discussion we planned.
Yeah. How have we not talked about agents? Also, you guys Well, we scheduled it.
We we we I was like I was like, okay. You know, like, let's have, like, cohesive sections or Yeah. I mean, there's a big news.
Right?
like, deployment of codecs?
Yeah. The term of uses everything. I mean, it uses cursor and we use this But that's that's a pretty big deployment.
Right? Like, that's tens of thousands of people. Totally.
Yeah. We we're curious about it. Yeah.
I mean, it goes back to the mosh pit of emails we kinda mentioned earlier or just the like, how fluid the org feels. So when there's new technology, people will just email it out, and everyone will try it. And if if it's making people's lives easier, it'll spread like wildfire.
A lot of times, Jensen will get it, and it'll be like, let's make this work across the company. Let's make this work right now. Honestly, if I was a startup, I feel like a cool hack.
If you have something that's gonna save an Nvidian's time, they'll spread it to a couple and the same thing. Right? It'll just spread like wildfire.
Like, careful before your email blows up from startups, by the way. Well, you gotta know the person. Right?
But, no, I I yeah. So, I mean, we I love using Codex. It's been a ton of fun.
Yeah. I've been using it personally. I've been using it at work.
It's been yeah. Don't know. It's been great to see the rollout.
Something really funny on the day that we got Codex and Claude Codexcess. I founded this person. His name is Carlos at the company.
He wrote an Outlook CLI. Oh, yeah. And just the CLI for email.
And this was I've been using that. Yeah. Maybe, like, four or five weeks ago.
And the site so once I got, like, Codex access, I installed the CLI. It had a skill, and I just asked it to go through all of my emails, which it's very messy. So if I don't respond to your email, I'm really sorry.
But I asked that they give me a summary, highlight any escalations that I should look at, put any thread that it thinks I should respond to in a folder, and then archive everything. And it did. So if I missed your email, that's because it didn't get picked.
So I should put a prompt injection in my emails to you. Yeah. What you should do is just FaceTime.
Shit. Yeah.
Yeah. Oh, my yeah. My SLA is highest on FaceTime.
But that was it was magic. And so sent it in a big email thread to like 500 people. A bunch of folks tried it out.
Gary team at NVIDIA is incredible. Like, shout out to them. They're they're they're trying to be.
We have that we have an amazing security team because they're progressive and they know that this is really important technology and we have to bring it in. If you think about, like, if you work at a big company, your laptop's usually very locked down. If you can only access certain things, NVIDIA engineers have those restrictions aren't there.
So you're expected to understand the risks when you try things out. And so very quickly, you know, made sure to chime in security on what we were doing. There's actually a lot that we've thinking about, especially with OpenClaw.
Right? Like, there's, you know, agents could do three things. Yeah.
Agents can do three things. They can access your files. They can access the Internet, and then now they can write custom code and execute it.
And you really only let an agent do two of those three things. If you can access your files and you can write custom code, you don't want Internet access because that's one is see for vulnerability, right? If you have access to internet and your file system, you should know the full scope of what that agent's capable of doing.
Otherwise, malware can get injected or something that can happen and so that's a lot of what we've been thinking about is like, you know, how do we both enable this because it's clearly the future, but then also, you know, what what are these enforcement points that we can start to, like, protect? And is there any directive of, like, hey, we have a company account or company agreement with OpenAI. We use OpenAI models here or, like, choose whatever.
No. No. So so I would never put any company data in a model that's not either that we don't even.
It has the most security of it. Yes. Yeah.
I like how to.
know, obviously, you could run your own models here at Nematron and and we we
an have an internal cluster so we we, you know, of course, are you in the abandon? Yeah. Yeah.
I think we're Dynamo's first customer.
there's a funny story about, like, how I got the experience that informed what we needed for Dynamo. At one point, there's a website called build.nvidia.
com and also for us, inference.nvidia.com that is allows people to try models.
It gives an API service. You can call the model with like a REST API and, you know, you get a response. I ran the model site for that and it was at one point the largest inference deployment, and still may actually be the largest inference deployment in video.
I've I've since like handed it off to some people and they're doing a wonderful This is an extremely under known or less known resource. Build.mv.
com, you can get any of these open source models, and it's rate limited, but it's free. So it's perfect for hackers because there's And and and the SLA on getting models, day zero models up is like a day. Yeah.
Like, they're they're incredibly good at like figuring out the right way to host the model to get it up there as soon as it comes out. You ran this? Yeah.
I ran I ran it a long time ago. It was originally called NVIDIA AI Playground, then it was called Oh, found the initial answer. Yeah.
And then it was called build out NVIDIA call. And I I ran the model side of it. So there there was a large multi organizational team.
How I ran how which models should we post? How should we host them? And like, what's the proportion of them?
And then, of course, there was like an SRE team that like made sure that things ran well and scaled the models as well. But I ran, like, you know, model how do we get the model to silicon?
which model also worked with our product team to determine, like, which models were important. A very long time ago. Yeah.
There's also, like, a middle ground in between there. This is, like, for the hacker try anything. There's the Brev console, then there's Dynamo.
There was also Nims. Right? Yes.
I remember it had its little moment like a year or two ago.
Is it still Yeah. No. Nim is, you know, inference Oil.
I I think it looks up for something. It's It is no longer an acronym. Yeah.
It's just a it's just a name. But, yeah, Nim is how enterprises can take our any of the any of this technology and run it with support and all of that. And so that includes Dynamo.
That includes, I don't know, all of our other optimizations that are packaged up for enterprise. Yep. Yeah.
So so you got you got a bunch of experience, like, running the sort of internal inference gateway of the grounds. Yeah. I got it.
like, Versus Code thing. You call it MB code? It's like the extension?
Right. Yeah. It a it was a Versus device.
First relates to fork Versus Code. We joke Absolutely not. It turns a while back to be like, we should have a fork Versus Code hackathon where you That's for It's just the best fork Versus Code.
We were we were we were doing a How do make a billion dollars? Someone from Versus Code was there, and he was, like, somewhat down to get involved, and I was like, oh, you should do that. That's all I was.
Then then the cool thing became for a Chrome hackathon. A Chrome. And now now IDs are not cool.
I am what's it called?
from Roboflow and Your partner in crime. We were talking about how with the new Alfa Romeo model. So NVIDIA just released an open source, the the Mercedes cars that you saw.
Drugs sound crazy. Yeah. Release.
Will you open source a autonomous driving model? I already yeah. So we were thinking, like, could we hackathon a driverless car?
Like, I have my old car. Let's just try it. Oh my We'll take it take it to, like, click trailer.
Yeah. With a treasure island in the middle of the day. Yeah.
That's why I just see. Like, everyone yeah. Like, how many how many cameras do we need?
Right? Like, one, two, three, four.
I don't know. We find space. I don't know.
I yeah. But I think we're gonna try. You just do with us.
Alright. We can see. We could even have a race.
It's like the first person to automate their driving. Let me over the weekend, we do have an autonomy track. It was fair.
And, Wemo was there. Like Yeah. NVIDIA did send people those for Groot.
No. No. Because he didn't have the driving thing yet.
Yeah. Yeah. That's that's cool.
Yeah. I think Comma Comma also has a version of this. Comma, yeah, they have open source driving.
They've they've done a fun hackathon on me. He and I also because I I really what I really want is a Tesla with Tesla level self driving Yeah. But as a smart car, like a two seater.
That's the bay basically a wheelchair with a roof. The only thing they make them, I knew. The thing is the the demand has been there.
Yeah. Think they don't waste They're this? No.
They they're like this for, like, five years. Yeah. Really?
Yeah. They were a different manufacturer.
Ways into playing. I thought it's one of those things we'll where we'll see someone buy the brand and it'll be revived. I I would buy it.
Like, I'm probably that'd be perfect. Go. Someone hears this.
Yeah. Yeah. That's crazy.
No. Because like Mercedes. Because that they're like, I think my old camera says, I'm Mercedes.
I'm And they're saying you're used to make them. Yeah. Don't know.
I feel like they own the brand. And you out. That's actually Your dream might come true.
Enough. Okay. We're we're time being here.
Fastenal. And and it's gonna and like, every time I I try to park in San Francisco, I I I have to buy a smart car because, like, 20% of the parking lots in San Francisco only fit smart cars. Yeah.
So I'm gonna take that. Really?
That's what I mean, was so small. Everybody was late here trying to put my own. This comes from someone that, like, basically does not drive.
Yeah. That's one of the the Vespa was a life hack. Yeah, exactly.
Yeah. You know what happened to the Vespa? I used have this yellow Vespa.
I left it outside the hacker house when we moved out. It's trendy. It's just it was always there and then like a month ago, it's not there anymore.
I've been meeting today. I don't know. You could've let lights off.
It's actually been it's like a TV hit. Does he you forgot about it. Yeah.
And left. Okay. Yeah.
No. It's it's probably hazard. And speaking of hackathons, I also wanted to give a brief shout out to the World Shortest Hackathon.
Let's go. You did twice? You're gonna it a handful of times?
Yeah. There's gonna be one at GTC. Oh, we're doing another one?
go through those channels. Yeah. That's like a zero the zero minute hackathon idea.
Because you just you just bring your I have noticed a night a long long time ago. You just bring your agent, and then you press the go button. You're not allowed to code.
It's just the agent doing hackathon. It's a good hidden email. Right?
Yeah.
like, drop it in. Because you don't you don't know you like supervise. Will it be a, you know, operate a browser, order a pizza?
Well, just see it like that. Snake game, you know? All the and you don't know what the task is.
Yeah. You do them. I don't know what the task is.
Like Or just, like, you don't even know what the judging categories are, then you give it the judging categories. Like, try and bring as much as possible. It's great though it turns into, like yeah.
So let's build something on DinoPod. It's a great person. Yeah.
It's all. Okay. Great.
Funny story, actually. We have a couple of people at NVIDIA. We've been working with security to, like, bring agents really close to compute.
So we now have, like, stuff where you can, like, tell Dynamo, like, go run some experience with Dynamo, like, on x cluster and just, try it right now. Like queue up. Once you get queued, like, send this request load and we've actually been able to like just like, you know, like one shot problems.
Like, we used have this problem where, you know, with with Dynamo, you have to like find the right configurations. And we sort of do it automatically for some parts of it. But you have to like a good initial configuration that you wanna use and we've just had like an agent just completely one shot that.
It goes, it gets the compute, it like runs a couple experiments, it's like, this is the best. This is this is these are part of the Pareto frontier, go run this. And then we just, like, give that to people and it's, like, faster than anything that they have.
Agent UX and agent marketing are super important. There's something we've thinking a lot about.
is, like, redoing the entire Rev CLI so that you can fetch all the different compute types that are available. I don't know. It's gonna be really soon, but then you can you can just browse what GPUs are available and then provision one, s h to it right there, and you can pipe all the commands.
But I think it goes back to, like, the ALLOC CLI. Like, if you coding agents are it's kinda funny. I feel like coding agents have been so much more effective than general purpose agents, and I think a large part of that is it just has access to the terminal, like you said.
And that means it has access to everything that you've installed into your terminal. It can run so, you know, it would write code and and it can compile the code. And if there are errors, it can fix it.
It can run your suite of tests because that's all just in your terminal. And so that, you know, then the for the idea or what have got me really excited about the Outlook CLI, we're now just churning through building CLIs for the entire, like, for the entire business suite. Slack building, Slack, also a workday CLI, CPU go.
I I've also done that for myself for some Really? Yeah. We're gonna we're gonna open source all of this and like, yeah, all the the, I mean, they're just they're they're yeah.
CLIs for the business applications. We would love for someone to run with this and like build like, I don't know, like, open CLI foundation or something. Yeah.
We NVIDIA would love to support anyone that's doing this. Like, every dev tool should really have good CLI support at this point.
accessible
by an LLM. Right? You want LLM good doc.
No. Every everything needs some CLI. Yeah.
It's kinda funny. Right? Like, we like, computing began with a terminal with a shell, but we said that it's not empathetic to humans, so we built these nice user interfaces.
And then now we have LLMs navigating our user interfaces, and ironically, we're not empathetic to the machine anymore. Yeah. Just give the the LLM access to the shell.
One thing that slightly makes me uncomfortable is, like, why do we have to build CLIs? Why can't we just expose APIs?
Like I I have I have an interesting answer to this. So that there are a couple reasons, like, there's there's, like, you know, portability is, like, one issue, like, know, like, sometimes APIs are not like discoverable or like reachable by by some, you know, types of things. There's some element of locality.
Right? Like like the CLI is like literally you interfacing with your like local system, which is a little bit different. You could still do it by API.
But, like, there's this highlighting of, like, what is the difference between, like, a CLI and an MCP. Right? Like, they kind of occupy at the same purposes and you call them, it does something on the system and and that's done.
I think that in pre training, there's just an enormous amount of command line data. Yeah. Yeah.
Like, even let's ignore let's ignore RL, like, you're doing no harness you're doing no harness post training. Just the amount of, like, CLI versus API documentation for just like navigating this world of the CLI in your file system through that is just enormous.
Yeah. Yeah. Right?
I think there's there's a couple of things too. Like, if let's say we wanna so one, I think your intuition is right. The CLI is just wrapping the API.
Right? So, functionally. Functionally, right?
Yeah. And I think it's nice because one, you're you're being very specific and pedantic even of what and that's really good because you're describing the problem space. So, you know what the, I don't know.
I don't wanna call it, like, what the the space for vulnerability. You know what network calls you're making. It's not arbitrary, that's undecided on the fly.
That's, like, predecided, which is important from a security perspective. But then if you were to write a bunch of API requests, you would probably do that. I don't know.
Would the model, like, use Python to do so? I kind of like that everything like, a CLI is just bash because it's ubiquitous. Like, it's just there, and you don't have to make sure that there's certain environment variables that are set up.
Like, if your Python versions of it, the MyPython version were using the same model to go do the same thing, is it gonna write, like, different code? It probably would. And so it's kind of a nice nice to do work.
Right? We are human. Yeah.
I think just, making those decisions happen ahead of time versus yeah. One last thing on this sort of agent, I guess, maybe colocation or whatever you call it. One pattern I'm tracking for this year.
I always try to think about what's the theme of this year gonna be. Last year, definitely coding agents. This year is definitely coding agents breaking out of containment into brothering.
Gabriel Waltz, I go definitely ask him. So you rent a human? Yeah.
Oh, yeah. Yeah. I'm on there.
Are you around? When I pay pay, sir. Yeah.
I'm like, $5,000. I'll do anything. Really?
I think so, I need I need the my powers from Costco. But I think the best part is only the agent can book me, you know? Yeah.
Yeah. Yeah. It's very usually like it's just like another labor marketplace.
Mechanical Turk was this. So it's this way. I have a weird story with why I did it.
So back to your example of just giving agent access to compute. Right? Yeah.
You guys are GPU rich at NVIDIA. Yeah. I hooked up He's not shy about you.
I have I have a twenty four seven agent running. I hooked up to RunPod. It doesn't shut down instances, and I'm like, I've tried prompting you.
I've given you instructions. Shut down when you're done. It's like, I need to keep it warm.
I'll need it soon. And it's horrible on time estimates too because, like, they realize it's like, yeah. I'll need it in forty five minutes.
Forty five minutes, I'll shut it down. Forty five minutes of human time is actually three minute of agent time. So it's like, I'm booting it up.
I'm waiting. I'll just leave it on all night. And modal's good at shutting down after some inactivity.
I had it on my local server, like a little dual GPU thing.
It just stays on. I have a little space heater at home now. But careful.
So, basically, you know, they don't care about the concept of money. Just burn it. I need it.
It's useful. And another thing where DGX Spark will be really nice. Like, I I think I'm looking at it as it's possible.
Super useful for agents because, yeah, you you buy it once, you plug it in, and they it can rip. I'm gonna make a I'm gonna make an NVIDIA ad here. K.
like, RTX 6,000 cards Pro. Pro are only, like, I think it's $8,000, slightly cheaper. Yeah.
Well, it's much it's much cheaper than the data center cards. Yeah. And it's got 96 gigabytes of URAM.
So if you and your your crew want to go, like, run a local agent for, you you know, you you in the home, I feel like it's got a significant amount of vRAM. I thought about purchasing this and running in my basement, except my neighbors didn't hate me. It's just a single, like, two, three slot GPU.
It's small. Yeah. It's a VCIE.
Yeah. It's a VCIE. GPU.
You can go by that. But, I mean, the big difference against, like, the RTX, like, gaming GPUs is it I mean, obviously, it's, like, BlackBell like, it's a pro GPU and has a lot of vRAM, which means you can run pretty large models on You can stack four of them for the max q in a system. That's that's a beast.
It's beefy. You can run what is that? 96 z or anything?
96? You don't know.
Also, they they are slow. They're not I mean, performance of speed will be somewhat slower.
like Oh, yeah. That that's true. So, again, the big learning, economy of scale allows you to do things that allow you to get both speed and throughput.
Like, you can run I'll give you an example. There's an optimization called wide EP. I'm not gonna go into it fully, but, like, it featured heavily in in inference max for DeepSeek.
And there's a there's a great set of stories from NVIDIA and from semi analysis about, like, why YEP is important. But for like MOE models, it's like basically essential and you run it like the a level of parallelism, the level of scale up parallelism used for it is like 32. So it goes beyond that eight barrier and it, like, really, really, really is important to have that m NVL 72, g b 200 NVLink to serve at scale.
And, like, it's, like, I don't remember, like, the, you know, cost improvement.
cheaper per token for, like, a lot of the curve Yeah. Which is crazy. Yeah.
And normalized per GPU, obviously, because part of the GPU's cost or the code the GPU's part of the cost. One thing I'm exploring is the sort of this year is also the year of the sub agent, where you have the main agent, but then that also kicks off tools which are in themselves agents that have limited agents. And so the Low contact.
Low good deals. Whatever. Right?
Different prompts. So for example, one thing that Cognition does is before you kick off a search, they do is, like, a fast context model where you kick off April or you just search across the code base. How's all that?
That is better than indexing a lot of the times. Not not all the times, and you should still index for some things. But, like, the idea that agents should be able to command sub agents and probably run them, like, maybe close to inference is why I don't know if that's, like, architecturally possible or even Yeah.
We're we're thinking about that for Dynamo. That's, like, our big theme for the year. Because, like, you know, like, if you can design that into your stuff, then a lot of people a lot more people use it.
Right now, it's, like, just kinda theoretical because you do pay a lot of, like, back and forth coordination costs. I I think you'll net speed up, though. Right?
Like, even at a basic level, speculative decoding, you're running a small model.
but it's not funny. That is one example. Yes.
Yeah. But this is like a little bit like different with like agents. Agents.
Yeah. This is not speculative. I I think I think there's like a summarization of that trend that I like to do or I like to say to my team.
It's like, this is the year so there are two things. This is the year system as model. Right?
Where like instead of having like a single model be a thing, you have a system of models and components that are working together to like emulate the black box model. So when you when you make an API call to something that's like like a multi agent in the background, it still looks like an API call to a model. You're still getting back to your Grids.
But under the hood. Yeah. Under the hood, it's like a billion different models.
and with other libraries and media where we're looking to help manage that complaint. Yeah. It's funny because we we actually for CES, we just released the model router Oh, yeah.
For DGX Spark, where you can have a local model that's running on the Spark, and then also a foundational model, and then the model router decides when to send queries to which one. So it's no longer this, like, either or. It's used the best of everything that's available to you.
You have a good post training bottle that's running on your If you are leads to the also the bread functionality of being able to manage the Spark. Oh, that'd be cool. Oh, yeah.
I did deal with it. Need to Jerry requested it up. Here we go.
I actually like a question, like, I I like to, like, extend and flip over. How much longer do you guys think, like, agents are gonna be running? Because that's one thing I've been throwing around.
Like, what happens when I mean, always are. It even affects the like, back to the prefilled decode.
Right? Like Yeah. Codecs is, I'd say, compared to Cloud Code, it's much longer at tasks.
Like Yeah. That thing will like to run six, seven, eight hours. I'll run it overnight.
Yeah. And I'll I'll go back and I have, like, a little crappy logging software I use, and there's just times where it wants to, like I'm gonna go deep on research, and it'll eat up 80,000 tokens, go on another, go on another. Yeah.
Just eat through tokens, and, you know, that's part of it. Like, at the end, it does it does hit a long task and I think you only see that. That expand.
Yeah. I yeah. There's insatiable demand for tokens and every improvement that comes kind of just makes our demand even higher.
It's kinda funny, right? Like, if you have, a teammate and you ask them to do a task and they're like, should I save some effort and not think too hard about this task? I'm like, fuck no.
I mean, my favorite was right. You can have four shots. Right?
Like, the original codecs before the app, you you why do one call? Like, give it four attempts. Just just use all the tokens.
Like, alright. Try more. Try again.
Try more. It's like it's like the the meta index, right, is the thing that tracks, like, how long models are able to run. I expect that we'll just see, like, log linear, if not log super linear growth.
We will see before the end of the year an agent that is capable of running for longer than twenty four hours with, like, self consistency the entire time. I I would also poke at different domains having different desires. Right?
I'm getting slightly frustrated at twenty minutes per basic query. Sure. You can optimize, you know, six, eight hour.
I don't see myself shooting off many one week agents. Right? Someone doing like, okay, GPU kernel research or medical or biological.
Like, you know, in in those domains, sure, shoot off a lot that take them out of hope. So, like, I think it will be somewhat domain specific because you also really need to train that in, right? You know, that's funny.
One thing was doing your taxes.
Right? Like, that's tax. Yeah.
It's got a month. Yeah. Yeah.
Okay. Yeah. Exactly.
Get it right. I wonder if, like, this may be school day. So that's sort of, like, speculative decoding is, like, your agent figuring out what you might be prompting it the next day at night and, like, prefetching.
Yeah. You know where you look at that. Yeah.
Really? Branch branch prediction. Oh, well, no.
That well, that's that's too that's too low level, but yes.
Sorry. Yeah. Yeah.
Yeah. One question I gotta get is, so, like, we actually did record a part with the the beta folks. Was that right here?
Their chart is the human equivalent work, hours of work rather than how long the agents themselves are are being autonomous, and that there's a huge difference. Right? Like, human work five hours, agent work thirty minutes.
Yeah. Five hours. Right?
So, like, that that that chart that you see is them estimating what the human equivalent replacement is. I think the I think, actually, Anthropic released a more recent chart that showed cloud code autonomy from their production traffic numbers, and that was twenty to forty five minutes. That's roughly where we are.
So Yeah. Yeah. That's that's the sort of realistic thing.
I mean, I I do think, like, there's experimental setups where we can just, like, Ralph Widom may just prompt it to keep going Right. When it stops, and, obviously, he can that can go arbitrarily long. I feel like from my experience, around yeah.
I guess twenty to forty minutes seems right for when I'm using, like, Codex or Cloud Code. But then, like, what I always try to just, like, if I wanna spin up, a new there's a net new project, I'll I'll often start with Replit. And, like, it'll be for the BladeBand.
Yeah. But, like, spin up like, they're they're new, like, from the v three agent, like, it'll spin up a web browser and, like, click around and discover new bugs and just keep churning. So I I think, like, my longest was, over an hour that, hey, I've been churning.
I think before we see super long running, I think there's gonna be a bit of an efficiency hit. So, sure, you can take an hour and go down paths, but you also want you wanna be more efficient. You wanna be smarter in your reasoning.
Right? So I think that'll actually go down before we go back up.
non optimized systems just for the heck of it.
you know, they are expensive. Like, going from dense to reasoning models, that's an added cost. Right?
You're paying for a lot of tokens, and it doesn't make sense to just scale stuff that's not optimized. So there's there's always that little balance. Yeah.
But, you know, yeah, I think you'll see both sides of it. Yeah. So 2023 was super exciting.
I think if you were in SF, you were like, okay, I know this is gonna be a huge world changing moment, but it seemed like, you know, no one had known yet. And maybe even before. Was it 2022 maybe?
Yeah. Yeah. I would say, yeah, like, Rune had this tweet where, like, everyone was in SF from, like, 2021 to 2023.
Yeah. Like, understood what it was like to be, like, early. Totally.
Yeah. 2021. That's when had my first OpenAI account.
Yeah. It would it was crazy. And I remember it was so funny because at the time, SF had not been doing well.
So pretty much what it felt like was the concentration of founders in the city had had risen because where my neighbors were used to doing a bunch of stuff, those people had all laughed. So the only people that were still in the city were people that really wanted to build it was cheap tech. It was yeah.
It was also way cheaper. I feel really bad about anyone who is trying to get rent now. But there was Cello was they had a huge office.
So Blockchain, it, took over the the old Casper Building. Yeah. They had the showroom and they had the, like, the what would it I think it was like the back warehouse.
It was and it was a huge office. And It's right across an opening eyes in New Orleans. Yeah.
It was in the original arena. I named the arena because of it. Yeah.
Yeah. And so it was really exciting because, like, Roboflow, I think I forgot. They're really Lify.
Yeah. Mint Lify. Yeah.
Brev was there. You guys were there. I remember that was actually it was there that you bought the ai.
engineer domain. Yeah. I didn't know what I was gonna do in AI.
I just wanna do something. But it was kind of this it was a really fun moment where we were kinda all in this solo space, and it I don't know. It was it was a really cool community especially being so early.
Yeah. And so it they will then you got me early cruise access. Oh, yeah.
And so there was a going period of time that both cruising at Weymos were just free. Yeah. Always singing.
If you had yeah. Mean, they're they're so back. Cellos opened again.
Yeah. So nature zooks? Zooks is doing.
Zooks at robotaxi.
Yeah. So Totally. But yeah.
And so it's actually really cool that you guys have this studio so close to Cellos. Yeah. This rock climbing gym right around the corner.
yeah. So, yeah, it's a it's an awesome block. Cool.
Yeah. Just and well, you've been a little a San Francisco ownership, but I do think one of one thing I try to do with the podcast is, like, bring, like, what it's like to be in San Francisco to the rest of the world. Yeah.
And also just, like, maybe give El Teppa Taqueria. Yeah.
My favorite tacos in the city. And Yeah. Stick and shrimp.
I know it's very good. Yeah. And I guess what it's like to be in San Francisco, I think, is just everyone seems to be super supportive.
Sometimes I feel like the city believes in you more than you do. And even I don't know if you remember, but I remember posting my first blog post, and I had met you on Twitter. And you gave me, like, an hour of your time super randomly, and you kind of coached me through writing content for developers.
And I was trying really hard not to come off salesy or plug myself and so I kind of stripped all personality out of the blog post. Yeah. And you you brought that out.
You're like, people don't. It's it's okay to talk about what you're doing. Like, you don't have to be weird about it and I remember just that that I think that really helped me kind of figure out what our voice is and not shy away from it and so, always really grateful for you.
Hey. You inject your voice into, like, everything now. It's actually Okay.
Actually a huge advantage to be, like, very genuine about what you care about. Yeah. Yeah.
Yeah. Like, imagine, like, some some of our interesting DMs you and, like, he's like, can you give me feedback on this blog post? And it's pretty boring, and you're like, fine.
Like, you know, he looks interesting. I'll just do a Zoom call. And then you meet this guy.
Yeah. Right? He's so energetic.
Like, how do Just be right there. That's And but and I I think people are trained to write a certain way in school, and Yeah. They never Totally.
See there's, like, a broader world. A lot to unlearn. Writing writing is thinking, and, like, everyone thinks differently.
So, like, you you might as well just, like Yeah. Yeah. Write your way.
Cool. Well, thank you for, indulging with us. Really broad breaking discussion, but I love like, you guys are, like, sort of, like, the sort of young faces on video with so much energy and but, like, also a of typical death, and I think people learn about for this session.
So thank you. Hell, that's awesome. Thank you, guys.
And thank you for everything that you've done, NJ. Good to talk to you. Yeah.
NJ, the podcast, all the above. And see you at GTC? Yeah.
We're looking forward to it. Yeah. Cool.
Thanks. Awesome. Thank you, guys.
Thank you.
Shared via Hopper