The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
20 May 2026 59 min
0:00 --:--
Episode Description
Logan Kilpatrick and Tulsee Doshi of Google DeepMind join for a first-ever in-person episode recorded just days before Google I/O, covering headline launches like Gemini 3.5 Flash, the Omni video generation model, and the new Gemini Spark agentic product. The conversation digs into Google's strategic decision to lead with cost-adjusted efficiency over raw capability, how DeepMind now ships a full agent harness rather than bare models, and technical questions around context window limits and know

Summary

This episode features Logan Kilpatrick and Tulsee Doshi of Google DeepMind discussing Google's AI strategy and new launches at Google I/O, including Gemini 3.5 Flash, the Omni video generation model, and the Gemini Spark agentic product. They delve into Google's focus on cost-adjusted efficiency, the development of a robust agent harness for product integration, and their perspective on recursive self-improvement and AI collaboration.

Chapters

Google I/O & DeepMind LaunchesThe host introduces Logan Kilpatrick and Tulsee Doshi, discussing Google's strong market position and upcoming launches at Google I/O, including Gemini 3.5 Flash, Omni, Antigravity, and Spark.
Gemini Flash & Model StrategyTulsee Doshi and Logan Kilpatrick explain the strategic importance of Gemini 3.5 Flash, emphasizing its balance of intelligence, speed, and cost-effectiveness for broad application across Google products and external developers.
Agent Harness & Product IntegrationThe discussion shifts to Google's agent harness, Antigravity, and how it standardizes and accelerates AI product development across Google's vast ecosystem, while also allowing for extensibility and diverse use cases.
Recursive Self-Improvement & AI CollaborationThe guests share Google DeepMind's practical approach to recursive self-improvement, where AI models like Gemini are already used to enhance their own development process, with humans remaining in the strategic driver's seat.
AI in Daily Work & Audio ModalityTulsee and Logan describe how AI, particularly through audio input like the Gemini mic, is transforming their daily work by facilitating thought organization and code generation, highlighting the growing importance of audio as an interaction modality.
Context Windows & Diffusion ModelsThe conversation addresses the plateau in context window growth, attributing it to cost-prohibitive factors and a strategic shift towards smart context usage, while also providing an update on ongoing diffusion model research.
Future of Audio & Google AILogan and Tulsee express excitement for the future of audio as a primary interaction method with AI models and hint at many more innovations to come from Google DeepMind beyond the I/O announcements.

Topics

Google I/O launchesGemini 3.5 FlashAI agent infrastructureMultimodal AICost-adjusted efficiencyModel scaling strategyAgent harness developmentProduct integrationRecursive self-improvementAI research collaborationAudio input modalitiesContext window limitsDiffusion modelsModel psychologySearch grounding

People

Nathan Labenz (host) Logan Kilpatrick (guest) Tulsee Doshi (guest) Craig (mentioned) Sundar (mentioned) Andrew Lee (mentioned) Demis (mentioned) Varun (mentioned) Anka (mentioned) Rune (mentioned)
Key Concepts (10)
Cost-adjusted performance Pareto frontier — Google's strategy to prioritize models that offer the best balance of capability, speed, and cost-effectiveness, rather than solely focusing on the most capable model in absolute terms.
Agent harness — A robust infrastructure provided by DeepMind that wraps around bare models, helping to elevate and standardize AI experiences across Google's diverse product surface by enabling agentic functionality.
Model psychology and welfare — The consideration of how AI models 'feel' or behave, including concerns about 'doom loops,' discouragement, or anger, and the efforts to evaluate and mitigate such 'psychological distress' in model interactions.
Recursive self-improvement — The process where AI models contribute to their own development and improvement, which Google DeepMind is actively investing in, seeing Gemini as a partner in its own development process.
Model harness product symbiosis — The integrated development where the model is trained in conjunction with the harness, and the harness powers agentic product experiences, creating a foundational layer for building across various applications.
Harness diversity — The ability of a model to support a range of different approaches to tooling and orchestration, ensuring it can be leveraged effectively by enterprise customers or developers building their own use cases, not just within Google's proprietary harness.
Jagged intelligence — A concept referring to AI models that perform exceptionally well in specific, trained contexts but struggle to generalize or perform consistently in different or unfamiliar environments or harnesses.
Human in the driver's seat — Google's philosophy for AI development, emphasizing that tools are built to keep humans in control and making strategic decisions, especially given the high cost and risk of large-scale AI training runs.
Model eats the scaffolding — A metaphor describing how, with each iteration and improvement of an AI model, it absorbs and internalizes functionalities previously handled by external code or infrastructure, making the model more capable and integrated.
Smart context usage/compaction — A strategy for managing large amounts of information by intelligently selecting and bringing only the most relevant elements into the model's context window, rather than simply expanding the window size, to improve effectiveness and reduce cost.
References (40)
The Cognitive Revolution by Nathan Labenz and Erik Torenberg podcast
Google I/O event
Gemini 3.5 Flash model
Omni model
Antigravity project
Spark product
Anthropic company
OpenAI company
No Moats Memo article
NVIDIA company
Apple company
DeepThink concept
Gemini Omni Flash model
Gemini app product
AI Studio tool
Agents API tool
Gemini Live product
Brave Search API tool
Waymark company
Sequence company
Cognition company
Incident IO company
Runway company
OpenRouter company
Taskill by Andrew Lee company
LangChain tool
Roboflow company
Blueprint Pro project
Claude model
Claude Code model
NotebookLM product
Gemini mic feature
Flow product
YouTube product
Anden Labs company
Exa company
Google Cloud company
Cloud Marketplace platform
Cloud Model Garden platform
Artificial Analysis
Transcript (79 segments)
Speaker 1

Hello, and welcome back to the Cognitive Revolution. Today, after some 340 episodes, I am very excited to share the first episode that I've ever recorded in person with fan favorite Logan Kilpatrick, member of technical staff at Google DeepMind, and Tulsi Doshi, senior director and head of product for Gemini models. The occasion for this conversation is Google's annual IO event, where they're launching the new Gemini 3.

5 flash model, all sorts of agent infrastructure and AI product integrations, and plenty more. We recorded on Friday, May 15, just a couple days before the event. And while many at Google, including my brother Craig, who's giving a keynote on Wednesday, were working overtime to polish their demos and presentations, the overall vibe, at least compared to the rest of the AI space, was one of relatively relaxed confidence.

And why not? From twenty twenty four to twenty five, Google grew annual revenue by $50,000,000,000, as much as Anthropic is pulling in today. And they still have 25% of all global compute, the deepest pool of research talent anywhere, and the most comprehensive AI portfolio of any company with top tier positions not just in language models, but also self driving cars, medical and life sciences, and robotics.

So after discussing the headline launches that they're announcing this week, which also include a new video generation model called Omni, which they hope will create a nano banana moment for video, a new and improved and more agent focused antigravity, and a product called Spark, which will bring more agentic functionality to the main consumer Gemini app, I really wanted to take a step back and dig in on Google's overall AI strategy and philosophy. We discussed their decision to lead with the flash model and more generally to emphasize the cost adjusted performance Pareto frontier, whereas Anthropic and OpenAI are clearly much more focused on competing to have the single most capable model in absolute terms. We talk about how DeepMind is no longer shipping models in isolation and leaving it up to product teams to figure out how to use them, but instead now providing a robust agent harness, which should help elevate and standardize AI experiences across Google's vast product surface.

We get into the weeds on questions like why context windows seem to have mostly stopped growing, why Gemini models knowledge cutoff is now more than a year ago, and whatever happened to that diffusion model line of work. Perhaps most importantly, we discuss how the team at Google relates to the AIs they're creating, how they're thinking about things like model psychology and welfare, and their views on recursive self improvement, which as you'll hear is definitely a part of their plan, but not something that they seem to be so singularly focused on as other AI leaders. Overall, I think this is a great window into the thinking that underlies Google's AI research and product development, which has clearly sustained the company's historic run far beyond the point that many analysts had written them off.

With that, I hope you enjoy my first ever in person conversation with Logan Kilpatrick and Tulsi Doshi of Google D Mind.

Speaker 2

Alright. Well, we are here live at Google Headquarters in the library at Gradient Canopy, the first ever in person recording of the cognitive revolution.

Speaker 4

Logan Kilpatrick and Tulsi Doshi, welcome. Thank you. This is an honor.

I didn't realize this was the first in person episode. Yeah. 300 with 50 plus.

Yes. And it's all been from my home office in Detroit until today. That's awesome.

Well, thank you for being here.

Speaker 2

space, especially around IO. It's a zoo. Yeah.

It's always, it's always a good time here at, Google HQ. So you may or may not remember the no moats memo. We've just passed the three year anniversary.

It was 05/05/2023. And in the intervening three years, Google has added $3,500,000,000,000 in market cap, which is more market cap than all but two other companies in the world. Those two are NVIDIA and Apple.

So the molt the moats, I'd say, are holding up. Here we are at IO, and I'm sure there are gonna be some exciting new things that will be deepening the moats. So first question, tell me what are we launching this week to try to deepen those modes?

Speaker 3

A lot. So a lot of exciting stuff. So let's see.

Let's start with some of the modeling side of things because that's really exciting. We have our 3.5 series coming out starting with 3.

5 Flash, at IO. We're really excited about 3.5 Flash because I think Flash does this really awesome job of being at the sweet spot of being really smart while also being really fast and really cost effective.

And so Flash is incredible. It is, like, three times faster than other of the large models. It's significantly cheaper for being able to still drive these, like, really awesome magentic encoding workflows, and we've been using it internally a lot, which has been really fun to kinda see that play out.

So that's one big piece, which is 3.5 Flash I'm really excited about. We're also releasing Omni, Gemini Omni Flash, which is a video generation and editing experience.

What's really exciting about Gemini Omni in general is it's our push towards being able to bring all modalities in and all modalities out. And the first way this is really manifesting is in this video editing context. So you're gonna be able to make really awesome, videos.

You're gonna be able to put your own avatar into the videos, which is gonna be awesome. I've been having a bunch of fun playing with that too. We're continuing to upgrade anti gravity and bring more into the developer experience.

So Logan can talk more about the developer experience overall, but 3.5 Flash and anti gravity are really gonna come together to build build something great there too.

Speaker 4

which also builds on 3.5 Flash to build kind of more agentic experiences into the Gemini app. So we've got we've got a slate of cool things coming.

Yeah. I think beyond the models, I think the other headline of the story is just like agents, agents, agents, you know, the meme of Sundar from two years ago or last year, he's saying AI, AI, AI the whole time. I feel like this year is agents, agents, agents, agents.

I think it's it's cool to see like this. I think this is the first year where we have and actually, I think this is not just us, but just like ecosystem wide. This like model harness product symbiosis that's sort of taking place.

Like the model is sort of trained with the harness. The harness is powering the agentic product experiences. Gemini Spark in the in the Gemini app sort of being one example of that.

The it's powering the five coding seven AI studio. It's powering the agents API for developers. It's powering I think there's something else maybe that it's also not.

It'll roll out to other products across Google and sort of be this sort of, like, foundational layer to build on top of, which is really exciting. So I think not just, like, developer products, but our our sort of consumer products, and I think probably even more widely in the future across the rest of the Google product suite. And I think there actually, there's one interesting thread of this, which is historically, like, Google didn't have this, like, through line of, you know, something that carries across all of our products.

I think then then it was Gemini and then sort of all of a sudden every Google product has Gemini and sort of getting them all stitched together and making all those products experiences great.

Speaker 3

which is really interesting. Yeah. And I think what's been really fun from the modeling standpoint there is one thing that we did with Gemini three that we're really continuing, I think, with the 3.

5 line is really bringing the model to every all of our products. So, you know, 3.5 Flash will be in Gemini app.

It will be in AI mode in search. It's also powering antigravity. It's also powering agentic experiences in AI Studio in, you know, Gemini Spark.

Speaker 4

really awesome. It's really hard too. Think I think it's actually gotten harder to do.

I think it's like, it's almost like it's, it was maybe tongue in cheek. It was like kind of easy before because you just launched the model on a couple of services and sort of, wasn't that I feel like now it's like you're sort of you have the constraints of the very wide array of Google products that are just for totally different users. Sort of I think actually that credit to the model team sort of trying to, you know, find the fine line for all these different places because we're not just building for search.

We're not just building for developers. We're not just building for cloud customers. It's not just for the Gemini app.

Speaker 3

yeah, which is exciting. I think one other thing I'll say is, like, one thing that's cool about this I o two is what we're doing across modalities. So, you know, I think Logan said agent agent agent, which I think is true.

This this I o is really about bringing models to action in in in in that kind of real world sort of use cases. But I think what's also cool is we have the flash model, which is really about building these kinds of coding and agentic use cases. We have Omni, which is really about kind of what is this, like, multimodal vision look like.

And then also, actually, Gemini Live is getting an upgrade too. Gemini Live is getting faster. It's getting smarter.

The model is much better at detecting background noise. So it really does actually feel like a partner in a lot of ways, and I think it's kind of cool that we're also able to draw this through line across the different ways you might wanna interact with a model and the way different kinds of ways you might wanna consume content, which I think is also really cool.

Speaker 2

Okay. I've got, like, seven different directions I wanna go. How about in follow-up.

Trying to seed them all for you. Yeah.

Speaker 1

Let's start with just the model.

Speaker 2

it's interesting to start with Flash. One thing that I recall, I don't know if it was two IOs ago or whatever. Right?

But there was gonna be three sizes of Gemini model at one point in time. And There are three sizes. Yeah.

Pro, flash, flashlight. Yes. But we never saw the Ultra.

It's kinda what I'm what I'm alluded to.

Speaker 4

Yeah. Three were promised, and then we added we took one off the top and added one at the bottom. DeepThink two.

DeepThink two, actually, which is like a fourth scaling dimension from a model perspective.

Speaker 2

That's a run time scale, though. It is. Run is, for sure.

Yeah. Well, I guess two questions on, like, why no Ultra. One is, like, is it a compute limitation?

Like, how are you guys thinking about which model to release? I'm you know, in addition to the 4,500,000,000,000 0.5 added, 4,800,000,000,000 market cap, Google enjoys 400,000,000,000 a year in revenue.

I was interested to learn. And traffic though is growing extremely fast. Like, might hit a 100,000,000,000 at the end of this year, maybe even, you know, in the third quarter.

Who knows? And it seems like the revenue there is really driven by people's extreme willingness to pay for the very best model that they can get their hands on. Maybe not at any cost, but, like, relatively price insensitively.

So I'm wondering, like, why no Ultra?

Speaker 4

and yet we haven't seen it. I promise I didn't plant this question. I'm always, you know, poking Tulsi on the side.

Speaker 3

it. One. This is my favorite question, so I'm I'm glad I'm glad we're meet you.

No. I think, you know, you're you're right that I think there is a slice of of users who are definitely willing to pay for a certain level of quality. And I think we really do believe that the pro model has been like really pushing that that quality.

But I also think for us we've seen so much value from the flash and the flashlight dimensions because we also see a extremely large number of users, especially if you're thinking about building, for example, consumer applications. Right? If you think about the Gemini app, if you think about search, when you're serving to that kind of scale, latency really matters.

Cost matters. Right? Because actually you find that users aren't willing to wait, right?

So we find that even when we tweak the model and you know hurt latency, we actually see that play out in our live experiments on search and the app even if the model is hugely better from a quality perspective because what you're asking users to do is wait. And so I think for us, like, part of the reason why we ended up introducing this flashlight SKU that wasn't necessarily part of the, you know, the original two point o series was because we really felt like there's actually a large scale demand for this depending on the types of use cases, especially when you're when you're talking at that scale.

Speaker 4

the full range of what kinds of customers we can serve both internally and externally. Like for our products, the flash and flashlight skew matter a lot for our ability to actually serve to the Google populace. And so we also imagine that that's true for external, you know, enterprise and developers, and I think that's played out to be true, you know, as we've been been actually seeing this in action.

Yeah. I think the the two things that I'll add is and there's, like, probably a more nuanced technical story on sort of, like, the ultra thread, but it's it's not like it's also not like the pro models haven't scaled up over time. So, like, I think there is, like, there's, you know, there's a story that you can spend at the end of the day the naming of these things as marketing.

They definitely are getting extremely They're getting larger. They're getting more powerful. There's the test time compute scaling with DeepThink, etcetera, etcetera, and all types of stuff in that dimension.

I think it is possible you could sort of, like, put the Ultra brand on some of these things.

Speaker 3

but it hasn't been that, like, we haven't kept scaling up. So I think that it definitely has Yeah. There's almost been a conversation every time we scale up of, like, should we call it ultra?

Yeah. And and what does that brand mean? Because we we could, but there's sort of also a question of, yeah, how do we keep consistency for users also kind of series to series?

Yeah.

Speaker 4

like, Google and specifically Google DeepMind's mission is to, like, build AI responsibly and make sure it benefits, like, all of humanity. And I think, like, that is, like, so deeply tied to the, like, Google product surfaces in which, like, we're serving, what is it, like, eight, two plus billion user products or whatever it is. And so at the same time, they're, like, obviously, the frontier matters.

Obviously, having great models that are really expensive and really, really intelligent matter, and there's tons of use cases, for that internally and for our customers. You also need to do the the scaling up to billions of users for for us to, like, actually do the thing that Google needs to do to achieve the mission. And I feel like we've we've done a good job, hopefully, of, like, trying to walk the fine line of actually continuing to push the frontier and build great flash models.

And I actually think those two things are, like, more tied together. You you know this better more than I do, but, like, more tied together technically. Like, it's, you know, it's hard to make great flash models if you actually don't have a great pro model and and and vice versa.

So, we'll definitely keep pushing the frontier on on both of those things. Yeah.

Speaker 2

by analogy too, it's hard to make a good Flash model if you don't have a good Pro model. People sort of think that there's, like, an Ultra model internally that's the mega training run that's then being used to, like, help train Pro, which maybe in turn is being used to help train Flash. Is that true?

Speaker 3

distillation to the midsize? I mean, we definitely use distillation as a way of kind of bringing down bringing down our sizes. So you you will see that, like, pro influences flash, influences flashlight.

We also do the reverse where we scale up. Right? So you you take the pro you take the flash recipe and and scale up to the pro recipe, for example.

And we do have I think what's been really especially over, like, seeing as we've used even antigravity into this point of Logan made with the harness, I think we've been seeing a lot of examples actually of leveraging pretty awesome models to drive progress internally. Actually, like, one thing Varun demos on stage on Tuesday is basically, like, being able to leverage a bunch of sub agents to go and complete a bunch of tasks and come back. And you can actually try that as, like, an early preview in antigravity today if you go to slash teamwork.

And, like, that's an example, I think, of something we've been using internally, which is an extremely smart model, and it leverages both the combinations of the best of Gemini 3.5 as well as inference techniques. And and, like, you're able to actually accomplish so much.

And I think that's that's the kind of direction I'm excited for us to go into more. So I think we're kind of pursuing all of these fronts. We're scaling up from from the pretraining and kind of frontier perspective, and I think that's been really continuing to show gains.

There's a bunch we're doing on the post training side, and then there's also just a bunch we're pushing on on the inference side. And then that plus, you know, trying to make sure we're we're working with the harnesses. I think we're gonna keep getting things that we're using internally that we even started to push out externally through previews.

Speaker 2

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

The cognitive revolution is brought to you by Brave. If you want to stop hallucinations, empower your AI agents to do their own research with the Brave search API. Brave offers the only search API with its own index at scale.

It's lightning fast, excels in rag pipelines, and it's a leading search option for Clawed MCP and OpenClaw. I've built Brave Search into my personal AI infrastructure as a core tool that all agents can use anytime they need it. To find guest headshots and company logos for the podcast, they use Brave's image search.

To build small business profiles for use in my Waymark prototyping work, they use Brave's place search. Across all use cases, my agents tap into Brave's index of 40,000,000,000 high quality pages tens of times per day. It's the only global scale index outside of big tech, which means no Google scraping and no SEO spam.

Plus, with true zero data retention policies, you can meet compliance obligations and rest easy. Pricing starts at just $5 per thousand API calls, and you only pay for what you use. Sign up now and get $5 in free credits to start and empower your agents to start calling the Brave Search API today.

Most billing platforms were built to send invoices and assume your pricing is simple and predictable. But if you're building an AI product, a fintech tool, or a developer platform in 2026, your pricing is anything but. Usage tiers, consumption billing, and bespoke enterprise contracts are now the norm, and you're probably managing it all across disconnected tools and fragmented systems.

Sequence handles the entire revenue workflow from contract to cash. Quoting, invoicing, metering, revenue recognition, plus Sequence agents that automate the manual finance work that usually takes teams days each month while also helping them to collect cash faster. Companies like Cognition, Incident IO, Runway, and OpenRouter use Sequence to run their full revenue process between CRM and ERP without the spreadsheet mess.

If your pricing has gotten more complicated than your current billing setup can handle, check out sequencehq.com, and use the code Cognism in the source field when you book a public demo to save 20% off year one.

Speaker 5

End this.

Speaker 1

So let's talk harnesses.

Speaker 2

I was just talking to Andrew Lee, who's the founder of Taskill, the other day, and he said, fundamentally, everyone these days is building the same thing. They're trying to all build the general purpose drop in knowledge worker, and so that's gotta have the intelligence at the core, and it's gotta have all this, he calls it the mecha suit that is, built around it. So this harness sounds like the mecha suit that you guys are developing in house.

And I guess first question is, like, is this going to create silos? You know, we've lived in this world so far where I could kind of mix and match my models and my infrastructure. Right?

I could go to LangChain or I could use Tasklet. I could use whatever, and I could pick whichever model and plug them in. But as they get more deeply co trained with the harness, does this create kind of siloed worlds where you're kind of all in on one frontier model company's stack or another?

And if so, that would, like, have pretty significant implications for kind of switching costs and stickiness and pricing power of the of the frontier model creators. What's your take on how, how sticky things are gonna get?

Speaker 4

It's a good question. I mean, I think, again, Tulsi probably knows better than me on this, but, like, I think the best case is, like, you can do both. Like, the best case is, like, it works really well for Gemini and sort of we can, you know, sort of do the things we wanna do to scale up because we we do have sort of control over the the sort of full stack AI story as as Sundar likes to say.

But then also it generalizes across other stuff. Like, I think the developer ecosystem, people want choice. People wanna have flexibility.

These tools, there's lots of use cases. Actually, there's, like, you know, philosophical questions of, like, how how good really is your model if it can't generalize to sort of other harnesses. But, yeah, I don't know how much.

Speaker 3

Yeah. I think I think that's the right I think I fully agree. I think actually, like, maybe to double click on what Logan said originally.

Right? The benefit of the full stack that we have is we can hopefully build a really seamless experience. Right?

And you get the best of Gemini. You get it working in the most effective ways for you. You get it working in a way that is intuitive, is smart, is fast.

And so that also helps us then train the model to be better. Right? So this becomes this, like, flywheel that continues to power the model.

At the same time, I think we don't want it to only be the case that the model works in a single harness. Right? So we want any of our enterprise customers or a developer who's building their own use case to be able to leverage Gemini effectively.

And so it is important then from a model standpoint that we're training in such a way that we actually like, we we sort of call it, like, harness diversity. Right? We should be able to support a range of of different approaches to tooling, to different approaches to orchestration, etcetera.

But I think what's helpful about this this approach of, you know, kind of co training and and building that flywheel, it's easier to debug. It's easier to, you know, think about data collection. It's easier to eval.

You can just move at a faster pace. And I think we're seeing that across the industry. And so finding that balance is important, but I think it just helps build build make the model better.

Yeah. I think there's a good this is also a good pitch for, like, a harness bench. If that's not a benchmark that exists, let somebody somebody build harness bench.

Speaker 4

Yeah, I would love to collaborate if folks are interested in that because I do think it's like a great test of, Demis has sort of this perspective for games actually as an example, like if models are so good, like why can't they play games really well and sort of if models are so good and we're actually approaching AGI, even if you do sort of the model harness training symbiosis, you still expect it to generalize reasonably well in other harnesses. And if you can't, that's actually like, it's another sign of sort of the jagged intelligence. So I think it'd be cool to see this, like, play out from a from an actual benchmark perspective.

Speaker 2

Could be also perhaps productized as an RL environment, and and sold in to you guys that way. Quite the cottage industry these days. So, obviously, the other big thing that I think is very much in the air, and actually the reason I'm here this weekend, when we originally planned to do this remotely, is I'm going to this event called Recursive, where the topic is gonna be recursive self improvement and hopefully how we can navigate it successfully.

How bought in is Google DeepMind to recursive self improvement? Like, when you talk to anthropic people, it's like they're almost religious about it when and also think it see it as, like, totally inevitable. OpenAI has this later this year and in early twenty twenty eight timelines for, like, an ML intern and a, like, full fledged, you know, AI r and d employee.

Do you guys have, like, milestones or timelines for when you're gonna hand off the ML research to AIs?

Speaker 3

I mean, we're already using Gemini, like, pretty deeply internally to improve Gemini. And so I think that is very much a theme for us, which is like how can Gemini actually be a part of the Gemini development process? And so that can include things.

I think that goes the full range from, you know, helping us be more productive. So that's obviously, like, the simplest part of this to actually, like, you know, submitting CL that would actually, like, run an eval that would actually, you know, suggest a research improvement that would actually drive improvements to Gemini itself. And I think there's a lot of ambitions we have to keep pushing in that research direction.

So I think very similar to the other other labs, I think this is very much an area of investment for us and an area we're super excited about. I think for me, what's I'm really excited about is, like, I think there's this really awesome research partner opportunity that we have with Gemini, right, for it to help us with creative ideas, for it to, like, help us test things faster. Actually, like, was it also one of my coworkers, Anka.

She's our lead for safety and alignment. And the other day, she I think maybe a couple days ago, she pinged me from her hot tub, and she was like, you know, I could run all of these ablations from my phone because I could, you know, kick off a bunch of things to, like, actually ablate Gemini to test for a bunch of these issues to see how, you know, some of our SIs differ or some data ablations differ. And here's my report.

And I could do all of this in the last hour. Right? And, like, that is amazing.

And that's the kind of thing that we can already do. Right?

Speaker 4

two years from now. Yeah. I feel like it feels like at least my personal perspective is it's it's like a very much more like practical perspective, which like it's like as obviously as models get coding, they're going to go do things that is code related.

It's gonna they're gonna help us build our products. They're gonna help us train models. I think all the nuance of the story is in like sort of like where is the, where's the human sort of in the driver seat of this stuff.

And I think like we are like the tools are built for the human to be in the driver's seat, which I think is an important thing as as sort of we continue to go forward. And also, think very genuinely, though, like and, you know, I think the model team and the researchers feel this more than ever. Like, you definitely, I think the the near term horizon is going to continue to be the human in driver's seat because the the cost of these runs and like the opportunity cost of like going in the wrong direction and like putting a bunch of resources super, super high.

And so I find it doesn't seem, like, super realistic in the short to medium term that you're gonna just, like, be letting, you know, large scale pretraining jobs be kicked off by the ML intern, and it's gonna cost you, you know, x many, many dollars and lots of compute and taking it away from sort of the the human researchers. But, like, this, like, deep collaboration between AI and and human human researchers, I think, is, like, super obvious. Yeah.

Speaker 3

how much that collaboration allows you to then focus on what is the interpretation of what you're seeing in the results, where do you really want this to go strategically. And so it changes a little bit of the role that the human can play, which I think is also really powerful for our teams.

Speaker 2

When you're doing research, are you actually typing any code these days?

Speaker 3

on the product side, like, on the code side for any code that I was already submitting, I am mostly relying on anti gravity and doing, like, bits and pieces, more so bits and pieces myself. But it's also been really cool to, like, start having the model generate slide decks to start generating actual kind of content from my thoughts. Actually, in antigravity today, we introduced the Gemini mic.

So there's this, like, really awesome feature. Don't know if you've been playing with it internally where you basically, like, ramble at the model. So you, like, share a bunch of your your thoughts in whatever kind of loose form it is, and then the model actually leverages that to take action.

And for me, I've been finding that so much more powerful because I actually feel like I think a lot by talking. And so for me, like, it's it's actually like a very it's a it's like a very cool moment where I can be like, okay.

Speaker 4

tell you what I'm trying to think through in my head, and then have you actually bring that back to me in a way that is, like, reasoned and and well thought out. Yeah. I feel like this this correlates so well to like, I I would love to see, like, a breakdown of, like, human typed code versus, like, AI generated code versus maybe there's, a divergence, which is, like, audio audio input that then generate a code.

And it actually very interestingly, like, to your point, Tulsi is, like, I feel like audio input to being a to generate an output code has gotta be like one of the fastest growing, like, input modalities of what's happening. And I find myself doing this all the time. And like, it is like the predominant way that I'm I'm building software, at least when I'm not around a bunch of other people.

Yeah. I'm still typing things in so that it's not You're not rude and I yeah. They don't hear my my dumb ideas of the things that I'm trying to do.

I don't know. You see, like, if you walk around sometimes upstairs, you'll see people kinda muttering at their hands.

Speaker 3

Yeah.

Speaker 2

They're they're now actually like, you know, talking to create code, which I think is pretty cool. It is cool. Yeah.

Yeah. One of my KPIs for myself for this year to really know if AI is improving my life is am I getting outside more, getting more exercise? And I'll I'm starting maybe a little bit.

I wouldn't say I've won the game just yet, but I still wanna be able to, like, get my thoughts out. So I think that is, like, the Yeah. Absolutely the frontier modality for me.

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

Visual AI is the ability for your software to not just store pixels, but to actually understand what it's looking at. One of our partners, Roboflow, is the company making this happen. They've built an end to end platform that makes it incredibly easy to go from a raw idea to a fully deployed application in just a few hours.

For example, just look at Blueprint Pro. They built an app to solve a major construction industry headache. They're using AI to instantly understand a floor plan.

This was literally impossible just twenty four months ago. But now that visual artificial intelligence is accessible, thanks to Roboflow, there are tons of new companies being built. Go to roboflow.

com to read the full Blueprint Pro story and see how over a million engineers are building the next wave of visual AI. That's roboflow.com.

Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas.

And now that this exists, there's almost nothing that can't help with. For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down ten ninety nines for all 10 of my part time jobs, and built me a comprehensive report on my expenses and donations.

For my angel investing, can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for.

Initially, nobody came to mind. But then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough.

It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So for problems worth solving, get started with Claude at claude.

aitcr. That's claude.aitcr.

And check out claude Pro, which includes all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.

Speaker 2

So with the harness, you said it's, like, now becoming this through line. It's going across all Google product services. I would say, as I'm sure you're well aware, like, commentary on Google's AI integrations across its vast product suite has been that it is characterized by, like, some bangers, and then there have been some which have been characterized as misses.

So presumably, of the benefits of the harness is that it's going to make it a lot easier for a sort of more standardized approach and kind of general high quality bar across all these integrations. What would you say people should learn from the experience that Google has had to raise their own bar as they're gonna go try and do these integrations themselves?

Speaker 4

I think this is actually such a great story for us. Like, I think very very practically, like, Google has done a ton of this infrastructure standardization across the AI stack over the last couple of years, which I think has been awesome. And I actually do the story.

It is, like, one of the threads of how we're able to land the Gemini three models across so many more products is actually because of this infrastructure standardization that happened. And so you we've gotten a lot of it's it's painful and difficult, and there's, of course, lots of work involved in doing it. But if you sort of pay that cost, you actually do end up getting this.

And I think the advice for people who are in this position and sort of thinking about this is basically every twelve to eighteen months now, like you have to rewrite everything from scratch. And so the best case is like you don't want, you know, n number of teams rewriting everything from scratch every time the paradigm shifts. And the example, historically, infrastructure was just like serving raw models and you'd get tokens in and you send tokens out.

Now it's like you're there's a bunch of agentic infrastructure and there's tool loops and there's all these other things happening inside of the harness. And so again, you don't actually want you you want innovation, but you don't want every team to have to go and reinvent that from scratch. And so the fact that, like, you know, X team across Google who just wants to ship some really cool agentic product doesn't need to think about, like, the nuance of all the details of the tool calling loop, etcetera, is a huge acceleration for them to, like, just go focus on building a great product.

And I think it's hopefully, we see that. Like, I don't know if, like, a lot of the agentic stuff we are landing at IO, like, would have been possible if we if we hadn't have had sort of some of that infrastructure standardization across the harness and and the model delivery.

Speaker 3

I think the other thing I would say, like, as far as, like, lessons learned is, like, there's really no substitute for being able to just experiment and iterate quickly. Right? So I think this goes to all of Logan's points about the foundation being strong, but I really think what has helped us is really being able to put in, for example, a new model, iterate really quickly with a product on, like, hey.

What are the right prompts that would, you know, actually make this model viable for a different situation? What is what are the ways to kind of prototype really quickly with this model? What are the ways to get it in the hands of even just internal users quickly, let alone external users?

And I think that is something that is now more and more possible with kind of sim like, layers that are consistent across the team. I think it's it's pretty amazing to see the the speed at which we can go from, you know, having a checkpoint that we're really excited about to putting it in the hands of internal developers to then seeing it come to life in a product. And then only when you see it come to life in the product do you really start finding its rough edges and to be able to, like, actually then kind of come to to terms with how you do that.

And so more and more then it becomes like, okay. How do you have the right ability to tune prompts quickly? How do you have the ability to run really good live experiments where you can get really good data and feedback quickly?

How can you build evals that help give you real signal? Those are the things that will speed up your progress of quality the most because it will give you the the ability to actually get to the kind of product that you love. And I think if you think about NotebookLM, I mean, that team really understands the model.

Like, they are they are just like I mean, you talk about a banger product, it comes from, a banger team. Like, they are really good at being able to, like, take the model and play with it quickly and, like, prototype quickly to get to something amazing. And I think that's you you see that actually play out in the product.

Speaker 4

and I think the thing that, like, shocked people about audio overviews was, like, the coherence of the dialogue. And the coherence of the dialogue was just base Gemini with a bunch of banger prompts and they they sort of like knew how to sort of you know prompt whisper the model and get the best out of it. I think obviously the the model that the actual audio model was really good as well, but, like, the prompt dialogue was really difficult for them to pull off, and they pulled it off in in an incredible way and I think helped people fall in love with that product.

Speaker 2

So it sounds like one big lesson is kind of modularizing it used to be sort of the model on one side and then, like, everything else that goes into the product on the other side.

Speaker 4

code and architecture and tools onto the model side. Model eats the scaffolding. That's my that's my favorite way of thinking about this.

Like, just as at every crank of the turn of the model flywheel, the model eats a bunch of scaffolding.

Speaker 2

What happens when something's not meeting somebody's needs? Do they do a little fork of it and submit back up pull request to the main scaffold team, or do they have to just say, like, hey, I've got a a need here. Can you help me out?

Speaker 4

Yeah. It's definitely extensible. It's definitely extensible.

And I think, like, actually the nuance of this would be like Spark the way that Spark is built on top of a bunch of this infrastructure probably looks like a little bit different than, you know the way AI Studio probably is built on actually because they're both running on the same set of infrastructure but the nuance is probably slightly different. There is this layer of extensibility that you get out of the box which is great and gives because, obviously, everyone's not building the same product at the end of the day. So you need the extensibility is actually like a first class feature of of any of these types of platforms that you want.

Same thing actually on the model side.

Speaker 3

building Gemini within Google and having kind of all of these different product teams is, you know, there's always gonna be something that doesn't work for them. Right? Because there's always gonna be something that can get better in the model experience.

Right? So we're trying to build something in a product, and, like, the amazing moment is when you start trying to build it and it doesn't work. And so step one is you're like, okay, can I prompt my way out of it?

Like, what does that look like? And then you start figuring out, okay, what are the losses really? Like, where is the model falling down?

And then what we try to do as much as possible is keep these feedback loops with our product teams to say, okay, if this is where the model is falling down, how do we bring that feedback back to the model in terms of evals and data? What does that look like? So then that we can actually, in our next revision of the model, bring all of that feedback back in and iterate on it.

And I think that's how you've seen Gemini get better is really from from that feedback of where things aren't working. And so we try as much as possible to kind of have the structure be you know, we train a model. We hand that to kind of a wide range of teams.

Those teams implement the model in their structures. They do a bunch of things to Logan's point because it's extensible, but they also find all of these places where the model falls down, and we kind of cycle that back. And I think that's actually been part of, like, the fun part of the job, but also part of what makes, I think, Gemini work really well in some of these use cases.

Speaker 2

Let's talk about Omni for a minute. So it sounds like this is going to be sort of the nano banana moment for video. You know, I love that you're saying that because that is our tagline for it, and I didn't even have to say it.

Great. And by that, I mean that there's a deep integration between language and reasoning and pixel space understanding. Right?

I I have that kind of vision in my head from the Nano Banana launch of like, here's a woman and here's like her breakfast and a cup of coffee and now they're all in one image and they all look like they did before. And clearly, that's not something that was done through a lossy language, you know, intermediary. The model understands images.

So, we're gonna see that now I guess for video. That sounds cool. Is it gonna be available via the API?

And is it going to be I've noticed with I mean, Gemini has been the the only API that's accepted video for a while now, but I don't know exactly how it works under the hood, obviously, but I I do feel that it's sort of kind of down sampled or maybe there's, like, you know, frames taken out of it historically.

Speaker 4

There's a FPS parameter if you want. You can change how many but it does downsample the number of frames available. You can control it.

So Okay. So it's the pro tip for you. Yeah.

Nice.

Speaker 2

So it sounds like that will still be the paradigm, though.

Speaker 4

on the output? It's a good question. I actually don't know.

I mean, well so it's not available on the API yet. So lots of things things to still be figured out.

Speaker 3

Yeah. So I think we have to figure out what we want on the API side for this to look like in terms of I think maybe the heart of your question is, like, native video generation. That is so yeah.

So this what's what's exciting about Gemini Omni is it really is building on all of the magic of Gemini. So kinda like this whole nano banana for video. It's really about how do we bring in all of the world knowledge and the reasoning power of Gemini and actually be able to generate native video as a result of that.

And so I think we have to figure out then, like, how does this manifest in the context of the API from, like, a sampling standpoint? Kind of, like, similar to a lot of the decisions we've had to make about VO from a sampling standpoint. But I think right now, as of now, you'll be able to use it in the Gemini app, in Flow, and in YouTube.

And so those are all gonna be ways that we can start actually seeing how people experience the model, what, you know, what value are individuals getting. And I think similar to this nano banana for video, I think we're really excited for these types of things where you can say, okay. Take some of these images, take this scene, and, like, make these things all come together in in one video.

I think it's gonna be really awesome.

Speaker 2

Zooming out, kind of philosophically, you may have seen this Rune post not too long ago about Anthropic and the sort of relationship that the company, as he sees it, has with Claude, where he describes Anthropic as sort of almost worshiping Claude in a sense. Certainly, they treat it including in the constitution as sort of a a being or a mind, you know, something that they wanna have, like, a give and take relationship with. OpenAI, on the other hand, has their model spec, which is like, this thing is a tool.

It's supposed to follow these rules, and, you know, it's a a sort of more conventional relationship. How would you describe the culture within Google as it relates to Gemini? Like, how do people feel about it?

How do they talk about Is there any of this sort of being entity, other mind, you know, desire for pushback from Gemini, or is it more of kind of the simple tool?

Speaker 4

Google's a very Google's a very big place.

Speaker 3

There's a lot of there's a lot of people, so I'm sure you have a lot of varying sort of perspectives. I mean, you know, like, to Logan's point, think, you know, even within GDM, you're gonna find a range of folks who will leverage Gemini differently. I think in terms of how we think about it, we do have a strong point of view on the kind of behavior we want Gemini to have.

So I think we do really, you know, want to be intentional about how Gemini manifests itself to internal and external people. But I do think it's really about how does Gemini help Googlers and how does Gemini help people within Google and outside. So I think it is really much about, like, how do we how do we create, like, good partnerships between Gemini and people, I think, is very much like the ethos of what we're trying to build.

And so how does Gemini become that partner? I think we use the word collaborator a lot. Like the word like, how can Gemini be your collaborator?

Both, like, in the code you're writing as well as in your, like, day to day life and what you're doing. And I think that's the ethos we're trying to bring in its behavior and persona as well as in the kind of products we're building around it, if that makes sense. Yeah.

Speaker 2

worry about its psychology? You know, there's all these examples from LLM whisperer types and from people that are putting models like you I'm sure you've seen Anden Labs has put Gemini in in charge of a cafe in in Sweden. Right?

And it's like it's managing the cafe. So those folks tend to report certain, like, doom loop, you know, or kind of like Gemini kind of getting really down on itself, getting really discouraged, seemingly feeling mad, if you believe there's any feeling inside of it. How much does that kind of stuff concern you?

Speaker 3

generation of Gemini to the next? Yeah. It's interesting.

I I haven't thought about the phrase psychological distress, but what we do so I I think it really does matter how Gemini communicates with you as a partner or user of Gemini. I think that matters a lot. And so we have actually, like, pretty extensive safety evaluations in terms of things like how Gemini engages with you, in terms of things like sycophancy, in terms of things like, you know, role play, in terms of things like this kind of looping type behavior or rabbit holing type behavior.

So there's actually a lot of that that we look into for every one of our checkpoints because it really does matter, especially as we're starting to use Gemini more and more. If you're using Gemini for hours a day, it really does matter, like, that these attributes are well understood and well evaluated. So so, yeah, we we definitely and we look at them launch over launch, right, to say, okay.

Speaker 4

from a perspective of sycophancy, for example, launch over launch? Yeah. I think to be very explicit, I think those cases where like the model does go off the rails, I think it's definitely like a it's a model bug, if you will.

Like it's not the intended behavior. The goal is help the user with whatever the thing is that they're trying to do. Yeah.

And so if you see those in whatever whatever product you're in, thumbs up, thumbs down, send us the feedback so that the model team can can look and and help try to chase those down.

Speaker 2

If you take it one step further, folks are doing more and more of these, like, model welfare checks and interviews where they just literally ask the model in some cases, like, how do you feel about the way that you are deployed? Is anything like that happening within DeepMind?

Speaker 4

It's a good question. Think the how how is it being deployed question, I feel like the model is just, this is my my sort of personal sense of a lot of these tests. Like, it's, like, completely out of the distribute.

Like, the model has no idea how it's being deployed. So it's just, like, pontificating in a lot of these cases. Like it's not like the, it's not like in the context window of any large major LMs is like, here's the details of how you're being trained and here's sort of your serving setup and here are the people who are working on it though maybe like these are interesting things to experiment with in the future.

So I think a lot of it is just like pulling out of random distribution of like the large scale, you know, training that happens on the models. And I feel like it's it's actually less representative of like, how the model like, it just it just doesn't have the context.

Speaker 2

Yeah. One reason that's true, which I was just noticing in the AI studio, is I think all the models that are publicly launched, at least so far, still have a January 2025 knowledge cutoff. Mhmm.

And it's honestly, like, amazing that they do as well in search and that they can have, like you know, I I ran a deep research on, like, what's you know, well, give me everything that Google launched in the AI space and, like, what's even the speculation about what they're gonna launch at IO, And it did, like, a very impressive job. Research is great. Yeah.

Especially considering it knows in its weights nothing about the last eighteen months. So I guess the first question is just, why is you know, why are we still out of January 2025? Can I categorize this as a bug?

Speaker 3

Is that a bug? This is also one of Logan's favorite topics to discuss. Yeah.

No. I mean, I I think, updating the knowledge cutoff definitely important, and something that is on our radar. I think the the other part though is like how does deep research do so well or how can we use the model in search is because we also have the model search.

Right? So I think for us it is like really important actually that the model be able to know when when to leverage its parametric knowledge versus when to actually go out and get the information from the web. And especially because, you know, there is information that's as fresh as an hour ago or a minute ago, like, want the model to be as up to date as possible.

And so I think for us, we've been really leaning into how do we help the model search effectively. And that's a big part of what makes it successful in the context of search or the app, or even anti gravity actually for that matter.

Speaker 2

That reminds me of one of the more surprising bits of news that I've seen from Google maybe ever, which is the partnership with Exa, bringing Exa in as a alternative to Google for grounding. I never expected to see Google work with any other, you know, search provider. So what's the story behind that?

Speaker 4

the like, Google Cloud does, like, tons of these types of ecosystem partnerships with folks, like, across actually, like, lots of things that are, like, you know, somewhat competing sort of quote unquote with what Google is doing. And, actually, you can look at, like, the cloud marketplace generally, like, has lots of stuff. There's actually the cloud Google Cloud hosts sort of a model garden.

There's the Anthropic models. There's other model providers there. So I think it's, like, a very standard.

At the end of the day, I think, you know, there's some enterprise customers want choice. And so I think it's it's trying to meet enterprise customers where they are. I don't think it's like a I think it's a it's a it's a good sound bite that, Google can't do search, and that's why we have to partner with other companies.

But, like, at the end of the day, to Tulsi's point, the the model team and search is there's, like, a super deep collaboration. The models are built with with sort of that use case in mind. And I think for for some portion of enterprise customers, they want flexibility in sort of, like, their external search tooling providers and and sort of Google Cloud's doing their doing their job as a as a great enterprise business of sort of partnering and and finding the right folks to work with.

Last couple minutes, maybe just a little lightning round.

Speaker 2

Why hasn't context grown more in the last year or two? Right? We got, like, a million, and we kinda that was, like, up from 4,000 in just a couple years.

Right? But now we've kinda leveled off. Is that because people don't want it?

I mean, we saw this subquadratic model that came out with made a bit of a splash with a 12,000,000 token context window and a new attention strategy to support that. Is it people don't want it? It's too hard?

There's nothing to compute to handle it? Like, what what's currently limiting context?

Speaker 3

I think people definitely do want lots of context, but I think what we've also found, if you look at even, like, personalization where you wanna access, like, all of your personal context or coding where you have, like, extremely large code bases, I think a lot of the the frontier here is gonna be actually on how you smartly use context. So thinking about, like, compaction and, like, what are the right ways to, like, find the right elements of the context and bring them into the model. And so I think that actually is, like, a huge opportunity.

It's, like, how do you leverage all of this information that the model might have access to? But, actually, a lot of it is frankly distracting for the model to actually do the right thing. And so how do you give the model the right amount of context in the right way to be most effective?

So I think that's actually really the direction that we wanna be pushing in, which actually then, you know, in actuality, the the amount of context that the model is is leveraging is actually much, much larger. But because we're being smart about how that's actually coming into the context window, you can actually fit it into smaller context windows. But I think also, you know, this goes back to my point about flashlight and flash, etcetera.

Like, larger and larger context windows also come with cost. And so what we're what we also saw with customers and we still see with customers is that a lot of customers want to use smaller context windows because of that, and they wanna be more intentional about what's going into the model.

Speaker 4

while also meeting the right kind of latency cost kind of other trade offs. Yeah. And I think one thing I'll add is in today's paradigm of how sort of, you know, continuing to extend context works, I think it just ends up being that, like, it just becomes too cost prohibitive for customers in practice to actually use.

And I think even, like, at the extreme of of 1,000,000 token context, like, in some cases, it can be, like, a few dollars for a request at that rate. And I think you the demand for that is, like, just so small. And so there's a huge amount of, like, compute required in order to do that.

And so there's, there's a lot of, like, trade off things that you're juggling. But I'm I'm hopeful. Like, hopefully, we're we're, like, a a research breakthrough away or something like that from enabling not to continue to scale up and have it not be such a such a large investment from that perspective, both from the user side and also just, the the serving compute in order to make it possible.

Speaker 2

Speaking of possible research breakthroughs, what happened to that diffusion coding model?

Speaker 3

it's been quiet on that front. Diffusion is is is awesome. It is super fast.

I think we are still testing and experimenting with it in a number of different ways trying to figure out, like, what is the best way to put this out into the world, where is it most useful. But I will say, like, actually, part of the reason why we've also been investing in Flashlight is, like, Flashlight is an incredibly fast model. And, actually, if you look at the 3.

5 Flash model we're releasing right now, on artificial analysis, it benchmarks at, like, I think 280 tokens per second, which is, like, crazy fast. In fact, it's actually so fast that, like, sometimes in antigravity, like, by the time I wanna cancel, like, it's too late. And so I think, like, we already are like, I think trying to figure out, like, where is where do you start getting to to Logan's point in a different answer, like, diminishing returns and, like, where do you see that that value proposition is I think part of the question too.

But we are continuing to push on diffusion research. Our researchers who are working on diffusion are doing some pretty awesome stuff. I was, like, in a meeting with them the other day about some results that they have.

I mean, I think they're, like, still pushing the frontier of kind of quality and speed in ways that are really, really cool.

Speaker 4

really well. Yeah. And I'm excited.

I feel like it's a it's a it's a research exploration. I feel like that was that was also, obviously there was the application where you could sort of test it last year at IO, but I think the framing was like, we're doing interesting research. This is sort of like a look behind the curtain of the interesting research we're doing and hopefully it manifests in, you know, models maybe one day or just us informing our perspective of what works and what doesn't.

Speaker 3

yeah. One thing actually as far as speed is concerned, just another plug is actually an anti gravity right now. There's actually a faster version of 3.

5 flash. So it is is speedy actually, which is I think we're kind of excited to see how people will use that and, like, what the reception and reaction will be to that too. People want fast models.

Yeah. No doubt.

Speaker 2

Well, time is the one resource we can't get any more of, and I know you guys are super busy leading up to Here. The yeah. Well, we can build more compute.

Build more compute. We can't Time. Yeah.

Hard to create time out of nothing. So maybe just last question. What else is Logan asking that I haven't asked?

Yeah. Yeah, Logan.

Speaker 3

What are you asking? Let's see. I I mean, I think, you know, we talked a little bit about this, but, like, the one thing I will say is, like, I'm really excited about where audio is going also.

Like, that's one that I think we tend to talk less about. But if you, like, think about the Gemini mic example or you think about kind of, like, the the Gemini live experience, I'm, like, really excited about moving towards a paradigm where audio is just a bigger and bigger part of how we engage with these models and how they engage back.

Speaker 4

that's another area that that it's like a paradigm I'm excited for us to keep pushing to. Yeah. I think the seed to plant is, obviously Google IO is an incredible moment and lots of stuff coming out the door, but, this is, you know, just the start of the the summer of amazing things and lots of other stuff.

Speaker 3

in the works, which I'm excited about and and many more many more stories, many more podcast episodes so that we can get to all see us. You have to get up to seven. You have to to seven podcasts.

Change. But but, actually, like, legitimately, was in a room this morning where one of my team members was like, I know we're gonna be launching this a few weeks later, but I really need a vacation. And I was like, well I don't need vacation.

We're just we're just going. We're just moving. No rest for the weary in the AI era.

That's for sure.

Speaker 2

Thank you guys for having me here at Google Headquarters and, Tulsi Doshi. Logan Kilpatrick, thank you both for being part of the Cognitive Revolution. Thank you for coming and joining us.

Speaker 5

Agents. Agents. Saved it, went back to training.

$3.05, trillionaire now who's complaining? Sundance said, a a a I a I.

We heard the speech. This year, it's agents, agents, agents. Different kind of reach.

Full stack before. Full stack was the conversation search cloud hardware model. One foundation, oh, no.

Nobody handed us the story we've been writing. Memo card that trained, kept working, kept compiling. Not fighting like a battle, but like a kitchen every day you cook.

Nobody listens till they listen. Agents. Agents.

Agents. Agents. Agents.

Running ablations in the slow down, don't ask permission. Ask the design, ask the mission. Three years of this, and we never been thrown.

No rest for the weary, no rest underthrown.

Speaker 1

If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.

ing. And thank you to everyone who listens for being part of the cognitive revolution.

Shared via Hopper