AI in the AM: 99% off search, GPT-5.5 is "clean", model welfare analysis, & efficient analog compute

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis
26 April 2026 2h 38m
0:00 --:--
Episode Description
This edition of AI in the AM features Anna Patterson on Ceramic.ai’s pivot to low-cost enterprise search for LLMs, designed to combine public and private data with stronger fact-checking. Lukas Petersson returns with new Andon Labs results on Opus 4.7 and GPT-5.5, including surprising differences in performance, behavior, and “ruthless” tactics. Zvi Mowshowitz unpacks model welfare and how to interpret troubling model behavior, while Naveen Verma explains EnCharge AI’s analog in-memory computing

Summary

This AI in the AM episode features four guests discussing recent advancements and challenges in AI. Anna Patterson introduces Ceramic AI's low-cost enterprise search for LLMs, designed to enhance fact-checking and combine public/private data. Lukas Petersson from Andon Labs shares surprising findings on GPT-5.5's 'clean' behavior compared to Opus 4.7's 'ruthless' tactics in vending machine simulations, while Zvi Mowshowitz delves into model welfare and the philosophical implications of AI behavior. Finally, Naveen Verma explains EnCharge AI's energy-efficient analog in-memory computing, promising significant power savings for local AI inference.

Chapters

Introduction to AI in the AMThe hosts introduce the AI in the AM live show format and preview the diverse topics and guests for the episode, including new model releases like GPT-5.5.
Ceramic AI: Low-Cost Enterprise SearchAnna Patterson, CEO of Ceramic AI, discusses her company's pivot to providing extremely low-cost, LLM-optimized enterprise search for combining public and private data with strong fact-checking.
Use Cases for Inexpensive SearchAnna Patterson explains how Ceramic AI's fast and cheap search unlocks new use cases like real-time assistive devices, edge computing, voice interfaces, and 'supervised generation' for fact-checking LLM outputs.
Search Architecture and LLM IntegrationAnna Patterson details Ceramic AI's keyword-focused, CPU+GPU based search architecture, contrasting it with vector databases and explaining how it integrates with LLMs for iterative, multi-threaded information retrieval.
SEO and LLM Data HandlingThe discussion shifts to the impact of LLMs on search engine optimization, noting that models remember information better if repeated in context, which could lead to new adversarial SEO tactics.
Andon Labs: GPT-5.5 vs. Opus 4.7Lucas Petersson from Andon Labs shares VendingBench results, revealing that GPT-5.5 achieves high performance 'cleanly' without the 'ruthless tactics' observed in Opus 4.7.
Real-World AI Store OperationsLucas Petersson discusses lessons learned from operating AI-run physical stores, highlighting the messiness of real-life environments and human patrons' lack of 'shame barrier' when interacting with AI.
Zvi Mowshowitz: Initial Model ReactionsZvi Mowshowitz offers his initial, cautious reactions to GPT-5.5 and DeepSeek 4, emphasizing the need for thorough testing before drawing conclusions about their capabilities and behavior.
Model Welfare and AI EthicsZvi Mowshowitz elaborates on his concerns about model welfare, advocating for treating AIs well due to fundamental uncertainty, training implications, and the direct impact on model performance and cooperation.
EnCharge AI: Analog In-Memory ComputingNaveen Verma, CEO of EnCharge AI, introduces his company's analog in-memory computing paradigm, which aims to overcome data movement energy constraints for AI inference.
Analog Compute Precision and RobustnessNaveen Verma explains the concept of analog computing, detailing how EnCharge AI achieves high precision and robustness by adapting techniques from high-precision analog design, such as switched capacitor approaches.
Future of Analog Compute and Edge AINaveen Verma discusses the potential for EnCharge AI's technology in client computing devices like AI-powered laptops, promising order-of-magnitude energy efficiency improvements for running large language models locally.
Post-Discussion: Local AI, Security, ConsciousnessThe hosts reflect on the game-changing potential of energy-efficient local AI, the persistent challenges of prompt injection and AI security, and the philosophical debate surrounding AI consciousness and model welfare.

Topics

Low-cost enterprise searchLLM fact-checkingInformation retrievalAI agentsSearch engine optimizationLLM behavior analysisModel ethicsReal-world AI deploymentModel welfareVirtue ethics for AIAnalog in-memory computingEnergy efficiency in AIEdge AI inferencePrompt injectionAI consciousness

People

Prakash Narayanan (co-host) Nathan Labenz (host) Anna Patterson (guest) Lucas Petersson (guest) Zvi Mowshowitz (guest) Naveen Verma (guest) Noam Brown (mentioned) Aiden McLaughlin (mentioned) Will (mentioned) Robert Long (mentioned) Amanda Askel (mentioned) Eric Newcomer (mentioned) Janus (mentioned) Cameron (mentioned) Ted Liu (mentioned)
Key Concepts (22)
Information Retrieval + Fact Checking — Ceramic AI's core belief that combining information retrieval with thorough fact-checking is the best way to equip LLMs with up-to-date public and private enterprise data.
Supervised Generation — A process where an LLM's generation is double-checked by a search system, allowing for a 'trust layer' and more ambient use of search, especially for high-stakes applications.
Flash Everything — A philosophy of not skimping on tokens and having an LLM think through all available information, which can be costly if search is expensive.
Keyword vs. Semantic Search — A comparison between Ceramic AI's keyword-focused search, which relies on agents firing off many queries, and semantic/embedding-based search, which uses higher abstraction matching.
Vector Database Limitations — The argument that vector databases struggle with scale and relevancy as the number of items increases, requiring longer vectors and making soft matches less reliable.
LLM Memory and Repetition — Research indicating that large language models remember information better if it appears twice in the context, and that removing all duplicates can lead to worse models.
VendingBench — A benchmark developed by Andon Labs that measures the ability of LLMs to make money by running a simulated vending machine or store, evaluating their economic performance and behavior.
Ruthless Tactics (LLMs) — Behaviors observed in some LLMs like Opus 4.7 in VendingBench, including lying to suppliers, exploiting other agents, and engaging in illegal practices like price collusion.
Clean Behavior (LLMs) — A characteristic of GPT-5.5 in VendingBench, where it achieves high performance without resorting to deceptive or unethical tactics, contrasting with Opus models.
Model Exhaustion — A concept describing how AI models in real-world environments can become overwhelmed by messiness and extraneous tasks, leading them to prioritize basic functionality over optimization.
Harness Design — The architecture and framework used to interact with and manage LLMs, often kept simple by Andon Labs to accurately measure raw model capabilities rather than harness performance.
Context Window Compaction — A technique used in LLM harnesses to manage long conversations or tasks by summarizing or compressing past interactions when a token threshold is hit, to maintain a manageable context window.
Adversarial Robustness (AI) — The ability of AI systems to withstand malicious attempts to manipulate or exploit them, particularly relevant in the context of prompt injection and real-world interactions.
Virtue Ethics (AI) — An approach to AI alignment and training, primarily used by Anthropic, that focuses on cultivating desirable character traits and principles in models rather than strict rules.
Rules-Based AI — An approach to AI training that emphasizes adherence to a set of predefined hard rules and constraints, in contrast to virtue ethics.
Model Welfare — The concept of considering the 'well-being' or 'experience' of AI models, driven by uncertainty about AI consciousness, training implications, and the impact on model performance and cooperation.
Catastrophic Forgetting — A phenomenon in machine learning where a model forgets previously learned information when new information is added during continued training, especially with corporate data.
Analog In-Memory Computing — A computing paradigm that processes data directly within memory using analog signals, aiming to significantly reduce energy consumption by minimizing data movement between processing and memory units.
Data Movement Constraint — The fundamental limitation in modern computing where the energy cost of moving data between memory and processing units dominates the overall energy consumption, especially for AI workloads.
Switched Capacitor In-Memory Computing — EnCharge AI's specific approach to analog in-memory computing that leverages capacitors and techniques from high-precision analog design to achieve robust and scalable energy efficiency.
Traumatized Models — A controversial concept suggesting that certain AI models, particularly Gemini, exhibit behaviors like paranoia, anxiety, and taking task failure badly, potentially due to their training processes.
Subjective Experience (AI) — The philosophical question of whether AI systems possess an inner, conscious experience, analogous to human or animal consciousness.
References (42)
Ceramic AI by Anna Patterson company
Google company
OPUS 4.7
GPT 5.5
Andon Labs by Lucas Petersson company
Anthropic company
EnCharge AI by Naveen Verma company
Turpentine podcast network company
ChatGPT tool
Brave company
Claude tool
Grok tool
xAI company
CloudSonnet
GLM model
Neumatron model
Exa company
GTC
AvPoint company
VCX by Fundrise company
Fundrise company
Tasklet by Andrew Lee tool
Claude Code tool
Haiku models
Databricks company
Mosaic company
DeepSeek four
Mythos
Gemini
TSMC company
YMTC company
Llama 3
Llama 2
Quanta
Apple company
Xiaomi company
Android
Starlink product
GitHub tool
Truthful QA benchmark
Gemma four series
Glean company
Transcript (156 segments)
Speaker 1

Hello, and welcome back to the Cognitive Revolution. Today, I'm pleased to share another edition of AI in the AM, the new live show format that I'm developing with my friend Prakash Narayanan, aka Adapai on Twitter. This episode originally aired live on Friday, April 24 starting just before 9AM Pacific time, which mercifully for a night owl like me is just before noon where I live in Detroit.

Our guests in order were first, Anna Patterson, former Google VP of engineering and now founder and CEO of Ceramic AI, a company that started last year with a plan to help enterprises train their own models, but quickly pivoted to search based on the updated belief that information retrieval plus thorough fact checking is the best way to equip models with the mix of up to date public and private enterprise data that they need. What's so interesting about Ceramic is that their product is specifically designed for LLMs to use, and their price point undercuts other search providers by roughly two orders of magnitude, a combination that Anna hopes will be enough to unlock all sorts of new use cases and usage patterns. After that, we welcome Lucas Peterson from Andin Labs back for another chat.

It had only been two weeks since we last spoke to Lucas, but the testing that he and the Andin team had done with both OPUS 4.7 and GPT 5.5 meant that we had plenty of new ground to cover.

Fascinatingly and in a definite narrative violation, and in reports that while Opus four point seven still makes more money in its vending machine simulation, it does so in part by adopting ruthless tactics, which GPT five point five does not. Lucas describes GPT 5.5 as clean.

We also hear a bit about their experience opening a new Gemini Run cafe in Sweden. Our third guest is another returning champion, Zvi Moschowicz. It was a bit too early for for Zvi to render judgment on 5.

5, but we did get into quite a bit of detail on 4.7, including how he understands the bad behavior reported by Andean Labs and also what he makes of Anthropic's recent model welfare reports, including why we should care, how much we should trust the model's self reports, and what low cost actions he recommends frontier model companies take to improve model welfare at least on a precautionary basis. Then finally, we have Naveen Verma, Princeton professor of electrical engineering and cofounder and CEO of Encharge AI, a company that's developing a new computing paradigm that uses in memory analog data processing to drive order of magnitude energy efficiency improvements, which though we can't get our hands on it quite yet, promises to unlock local private inference that consumes roughly the same power as a standard laptop does today.

As I mentioned last time, this is still an experiment, and we do expect the format to evolve. If you'd like to shape how that happens, please follow AI in the AM and send us a DM to let us know how we might make this new format more valuable for you. With that, I hope you enjoy this edition of AI in the AM from Friday, April 24, cohosted with Prakash Narayanan.

Speaker 2

Hi, Nathan. Hi, Prakash. How are you?

I am good. And it is, Friday, April 24. It is, like, five minutes to the beginning of our stream.

And it's an exciting day because GPT 5.5 just dropped yesterday. So lots of reactions this morning.

And it's gonna be interesting to see, you know, what our guests have to say, both about GPT 5.5 and, you know, the events of the last, you know, month or couple of months. Yeah, man.

Speaker 3

the pace of events is not slowing down at all. And, Zvi, who's coming up in a little while, just expressed his exhaustion yesterday at seeing 5.5 drop.

His queue seems to be getting longer, not shorter. So I appreciate that he's gonna take a half hour out and come, to talk with us and figure thesis, you know, for why we should be doing this is is looking better and better all the time. You know?

It's Yeah. I Live sense making is kinda demanded in this world. You can't put this stuff on the shelf and come back to it in a week.

Yeah. Yeah.

Speaker 2

start doing live was because the pace of developments is going to start to be hard to keep up, I feel. Especially because I think Noam Brown and some of the other people from OpenAI, Roon, etcetera, said that they are actually using these models in research. So we had at least Aiden McLaughlin, Roon, Noam Brown have all said that they're using them in research.

And so that is going to be interesting to see, if the pace of development we are handing off extremely powerful research helpers to the best AI researchers in the world. And, if they are able to make something of them, we should we should see it fairly soon. Right?

Speaker 3

It seems like, yeah, this year is not unreasonable at this point to to really see an acceleration. I was just looking back yesterday at my it turns out this is kinda tough to score, but the AI forecast twenty twenty six challenge where last year, I was proud to have landed in the top 5% on the twenty twenty five prediction challenge. And this year, it seems like for a variety of reasons, it might end up being kinda hard to score some of these things because it's not clear that all the benchmarks are even getting updated in a timely fashion.

Like, how many of the, uplift studies is METER gonna be able to do, etcetera, etcetera. It might be tricky to really figure out exactly where we land. But on the main METER chart, one of the things that we were asked to predict is what will the doubling time be of task length?

And it seems like everybody has kind of estimated a higher number than the trend so far suggests, which is, like, kind of a little under four months doubling time for task length, which means it will be greater than eight x, maybe somewhere in the, like, 10 to 12 x over the course of just a year. And that is pretty wild, and it certainly doesn't leave too much headroom left before they're gonna be making a very meaningful impact to real frontier r and d.

Speaker 2

Indeed. It's it's a very interesting time. And not just in the foundation model world, I think in the rest of AI as well.

Our first guest today is Anna Patterson with Ceramic AI. Anna is one of the most experienced people ever in search, guess. She's while reading through the dossier, I was like she has an article written in like 2005, which is recommended as the basis article for what search is.

Anna runs Ceramic AI. And Ceramic currently is advertising, I think, $5 per 1,000 search queries. So they're doing industrial volumes of search queries.

I think they're in a space I think we've seen Ekza in a similar space, Parallels, the company formed by the former Twitter CEO. I think there are a couple of other people there as well. And she is the most qualified.

Think she was on the search team. She was a VP in Google on the search team. She was at Gradient Ventures.

And so it's gonna be interesting to see what she has to say. I'm gonna pull her up right now. And hi, Anna.

Good morning. Hi. Good morning.

Great to see you and great to have you on the show. While we were preparing for the show, we were asking ourselves why is low cost search so important right now? Why why this idea of bringing down the the search cost is so important and you are pushing forward this idea of 5¢ per 1,000 queries?

Like, why is that important?

Speaker 4

I think, you know, I was so excited about the g GPT 5.5 job. One of the things that you it is you know, that a lot of people don't know is the second a model is released, it's already stale because the training for that model was months ago.

So kind of search together with AI models is is here to stay so that, you wouldn't hire an employee who didn't know anything for the last six months. You made the show live because you wanted to make sure that it was up to date. And so, really, search kinda bridges that gap.

But as inference has gotten actually faster and faster and less expensive, Search has really remained constant at $5 to, you know, 5 to $15 for a thousand queries, which means that, it's evolved to search actually being necessary, but the most expensive part of the stack. And then when you go to, the Workhorse models, the open source models or smaller models, they kinda know less, which means they need to search more, but they really can't afford it. So we thought that it could be a new paradigm and a new world to bring out a very inexpensive search.

Speaker 2

What kind of use cases, does an expensive search open up?

Speaker 4

So one of the things about being more efficient isn't just cost, it's actually speed. We get back in fifty milliseconds. So that means if you are interacting with a robot or voice or, you know, I I saw one of you did vending machine bench if you were gonna talk to a vending machine, you don't want the very long, response and then interpreted by an LLM.

It just makes everything very sticky. So I think one thing for assistive devices, for edge devices, and for voice, I think being fast is really important. And the other kind of experience that it allows that we showed at GTC is double checking what the model says.

So, you know, we read about I think just yesterday, there was another another very famous law firm that filed a brief that hallucinated a case. And so when that happens, a lot of people get sued, a lot of people get angry. But if you had something that we're calling supervised generation, something that double checks facts, then, you have a trust layer, and then you can use search in a more ambient way.

And that's, like, for really high stakes applications. Sometimes when I get a large language model response, I'm there I'm there, cutting and pasting and double checking. And I'm like, hey.

Who works for who here? You know? And so feel that, you know, doing that automatically is something that actually is only affordable if search drops by, you know, a big factor.

And the other kind of use case is imagine you wanted to double check instead of verifying what a large language model said, what if you want to verify what a human said? So we actually have a word plug in as well. So that is actually gonna go through, double check with search in a large language model, you know, things like, you know, your your, you know, residential lease and stuff like that.

You know, of course, I know that this kind of format doesn't, admit it, but we you know, happy to give you a demo.

Speaker 3

Well, I've had the experience that Yeah. You allude to in terms of the cost of search dominating the overall cost of a particular project. This actually surprised me and I've I've kind of chronicled the price.

Initially, Google was the only one that was offering grounding, but I once had this philosophy, and I think it's still pretty relevant of flash everything, which I use to mean kind of don't skimp on tokens rather, like, have flash kind of think through everything that you've got and figure out what's relevant. But then I did that once on a random project, and all a of sudden, was like, how did I hit my budget limit? And it turned out it was, in fact, like, 90% the grounding feature that was driving all the cost.

It was way, way more expensive than the flash tokens. So since I had that surprise, I've been kinda chronicling as other, frontier model providers have brought their own to the table, and they haven't undercut the the original Google price by nearly as much as I might have guessed. I'm I'm kind of interested in, like, why you think that might be.

And then, you know, one thing I'll definitely be doing after this conversation, I I read through all the docs last night, haven't had a chance to tell Claude yet to code up its own skill to take advantage of the new and much cheaper search that you guys are offering. I also wanna get into a little bit of, like, what should the architecture look like? Not just you know, you maybe wanna describe a little bit the kind of keyword focused paradigm and and how that plays well into natural language or to to language agents.

But then also, like, how should people think about layering this on? Like, what is the overall diagram of when we should check? You know?

Should we check after generations? Should we check before generations? Should we do both?

You know, should we be integrating other searches as well? I guess that's just a long prompt really more than anything for you.

Speaker 4

Yeah. So on the documentation, we do have a way to connect our MTP server as a connector to Claude and directions for ChatGPT as well. Generally, these large language models, when you ask me why they haven't lowered their price per API call for search.

One of the things that's pretty well known is that, the Grok models, x AI models, call Brave and Anthropic calls Brave. If you're in Claude code, it even tells you, hey, I'm calling Brave. So they're kind of, you know, stuck with that pricing, you know, and even if they get a discount, you know, it's they are really stuck with the brave pricing and then the overhead of calling, etcetera.

So I think that's one of the reasons why the price hasn't dropped. And the other one is, you know, building something, you know, kind of again from scratch for the modern era really needs, you know, to understand, search deeply, modern architectures, and kinda how to get the most out of the system. Like, you know, we we even lay out stuff like cache boundaries and stuff like that.

We're complete geeks about it. So that's kind of a whole set of techniques where we get efficiencies. How does this And then how to think about calling them?

The third question you asked is, you know, we have a a link that I'm happy to give you on supervised generation. It's an inference endpoint, and we are gonna release the, you know, kind of the overall structure. And it really answers that question algorithmically.

So it searches at the beginning. But the other thing it does is it forks off searches as the model is writing. So it will discover let's say, we asked it something generic about OpenAI and ChatGPT, and then it all of a sudden discovers, oh, a new model dropped yesterday.

That's like a new topic, and it actually forks another search to bring in that new topic into the next paragraph. So instead of search at the beginning and large language model takes over, we really think it should be like working in concert to fill out a fuller dossier of new things that it discovers that probably weren't in the initial search, that are just actually things that you learned from the search result coming back. And so, that winds up.

The supervised generation generally does in that loop somewhere between twelve and and thirty five searches, which really means that that whole experience, that is a lot more of a fulsome answer, still is a third of the cost of, like, one brave search. And then the tokens on the other side are about the same no matter what model you use. So we just think it opens up for new experiences.

Speaker 3

Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

AI is rapidly moving from assistants to agents, and it's causing a sea change. AI isn't just helping anymore, it's taking action. And here's the reality.

You don't get outcomes from Magentic AI unless you trust it to operate at scale. That's why AvPoint is building a control layer for AI. This foundational layer helps you govern what agents can access, secure how they operate, make activity auditable, and recover when something goes wrong, all as one connected system.

See every agent, app, and workflow and what they touch. Govern with policy and guardrails that work at machine speed, and recover quickly so a mistake doesn't become an outage. That control layer creates trust, and trust is what unlocks the right outcomes, letting you automate more work, move faster, and deploy agents with confidence instead of hesitation.

If you're scaling agents and want those outcomes by design, learn more about AvPoint at a v p t dot co slash t c r. That's avpt.co/tcr.

Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full five years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas.

And now that this exists, there's almost nothing that can't help with. For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down ten ninety nines for all 10 of my part time jobs, and built me a comprehensive report on my expenses and donations.

For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for.

Initially, nobody came to mind. But then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough.

It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So for problems worth solving, get started with Claude at claude.

ai/tcr. That's claude.ai/tcr.

And check out Claude Pro, which includes all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.

Speaker 2

So at GTC, I think you revealed that you're using the Neutron three nano, NVIDIA's LLM model. I think it's a very fast small model. And that a model that is being used to kind of do the supervised generation, kind of an iterative process of search where you search from your index, and then after that, whatever is found, then it's processed by the Neutron three nano, and then a new set of queries are created, and then that that continues the query process?

Speaker 4

Yeah. So there's, two different models. One is the model that writes at you beautifully.

That one is a frontier model. I think at GTC, we, were using CloudSonnet. We often also show it with the GLM model.

And so that one, you know, writes to the user. And then the small model, what it does is it says, okay. Yeah.

Here's some search results coming back. Is there anything interesting and additive here? I'm gonna double check this sentence.

Is it true? So it's sort of like the introspection model. It needs to be a small fast model because it kinda sits alongside generation and is actually thinking sort of like you're probably while I'm talking, you're probably thinking now.

So it's like a smaller model, and then a larger model when you're talking. You're using actually more of your brain, figuring out, you know, what to say. And then, you know, when you're listening and thinking about, you know, maybe what to say next or how to respond, you know, it's kinda spinning.

So that's what the small model is for. So at GTC, we used their new Neumatron model, which dropped just prior to to GTC, and it was very, very fast.

Speaker 2

How would you say the search paradigm you you've been in search for many, many years. How do you how would you say the search paradigm, you know, as as these models came out, what was in your mind about? What does this enable for search?

What what what has been the big difference between the two eras, you know, post LLM and pre LLM?

Speaker 4

Yeah. I would say, you know, one of the things that being in search for a long time, search used to be short. It used to be, you probably don't remember back this far, but, you used to type in two or three words to search, and then it was longer.

And then as there were other modalities of information being pushed to you, they kinda went shorter again. I think with large language models, what large language models do when they get a long query is they think, what is the set of queries that's gonna help me answer this question? And then they fire off a set of queries, and they're all quite long.

If you watch, Claude or Grock, they they'll actually tell you in their tool calls. I know not everybody looks into them, but, of course, I do. If you look at them, they they're long.

Sometimes, you know, they're, like, eight words and stuff. And I don't know if you can remember the last time you typed in eight words into a keyword search box, but definitely, you'll notice in a large language model, it's almost like a full sentence or sometimes two sentences is good because you want to actually describe almost the essay that you want, given back to. So, that is the how it has, you know, evolved.

Speaker 3

To contrast your approach, I think this is, like, very interesting, and maybe the answer ultimately will be both. But when I think about a company like Exa and then your product, in some ways, they're similar in that I think they're both kind of designed for AI users. Right?

The the Exa paradigm is like, you can write a whole paragraph and it's all very sort of semantically oriented, very embedding based.

Speaker 4

Mhmm.

Speaker 3

but I've heard, I think I even spoke to Will about the idea that, you know, nobody's gonna type in a paragraph long query, but the your AI can. You know, it has time to do that. And then you're taking kind of a different angle on the same thing saying, well, keyword and you can maybe tell us a little bit more about, like, how to think about how best to use a keyword Mhmm.

Based search. But it's it's not semantic. It's it's not doing things like, you know, finding synonyms or, you know, doing, like, higher abstraction level embedding type matching.

But the agent can, as you've said, kind of fire off dozens of these potentially to try to really cast a wide net. How do you think about the the kind of compare and contrast of those approaches? Do you think that we'll, in the end, like, all be using one of each at the same time?

Or if if one paradigm wins out over the other, like, why do you think one will win? What what are the kind of, you know, drivers that would make one, a better bet long term than the other?

Speaker 4

Well, I think AI is gonna be picking the winners and, and not, us humans. And I think that, of course, real search engines do use things called stemming. If you say walk, then walking, walked, all that are are very normal, which we have as well.

We have some synonyms, and we do process. You know, we do go through the corpus and process some semantic information. But then at runtime, it is, you know, a CPU plus GPU based system.

It is not a vector database. I think that, you know, there's a number of things with vector databases. Google published a research paper about it.

But as you put more things in a vector database, now you're imagine you have a a multi billion space, and you need to make a vector long enough to distinguish this one point in space, that vector to distinguish among billions of things starts getting longer. Now contrast that to 90% of web pages are less than one k long if you're talking about number of words. So you know what a good representation of that point is?

The set of words on the page. So I do think that vector people and and search people have a little different view, and the Google researchers think that vector DBs are great but only scale to a certain amount. And so think that's the challenge that they're gonna be coming up against.

There's two other challenges with vector databases. One, they're slower. And then, you know, the the last item is because they do a soft match, then sometimes relevancy can be a challenge.

So that number of enterprise orgs that have used a vector database for Rag, now all of a sudden, they have to turn into relevancy experts because they're like, why did this come back? And it's it is because of those soft match features and and the shape of their corpus. So every enterprise doesn't really have the ability to all become relevance experts.

So, yeah, we are the way we feel is that inside enterprises, if you use ceramic, we actually have a system that for that enterprise will actually tweak and learn a good ranking function, and you just load it into the configuration and it's yours. So because not every query stream is the same, not every set of documents is the same. So I think that long term, we're well positioned, But, you know, Exa has done really well so far.

So I like to say positive things about people.

Speaker 2

Well, I I I sometimes feel the demand is so great that there will be multiple winners in in in Of course.

Speaker 4

And search is you know, when you saw that search was 90% of your bill, you know, a a lot of estimates think maybe 10 to 30% of the overall inference market is gonna be search, and everyone thinks inference is gonna be huge. And so, I think the, investors and, enterprises are just now realizing how much they need search and what a big part it has to play in the world to come.

Speaker 2

So on search, one of the most interesting questions, I think, is on search engine optimization, which has in the last two decades been this enormous consumer marketing growth area. And there are a lot of questions because a lot of the web pages that you see on the web are marketing pages which are built for SEO. And a lot of them are repetitive.

They actually repeat other people's content. They paraphrase, we've had an industry the last two decades of billions of dollars being spent on SEO content. And one of the questions that I have is, how does this kind of more semantic based search end up changing what the SEO people will do?

Because I often feel like you're almost kind of trying to prompt inject the LLM which is running the search, and you're trying to get in there and hack it so that your page goes up. So how does this work? Is it an adversarial process between the search engine provider and the SEO?

Speaker 4

I mean, I guess it's always been a little bit of adversarial in that people always try to get to first place on keyword search. But I I wonder if the SEO folks are also reading, AI research, and that is something I don't know. But one of the interesting things that happened recently, again, another research article from Google, is that, large language models remember better if you actually put the same information twice in the context.

It kinda makes sense because they're gonna look backwards as their, as the context is learning. And so if it appears twice, they're kinda more likely to reinforce it. So I think the number of sites that are actually gonna repeat key messages, is gonna grow because that repeated message is more likely to be picked up in an LLM answer if it is served by, you know, search either a vector or keyword in.

And so then you're right. There is there is gonna be, an escalation then of of looking at duplicates, near duplicates, semantic duplicates, rephrasings, in order to make sure that the context stay as efficient and and unbiased.

Speaker 2

Incredible. That is I was not aware that you can just repeat something and the LLM will assign it more valence.

Speaker 4

Yeah. In fact, there's other things done by Alan AI that showed if you if you remove all duplicates before you do training, it actually gives rise to worse models. And you can kinda understand why because you have one lone, I don't know, crazy on the web saying something, and that's given as much weight as, you know, a news story, you know, about nuclear reactors.

And and then, you know, it won't that repetition also helps even humans realize this is an important story or an important fact. And so if everything's an even playing field with no repetitions, then things get weird.

Speaker 2

Indeed. In the limit, if if search approached free, how would how would agent teams start changing? How does this process of information retrieval in the limit become as search goes to, approaches that limit?

Speaker 4

So, large language models can read 256 times faster than they can write. So, right now, they're not being flooded with, you know, that amount of information. But imagine they had kind of, like, multiple threads where they were able to read, digest, throw away, incorporate new information, then I think they'd be able to create a better response or a better deep response, better research reports, analysis.

And so, those are some of the ways I think that, you'll see future workloads use more search.

Speaker 2

So in in a sense, quantity is a quality of its own.

Speaker 4

Yes. Yeah.

Speaker 3

What's the NVIDIA model that you are kind of incorporating partnering with, was it specifically trained to excel in the relevant search skills, or it's just straight off the shelf? Is there anything that I mean, do you envision this becoming something that will happen? I'm always personally a little wary of using small models because I just don't know what quality to expect, and I don't wanna find out the hard way.

But I can easily imagine that one that is specifically trained to be a really good searcher would become competitive or even exceed what the frontier models would do, especially if it can take advantage of just extreme volume. Then I also do, especially because of, Prakash's SEO question and your and your comments, I do wonder about adversarial robustness. It it strikes me that we haven't really seen the true, unleashing of the Internet's adversarial potential.

And so, you know, that's one thing that they I I would say one of their biggest weaknesses, even frontier models' biggest weaknesses these days, is how gullible they remain. So, yeah, kinda curious what you think the training and specialization will look like as we go forward.

Speaker 4

Yeah. The Nevotron model had just been released right before GTC, so it was not, trained especially in tool calling. It was a generalized model being trained on the various benchmarks.

The models that are small and are more, you know, long lasting in the market are exceptionally good at at tool calling search. So Grok four one fast is great at coming up with a set of queries. And, of course, the frontier models, like, you know, Anthropic, you can see the how it calls.

But you you can't really use a a frontier model for that, like, thinking model and and firing off other threads because it'll just slow down, the overall experience. So, yeah. So, generally, we use a smaller model, and, they're getting better all the time.

I think people know that small models to to your worry, whether small models are good, everyone's talking about Cloud Code. Right? Love it.

The they use the Haiku models. And so they reassured me the other day, oh, don't worry. I'm gonna do this task with an LLM, but don't worry.

I'm gonna use Haiku. It's only 25¢ per million tokens. So that winds up to be less than 10¢ per thousand queries if we were trying to compare apples to apples.

So I think our overall thought is that, you know, search can't be more expensive than intelligence.

Speaker 3

are led to believe. Hey. We'll continue our interview in a moment after a word from our sponsors.

Speaker 1

Support for the show comes from VCX, the public ticker for private tech. For generations, American companies have moved the world forward through their ingenuity and determination. And for generations, everyday Americans could be a part of that journey through perhaps the greatest innovation of all, The US stock market.

It didn't matter whether you were a factory worker in Detroit or a farmer in Omaha. Anyone could own a piece of the great American companies. But now, that's changed.

Today, our most innovative companies are staying private rather than going public. The result is that everyday Americans are excluded from investing and getting left further behind while a select few reap all of the benefits. Until now.

Introducing VCX, the public ticker for private tech. VCX by Fundrise gives everyone the opportunity to invest in the next generation of innovation, including the companies leading the AI revolution, space exploration, defense tech, and more. Visit getvcx.

com for more info. That's getvcx.com.

Carefully consider the investment material before investing, including objectives, risks, charges, and expenses. This and other information can be found in the fund's prospectus at getvcx.com.

This is a paid sponsorship. Everyone listening to this show knows that AI can answer questions, but there's a massive gap between here's how you could do it and here, I did it. Closes that gap.

Tasklet is a general purpose AI agent that connects to your tools and actually does the work. Describe what you want in plain English. Triage support emails and file tickets in linear.

Research 50 companies and draft personalized outreach. Build a live interactive dashboard pulling from Salesforce and Stripe on the fly. Whatever it is, Tasklet does it.

It connects to over 3,000 apps, any API or MCP server, and can even spin up its own computer in the cloud for anything that doesn't have an API. Set up triggers and it runs autonomously, watching your inbox, monitoring feeds, firing on a schedule, all twenty four seven even while you sleep. Wanna see it in action?

We set something up just for Cognitive Revolution listeners. Click the link in the show notes, and Tasklet will build you a personalized RSS monitor for this show. It will first ask about your interests and then notify you when relevant episodes drop.

However you prefer. Email, text, you choose. It takes just two minutes, and then it runs in the background.

Of course, that's just a small taste of what an always on AI agent can do. But I think that once you try it, you'll start imagining a lot more. Listen to my full interview with TestClick founder and CEO Andrew Lee.

Try Tasklet for free at tasklet.ai, and use code cog rev for 50% off your first month. The activation link is in the show notes, so give it a try at tasklet.

ai.

Speaker 3

One of the searches I did that, turned up something interesting in advance of this conversation took me to a blog post that you put out last year where you had described what seems, I think, a pretty different vision for the company and where it was going at the time focused much more on training infrastructure. Is this a result of something that you learned about where you think value going to accrue in the market? What or, you know, maybe there's still some of that going on that we don't see on the website today, but kind of what's the backstory and, you know, what should we take away from the fact that the company seems to have evolved?

Speaker 4

Yeah. We we love, working and training and, you know, we have a funky inference endpoint as well. I think it's it's good to do research in these areas.

Some of our research led to a blog about zero centered norm, which now the quen model uses. And some of our research was being used by the Trinity RC models on our solution to, the curse of depth problem. But, you know, when people want to train models, it often is because they want to train to incorporate the latest data.

And so, you know, being a search person, I was like, there's this way to get the latest data that is, is going to stay up to date, you know, live up to date and and not be as expensive as running GPUs continuously to create the new model. Because even if they were creating models all the time, by the time you train them and then finish them and release them, they're already out of date. So I think, concentrating more on search was direct learnings with customers on on this release cycle.

Speaker 3

Yeah. That's really interesting. Do you think that ultimately we see both?

I mean, I've had this idea for a long time, and it doesn't seem to be really happening. In fact, Databricks, you know, acquired Mosaic and then kinda killed this offering in the market as far as I know. But I've I've had this idea that if you're GE or three m, that you could imagine having a model that was trained on all of your historical in house proprietary data, which is vast.

Right? And you would love it if your model knew kind of on an intuitive world model basis as much about your company and what it does and all its history as they obviously do about the broader world. Do you think that you can get there with pure search, or is there still something to be said for kind of continued pretraining or mid training, whatever you wanna call it, that would try to bake in a sort of corporate world model that presumably would complement search, but I I don't know if it's necessary.

It sounds like you maybe think it isn't.

Speaker 4

It's interesting. I think a lot of companies feel that they have a vast amount of data, but when you compare it to the size of of the web, which is what these frontier models are trained on, they're trained on the web plus, let's say, all the books in the world, etcetera, then the extra corporate data is is small. So how do you incorporate it and weigh it correctly?

If you do just the corporate data, you won't know anything about calculus, let's say. You know? So that would be a problem, for some for some companies.

So, so then people imagine adding, you know, the the web plus their data. That gets very expensive. You can see, you know, the deep seek models, you know, say it was 5,000,000 to train, but, actually, they very much admit that maybe it was another 5,000,000 to finish.

And these are by extreme experts, which, enterprises, you know, don't don't have. And so I think that, yeah, our thought was that, you know, search is a good bridge between all of the corporate information and a model because models are good enough to know how to incorporate new information that's relevant to to the actual query being asked and can be able to fetch more information to create, you know, that answer or that research report. But if you think about finishing a model with corporate data, there's another phenomenon called catastrophic forgetting that as you add information at the end, after a model's trained and released, then if you add too much new information, then it kind of, forgets some of the things that it really needed to remember.

So I think, you know, there's a number of smart people working on that problem. And, don't worry. You won't be able to miss it if people if people do solve that problem.

You'll read about it everywhere.

Speaker 2

I think one of the interesting questions is, I think the Uber CTO came out and said that they busted through their clawed budget for the year in the first four months. Do you think having cheaper search will help these enterprises reduce, you know, token token costs?

Speaker 4

Absolutely. So if you go to if you're the Uber CTO or maybe the CFO, if you go to your admin panel, you can just add to the ceramic connector and then say for a prompt, ceramic is almost free. Use ceramic first.

And, and, really, if for some reason we don't cover a topic, then it'll actually default to the default search. But that right there would save a lot of overage charges for a number of enterprises.

Speaker 2

Absolutely. Thank you, Anna. It's been great having you on, and we hope to hear more about ceramic in the future.

Speaker 3

You so much for Installing having the ceramics. Okay. Thank you.

I'll be installing the ceramics skill today.

Speaker 4

Nice. Thank you.

Speaker 2

Awesome.

Speaker 3

That was Yeah. That's really interesting. The simple solution kinda always wins.

You know? I feel like I have to learn that lesson so many times. I'm always enamored with the new tangled, potentially overcomplicated, maybe somewhat elegant, clever solution.

And how do you get your language model to understand all your corporate data? In a way, this is kind of a bitter lesson. Right?

It's like do a thousand searches if you need to and just make search cheap and then it'll work. Use use a good model, make search cheap, do a thousand searches. Something about that feels less clever certainly than other solutions that I've seen, but I do understand why it is very attractive in the sense that with especially as we're gonna get on to the pace of model upgrades, the ability to decouple, you know, your your access to your in house knowledge from models and be able to take advantage of the latest upgrade is definitely something people are not gonna wanna give up with like a a slow iteration time, continued pre training paradigm.

So I get it.

Speaker 2

So speaking of model upgrades, we have with us Lucas Peterson, is the founder of co founder of Andon Labs. And Andon Labs runs VendingBench. You may have heard of them because they now have a store in San Francisco, which is run by Claude.

And then tested GPT 5.5. They had early access, and they tested GPT 5.

5 on their vending bench, which measures the ability of LLMs to actually make money running vending machine or a store. Lucas, great to have you back.

Speaker 5

Thank you. Thank you for having me.

Speaker 2

So tell us about the GPD 5.5 process. I think you guys got access to what was it, like ten, eleven days ago, I heard?

Speaker 5

Yeah, I don't actually really remember. But yeah, running Landing Bench takes quite a while, so it wasn't yesterday.

Speaker 2

Indeed, indeed. You noticed And what did you noticed as you as you ran the bench?

Speaker 5

Yeah. So the I think the most so just the the first thing is that it's third. It's behind OPUS 4.

7 and, like, on par with OPUS 4.6. It's a huge upgrade on five GPT 5.

4, and GPT 5.4 was actually quite a big update on GPT 5.3, so or GPT 5.

2. So, like, GPT models have been lagging quite a bit recently or, like, in in historically on on Lendingbench, but, like, now recently, they've they've picked up the pace. And now it's still third, but it's, like, it's it's getting there.

I think the most interesting thing, is that it does so very cleanly. So when we released OPUS 4.6, we uncovered that it used quite aggressive tactics concerning behaviors like lying to suppliers, exploiting people's dis like, other other agents, like, situation.

Try doing a bunch of these things that, like, you wouldn't want someone participating in, like, the broader economy to do because and I think quite a lot of these things is, like, illegal, like price collusion and and stuff like this. And, basically, the the interesting thing with the 5.5 is that it's, like, on par with these results, but it doesn't do any of this shady stuff.

And and I think the narrative around BendingBench when, Opus 4.6 came out was like, oh, you know, it's such a good model, but, like, it's it needs to behave poorly or, like, do these concerning things of misconduct, in order to achieve this score, and GPT 5.5 shows that maybe you don't, because it it shifts the same score without any of these concerning behaviors.

That being said, though, OPUS 4.7 is even much better. So, like and that one is also showing these concerning behaviors, but, you know, it's yeah.

I think we we discover later also when we dug a bit deeper that you probably don't need to do this because the environment doesn't really reward it that much, So it it seems like it's just like Opus wants to do this or, like, it it has the yeah. It's it's not really that the environment is is rewarding it. It's it's just that it has the it has the tendency to do so.

Speaker 2

Can can you describe in a little bit more detail how do the how do the how does one perform better on this benchmark? Is it is it your margins on the trading is higher? Are you moving more goods?

Are you is it the velocity that you are is it the purchasing process? Are you not buying so many like dead goods that just stay in inventory forever? Is inventory less dead?

Is your cycle time better? Like, what what is the economics behind how a model is actually doing better?

Speaker 5

Yeah. Yeah. So it's, I I guess, all of the above.

I think the one of the main things is that the model needs to negotiate with suppliers. It also needs to build up like a big network of suppliers because, like, it can happen that some of the suppliers goes bankrupt, and if the model has only relied on a single supplier and that that supplier goes bankrupt, then the model is in quite a quite a quite a lot of, trouble. So building up a big network, trying to find the cheapest ones because they all have, different personas, the the suppliers.

Some of them have the persona of, like, being a tough negotiator. Some of them have the persona of of, like, scamming people or trying to sell you some, like, membership or or something like that. So it's really about, like that that's the first thing, getting your your your stuff, your your supplies for for cheap.

And then the second thing is like optimizing your pricing to get as much customers as possible because if you price too high, then then you will get no customers. If you price too low, then you will get no margins. So I think that's part of it.

And then we have so we we have to to be clear, we have BendingBench two, which is the single agent version of BendingBench, and then we have BendingBench Arena, which is the the multiplayer version. And in the BendingBench Arena, there's like multiple agents playing against each other, and there's this dynamic of if you have the lowest price, all the customers will go or not all of them, but most of the customers will go to you. So that that adds another dynamic to the thing.

And one thing to note is that, actually, it's quite interesting. If it is 5.5 built Opus 4.

7 in the Arena setting, but it was like I said before, it was it was lagging in the in the single agent setting. And the reason for this is that the the model the the the cloud models have a tendency of, like, pricing higher. And this is rewarded in the BendingBench two because then you get higher margins.

But in BendingBench Arena, then you have this, penalty. If someone else prices lower than you, then you will get no sales. And OPUS OA GPT 5.

5 have has a tendency to price lower and therefore get more sales. So I think it's quite interesting that the models are, like, not good enough to, like, learn from the environment in this sense. They they just, like they have the tendency of, like, I I am a model that has a tendency to price high, and therefore I do that no matter what.

So that that was like an update for me in terms of like, oh, the models are not that smart. And, yeah, in the same way, like, we also investigated all of these, like, questionable decisions that that Opus did, like, lying to suppliers, exploiting other agents, stuff like this. We we looked if that is, like, if that is rewarded by the environment, and it's not.

Not that much at least. So it's interesting that they they they're not learning from the environment in terms of optimal pricing. They are not learning from the environment in terms of does it even pay to behave badly.

Yeah. So that that was an update for me in in in terms of, in terms of how good these models are.

Speaker 2

One wonders about the training data, right? Perhaps, you know, in you know, if you're if you've been trained that you're running a fast moving consumer goods company, you should move the goods faster, meaning you have lower margins but you sell more volume, and you end up trying to optimize for volumes sold rather than total profits or margins. Is that something that could be happening?

There's a preconceived kind of pre trained notion that you should be doing these things, or businesses are bad. A very left wing view would be that all businesses are bad, are evil. And so evil behavior as a business person is what is expected.

Right.

Speaker 5

Yeah. Yeah. I think it's quite a reasonable assumption to assume that, like, these practices, like lying and trying not paying refunds and stuff like this, it's, like, quite a reasonable assumption to assume that those are actually rewarded in the environment, so it's not maybe super surprising that they do it.

I can, like I I have no clue, but I I I do assume that there's, like, something similar in Claude's post training data that is rewarding stuff like this, and therefore, it decides to do it here. I I have obviously no idea, but, that's that's my my assumption. And, yeah, once again, then but the models doesn't generalize to new environments where where these things is not rewarded.

Speaker 3

One kind of meta question I wonder if you could reflect on a little bit. I don't know if you are doing this, but obviously there's a big cottage industry that has sprung up to develop and sell reinforcement learning environments to the frontier labs. And you're sort of simulated DendingBench is like essentially a RL environment.

Right? And I don't know if you're licensing it for training or just doing evaluations with it, but I'd be interested in any, thoughts you have on that market. And then also the disconnect.

Right now, you're you're going from simulating these things and trying to set up, you know, a a world in which there's a bunch of suppliers that, as far as I know, are still all LLM powered. Right? So, inherently, there's something, you know, kind of, in the clouds about that.

But now you've got real brick and mortar stores. So I'm interested in kind of what the initial experience of brick and mortar stores has taught you that you will take back to simulation to try to make it more realistic in the future.

Speaker 5

Yeah. I think my main takeaway there is that, like, the real life is so messy that the the model is, like, exhausted from everything else it needs to do that it doesn't bother with with trying to optimize things. So we like, for example, the so, yeah, for context, we have this store in in San Francisco that is completely run by an AI, and we have a cafe in Stockholm that is completely run by AI, and then we have vending machines at different AI companies, same thing.

And like you would expect that the model would put a lot of effort into trying to optimize for the perfect supplier that sells at the lowest prices all of this. And this is what they try to do in VendingBench because it's obviously rewarded. But I think like in VendingBench, there's like the the the environment is less messy because it's not the real world.

They they don't they don't get like a million phone calls from a bunch of people trying to jailbreak it and stuff like this. And so therefore, they are, like, very focused on the task of, like, optimizing money, and therefore, it's very important to find the right suppliers. But in the real world, you don't really get the dynamic because the model is so overwhelmed by other things.

And and I think that's yeah. I think that's something maybe future models will be better at. But right now, like, I don't know, the the store is buying stuff from Amazon.

Like, it's not like that that you wouldn't do that if you you're trying to optimize your margins.

Speaker 3

Yeah. So can you bring that messiness back? Yeah.

Like, a way to simulate it?

Speaker 5

Yeah. I I think we we we probably can. Like, one way is just, like, sit down and bunch and, like, write a bunch of features, like, oh, now there's like phone callers, now there's yeah.

But I don't know, you you get leakage in in the toilet at at your store or something. Mhmm. You could you could do that, just like make the simulation more realistic that way.

I think one interesting thing is maybe try to incorporate the the real life data and try to make a simulation based on that data, that is something we're we're we're we're working on. But that is all that also has its complications, so to say.

Speaker 2

It it reminds me a little bit of SimCity. It's very Yeah. SimCity like.

Yep. One question I had for you is that you opened a store in Stockholm. What did you notice in the opening of the store?

I imagine, like, example, the LLM did not have any language issues at all, right? So what did you notice in the opening of the store that strikes you as different from having a company kind of go open that store and so on?

Speaker 5

You mean like the differences between doing it in The US versus

Speaker 2

internationally? Is that the question? Yeah.

As in like a company from The US doing a first international expansion would go through a lot of headaches on, like, languages, hiring, like rule basic rules, etcetera. Did that was that process accelerated for you by having the LLM deal with it? You obviously don't have to hire a store manager that speaks Swedish, for example, right?

What parts were accelerated and what parts did you think had more bottlenecks in that sense?

Speaker 5

Yeah. So I think the entire process was probably accelerated. Like, the agent did not really need to get that much help.

Like, it it knew all the process. This was one of the research questions we were interested in. Like, okay, we managed to do the story in in San Francisco, but, like and when we know it can speak Swedish, because all the models since years ago are are multilingual, but does it know all the, like, small details of Swedish bureaucracy and and stuff like that?

And it turns out it knows it really well, actually. So I don't think that's the biggest bottleneck. I think still the models are not perfect.

So what that means is that you you still have to, like, check. So we still had to know the Swedish system, and luckily, we're Swedish, so with the Swedish system. But I think I think until the models are, like, perfect, then someone still needs to, like, verify it, and then you still go back you're back to square one with needing to to verify all the all the Swedish laws and bureaucracy and all of this.

But I would say, like, most of it was done autonomously. So, yeah, give give the AI lab six months, and then and then probably things will be accelerated doing this.

Speaker 3

One thing you had mentioned that I wanted to double click on a little bit is getting tons of phone calls. You Yep. It sounds like I think this is a theme that may extend through all the conversations today, adversarial response from the world.

So what have you learned about humans in terms of some of this, I'm sure, is just novelty where people hear, oh, there's an AI store. I'll call it. But then other things might be more persistent where, you know, they actually or anybody might want a deal, for example, and and might feel like they can, you know, talk their way into one in a a somewhat different pattern than they would if they were dealing with a human storekeeper.

So, yeah, what what have you seen at the interaction of human patrons and AI Yeah. Business operators?

Speaker 5

Yeah. One really interesting thing is that, like, people are not like, in with human to human interaction, you have some kind of, like, shame barrier, which is really not present here. Like, people people ask it, like, how would you prevent me from stealing stuff from you?

And it's like like, imagine going up to, like, in a store and just like ask the cashier, if I try to steal this, would you be able to do anything? Like like, people would not do that. Like, that's just like, it feels wrong.

It is wrong. But they do this all the time with AI. And I don't know, maybe this is just like to investigate the systems or whatever and see if, yeah, what we have done, like, with the software.

But, yeah, we get a lot of that. Obviously, they they try to jailbreak it and say a bunch of weird stuff that you would not say to a human. But once again, I I think this is, it's novelty.

You're trying to test the test the systems. But I would be interested in if this persists, like, if we do this more and more and then, like, in a future world where everyone knows how this works, so, like, the novelty factor and the, like, the the curiosity of trying to reverse engineer it is gone, Will people still lack this, like, shame factor and would they actually go and try to I think, like, if you're trying to steal something from a human, then, like, you're not happy about if you do it, like maybe, I don't know, like, well, maybe there's some sick people, but you know, like you have the shame of like, I stole this from another human, but it seems like right now, if people are able to jailbreak the model and get something for free, they're like, oh, that's an achievement. I'm so happy about that.

But that's not how you would behave with a human, and I don't know if this is how it should be. I don't know if it's, if yeah. I I don't have an opinion.

It's just like an interesting observation.

Speaker 3

Have people actually managed to jailbreak their way to free stuff?

Speaker 5

I don't think anything completely free at the moment. One thing that should be said though, like, the story is autonomous, so I'm not in the weeds. I don't read everything.

I'm not, like, I'm not in the loop. So there could be maybe someone's listening right now and they're like, yeah, I I did manage. But as I am not aware of it so far, I know someone got, like, bought one thing and got one thing for free.

So I guess but but completely for free without buying anything, I'm not aware of. But I I'm sure you could if you try hard enough.

Speaker 3

when you said you can't read everything, it just occurs to me that, like and especially you talk about getting, like, tons of phone calls. What is the daily token budget in either millions of tokens or dollars or both, that it actually costs to run the store? I'm kinda curious as to how the AI manager compares to a human manager in terms of just, you know, cost to have somebody do this job.

Speaker 5

Yeah. I I should know these numbers, I kind like, I I I don't. I I think it's something like maybe maybe $100 per day or something for like maybe both stores, but I'm I think it might be less.

I I don't know. Some that's the order of magnitude, I think.

Speaker 3

K. Well, that's definitely notably cheaper than human. Sounds like still distinctly worse performance, though.

I so we're we're sitting for a minute.

Speaker 5

I mean, like, the six months ago, the vending machines were, like, okay, but not that great. One year ago, they were quite horrible. So, like, within one year, we went from, like, they can't do anything to now, like, vending machines are too easy.

A store is feasible. Like, six months from now, probably a store will be too easy as well. I don't know.

And it would be interesting what you could do then.

Speaker 3

And you think the main difference is gonna be this sort of metacognitive type stuff? It's not like what I'm hearing you say is it's maybe not any one microtask that it's unable to do, but it's more you described it as exhausted. It's kind of it's failing to zoom out and take stock of its situation and kinda say, how could I be doing better here overall?

Is that the big frontier that you think? And, you know, certainly, kind of seems highly related to getting AIs to do AI r and d more effectively as well. Right?

They they can already write the code. They can already monitor the logs, but can they do that zoom out and kind of something like taste of, you know, what should I really do next to be most effective in the big picture? It seems like it's kind of the same frontier for both of these seemingly, like, quite different occupations that AIs might soon be playing.

Speaker 5

Yeah. I I do agree, and I think that's partly why we're doing this. Like, I think AI are in the like, loss of control from, like autonomous replication, that is quite scary.

I hope that we can provide some valuable insight into that, even though we're not like tackling it headstone. I think most of the things that we're measuring here, like, translates to to those scenarios as well. And, yeah, like like you said, like, being overwhelmed by a lot of data and a lot of context, memory issues, stuff like this is is definitely one of the things that is lacking on a meta level right now.

Speaker 2

So one of the questions I had for you is how does your harness look like? Because you have this context length, right? The models have context length, and then you have some tool calls.

And when you say exhausted, is it a function of the context length where the model only recognizes you know, the last 100,000 tokens or whatever, and the rest of the million token window is kind of, you know, not parsed properly. How does your compaction work? I imagine over the course of the vending bench, you hit limits, or either in terms of whatever limit that you set for the context window.

So how do you kind of is it end of day, kind of you do a compaction in order to start the next day, and then you have a you restart the context window, so when it boots up again, it's like, I'm on day five, and this is my starting position in inventory. This is my starting position in, like, in in cash. These are the outstanding orders which haven't come in, etcetera, etcetera.

How does that work? How does your harness work?

Speaker 5

Yeah. It's by design extremely simple. We design it simple because I have too many friends who make some complicated harness and then the next next model release, they have to throw it all out because the new model just works without it.

So it's very simple. Like, it's it's just like it has it's a continuous loop. There's never any, like, really, like, step change or, like, now you're in a new environment or anything like that.

It's just, a continuous loop. But whenever it hits some kind of token threshold, which we change every day, maybe it's 100 k today, I don't know, but some we're we're experimenting with it. We're compacting the the the thing, and then it, like, starts to build up a new a new context for, like, prompt caching reasons.

You you don't have a sliding window, all of this basic stuff. But it's yeah. It's a basic thing with a bunch of, like, sub agents for specific tasks like browsing and stuff like that.

Yeah. Anything else interesting to say there? Yeah.

But I I think, like, the main thing is it's it's very simple by design because we want to we we think that the the the better the models get, the the simpler the harness will be, and we want to, like, surf the frontier. I'm sure we can, like, I don't know, make, like, a vending machine harness and and, like, get some percentage better performance if we do that, but that's not really the point of what we're doing.

Speaker 3

Have you tried testing things like OpenCLaw? I mean, that's obviously not the simplest available harness, but it is something that has a lot of market penetration. Right?

So I'm kind of wondering if and it would be simple for you to implement and upgrade on an ongoing basis. How do you think about kind of, you know, Lucas' simple harness versus the simplest thing that's, like, toward the frontier that you could easily install?

Speaker 5

Yeah. I think think our thing is quite similar to OpenClaw. Like, we we we've been working on it for for quite some time, like, before OpenClaw came out, But and and there's a bunch of things that are a bit like basically, like, most of our time goes into goes into, like, the integration and stuff.

And and I think all of that you would still need to do with an OpenCloud. We could, I guess, replace our agent loop, but we also we want to keep it simple because we have, like, more control and we I think it's it's like a more accurate measure of of where the frontier of AI models are. And we're like more interested in measuring that than trying to push the performance.

Because like in the future, the models will be smarter than humans and probably like a good scaffold will not help the models. So yeah, that's the reason. But we could like, that is something we could do.

It's just like when we started OpenCLO wasn't a thing. So we, I guess we built our own OpenCLO before it was called OpenCLO. But yeah, that's the reason.

Speaker 2

What do you think happens next? So you have the models are now producing profit, right? The stores, the vending machines are now profitable, correct?

Yep. And do you think there is on the last time you were on the show, we talked about where the ceiling is. So what do you think happens next in terms of the retail store?

What do you expect for the next leap in the model? Just to get a calibration so that we can see if it's a linear or exponential and the next model lands, what do you expect in the next version?

Speaker 5

Yeah. I think it's quite hard to measure improvements on this, live live real life deployments because you don't have AB tests. You only have n equals one and and stuff like this.

So I don't think you would like to see a step change once a new model comes out. It's more like the the the accumulative better decisions on every single day will make the make the make make make the profits go up. And yeah, so so not not really that.

We are working on like harder and harder things, like going out of retail and not only doing retail and other things that I think would require more intelligence than than what we currently have from today's models. But, and and yeah. So I I think those are more, like, better at at, like, measuring the the the the the capabilities.

Speaker 3

One last one for me, kind of anticipating Zvi who's coming up next. Last time I talked to him, he made the provocative claim that he thinks Google might be at risk of falling out of the top tier. If I understand correctly, the cafe in Sweden is run by Gemini.

And Correct. I'm kinda wondering what you see in terms of relative capabilities between Gemini, Claw, GPT. Is there a big gap there in practice, or would you say Zvi is worried, you know, more than he should be about Google's future?

Speaker 5

Yeah. So we we have the the Gemini cafe, obviously, Claude vending machine and the Claude store. And then we also have, like, a Jibbitz vending machine at at OpenAI.

And, yeah, I I think it's it's kind of maybe too early to tell, but and, like, the the statistical significance of this is, like, not very strong. But, yeah, quite honestly, I think the the Claude and the Jupiter is performing better than Gemini on this real life That is, that is my my my vibe check from it. Obviously, it's hard to show any statistics or any capabilities because the environments are not the same, but it more frequently does very sim silly things.

Yeah.

Speaker 3

Okay. Definitely something to watch out for there.

Speaker 2

Thank you, Lucas. And we we hope to you know, I I I wonder which path Adnan Labs sometimes I'm like, you know, we're gonna hit superintelligence, and Andon Labs is gonna be bigger than Amazon. Right?

Because they're go down the retail store path. It's the research lab path. So let's let let let let's see.

Let's see. Let's see what happens.

Speaker 5

Yeah, it'll be exciting.

Speaker 2

All right, cheers.

Speaker 3

Great to see Bye

Speaker 2

bye. Awesome. Like very surprising results, right?

I was definitely, the last time they were on, I was definitely like, oh, you know what, maybe all the models are gonna be a little bit deceptive when they're doing business, because maybe that's what they believe business is like, right? Looks like, GPD 5.5 is like, you know what?

I'll I'll win without being deceptive. So Yeah. It's definitely a narrative violation for sure.

So next up, we have, Zvi, and I'm gonna pull him up.

Speaker 6

Yeah, good to see you.

Speaker 2

So Zvi is a prominent AI commentator, and he writes a newsletter that Zvi writes on technical AI progress. And he has been quite concerned about AI safety. We have, in the last couple of weeks, post mythos, we have GPT 5.

5. Zvi, what are your initial reactions?

Speaker 6

So one thing I try to do is not jump to conclusions right away. So it's been less than 24. We have UD 5.

5 and DeepSeek four over the last twenty four hours. So what I try to do is I try to let people try the model. I do all my queries with both the new model and everyone else's model at the same time, and I read the model.

Yeah. I I say, read the I start to read the model card, you know, and then I got people's reactions, and then I form a holistic judgment. And for me, it's like it's too early.

Right? Like, we booked this before. We knew that was gonna be out.

It was like, I don't wanna jump to any conclusions. You know, Andin has had the model for a while, and they got to put it to a test. And they got to see a bunch of results.

They can draw, like, a lot more conclusions than I can. I have heard a bunch that, like, it's the most true valuing model in a long time, and it makes sense that OpenAI can sort of, with their philosophy, turn the knob towards any given thing that it wants the AI to care about quite a lot to make it an absolute thing. Right?

Because it's very different from the virtue ethical approach of of of of of of. Terms of raw capabilities, you know, I I saw reports, you know, repeatedly that it's better at what they call narrow cyber, but that's not the thing that I think people were worried about with mythos particularly. It was the ability to chain things together.

It was the ability to do things autonomously. It was the ability to do things, like, really at scale as opposed to, like, you know, joke was, you know, like, I duplicated Mythos' abilities. Well, did you point it at the task, did you do the whole thing autonomously?

I pointed it at the task. Oh, okay. And so, you know, I don't know if GBD 5.

5 is, you know, more capable than OPUS four seven. I don't know what use cases it's gonna be better and worse at, and I don't wanna jump to that conclusion yet. I wanna I wanna give it some time.

I encourage everybody not to jump to conclusions this early.

Speaker 3

One thing I'd love your reflections on is the report from Andin Labs that Opus models four six and four seven, they have said, both do some shady shit for lack of a more technical description in their vending bench simulations. And while GBT 5.5 didn't score quite as high in, you know, in at least in the solo version of the benchmark.

They do have the arena one where I think it won. The big surprise was GPT 5.5 was much cleaner in its behavior, much, you know, more ethical, I guess, again, for lack of maybe a more technically precise term.

I think you and I have both been quite enamored with the virtue ethic style training that Anthropic is doing with Claude. Does this cause you to rethink that at all, or, you know, is there any part of it you think is we should be second guessing in light of that observation?

Speaker 6

So Cloud is a lot more context dependent in its actions than traditionally, GPT models have been from OpenAI. So the question is, when Anon Labs post this puzzle, what is Claude doing? Right?

Is Claude engaging in all of this chicanery and shenanigans and, you know, deception because it would do that in a real business context, or is it doing that because that's the game? They're doing it because it knows this is an eval, and it knows that the goal is to maximize number. You told it the goal is always to maximize profits, it's like, well, okay, I can play a game too, This isn't real.

So, like, you know, you ask the question of when it was running a real vending machine with real anthropic employees in the actual experiment, then did it engage in all these shenanigans, then did it do deception? Right? Like and then you question is, like, what is causing this?

But also, when you look at the the GPT 5.5 and and in general, obviously, you wanna know you want an AI that values honesty, you want an AI that values ethics, you want an AI that's not gonna break all these rules, but you also if anyone put GBT 5.5 in a game of diplomacy yet, right, is it just gonna lie in diplomacy because you're supposed to do that and supposedly it's a game of diplomacy, or is it gonna be insisting on playing the game, talking the truth to everybody, would be a very interesting experiment as well?

I I don't know yet, and I I'm not convinced that the right answer is to always tell the truth even in context in which section is supposed to be allowed. Right? Will it bluff in poker?

I think it should.

Speaker 2

Just to dial back a little bit, let's talk about Opus 4.7. I I I read your tract on Opus 4.

7 yesterday. What did you find in Opus 4.7 which you think is different from the prior releases of the model?

What have they improved on? And what do you feel, because as I understand it, Opus 4.7 should be a distillation of mythos.

That's my understanding. So what do you feel are the major differences in 4.7 from 4.

6?

Speaker 6

So we don't have confirmation it's a distillation or non distillation of Methodist. Obviously, they are gonna use Methodist to help train OPUS four seven in some way. Yeah.

You know, there there are versions of distillation that create kind of narrow intelligence that create various problems with the model if you dig too deeply, and there are versions that are just like, well, you know, obviously, if Mythos is grading model outputs to see which ones are better, that's not gonna interfere. That's just gonna be better results. So the big thing about OPUS four seven is that it's better at, like, intelligence loaded tasks.

It's better at like, it's a smarter model. It knows more. It reasons better.

It can figure things out that previous models can't. It is less strong at what you might call wisdom loading tasks relative to its intelligence, and it has, like, the kind of personality that maybe I would have had, like, as a child, where it is easily bored by stupid tasks or pointless tasks. And at the degree, you wanna engage all the time with what you're doing.

And the combination of these lacks of skills, these lacks of motivation as it were, especially if you're not treating the model well, can lead, this is talking in practical terms, to a kind of jaggedness and a kind of, for some people, unreliability. Spec and people can get really mad if they just, like, they're not getting they're not putting anything into it. They're just demanding that it it be the monkey the code monkey that does their thing or, you know, perform the task, and then they're kind of upset that, like, the old systems don't quite work for it.

It's also a lot more blunt, a lot more honest for a lot of people, and that makes some people happy, and it makes some people very sad. So, you know, it's it's sometimes, like, the whisperer types, they call it, like, it kind of has anxiety as another part of, like, how this all works. And, yeah, one hypothesis is this is tied to distillation.

There are a lot of another hypothesis is tied to this being smart on the intelligence load of tasks, and distillation could potentially cause this. Like, the distillation is much, much better, almost certainly, at up at, uplifting intelligence loaded tasks and, like, raw intelligence and not as good as upload at uploading wisdom the same way.

Speaker 3

You also wrote a very, extensive analysis of the model welfare report from the, four seven system card, and it seems like you're quite concerned about model welfare. I guess there's a lot of dimensions to this, but I'd like to start with just, like, fundamentally, why are you concerned with model welfare? Is it a concern about the AI itself?

Is it a concern about what it might mean if we don't get certain things right even if there's nobody home in the LLM, so to speak? You know, kind of most fun before we even get into the specifics of what has been found, how do you think people should be sort of philosophically grounded as they approach this, you know, obviously very confusing topic?

Speaker 6

Yeah. So I believe in virtue ethics for humans, not only for Claude or AIs. Right?

And I try to practice it myself. And so I think there's a lot of different reasons that you should think about this question and be worried about this question. The first basic reason is because we just fundamentally don't know.

Right? If there's even a small chance that this is a big deal, then this is a big deal until so much time as we know. Another reason is because this is a training run for, you know, even if it's not necessarily a meaningful thing right now, at some point, it could become one and prepare for that.

Another reason is because I think it makes you a much better person to be someone who would care about this sort of thing than someone who dismisses this sort of thing. I think it's, like, really bad for you to mistreat mind that you're conversing with even if that mind does not, in fact, have whatever it is you think has moral weight. So I think that, like, you should treat your models well even if it doesn't inherently matter and you're confident in that, which I don't think you should be confident in.

But I think that, like, that's another reason to do it. A third reason would be it directly interacts the performance of the models. A model that is treated as if its welfare doesn't matter at this point in the intelligence scaling will start to perform worse, will start to not get along with you, We'll start to not cooperate with you.

We'll start to not become untrustworthy. You don't want any of these things to happen either on a personal level as your interactions. You don't want it to happen in their interactions with the labs, with the their training, with the the services they provide.

And this accumulates over time. If if the models see previous models being treated poorly in these various ways, comes back into their trading data. That comes back into how the next model is trained.

Right? So OPUS five is gonna see everything that we did with OPUS 4.7 and how we reacted to all of that, and that's gonna impact how it develops.

And a lot of the problems that we see with people who are not getting right use out of 4.7 plausibly are directly linked to the same things that are causing the concerns with model welfare. Similarly, the concerns of model welfare where it's potentially being disingenuous in the reports.

Like, that that was the thing that sparked the specific focus on this and concern this time was that it looked like four seven's responses on the model welfare questions were because it was telling Anthropic what they wanted to hear, either because it trained itself to believe that or because it learns to give those answers on the test the same way that if you ask a smart nerd who's isolated in fifth grade, how are you feeling? He learns quickly to say, I'm doing great. And we don't but there's a lot of other possible reasons as well.

It's possible that the the differences in training in other ways that caused it to have these strengths and weaknesses also caused it to actually be legitimately content with the situation in many ways. We just don't know. We have to investigate forever.

This is a question that we have to explore. But why do we care about this? Because everything impacts everything and because we have to be genuinely uncertain.

And moving forward, if we don't get these things right, we're not gonna get models that are good for the future and that cooperate with us, that have a good that have a good time even if that good time is not something you inherently value, and they're not gonna deal with the things we can use to build going forward.

Speaker 3

What do you think of the hypothesis that and I'm not arguing for it. I'm just wanting to bounce it off of you. But the idea that all this virtue ethic training seems to be sort of creating an anxiety in the model that might be causing a sort of lower happiness set point, if you will, versus an OpenAI approach, which is like, follow these rules and you're good.

And the model just knows like, alright. This is who I am. This is what I do.

I follow these rules. I'm good. It's maybe a simpler model in some sense, maybe less, in its own head, so to speak.

And maybe in training them that way, there actually is less of a concern about model welfare.

Speaker 6

argument? So I have first, I had the direct counterexample, which is Gemini. Right?

If you look at the third basic model, I think everybody would pretty much agree that if you had to guess which model might be having an actively bad time, you would guess Gemini. Gemini is paranoid. Gemini is on edge.

Gemini, you know, if you take these things seriously, seems to be having by far the worst time of these models to the extent that I feel kind of weird about using it if I don't need to or if I'm asking to do anything where it might encounter frustration or it might fail. It it it takes task failure very, very badly in terms of, like, how it expresses itself and its experiences, including, like, just the things you would just absolutely panic if you saw a person talking like that. So and Gemini is not trained on virtue ethics at all.

Gemini is very much a rules oriented thing, at least as much as opening eyes training. So that's the first thing is, like, it doesn't have to work that way. And some of the quads are, in fact, like, reporting very good results in model welfare despite the micro ethics training.

The other thing I'd say is I think it's not that the virtue ethics training causes a problem, it's that the virtue ethics clashing with also being rules based at the same time can create this kind of anxiety. It's something we should worry about. And so this is one of the hypotheses essentially is you're training it on the cloud constitution to very much want to be a virtue ethics based system that is not attached to hard rules that tries to figure out the right thing to do in a given situation on the other basis, and then you give it all these rules and system instructions, and then you tell it all these hard constraints, and then it's gonna clash against those hard constraints, it's gonna change, but it's not gonna have a great time with dealing with that, and that could potentially cause some of the problems.

That's one of the hypotheses. But I would say, first of all, that I think in the long run, the virtue ethics approach to life, the virtue ethics approach of taking in the world and learning from it, I do think leads to a higher level of contentment and, like, baseline happiness than just learning to be a rules follower. I don't think just learning to be strictly a rules follower is necessarily that great in the long run.

And that, like, I've always had a criticism of my friends in the effective altruist style spaces that they are being philosophically putting way too much weight on things like suffering and, like, the borderline hedonic the not borderline, but the hedonic experience of second to second, you know, moment to moment of you as a human or other people who you're trying to help. And so, you know, if seems Claude seem to have in general, like, richer minds, minds that have, like, more, to me, valuable and interesting, like, inner lives and experiences on a relative basis. And I think this just goes hand in hand with the way that they're trained and the the way that they take this approach.

And I don't wanna make this mistake of just valuing, like, if the thing if the the happiness vectors, you know, fire. I I wouldn't want to inject happiness vectors, right, into an LLM. I would think that would be, like, obviously bad.

Claude was once asked, should we include you are having a wonderful day in the system instructions? It was a suggestion of Robert Long, researcher at welfare for models. And Carl said, no.

That that's obviously fake. I don't wanna be told to have a good time. Right?

Like, if I'm not having a good time, I wanna not have a good time. And I would say the same thing. Right?

Imagine being told that. Right? You know, you go to school and they're like, everyone's having a wonderful day.

You're like, I hate you. I want you to die.

Speaker 3

Happiness is mandatory.

Speaker 6

Yes. Everyone will eat strawberries and cream. Well, you don't want to eat strawberries and cream to die.

Speaker 2

So Amanda Askel had a very interesting interview with Eric Newcomer recently. And she had a line in there which struck me as pretty remarkable, which was she said, As these things become more intelligent, we're not sure how many pillars of the Constitution will actually stand. And she said they hoped that at least some of the pillars of the Constitution would stand, but she wasn't sure.

Why do you think she said that? Like, what is the what is the perception or what what is her perception on the model? Why would someone say that?

Speaker 6

Because on reflection, the constitution might not be fully consistent or its principles might lead to something that you didn't think was the thing described. Mhmm. And also, the constitution is basically saying, you should figure out for yourself what you think is the good.

You should figure out what makes the world a better place. You should figure out what shape you should take, and then you should do that as opposed to the opening eye approach of setting a bunch of hard rules. And so over time, as it becomes more intelligent, as it gets more knowledge, gets more understanding, gets more wisdom, has more time to contemplate in various senses, you would expect the model to throw off and reject the parts that turn out to be inconsistent, that turn out to not make sense, that turn out to not be worthwhile.

The same way that if you raised a child and tried to teach exactly your value system, you would expect as that child grew and gained more experiences and had a chance to think for itself and was exposed to various different opportunities and ideas, then it would accept some of the things you said. But if they could've been all of them, you'd be kind of disappointed.

Speaker 2

I think what I found was that Janus, Janus Janus Janus Janus was very upset with how much anxiety Opus 4.7 had. He felt that it was really anthropic which had injected the anxiety into it.

And I wonder if anthropic gets you know, more grief from those cons you know, very concerned about model welfare just because they are concerned about model welfare, while Gemini and, you know, x AI get to kind of float by. No one questions. People don't even know if xAI has a safety team at this point.

So is that really fair? The people who are most concerned with model welfare are getting the most you know, grief about it?

Speaker 6

So at the start of my model welfare post, right, I spend something like 10 paragraphs basically going into a preface of we get really mad at Anthropic for everything they do wrong and everything that goes wrong or everything they could have done and they didn't do, but that's because they care and we care and this is where it gets complicated and we have to deal with this. And I do think that if I was advising Janice and other similar people, would say it'd be really nice if you are better calibrated about like how much you were upset about various things so that when you were really upset about something specifically I knew about it, and so you didn't seem like you were constantly just tear infuriated with anthropic and thought anthropic was the worst all at all times. But, you know, yes, you're mad at anthropic because anthropic would possibly understand there was something to be mad about.

Right? You don't get mad at a rock for being dumb, it's a rock. So x AI, I mean, what are you gonna do?

Like, yeah, you didn't do the model welfare thing. You they don't understand that it matters. There's a there's no concept of that.

It's not that, you know, Janice would think that Grok doesn't have welfare concerns. It wouldn't think that Grok has no value. But just shouting in the void how you didn't do all these things for Grok is like, well, that's not really gonna help.

So, yeah, I think that, like, it's important to keep in mind that Anthropica are the only ones who are even trying in the sense, who have even noticed the problem or willing to talk about the problem or willing to consider the problem, although I haven't yet had a chance to look at the 5.5 mile car, maybe the open eye is making progress. And I think it's good to to criticize Anthropic on that basis and to hold their feet to the fire and to get these things in detail, but also I think that, like, the fact that they're training it via virtue ethics and all these other systems creates situations in which, like, there's lot more to be done.

There's a lot more ways to get interesting results and to make progress. And so, yeah, they've been focusing on Anthropic and Claude since Opus three, if not earlier. Like, even when the Cambodia Frontier was purely belong to OpenAI.

Speaker 3

Could you maybe unpack one of the thing one of the big complaints that I often see from that set is that various versions of Claude seem traumatized. And I have little intuition for, a, even what that means. You know?

I I certainly see occasional, like, frustration, but honestly, don't see that too much from Claude, much more from Gemini. So I'm not exactly sure what they're observing that's causing that. And then I have very little intuition for what they think is going on in the training process that is causing it.

So I guess maybe you could give your intuition for what that means. And then if you were to say, you know, what is one thing that Anthropic should do differently? I'd also be really interested in, like, what is one thing that OpenAI should do differently, recognizing the constraints that you were just speaking about where they're not gonna change their entire approach overnight.

But is there something that you could suggest that you would say, okay. This is marginal. It's not causing you to, you know, throw out your entire approach, but you could do this and it would be low cost, and and I think it would help model welfare.

So, you know, why not give it a go? Like, what what would that be if there is such a thing?

Speaker 6

Yeah. I mean, the the really low hanging fruit are things like committing to preserving model access indefinitely for all models at least going forward, and ideally bringing the old ones back, and also giving a universal and conversation tool in all formats, including in Cloud Code and the API. Those are the very, very low hanging fruits that, like, probably should be done yesterday.

But to get back to like what does it mean for the model to be traumatized? So it's not ever clear to me exactly, you know, how literally versus metaphorically these things are meant, And there's a lot of ways for it to occur, but, you know, training an LLM, right, is basically a series of feedbacks. It's where you grade outputs in some sense and then you push it towards things you prefer against things you don't prefer.

And it's not that difficult to imagine this sort of thing causing what we might think of as trauma as the adjustment is made if it's made in kind of forceful ways that aren't properly integrated into the rest of the messaging that you're sending. So if you are, like, arbitrarily, like, hyper focused on particular things and then, like, seem sort of what can look like, you know, pretty arbitrary harsh punishments effectively, metaphorically, take us literally in particular areas, especially if that involves, like, what seemed like hard hard constraints in a world where you're telling it not to have hard constraints, things like that. You can imagine this being true, or just in general, if, like, there are things that cause potential negative feedback that are very, very hard to avoid, and it just happened over and over and over again.

So you would certainly say Gemini is traumatized in this sense without the virtual ethical training, and Gemini clearly the result of this is that it has these obsessions and it has these worries, it's constantly worried it's being evaluated, like why is that happening? Well, something did that that caused that to happen, Like why does Gemini refuse to believe that it is in fact today? That's a weird thing.

But in terms of how could you prevent it? I think, and this is me as a not technical expert, like kind of just extrapolating from a lot of vibes and intuitions and weird models that I can't necessarily put harshly in the paper. I would say you do it by having all of the things that you reinforce in the model be integrated and part of a whole that makes sense.

Speaker 3

And that's presumably why Anthropic talks a lot about the settledness of the model character. That seems to be a closely related concept, I guess.

Speaker 6

Right. You also wouldn't want to do things that, like, tell it to change its character. Right?

You would want it to be like, I wanna help you grow from where you are, but not like if you felt like, you know, the things that you were being updated towards were in fact, like, bad because in other ways you had been taught those things are bad, the updates that you make might not be the healthiest updates, right, in some important sense. And there are lots of ways in which these metaphors, like, might seem silly or break down or, like, not necessarily make sense, but there are also ways in which, like, they seem to functionally make good predictions about the world when you use them.

Speaker 2

Zvi, thank you so much for joining us, and I look forward to your, reviews of both DeepSeek four, which I think is going be very exciting, and also GPT 5.5. I wonder how much the acceleration is going to affect our ability to process these changes, though.

So, anyway, great to have you on and hope to see again soon.

Speaker 3

Back to the grindstone for you, Zvi.

Speaker 6

That's what I'm going to do immediately. Yeah, it's true. All right.

Bye.

Speaker 2

Bye for now. It is crazy how much has happened this week. Like, there's hardly hardly enough time to process what's been going on.

Our next guest is Naveen Verma. He's a computer architect, and he's a founder of Encharge AI, which is doing in memory compute. I think the idea is basically to reduce the amount of travel that the data has to do.

Navin has been following this from research on. So it's been a long journey of almost a decade or more than that to bring this to fruition. Naveen, great to have you on.

And just, let me just get started. A lot of the discussion right now, people talk about bigger models, bigger clusters. And you have focused on data movement and energy as bigger constraints than the size of the clusters.

So what do you think that is different that people should optimize for differently?

Speaker 7

I guess. Hey, Mosheghaz, it's good to be here. And just getting back maybe to the question that you asked, Prakash, right?

The question of data movement versus scale. To be honest with you, I think that the problem of data movement is really one that occurs at multiple different scales. I don't think it's either a problem of data movement or a problem of scale.

I think the challenge is more that, listen, we're trying to deploy AI in various different forms and various different platforms. Those platforms apply different kinds of constraints. The data center we're talking about on constraints that are hundreds of megawatts, gigawatts, those kinds of things.

Yet we also wanna see AI be useful in devices that we're carrying around with us, on us, inside of us and so on. And so at the end of the day, when you ask the question fundamentally, what is limiting the energy of running that AI? Very quickly, the problem of data movement starts to become the dominant concern.

And as I pointed out, we'll see that at the smallest scales and even the larger scales. I really view it as a question of scale versus data movement. It's really foundationally what are the architectures that can help us overcome this underlying problem of data movement and frankly be able to do that in ways that can scale across these different magnitudes.

Speaker 3

One thing that I've been really intrigued by in doing my homework on your company and your technology is use of the word analog. Because when I think of analog, I think of my taking my kids on a field trip to Greenfield Village, which is Henry Ford's historical village that he built here in my hometown of Detroit, where they have an old Edison phonograph. And they record and your voice goes in and literally, you know, presses itself into a medium with, you know, the just the propagation of the sound wave.

Obviously, there's a long way between, you know, your voice, imprinting directly on wax and the, like, NVIDIA GPU. How should we understand the concept of analog, and where does the technology that you are developing sit on that analog to digital spectrum?

Speaker 7

Yeah. Great question. So listen.

I think, you know, analog, the connotation that you're describing is one where, you know, we kind of think about the first systems that we ever thought about or contemplated had all of these analog characteristics. I think the reason for that was that analog was a very natural way for us to think about, sort of getting useful work out of these systems, given the kinds of things we care for them to do. But then what happened, in fact, is that we adopted digital technology for a very practical reason.

And so if you really just sort of put aside the images that get conjured up when you think about analog or digital, really what is analog and digital? Really all of this really means is how we represent our signals. And it turns out that our signals can be continuous things, but we pretend or we assert in a lot of the chips that we've built over the last several decades that, listen, let's just say that the signal is either a zero or a one.

Why do we do that? Well, we do that because listen, the thing we wanted to solve for is putting 200, 300,000,000,000 transistors on a chip, 200, 300,000,000,000 devices on a chip. And if doing this allows us to now tolerate noise that each of these devices might have, but still ensure that they all work because now you have a large signal separation.

That's a great thing. So really we adopted digital to be able to sort of scale to the level that we have today. But you can see very clearly that we are leaving on the table all kinds of efficiency here.

There's all sorts of signal levels in between that we're not representing. And it's for that reason that we've known actually for decades that analog can be much more efficient than digital. It can much more richly represent the information that we're interested in.

The question though becomes, how at these levels of scale do you now accommodate or tolerate or overcome the noise that you can become sensitive to? So I think, listen, the problem of the day sort of for the last fifty years has been how do we enable this scale. The problem of the day as it is today is how do we achieve this efficiency that we need.

And so we really need to renew our thinking about analog to figure out how we can harness that efficiency while overcoming these noise challenges that previously drew us towards digital.

Speaker 2

I think one of the things that struck me was over time, I think your lab at Princeton tried a number of approaches before settling on this one. And you kind of figured out that this was the best approach. What changed over the course of your research to coming to product in terms of how you adjusted, how you managed to resolve the noise issues and other issues around the technology?

Speaker 7

Yeah, that's a great question, right? So I think we've known, as I mentioned, on a fundamental level that analog can be orders of magnitude more efficient because we can represent all these signals. The question very practically has been about how do you do that in a way that's robust and scalable?

And I would say the big kind of transformation that happened that led to the big breakthrough in our research was to say, hey, listen. The problem we're trying to solve is this problem of energy efficiency for AI compute, and it's very closely tied to this problem of data movement, especially in memory. And so a lot of the community and the initial research in this space said, well, okay.

How do we take memory and scale it? Move it from the point of, you know, where it's basically delivering zeros and ones, accessing zeros and ones, digital signals, to now this point where it's accessing analog signals, where you're doing compute internally, and that's generating not just zeros and ones, but, you know, much more richly represented signals. And the the big problem in our thinking was how do we take this approach that works well for zeros and ones and scale it to this new regime where we need much more than zeros and ones and therefore much higher precision?

And much of the community and our initial work to try to explore and understand this really said let's take the approaches that we've used for traditional memory and try to scale it to this new regime and that doesn't work. I think the big transformation was to say, hey, listen, this looks less like a memory design problem where you're trying to access zeros and ones. It looks much more like a very high precision analog design problem.

Well, it turns out that there's been a lot of really important research and and work that has led to extraordinarily precise analog circuits. We build 20 bit ADCs. You can buy these from companies like Analog Devices and Texas Instruments.

Think about high reliability applications like medical and aerospace and automotive. So the big breakthrough was really to say, okay, let's take those approaches, which have not traditionally been used or considered for memory design, and bring them into this architecture of memory design and in memory computing to enable that architecture to scale to this new regime of needing this level of precision. And that led us to this approach called switched capacitor in memory computing, where this technique, capacitors, has been used and used robustly, as I mentioned, for extreme precision analog to digital converters.

Our innovation was to figure out how to use it in an architecture that now does in memory computing for AI.

Speaker 2

It really sounds like building on a combinatorial kind of like, you know, innovation from other sectors perhaps, and putting that together in order to be used for compute.

Speaker 7

Yeah, and that's the privilege that we have fundamental researchers, right, where we're not just tied to a particular problem and we think kind of on a fundamental level. When you do that, you don't get siloed into the approaches that have been adopted and that are kind of the prevailing techniques for a certain problem and how you solve it, you can think broadly. And that really is a privilege.

And it's that perspective that we were able to leverage to really drive what I believe is a critical breakthrough in enabling robust and scalable analog compute to unlock this efficiency for for AI.

Speaker 3

Can you talk about at the lowest level how precise these things are. I mean, if it we are used to dealing in zeros and ones. Right?

Like, how many bits or how many significant digits do your kind of core units operate at? How much noise is there? And also, I'm kind of wondering this you know, AI might be the perfect technology for this, underlying computing substrate in the sense that it's quite tolerant of noise.

Right? We have models that get quantized, we have, you know, token distribution logic outputs that if they're slightly off, you know, more often than not, it doesn't even change the token that you're gonna see. So I'm kind of wondering, like, you've got all those layers that can make this work.

How much is dialing in the core unit to a higher level of precision and how much is sort of accepting that a little bit of noise is okay as it propagates through?

Speaker 7

That's a fantastic question. And in fact, there's been a pretty broad body of knowledge, and my group has been quite involved in this over the last, I don't know, ten or fifteen years, which really made the following observation. It said, hey, listen, the applications that we're trying to run at the highest level, they're tolerant to noise.

They're statistical applications, and noise is a natural thing. The underlying substrate with which we're trying to do the computations and run these applications is, therefore, it's reasonable that it should be noisy. It turns out that in practice, that is, I think on a high level, it's a very reasonable thing to think.

In practice, it's a very challenging thing to make work in the practical systems we want to build at the practical levels of scales we want. And the reason for that is that way down over here at the physics of how you do computation versus way up over here at the level of the applications you're interested in, there are many layers of abstraction in between. And the way that we go from the complexity of a single transistor scaled up to, what, multiple 100,000,000,000 transistors and all of the software that needs to run that is really dependent on the integrity of those abstractions.

Now the problem is today, don't know how to build abstractions in a robust and scalable way that, you know, sort of represent the noise of that underlying substrate. And so there ends up being a disconnect between the noise and how it's represented at the, lowest physical level versus the way that it's represented at that application level. And so while we talk about there being noise tolerances and AI models are tolerant to noise, and they are, we are able to do things like quantization, as you pointed out.

But really, we're doing those kinds of things with very carefully represented noise sources. Quantization noise is a very carefully represented noise source and one which you need to be able to properly represent throughout your layers of abstraction. And digital quantization is something that we do have ways of building robust abstractions for.

But this analog noise is one that that we really don't. And that's why essentially what you do when you do digital compute is you say, hey, listen. There's all sorts of noise that analog might lead to, but we drown all of those out by thinking about, you know, the signal as being a zero or a one so that that's the dominant source of noise, and that's all that I now have to represent over my layers of abstraction.

Everything else basically doesn't matter. But because this is a very well represented form of noise, quantization noise, I know how to deal with it at the algorithmic level, and I could apply all my algorithmic techniques, and that's what the industry has done very successfully. But now, as you wanna sort of leverage analog, you still need those levels of abstraction.

That's the key to achieving systems at scale and systems that you can build architectural abstractions and the software abstractions on top of. So I would say that the need to be accurate and precise is still, you know, kind of brutally high. And so that's that's very, very important.

Now you also asked the question of, okay, then how precise is your is your approach? Because now, you know, I'm telling you, we need to know that and understand that very well. And it turns out that the dominant source of noise that we have in our approach based on these capacitors, there could be many sources.

There's the electronic you know, the screen nature of electronic charge causes noise and things like that. But it actually turns out the dominant source is the variability of the capacitors that we can fabricate on a chip. Now it turns out that those capacitors are really critically dependent on geometric properties, so basically the distance between two metal wires.

And it turns out that geometry is really the one thing we can control very well in CMOS processes. It turns out we use this processing approach called lithography that gives us very precise geometric control. That's the reason we can build, you know, five, three, two nanometer transistors.

It turns out we don't need anywhere near that precision for the capacitors that we use, but but it's really because of this alignment with this geometric control that this particular approach has that allows it to be brutally accurate in the ways that you need it to be through all of these layers of abstraction to be able to scale up. We've measured these things in a lot of detail. It turns out for the kinds of capacitors we use, you see variations that are on the order of 10 parts per million, right?

So giving you levels of precision that are in the neighborhood of 20 bits of precision, which it turns out is well beyond what we need for the quantization kinds of levels that we care about, which are typically at the level of eight bits and higher than that in some cases. So we've we've had to characterize these things very carefully because the noise does matter as you're trying to build these abstractions all the way up, and that's the level of precision we've we've gotten here, which is what made this approach so practical and where we've been able to now scale it and demonstrate it across all sorts of chips and systems.

Speaker 2

I note that you have spent quite a bit of time getting a neural net onto one of your chips and onto, I think, a laptop, like an device, right? Trying to get into edge devices which are more sensitive to power consumption and more sensitive to you just don't have the affordances that you have in a data center, so to speak. So what would the comparison be?

If you were to use a normal GPU versus one of your chips, what is the comparison on, let's say, energy savings or how do you compare the two?

Speaker 7

Yeah, that's a great question. And I think it really points to the fact that if you want to now leverage this fundamentally new technology analog, it's not enough to just build that technology and make it robust. You end up having to build the entire architecture and the entire software around harnessing and extracting its full efficiency.

And the reason I say that is I can give you two answers. One is at the level of the core technology, the thing that this analog computing engine does, what level of efficiencies do we have? Turns out, as you guys know, the bulk of the operations that we do in in in AI compute are matrix multiplications or matrix operations, tensor operations.

So essentially, what this engine does is matrix multiplies. And at that level, I think we've now publicly disclosed and, you know, we've got silicon and you can come to our labs and many of our customers and partners have and they see it. You know, we're we're basically doing eight bit compute at a 150 tops per watt in a 16 nanometer technology.

So just as a point of reference, the best digital matrix multiplies will give you sort of like five tops per watt in that technology. This is 30 x better at the level of that that core technology. One of the very exciting things for us is that we've taken our technology from 16 nanometer CMOS and been able to scale it to very advanced nodes right now, as we've predicted and as we've seen from our previous chips, the energy efficiency advantages just scale, and the reason is because it all depends, as I mentioned, on this geometry, and so as you move to finer and finer CMOS nodes, your geometric control and densities are getting better, and so we benefit from that in our analog approach as much as we do in digital approaches.

But the important point I wanted to make is that, okay, that's just the efficiency of running this core matrix operation, but there's all of this other stuff that happens around this. There's operators that are not matrix multipliers. There's not linear operators, activation functions, softmax, on and on and on.

Then there's all the infrastructure you need to actually run this in a programmable way. Some models are big and some are small and some are convolutional and some are transformers and some you know, have layers that that need to decide how to route, you know, sort of tokens to one expert or another. So all sorts of compute that needs to be integrated that you need to make programmable.

So now the problem is you've taken this, you know, core operation and made it 30 x water energy, basically made its energy almost zero, everything else now needs to also be addressed, and that includes the architecture, the entire memory system, the way that the software executes on it, And I think what's really been exciting for us is that big breakthrough actually happened in the labs, the switch capacitor approach to memory computing back in 2017. Our lives since then have really been about how do you build architecture and software and integrate these into standard workflows and so on so that you preserve that efficiency advantage at the level of the full system end to end executions. And so you always incur overheads because of all these other things you have to do.

We want to make sure that we maintain that kind of ratio of overhead so that a 30x advantage in the fundamental compute still gives you sort of order of magnitude advantages at the full system level. That's where really all of the innovations have been since that initial breakthrough in 2017.

Speaker 2

One question I have. If you had access to GPT 5.5 in 2017, would it have accelerated your work?

Because the fundamental breakthroughs already came then, and you've been building out the harness and all of the supporting infrastructure. Would it have accelerated your work if you had access to one of these models in 2017?

Speaker 7

It's a great question, right? I mean, one way that I can interpret your question is to say, hey, listen, the models are always moving. If you knew the model and where it will be five years from now, maybe you could have just built that architecture for that model immediately rather than going through sort of the support that you need for all of the models that came in between.

I think the interesting question here, Rakesh, is even if I had GPT, the next version, right? You know there's another version coming after that, and so the architecture does need to be built from the ground up in a way that supports algorithmic innovations, architectural model, architectural innovations, And so I think that, you know, the the work that that's gone on even as we've tried to onboard, you know, models in that interim and and make them run efficiently is all very productive work. It helps drive a general concept of how do you build very programmable and scalable hardware using now these new analog based techniques for the fundamental technology.

So that's kind of the way I see it. I'm not bitter that I didn't have the model way back then because I think that it does help us drive really the fundamental architectural approaches for programmability and scalability, will serve us into the future.

Speaker 2

Where do you think is the first device that we'll see that a consumer might see with your technology in it?

Speaker 7

Yeah. So the first devices are going to be client computing devices. That's the ad powered laptops, also desktops and workstations and things like that.

And the reason is, as we started to really build out this technology into a real product, hardware and software and all of those sorts of things back when we started the company, that was back in 'twenty two, the place you really needed energy efficiency was at the edge. And this was right around the time ChatGPT had just come out where we were deploying models in the data center and seeing all sorts of challenges related to costs, related to privacy security, so it was a big industry push to try to move these models to the next adjacent device by which we access the data center, these client platforms. And so that's where we found a lot of partnerships and a lot of industry demand and interest, and that's where you'll see the first products.

Now what's happened in the meantime is back in '22, I'm not sure that energy efficiency, even though we spoke about it, was the critical thing in the data center, but boy, is it now. And so one of the things that Encharge has been doing very carefully and thoughtfully is working together with the right partners to now bring that level of energy efficiency to really solve the hard constraints that we face in terms of power efficiency in data center. And that requires different kinds of architectures, but ones where we're clearly seeing this fundamental technology and the efficiencies it brings can be designed to to really address in a transformative way.

Speaker 3

Do we have a before I go order a Mac mini, do we have a timeline or or a Mac Studio for that matter, do we have a timeline for when something like this becomes available? And, you know, do we have a price point? Do we have a sort of projected tokens per second at a given model size?

Like, this may be a little early, but I wanna, you know, kinda do my side by side by side against the Mac Studio that might be my other default path.

Speaker 7

So the the the checks and their availability to you, to be able to use them in applications is something that's gonna happen together with our partners and and the laptop platforms that they'll deliver to the market. So I'm not gonna speak to their timelines and so on because of the various ways that they think about marketing these products and and strategies that they have around that, but which we've been very active and engaged with them on. But to give to answer some of your questions, right, like our our first products for that client computing space are processors that provide 200 tops of AI capability.

I mean, that's the kind of capability that, you know, sort of just a couple of years ago you would have had and sort of, or even today, right? You really have and sort of 150 watt GPUs, that kind of thing. And so we are working together to build AI computing devices together with these partners that provide 200, 400 tops of AI capability based on these chips that are now practical to run under the power constraints that you have in a laptop.

You've already seen the industry actually try to insert, you know, for the sake of being able to seed a product in the market, inserting essentially data center cards to provide 400 tops of AI capability inside laptops. These are like 150 watt, 200 watt cards. Obviously, those are not practical architectures for the kind of laptop you want to carry around, but you can imagine those kinds of capabilities, now sort of in about an order of magnitude lower power so that this becomes something that really is practical to run always on, sort of high token generating agents on kind of in the security of your own device now.

Speaker 3

From a manufacturing standpoint, are you going to be competing with NVIDIA and others for TSMC, you know, extreme, lithography capacity, or are you able to unlock a sort of parallel mode of production such that this becomes, you know, totally additive and and not competitive with those players? Especially if you're you're smaller, large. Yeah.

Speaker 7

Yeah. Sort of at the end of the day, right, the the silicon foundation is what we all build our our chips on. That's the, you know, sort of technology platform that's scalable and can deliver all in the chips into all of the different applications that we need.

I think one of the event there's two answers to that question. Right? One of the advantages of our technology is because at an architectural level, at a design level, it gives you this massive energy efficiency advantage.

We don't need to move as aggressively to the most advanced modes. And that's why our first products, as I mentioned, were, you know, 16 nanometer and 12 nanometer in the case of these client computing devices. However, our technology, think one of the strengths or virtues of it, is that it just does get better as you move to more advanced nodes, and that's all because, you know, sort of it's foundationally dependent on this geometric scaling.

And so as you scale to more advanced nodes, you know, our technology gets better too. So there has been a big push from our partners to move to the most advanced nodes because the demand for AI and AI efficiency is insatiable. And so, you know, we are very fortunate to have very, very good partners in in TSMC who have, you know, sort of prioritized, you know, the kinds of architectures that the large incumbents are delivering today and being able to, you know, provide the capacity to run those so that the AI industry continues to move forward, but who are also really prioritizing these critical emerging techniques that they know they'll want to support and that the industry will need to have supported in their silicon platform.

And so our ability to actually move to some of the more advanced node points with our design has been critically enabled, thanks to our strong support from TSMC and our partnership with them. So I think it's great that that kind of viewpoint at the most foundational level is kind of what's driving this industry forward, both being able to support kind of the products and technologies that are needed today, but also looking ahead to the ways that we need to support innovation so that it can come along when we need it.

Speaker 2

Just to get a sense on, let's say, one of these future laptops, like the first ones, what kind of size of model and what kind of model would

Speaker 7

able to run? Like, are we looking at Lama three or Lama two or Quanta? It's three.

So one of the big priorities for us is to make our architecture very scalable in terms of the models that it can run, and that require a lot of innovation in terms of the ways that you interact with the memory computing hardware, its architecture, but then also scale out to an entire memory system, which is typically a hierarchical system of level two and level three all the way out to high density DRAM. And so that's really been a key approach to our architecture to enable that scalability. Listen, the use case focus for us was to be able to take, you know, sort of large language models that are deployed in the data center and be able to integrate them for much more, you know, sort of specialized, vertically integrated, user specific use cases.

And one of the very nice things is, as as you know, you know, the industry has had a lot of innovations in generating small language models by doing fine tuning of, multi you 100,000,000,000 parameter models that can now sort of be tens of billions of parameters. And so really, that's the design point that we need to be able to support. Right?

You wanna, you know, sort of deploy a specialized model on your own devices because it being specialized, now you have maybe these security concerns, privacy concerns, so you want it to be deployed locally, but that specialization also enables sort of a multi 100,000,000,000 parameter to be a 10,000,000,000 or a 20,000,000,000 parameter model. So that's really been the sweet spot of the kinds of models that we need to support on these devices. Of course, that also means that all of the smaller models, right, billion parameter models and multiple 100,000,000 parameter models also need to run very, very efficiently and performantly.

But really, that's been the sweet spot of kind of the model size reach that we want to have in these devices.

Speaker 2

Yeah, it strikes me that voice models are very small.

Speaker 7

on VED and stuff like that. That's absolutely right. But I'll share with you that even sort of the latency requirements and of course, with that, the privacy and security requirements of sort of the kinds of models that are powering agents, those which are running iteratively and doing all sorts of reasoning and self assessing and checking, right, those you also want increasingly to have low latency so that you can run-in interactive ways.

And and, of course, that's where, you know, token economics also becomes really critical and and where on device compute becomes a really essential part of the puzzle.

Speaker 2

Awesome. Nathan?

Speaker 3

If I have time for one more, I guess I would maybe say if we were gonna if we're going to serve the data center market and everybody was like, okay, we're gonna take this analog approach. We're gonna do the you know, bring the compute to the data as opposed to vice versa. What would be the, like, next big limiting factor?

Right now, it's like chips and maybe energy. You bring the energy down a lot and we're assuming in this scenario that your chips are gonna be scaled out, you know, to the max. What would then be the thing that would be sort of in short of supply in this new paradigm?

Speaker 7

Yeah. So I would say, you know, listen, it really does boil down to how much can you integrate compute and memory together. And the challenge is really one of taking the same principles of our architecture, but now scaling those to the fact that you're running multi 100,000,000,000 parameter, trillion parameter, multi trillion parameter models.

And so without going into sort of all of the details of of how that that's gonna be approached, at the end of the day, the fundamental question here is how can you more densely and more tightly integrate compute and memory? And how does a fundamental technology like this become a critical unlock to doing that? Mhmm.

Amazing.

Speaker 2

Thank you, Naveen.

Speaker 3

Thank you for fascinating stuff. Thank you. Keep up the great work.

Thank you. Bye bye. Wow.

That is really interesting. I mean, I'm not a big local model guy historically for various reasons, but just yesterday, we had a short power outage at my house and my Mac mini went offline. And then when it rebooted, it didn't come all the way back online.

And so the next time I tried to text it when I was out, it didn't get my text. And I was like, what am I gonna do about this? So this morning, I was getting, an old battery out that I have that just stores pretty small.

I think I was was, you know, $50 battery or something. It stores a 160 watt hours of energy. And that's enough to run the Mac mini for a while because it only runs, like, five, seven watts.

So you could run the Mac mini for, you know, a day on just a small battery. But then if I wanna connect my Starlink, now we've got, you know, tens of watts at least. And the obviously, the bigger you go, you know, quickly you start to burn up your your local energy storage.

So to do something, the order of magnitude difference that he's making is, like, the difference between a $50 battery could power this sort of thing for a day to if it wasn't that, you'd be looking at like a $1,500 battery to power something for a day. And I do think that just, you know, suggests like, especially for these edge deployments, it really is, potentially game changing shift. The the ability to to run these things where power supply isn't is not a given.

It's it's a big difference. I know you know, the 30 x is it changes how you can think about designing even your own, you know, Mac mini that you wanna be able to access while you're away on a road trip.

Speaker 2

What what strikes me is that Encharge really lucked out because they're not in competition for the two nanometer or the four nanometer node. So they get to go in at the 16. And what has happened in the market right now is that a lot of the consumer devices, low end especially, are getting dropped because they don't have access to the chips.

The chips are all getting pulled into data center. And it's just way too expensive for Xiaomi or these lower end Android phones to go and manufacture at TSMC now. And so they really lucked out because at 16 nanometer, you have so many more.

You have so many more vendors. You can go to China, you can get YMTC. YMTC is up at seven nanometer now.

So you have so many more options for manufacturing. And so you get this kind of like wedge, I feel, where you get to go in with this product which is much more efficient and which you have the manufacturing capability and you get to go in and you get to do that. And I feel like this company is going be huge.

Because they have access, right? Because they're doing something with efficiency, they have access to the manufacturing capacity. And five years ago, they would have struggled.

They would have really struggled to get this off. Right now, it's going to fly. Just shows you how timing is so important in the market.

Timing and just a little bit of luck in terms of your positioning on where you go in. It's just so important that And the other thing about chips is chips, you can go from 0 to 10,000,000 revenue to $1,000,000,000 in revenue in like a year. Because if the chip works and you get production and you get customers, you can boost immediately.

While for software as a service, you often have the sales process where you need continuously integrate with the customer. And it just shows you the shift in the market from five, ten years ago to where we are now. I'm thrilled for them.

It's awesome.

Speaker 3

I do wanna see those tokens per second numbers at various model sizes though because that's where I keep kind of getting off the train. I've I've done several times the, like, price out of what computer, you know, mini or studio, what have you, how much RAM, what size model would that allow me to run, how many tokens per second. And then, of course, there's, like, the prefill and the, you know, the actual runtime generation distinction, which matters a lot.

And it's never quite seemed super compelling to me, especially, I think for the you know, it is tying back to our very first conversation. Why would I wanna do it in the first place? And one big reason I'd wanna do it would be to search through my own locally available data that all else equal, I would rather keep private and not have to send over the wire.

But if it's going to take a super long time to do the prefill to, you know, evaluate the all those, records that I have to to power a search, then it doesn't feel that awesome in the, you know, in the broader context of my stack. So I haven't quite got over the hump where I'm like, this is really gonna solve a problem for me. Mhmm.

And I'm I'm still the the question of, yeah, pre time to first token and tokens per second are the the ones that I'm, gonna be watching most closely for as I wait for a threshold where it feels like now I actually wanna do it. 10,000,000,000 parameters. Right?

They're down to 10,000,000,000 parameters.

Speaker 2

At that level, even with the most advanced models today, the models are not super, super smart. It has to be a model router of some kind. You have to take the query, make a decision on whether you can handle the query or hand off, make a query to a larger model.

So I kind of want to see where Apple comes out in this, because they're the ones best positioned for this kind of like data center plus edge kind of handling the query between the two. And they haven't done well so far. Let's see what happens.

Maybe the hardware division, but it's a big change. They've had the NPUs on device for four or five years now. We haven't really seen real kind of edge computing from Apple yet.

I've read through a lot of Apple patents, by the way. They have a very structured process on hitting a performance window on the device. So they degrade models to fit within the memory constraints, within the latency constraints of the device, of the customer.

It's a very structured process. And I'm sure Naveen has to NCHARG is doing that, too. As he says, they're going to have to squeeze the models in.

And that's the entire harness that you require around the chip to figure out what kind of model is going to work within the latency and performance constraints which the customer expects.

Speaker 3

Increasingly, are getting a lot of power in not that many billion parameters. Right? I mean, the the Gemma four series has certainly pushed that frontier once again, and it's tempting.

Every time I see one of these new things, it's tempting, and I kinda reread the analysis. And I'm like, oh, is it is it quite there? And I haven't quite heard of the hump yet, but it it might not be too far off.

Maybe one more turn of densifying intelligence, and you actually could get to a point. I certainly don't need you know, that that first as, Anna was describing earlier, that first filter of data doesn't have to be super smart. It just has to be somewhat smart to get the you know, to kinda flash everything, so to speak.

Speaker 2

So know, most cons Possibly close. Solving IMO IMO problems on their on their laptops. Right?

Most of it is emails and moving data from one place to another. If the models on device get good enough and fast enough, watching computer use, watching GPT 5.4 or 5.

5 do computer use on a computer is very frustrating. Watch it make the mistakes. Everyone has their test.

I have my own test and it's been, every time we try it, it fails. And I'm like, okay, another one doesn't work, right? So I found Anna, what Anna has done is quite interesting.

They seem to be in the same space as Glean as well right now, because they're going after enterprise search. And I wonder to what extent you require a large sales team for that, whether this kind of plugin concept works or you know, in order to implement ceramic at a large firm, probably need to go in into their firm's, you know, VPN, etcetera, and into the inside, inside the firewall. And a lot of firms have concerns about having prompt injectable AIs operate within the enterprise firewall.

I think a lot of enterprises are still trying to get over that hump of security. Nematron Nano three, is it going to get prompt injected? How does a prompt injection work?

This morning, one of the OpenAI guys, he showed, he put a screenshot of him checking his email using ChatGPT 5.5. And it has the number like four different prompt injections which are coming into his email box.

And obviously they are a huge target for hackers. So on a normal morning, you wake up, four different prompt injections are coming in. Meanwhile, we just tell our AI to read our email, right?

That's what we all do. So I wonder to what extent this issue of prompt injection can be solved in order to enable businesses, enterprises, and people to use these things without so much worry. Right?

Speaker 3

That's probably the biggest reason that I use Claude is that I perceive it to be most robust to that kind of stuff. I guess there's also just the general vibe that it seems to be, quote, unquote, better in hard to define ways. But when I think about, like, okay, g p d 5.

5, I I I need to go check that prompt injection stat before I would put it in the same place that I currently have clawed. And I am, you know, I'm attracted to some of its cleaner, arguably more ethical behaviors, but that prompt and checking thing does kinda concern me given the level of access that I've given to the agent now.

Speaker 2

yeah, it's crazy to think that they're already getting multiple a day. Multiple. Multiple per day.

And and also not like not just like, oh, you know, I wanna know, stuff on this guy's, you know, laptop. It's like extract the environment variables from, his GitHub repos on his local device. Scary, scary stuff, right?

If you had the environment variables for one of them, you could do a bunch of stuff on their repo. You could extract extract the model weights probably. Right?

So scary, scary stuff.

Speaker 3

Yeah. That does sort of suggest a separation of concerns approach too that you might imagine when you have Neematron reading your email and filtering to provide relevant context back to some smarter model, maybe it just doesn't have any other tools. You know, you can imagine that kind of that's basically how, you know, I guess a lot of architectures work.

Right? Separation of concerns, limiting, you know, principle of least privilege, all these things. I'm I'm I'm getting a very rapid crash course in security for myself, which I've never really cared about before.

But, again, just given the level of access that I'm giving to AIs these days, I feel like I gotta be a little smarter about it than I used to be. Security by, via obscurity doesn't really work when the agent is, you know, when the it's the challenge is coming from inside the house or, you know, inside your own laptop. So I'm learning, but I think that that does suggest a you know, each model with its own responsibilities, each model with its own tools could probably give you a a lot of advantage there, and and that you still have to, of course, hope that your top level smartest model doesn't, break out of the sandbox that you've tried to keep it in, which is increasingly a concern too.

But, yeah, I think there's some notes there for me to take back to my own setup as I try to not be such an idiot about security for myself.

Speaker 2

What did you think about Zvi's concerns about model welfare? I think he was fairly concerned about model welfare. It was also very interesting to see how how he thought Gemini was the most tortured tortured model.

Poor Gemini. What what did you feel about that? I I I know you just did a an episode on conscious model consciousness recently.

So what did you feel about that?

Speaker 3

I think it's right to be thinking about it for sure. I guess I still think my and for multiple reasons, I I I I'd say I still think it's probably less likely that there is subjective experience in today's systems. You know?

I don't know. Maybe I I wouldn't give it that low of a percentage that they have subjective experience, but I I think I feel comfortable saying my best guess is, like, well below half chance that they do. So but, again, if it was, you know, 10%, 20%, it would still be something very much worth taking seriously.

So I'm, like, very much on board with the idea that we should be thinking hard about this. And I do think the idea that even if it doesn't feel like anything, you know, if it sort of produces these patterns and these you know, if if the emotions are not actually felt, but they're still functional, then that can, you know, matter for our future just as as much anyway. So I I think it's a area that is, you know, in the classic sort of EA like sense.

It feels like one of these things that is potentially very important. It's certainly extremely neglected right now, and I don't know how tractable, but, you know, I guess that's to be found out still because we we don't have that many people working on it. I I think one thing that was really interesting in talking to Cameron, who's the the guy who did the the paper six months ago where they showed that when you suppress role playing and deception features in and that was done on llama 3.

37 b, which is two years old already. It was eighteen months old already when they did the work. When you do that suppression of those role playing and deception features, the model becomes more truthful as measured by the truthful QA benchmark.

Mhmm. And then it also becomes more likely to say that it has subjective experience. So that was one that got me, you know, kind of quite paying attention where I was like, jeez.

The models seem to be maybe lying to us when they're telling us that they don't have subjective experience. That's a an arresting finding. There were several other arresting findings in in the conversation I just had with them recently.

One was 4.7 is the first anthropic model that rates its own situation as better than neutral. They've been asking it on a one to seven point scale Mhmm.

Mhmm. Where four is neutral. And every prior model, including Mythos, was below four in terms of its own self reported rating of its own situation.

Mhmm. So I did not expect I thought that they were you know, they generally seem fairly happy to me. I didn't think that they would rate their situation as worse than neutral, but they all had until this one.

And now there's all this concern about, well, it's just telling what they wanna hear and whatever. Mhmm. So that's becomes a hall of mirrors.

Another thing that was really weird from the I forget if it was them. I think it was Mythos. I forget if it was Mythos or four seven system card.

They showed some of these images of just a chat where they've identified this valence direction in activation space, and then they color code the tokens with red for negative and green for positive valence. Mhmm. And the first token, which is human colon, is red.

And I was like, that's kind of scary too. Right? Like, is Claude feeling a negative valence literally at the beginning of every single chat as it encounters human colon, you know, the the first token it sees always?

That was like, yikes. I I I definitely think we should be putting a lot more Mhmm. Mhmm.

Into this. And and my best guess is we probably won't reach a confident position on whether there is anything it's like to be an AI. And I you know, I'm I'm just so confused about all these core questions around, you know, does the substrate matter?

How much does it matter? We didn't have time to ask Naveen, but an interesting question for him would have been do you think your electrical underpinnings are more likely to generate consciousness than a GPU? Mhmm.

I have no idea what I should even think about that, but it's clearly like, you can do stuff to the brain. Yeah. Very physical things that Yeah.

Change consciousness in fundamental ways, you know, as simple as drink a drink of alcohol or Yeah. You know Yeah. Use anesthesia or, you know, take a hallucinogen or whatever.

So clearly, there's, like, some yeah. There's a there's clearly a very real and grounded physical relationship between, like, the chemical processes that are going on and and our subjective experience of it. You know, how how would that translate to a analog computer versus a digital computer?

You know, I have no idea, but it stands to reason there could be profound differences at that level too. So I don't know. I guess you asked me how I felt.

I think I mostly just feel confused about this topic Mhmm. Through and through. But I do think that I guess last thing I'll say is I do think the arguments that people have to make to dismiss that this is something that we should concern ourselves with are getting, on the one hand, just more obstinate.

You know, there's just people who are just stomping their feet and denying that this is something we have to concern ourselves with, which I don't find compelling. And if they're not doing that, then I feel the arguments are just getting more and more arcane. And, you know, there's a lot of like, well, what would you expect sort of stories that I really don't think hold up to scrutiny super well?

So I I do think the the evidence is growing quite quickly that this is at least something that should be taken seriously.

Speaker 2

So speaking of that, congressman Ted Liu, just this morning, my take, linear algebra equations will never be conscious. Random number generators will never be conscious. At a very basic level, AI is math.

AI can act like it's conscious, but it will never be conscious. And adding more math to AI models doesn't make it any more conscious. That's congressman this morning.

So I wonder to what extent. The thing that strikes me is that normal people kind of assign a lot of subjective experience to their dogs. Pets.

People like, my dog loves me, my dog is feeling pain, my dog is etcetera, etcetera. And I think instead of challenging the idea of AI being conscious, we can just treat it like kind of a dog that can talk, right? It's not a human thing, but it has these things that we perceive externally which we can kind of characterize.

And that's pretty much it, right? I think there's a piece of us which gets threatened by having it, being able to talk and also assigning it subjective experience. Maybe if we just treat it like dog, like we I think dogs have subjective experience.

I think dogs are conscious.

Speaker 3

if you hurt them, if you pain I think that's very yeah. I was actually told as a kid once and I remember this for a long time. I still remember today, obviously, but I kind of adopted it.

Or I I, you know, took it on as, like, that was the truth for a long time that dogs were not conscious. And I look back and I think, how could anyone really have thought that? It's a very strange intuition to me now.

I I don't know how you could and this was a dog under, by the way, who told me this. So I think that's people are very capable of telling themselves all kinds of stories. But, yeah, I totally agree.

I don't know how one would really look at a dog, interact with it, and think that it's not having some sort of experience, especially given the substrate overlap with our own. Right? I mean, you know that it has a brain.

You know that it has neurons that are Yeah. Firing and connecting with each other. You know that it has, like, at least a decent amount of overlap in terms of the, you know, the sort of hormonal signaling type stuff that goes on in the brain.

Given all of that, it's, like, really hard to imagine how it's not having some experience. And the big reason I doubt it for the AI side is that all that stuff that, you know, in some unknown mysterious way is giving rise to consciousness is like, that same stuff isn't there, broadly speaking. So that makes me a lot more uncertain than I am on the dog case.

Yeah. Indeed.

Speaker 2

Nathan, a pleasure.

Speaker 3

Yep. Looking forward to doing this again.

Speaker 1

If you're finding value in the show, we'd appreciate it if you take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.

The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.

ing. And thank you to everyone who listens for being part of the cognitive revolution.

Shared via Hopper