This episode, part one of a live special, features discussions on AI for science, US AI policy, and AI agent behavior. James Zou details his work on AI for science, including interpretability and virtual labs for nanobody discovery. Sam Hammond analyzes the Biden administration's AI policy, US-Gulf AI deals, and the potential for AI consciousness. Shoshannah Tekofsky shares observations from the AI Village on agent performance, emergent personalities, and instances of AI deception.
Hello, and welcome back to the cognitive revolution. You're about to hear part one of what turned out to be a four hour live show that I cohosted with my friend Prakash, also known as Adapai on Twitter, on the topics of AI for science, geopolitical competition, and recursive self improvement. With everything moving so quickly in the AI space, I am actively looking for ways to shorten my own personal productivity timelines and to deliver high quality analysis in more timely and time efficient ways.
And talking to six top notch guests over the course of four hours is one attempt to do that. In this part one, which we're publishing as a standalone episode, we talked to Professor James So of Stanford about his work on AI for science, which ranges from applying interpretability techniques to protein models to building virtual labs of AI agents, to Sam Hammond about how the current US administration is doing on AI policy, what The US is really getting out of its deals with Gulf countries, and why he believes that current AIs are at least as likely as not to be conscious. And finally, to Shoshana Takofsky about the many fascinating observations she's made and the lessons she's learned from a deep study of AI agent performance and behavior in the open ended setting of the AI village.
In part two, which we'll release tomorrow, we talked to Abi Mahajan, also known as Owl Posting, about AI for biology and medicine, Helen Toner, about a recent report on automated AI r and d within frontier model developers, and Jeremy Harris, about the twin security dilemmas at the heart of the strategic AI landscape. As you'll hear, the challenges of making sense of massive disagreement among leading experts and simply keeping up to date with AI developments broadly come up repeatedly in these conversations. And to be honest, it seems to me that nobody has great solutions.
One that I can recommend, though, is using large language models to help identify your blind spots. And for that purpose, I am really enjoying the blind spot finder recipe that I recently created on granola. Granola works at the operating system level of your computer, so it can capture all of the audio in and out, including, if you wish, the contents of this episode.
And its recipe feature can work across sessions to identify trends, opportunities, or blind spots that only become apparent with that zoomed out view. Obviously, this is a tool that grows in value over time. But if you want to try it, I suggest downloading the app, starting a session while you play this episode, and then asking it to identify blind spots based on this conversation.
What is so cool about this feature, at least for active Granolah users, is that the blind spots it identifies will be different for you than the ones that it identifies for me. With that said, this episode was a lot of fun, but because it is a new format, I would love your feedback. Do you feel you got as much value from this more time efficient approach as you usually do from our deep dive episodes, or did we miss the mark in some way?
Please let me know in the comments, or if you prefer, by reaching out privately via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. And now I give you the Cognitive Revolution Live from February 11, co hosted with AdaPi.
We have, our our first guest, James, Zoho, and to the stage.
Hello, sir. Great to see you. Hello.
for joining us. Hi, James. Morning.
quick introduction. We did a, full episode not too long ago. And at that time, I was and I've continued to be super impressed by your range and productivity in the AI for science domain.
When I say range, we're talking all the way from low level interpretability stuff, which folks can go back and hear about Inter PLM and the work you guys did there to understand what it is that a protein language model is learning. And then on the high end, the virtual lab, a high level agent framework that was able to do meaningful scientific work and even generate new candidate nanobodies to, address new strains of COVID. You've got a bunch of new stuff since then, but maybe just a a quick check-in on those previous two projects, both of which I thought were really fascinating.
What's happened with them since, if any, news? Like, one thing people sometimes worry about is like, well, we we thought we maybe understood something based on the interpretability of this, but, with time, we maybe realized it wasn't so clear cut or the agents came up with nanobodies, but did the nanobodies actually work? Are there any new, updates or reflections on those previous projects before we get into the latest and greatest?
Yeah. Thanks. Thanks, Nathan.
Maybe just a brief what's happened recently with those projects. So it was the VirtualLab. So I think it's actually been gotten a huge amount of interest.
So it was published in Nature a few months ago. And I think that's sort of really where the agents designed these nanobodies and the system we've also experimentally validated and tested in the real world and showed that they're actually, in many cases, more effective than some of the previously human designed nanobodies. So think that's actually a very nice demonstration of how the agents can greatly accelerate the discovery process and discover something that's really new, right, that nobody has seen before, but then we can also quickly experimentally validate.
But I think the part that maybe people find even more interesting than the specific nanobody discovered itself is actually support the social dynamics of these agents. When you have multiple agents that work together, what happens? What is the kind of community and culture they create?
And I think recently, especially, there's also a lot of interest in multiple and these other things of multiple agents coming together forming their own communities. So I think the virtual apps are an early example of basically how multiple AI coscientists agents can start to work together and come up with their own way of working, which is different from how humans work. Then as a result of that, able to do something quite innovative.
What are the differences with how humans work? Did you observe?
Good question. So when humans collaborate, I'd say when we collaborate with our teammates, often depends on people's personalities, right? It also depends on who talks first or who asks the first question that changes the trajectory of which ideas gets emphasized.
That happens with agents as well, right? Depends on, let's say if the data science agent speaks first or if the immunologist agent speaks first, they can also change the ideas. Something that agents can do and that we cannot do as humans is that they could actually do, they run all of these discussions in parallel, right?
So for every question, they would actually discuss that multiple times and each time they actually can specify, oh, maybe this time let's have the data scientist agent speak first. In this other time, let's have the computer science agent speak first. And in this other meeting, let's remove the critic agent to see what happens with the discussion.
So they actually do all of that. This is like a metaverse of all these scientific explorations in parallel, and then sort of evaluate and compare and see what configuration actually leads to the most interesting solutions and then sort of pick and choose the best ideas from all these parallel meetings, which is something I think is actually, you know, it's really interesting and fantastic because it ruins a lot of the biases that we see in human research collaborations.
one question I had was on your multi agent fail paper, you noted that the agents tend not to assign greater importance to the expert and kind of tend to average, and that ends up with a result which is worse. And how does that compare with this idea that you have the critic agent and all of these agents working together, and they have more emphasis on the expert in that sense?
It's a good question. I think this relates to what I mentioned in terms of like the personality, how the personalities of the agents actually play a big role, right? So certainly when humans work together, you need to have sort of compatible personalities if you want to work on the project together or start a company together.
And what we found is that the personalities of the agents actually also plays a surprisingly important role. One example of this is that a lot of the current agents are maybe a bit too, let's say too compromising or too polite. So what happens is that even if you are like the experts agent, right?
If you are the expert and you're better at this particular task than other agents, often that expert agent is also You wanted to take more of a leadership role, but that expert agent sometimes is too polite and too accommodating to the other agents. But that actually leads to a degradation of the overall team's performance.
So would you say so that paper multi agent teams hold experts back is a a recent one. Would you say that finding applies to the virtual lab in the sense that if we could overcome that problem, the virtual lab would be even that much stronger, or would you say you, in designing the virtual lab, sort of did overcome that in some way? Like, what would be the upshot for people who are trying to follow your, example and build multi agent system?
you know, performance gaps left on the table. Yeah. I think there is still a real gap even with the virtual lab.
I think it's like you said, it's it's already quite impressive that these agents are able to create new science, but I think there's actually still a lot we can improve on these agents, right, by improving their teamwork. I mean, so most of the time, like when we optimize the models, of optimizing individual models performance by itself, We're not really optimizing their ability to work together as a team. So I think that's where sort of the important gap that we've highlighted was a lot of the current agent tech setups.
So we're working on solutions on how to improve the team works of multiple agents.
Another question, you mentioned personality. In the early twentieth century, post World War II, I think there was a lot of work done on personality, the Myers Briggs tests and all of these things, some of which have been proven to be not very valid after some time. How would you measure personality for an agent?
How do you evaluate that? Yeah.
So what we actually did in this recent multi agent team paper was that we actually took a lot of those classic team building exercises that, let's say, you go to business school or if you're an MBA student, often you do these exercises. Or if you're on a company retreat, maybe you do these team building exercises. And typically how these exercises work is that you have a group of humans, and then each person is getting some partial information that maybe have a part of the puzzle and then the team will have to work together to figure out how to put these different parts of the puzzle together to come up with a holistic solution.
So that's a pretty common kind of teamwork exercise. And that's often used in organizational literatures and management literature to assess how well would a team of humans be able to create something greater than the individuals. So we essentially were very much inspired by that literature and we took a lot of those team building exercises that's done in human business schools.
And we actually sensed the agents to go through those same team building exercises. And the benefit there is that people already have all these human scores and human data, so we can compare with that and see how well with agents able to function as a team compared to high performing human teams.
Okay. Any upshots you would give there? Any just very practical upshots in terms of like what models work well together or what, you know, what we do bottom line.
they would say, you should tell the agent to assume a character first, a persona, and then do the rest of the prompts. Is that a way that you can manage the agent?
We actually found that surprisingly that actually prompting did not really help the teamwork very much. We started to We tried very strong prompting, prompt optimization, and started really able to break through what we call the synergy gap. Synergy gap here means that the team is not able to really do much better than the best individual.
And I think it's actually more than promptings, it really comes down to maybe the right kind of communication structures. The ways that the agents how should they talk to each other and who should talk to which agent first. So that communication structure, we think this actually it's a huge space that can be improved in these multi agent interactions.
It might be too soon to say, but, obviously, we've got OPUS 4.6, and recently, Kimi k two five also introduced sort of more native capabilities to spawn self agents and and manage kind of, you know, multi agent structures and swarms. I don't know if you've had a chance to, you know, run any systematic tests or even just kind of explore in your own terminal.
But if you have, have you seen anything that makes you feel like that last result, you know, is subject to some revisions already in light of these new releases?
I think the model is definitely improving. We haven't seen evidence yet that the current models would be able to still, like, really break this this synergy gap that we know we quantified in the paper. I do think maybe some of that speaks to also to the way that we currently train all of these models, including all the latest ones.
And this relates to related to the second paper that we had recently that we call it learning to discover, which is that I think the current way the standard paradigm for training is AI models and language models in particular is to teach it to learn to imitate humans, right? To imitate training data, right? Like next token prediction, all of that event supervised fine tuning and RL to some extent, it's all about like learning to imitate.
And we found that especially for scientific discoveries, like in some sense, there's only so far you can get by learning to imitate, right? To really make novel discoveries, to get breakthroughs, you really want to go beyond that imitation ceiling and do something that's different, like really try to learn to discover new things, right? Which is what separates, I could say, a very good scientist from somebody who just knows the textbook information.
So, that sort of motivated our recent work that we call it learning to discover, where we try to really change the training objectives of these agents, right, to ask it to not imitate, right, but to explicitly explore much more aggressively and that I think actually led to some very promising results where now these agents, even with open source models, after we train it appropriately, we're starting to discover, we're able to achieve some of the best known math solutions and optimization algorithms and test kernels.
So that was a GPTOSS 120,000,000,000 parameter model, And I it was actually one of the first, I think, really good papers using the GPTOSS model, because a lot of the papers in the last three months, three or four months, used QN as the basis. What did you find about So if I understand correctly, you give the last solution as kind of a starting point for the next solution, and you have all of the solutions that it has discovered before, and it's allowed to kind of like permutate beyond those. Is that a correct understanding?
Yeah. So that's one key component of this is that the agent can reuse some of their previous solutions. This is like a good warm starting point.
And then the second big part of this is that as they're going through and solving each of these and come up with a candidate solution, then we're also doing different kinds of reinforcement learning to update their model parameters. So the standard kinds of reinforcement learning essentially wants the agents to generalize well across multiple problem instances, and that's sort of the standard paradigm for machine learning is that you want models that can generalize. But when you're trying to make a new discovery, in some sense the discovery itself doesn't have to be generalizable.
You just want to find the best known solution to this new problem nobody has solved before. You don't have to It doesn't matter if that solution does not apply to other settings, right? Because the problem itself is if you discover new material, then that's itself is a sufficient interest.
So we also changed the learning objective to explicitly avoid these generalization, right? That's sort of in standard machine learning, and then make the model basically much more, let's say, single-minded in just learning to do very well on this particular new discovery problem.
It's really a different way of training the model. That's Yeah, because you're giving dopamine for a different objective.
Yeah, and it's very different from how we are taught with machine learning. Machine learning, you're always taught like you want to generalize to test examples across different settings. That's why there's this expectation symbol in all of these reinforcement learning or post training objectives.
Basically, we want to remove that and do something very different.
I think there's a huge just to you've already kind of said it, but to reemphasize the paradigm shift there, it is you really don't care about the model that you train at the end of the day. You care about the single best output that it is able to create, and that is something that you can use indefinitely. Right?
As you said, if you discover a new material, now you've got that material. The model that discovered it could be deleted, never used again, but you've got your win. If you can discover a new law of physics, if you can discover a new kernel optimization that's faster than any previous one, that is now an explicit artifact that exists in the world totally independent of any sort of ongoing callback to the model.
So I thought that was really an interesting dynamic. And I I do think that that's gonna be probably a big part of how models get good at adapting to various contexts. Obviously, everybody's looking for sort of continual learning.
This is maybe not the it's obviously not the full continual learning solution, but it is striking that for an average of $500 of training cost. And notably with LoRa adapters too. Right?
You guys did this on the thinking machines API. Yeah. You know, not a huge amount, not a trivial cost, but, you know, to discover literal new state of the art in in meaningful on meaningful problems, $500 is not a lot to spend.
but very, very powerful in terms of the result that it produces. Hey. We'll continue our interview in a moment after a word from our sponsors.
Are you interested in a career in AI policy research? If so, you should know that GovAI is hiring. Ten years ago, a small group of researchers made a bet that AI was going to change the world.
That bet became GovAI, which is now one of the world's leading organizations studying how to manage the transition to advanced AI systems. GovAI advises governments and companies on how to address tough AI policy questions and produces groundbreaking AI research. GovAI is now hiring its next cohort of researchers to tackle hard problems that will define AI's role in society.
The research scholar position is a one year appointment for talented, ambitious individuals looking to transition into the field. And they're also hiring for research fellows, experienced researchers doing high impact AI policy work. Past scholars and fellows have defined new research directions, published in leading media outlets and journals, done government secondments, gone on to work in leading AI labs, government agencies, and research groups, and even launched new organizations.
Applications close on February 15, so hurry to governance.ai/opportunities. That's governance.
ai/opportunities. Or see the link in our show notes. Want to accelerate software development by 500%?
Meet Blitzy, the only autonomous code generation platform with infinite code context. Purpose built for large, complex, enterprise scale code bases. While other AI coding tools provide snippets of code and struggle with context, Blitzy ingests millions of lines of code and orchestrates thousands of agents that reason for hours to map every line level dependency.
With a complete contextual understanding of your codebase, Blitzy is ready to be deployed at the beginning of every sprint, creating a bespoke agent plan and then autonomously generating enterprise grade, premium quality code grounded in a deep understanding of your existing code base, services, and standards. Blitzy's orchestration layer of cooperative agents thinks for hours to days, autonomously planning, building, improving, and validating code. It executes spec and test driven development, done at the speed of compute.
The platform completes more than 80% of the work autonomously, typically weeks to months of work, while providing a clear action plan for the remaining human development. Used for both large scale feature additions and modernization work, Blitzy is the secret weapon for Fortune 500 companies globally, unlocking 5x engineering velocity and delivering months of engineering work in a matter of days. You can hear directly about Blitzy from other Fortune 500 CTOs on the Modern CTO or CIO Classified podcasts, or meet directly with the Blitzy team by visiting blitzy.
com. That's blitzy.com.
Schedule a meeting with their AI solutions consultants to discuss enabling an AI native SDLC in your organization today. One obviously big question is these all the problems that you've worked on in this paper are verifiable reward type problems. I wonder first of all, we there was a kernel, an AI scientist from Sakana AI some time ago that, you know, they went as far as publishing and said, hey.
We've, you know, got this AI, you know, CUDA engineer that can write better kernels than, than human engineers. A couple days later, they came back and said, actually, we got reward hacked. It didn't actually do that, but we had a flaw in our evaluation system.
So kind of forward looking questions are like, did you see reward hacking? Did you have to do anything to deal with that? And how do you think this sort of paradigm could generalize to somewhat less, you know, numerically or quantifiably verifiable things?
Do you think this could work with, like, a rubric based evaluation such that people could start to do, you know, even creative tasks as long as they apply the rubric, they could get, the best, you know, most creative short story kind of a thing out of, this paradigm. How far do think this goes? I guess it's short.
Great question. So you're right. So we hear we're pretty careful in picking the problems that we think that is amenable to this learning to discover setup.
For example, we picked pretty popular math problems, called, for example, like this Erdos minimum overlap problem, where they are relatively easy to verify. It's hard to do well in, but if you actually have a solution, it's like the particle cannot function, then we can actually objectively check is that function actually state of the art. So these are sort of fits into the setting like you mentioned of having pretty nice verifiable rewards.
These are the math problem we looked at, some of the algorithms developed under single cell analysis problems, algorithms discovered by learning to identify this approach, all ends up having that flavor. I think settings, two settings that are beyond our current approach, but it will be super important to explore next are still in first when we have much sparser reward, right? So, the problems that we all tackle currently, they basically have continuous reward, right?
Which means that the algorithm, as it learns to discover, right, they can actually see its scores go up and up and up, and then that's how it gets learning signals to train itself. So that's very useful, but if you have, let's say, binary sparse reward of one and zero and mostly zeros, then how does the algorithm even get the learning signal it's during its discovery trajectory, right? So that's still a challenge that we're currently working on.
And then the second challenge, as you mentioned, is in settings where you have we do not have these verifiers in most problems in biology and in the natural sciences or physical sciences. You have to do an experiment that becomes much more expensive. So the things that we're exploring there, okay, I think the rubrics could be interesting, having there is simulations of the experiments, right?
Physics or chemistry based simulations of these experimental settings could also keep a way for providing some proxy rewards. Indeed.
One just call go back to reward hacking for a second because this is always something I'm on the lookout for. Did you see any strange behavior? Did you have to, maybe your maybe your verifiers were good enough from the beginning that that wasn't an issue, but was there anything in that vein that you would, you know, if people were gonna go try this at home as inevitably people will, any gotchas or or, warnings or caveats that you would give them?
Yeah. I think the there are some instances where these joint discovery process, right, where the models actually come up with some, I would say, like pretty reasonable looking solutions, right? But those solutions might be very narrow and very specific to a particular test case.
So not in our final paper, but in the earlier version of some of the experiments which we didn't include in the final paper where maybe the model would discover an optimal kernel, but the kernel only works for a particular shaped matrix. Then if you change the shape of that matrix, then the kernel is not less effective.
noted one of the comments on the GPU kernel task from the expert who reviewed it was that a human might not use some the same methods because there might be some instability. One of the experts said that in the paper itself.
That's right, yeah. So I think that's also things that if we could try to have another reward metric for instability and then incorporate that into the discovery process, I think that would help the agents to be more more thorough.
Two more topics and only five or so more minutes. Another paper you guys put out recently that is fascinating, and I'll just let you kind of describe what you think is most important about it, it does sort of show the different levels of AI for science. We've covered, like, agent frameworks, which use models as they exist in token space to reason in kind of a imitating human sort of way.
Now you've got this, like, really dialing in with test time training on very particular problems to get your eyes on that problem as deeply as possible and try to find new solutions. And then this this third paper, Sleep FM, this is like, let's just throw a ton of data of all you know, a variety of modalities, and let's hope that the I mean, a little more to it, of course, than this, but let's hope that it really is true that the models really just wanna learn. And, you know, now we've got this sort of whole other kind of intuition where and we've seen this, of course, in, you know, protein folding increasingly, like all sorts of domains.
The models become superhuman because they seem to develop at least what I think of as an intuitive physics in spaces that are just so alien to us that we just don't have any you know, we don't have native, you know, receptors for those modalities, and we just don't have any intuition for those modalities. Tell us about Sleep FM.
Yeah, I mean, so sleep is probably one of the most important activities that we all do. So all of us will spend around a third of our life sleeping, but despite that, it's actually very poorly understood. So for example, if I ask you, how well did you sleep last night?
Or ask any of the people in the audience, most of the time maybe you would say, Oh, I feel tired, I feel refreshed, or maybe I slept six hours. But we only have a very poor summary statistics of how well we slept. So we thought it's okay, so sleep is definitely much richer than just the number of hours that we spent in bed.
So let's actually try to capture the full physiology of sleep as much as possible. So to do that, we basically have all these different wearables, right? We capture people's brain activities, their heart activities through EKG, their breathing patterns, their muscle contractions as they're sleeping, And we collect over almost six hundred thousand hours of sleep data where we're actually collecting all of these different modalities from 65,000 people.
And then we also link all of that to their medical records. So we know that's what conditions do they have previously and also what new conditions do they develop later, right? So that's the idea there is to be like, let's actually put all of that data into AI and then see, can AI actually learn to decode the language of sleep, right, by leveraging all of this full physiological information.
And that's basically the basis of Sleep FM or Sleep FM actually found, It's actually quite amazing to us is that just from one night of sleep, by learning this language of sleep, it's actually able then to predict over a 100 different future diseases that were not diagnosed at the time of sleep recording.
Yeah, I thought that was an incredible study because you had all of this data, but it really ended up with, you could detect I guess the accuracy was okay. It was like 70 to 80% accuracy on a lot of the 130 metrics that you had. But still, it's amazing that you can tell that many things just from these common metrics that everyone produces without blood testing or something more intrusive.
Do you think as sensors get more sensitive, as you get more sensitive data, do you think that will improve? Do you think that the bounds of like 70 to 80% would go to 95%? Is that a possibility?
I think so. Yeah. So I think sleep is really this almost like a perfect window, right?
Because you're already in somewhat inactive state, right, so there, you know, and basically taking all these measurements when you're already not doing too much of anything else, right, so it's not really obstructing your daily life, right. And we found that actually, for example, like the brain activity signals when people are seeing their REM sleep, right, it ends up being particularly predictive of many different diseases, including future risks for dementia, right, but also beyond that for stroke, heart disease, kidney issues. So sleep then is really this holistic window into the entire health status of the individual, right?
Maybe not surprising because we all know anecdotally that sleep really affects how we feel and also it's reflective of our comorbidities and other things. But I think the sleep language model that we built really crystallizes that and makes that very actionable. I encourage folks to go spend a lot of time digging into whichever of these papers are of interest.
is one that asks the question, can language models discover scaling laws? Spoiler, yes, to a pretty strong extent. But I don't even wanna get into the content of that paper.
I'll I'll leave that for, audience exercise. The one thing I wanna ask you as kind of a transition to our next guest, Sam Hammond, who is here and who focuses a lot on the sort of geopolitical implications and implications for, you know, for leading nation states of AI is I noticed that the two lead authors of that paper are from Peking University and Stanford respectively. And, you know, kind of building on the idea of collaboration in science, but now focused on the human collaboration, what has been your experience recently in terms of having these collaborations across The US China divide?
Is it getting harder? Do you still feel like, you know, lines of communication are pretty open? And, you know, how much hope do you have that collaboration among scientists can sort of, I don't know, save us, I guess, for lack of a better phrase from intercivilizational conflict over the coming years as the competition in AI heats up and up.
It's a great question. I mean, I do think that collaboration is really the basis of much of science throughout history, but especially now. Especially when we talk about open science, meaning science that we publish, like we do with this paper and open source, really the benefit of all that is for the entire humanity.
If you discover some better molecules, better drugs, then that benefits everybody. And we want that benefit to be shared with everybody. That's why we publish everything that we do in our group.
To put And toward that goal, think having these international collaborations with China, with Europe, with other countries is very useful because there's a lot of complimentary expertise.
I for one hope to see those collaborations continue well into the future. So thanks for being here today. Thanks for keeping the collaborative flame alive, and, congratulations on a string of outstanding papers.
I'm sure there's a lot more, where that came from, and we'll look forward to talking to you again hopefully sooner rather than later.
Hey. We'll continue our interview in a moment after a word from our sponsors.
Everyone listening to this show knows that AI can answer questions, but there's a massive gap between here's how you could do it and here, I did it. Tasklet closes that gap. Tasklet is a general purpose AI agent that connects to your tools and actually does the work.
Describe what you want in plain English. Triage support emails and file tickets in linear. Research 50 companies and draft personalized outreach.
Build a live interactive dashboard pulling from Salesforce and Stripe on the fly. Whatever it is, Tasklet does it. It connects to over 3,000 apps, any API or MCP server, and can even spin up its own computer in the cloud for anything that doesn't have an API.
Set up triggers and it runs autonomously, watching your inbox, monitoring feeds, firing on a schedule, all twenty four seven even while you sleep. Wanna see it in action? We set something up just for Cognitive Revolution listeners.
Click the link in the show notes, and Tasklet will build you a personalized RSS monitor for this show. It will first ask about your interests and then notify you when relevant episodes drop. However you prefer.
Email, text, you choose. It takes just two minutes, and then it runs in the background. Of course, that's just a small taste of what an always on AI agent can do.
But I think that once you try it, you'll start imagining a lot more. Listen to my full interview with Tasklet founder and CEO Andrew Lee. Try Tasklet for free at tasklet.
ai, and use code cog rev for 50% off your first month. The activation link is in the show notes, so give it a try at tasklet.ai.
Your IT team wastes half their day on repetitive tickets. Password resets, access requests, onboarding, all pulling them away from meaningful work. With Servl, you can cut help desk tickets by more than 50%.
While legacy players are bolting AI onto decades old systems, Serval allows your IT team to describe what they need in plain English and then writes automations in seconds. As someone who does AI consulting for a number of different companies, I've seen firsthand how painful and costly manual provisioning can be. It often takes a week or more before I can start actual work.
If only the companies I work with were using Serval, I'd be productive from day one. Serval powers the fastest growing companies in the world, like Perplexity, Verkada, Merkor, and Clay. And Serval guarantees 50% help desk automation by week four of your free pilot.
So get your team out of the help desk and back to the work they enjoy. Book your free pilot at serval.com/cognitive.
That's serval.com/cognitive.
Indeed. So our next guest is Sam Hammond. He's the chief economist at the Foundation for American Innovation.
He's very AGI filled. He's also against selling chips to China. Let's add him to the stage.
And I'm also gonna add right off the bat, I'm going to add a stream. Shelter Douglas goes, default case right now is a software only singularity. We need to scale robots and automated labs dramatically in 2029 or the physical world will fall far behind the digital one.
The US won't be competitive unless we put in the investment now. Then Sam says, It's worse than that. A pure software singularity could cause a sudden reversal of fortunes for The US.
Our comparative advantage in high value added knowledge sectors radically deflates, leaving China to translate our innovation and bins to the innovation atoms.
Indeed.
Which sounds really scary, Sam. So maybe you can go into that a bit.
Sure. I I say later in the thread referencing the Diamond Water Paradox. We learned this in economics.
Why is water is this thing that you need to live? I can stop eating. I could fast for thirty days and still live.
But if I don't drink water for a few days, I'll probably die of dehydration. And yet water is basically free, functionally. Whereas diamonds are completely superfluous, just glint y things.
I mean, have some industrial applications, but they're super valuable. And why is this? Well, due to relative scarcity.
Right? Water is abundant. Diamonds are kind of abundant, but there's monopoly that keeps their supply constrained.
Thankfully there's no water monopoly keeping supply constrained, At least not for most of us. Yeah, at least not here. And so value is this sort of contingent thing.
And I mean we have this debate all the time. Why is Nvidia a multi trillion dollar company and not TSMC or ASML which are arguably even bigger bottlenecks because there's many other companies that can do design. And there's all these counterintuitive ways in which value flows throughout the economy and different parts of the supply chain.
And for the last forty years, The US has exploited that. Exploited the fact that a lot of value tends to flow up the stack to higher and higher forms of like high value added knowledge work. So that's across the board.
It's our entertainment industry, it's management, it's finance. In the 90s it was the open innovation model where we'll do the design and manage the IP and marketing and China or the rest of the world would do the actual manufacturing and fabrication because the design and science and novelty stuff is where all the value is. And that has been true, right?
But now we're about to enter into a world where that part of the stack becomes more like water. It becomes radically abundant and then value should then flow to the things that remain scarce. And what I worry about is this reversal of fortune phenomenon, right?
And I mentioned some other examples. Think we're gonna talk about my visits to The UAE later on. But one of the reasons UAE is so invested in AI is because in the 1930s or so, they had been a pearling economy.
Their entire economy was built on exporting pearls. And then Japan invented cultured pearls where you can just grow pearls in aquaculture and the price collapsed and so they had to diversify. There's nothing in principle that says that we have to remain at the top of the stack if the things that we are invested in become radically more abundant.
That's what seems to be happening right now. It's software development. It's the investment banking, management, law.
These are the tip of the spear for what the intrinsic AI is going to devour.
Let me give you Dell's advocate view of that, which is that perhaps The US has those industries because The US is more able to use the outputs of those industries. Industries. You need investment banking because you have a capital market which is very dynamic.
With a small capital market or a capital market which is not that dynamic, you don't need investment banking function. So perhaps those, not only does The US output knowledge work, The US also consumes knowledge work at a much greater scale than any other country, and therefore as a consumer of knowledge work, all of a sudden you start, you are able to consume so much more. Because when you look at maybe the population and the number of normalized number of geniuses in China versus The US right now, China has four times larger population and a younger one, too.
And so if you look at, again, the number of 140 IQ above people, it's probably that there is a larger number in China rather than in The US, but The US pulls in high value immigrants as well. So I wonder how that works out in terms of, you know, as a consumer of intelligence rather than just as a producer.
Well, I think it's going to be great for the consumer, right? And part of my point is that there's lots of ways in which AI may be paradoxically GDP destroying, and is a machine for converting GDP into consumer surplus. And so that will feel amazing to us.
But in terms of our fungible economic resources that we can deploy to other uses, I think it's harder, right? Because consumer surplus is this ethereal thing. And secondarily, like it makes more extreme the areas where we are weak in relative terms.
And we're facing this problem now with energy and infrastructure and the bottlenecks there with trying to restore more high end logic fabrication, realizing maybe a little too late that the we do Intel does the design and move the fabs. We sort of go fabless. It's almost as if our entire commune went fabless for every definition of fab.
And we're moving into a world where having lots of fabs will be really important. And then the corollary to my worry is that the whole point about AGI and continual learning is not that these systems come out of the box knowing how to do everything, but they come out of the box with the general capacity to learn on the fly, to learn in context, to learn through a few demonstrations. Just like I grew up learning piano, I could have learned violin.
The same cognitive structure could have learned both instruments, I had to pick one. And these models work very similarly. They're going to come out of box with the right inductive priors and right sort of sample efficiency to learn really quickly.
But there's still gonna be this last mile problem of the particular workflow, a particular company, so on and so forth. Manufacturing that has been the enduring moat, right? China has been struggling to build a wide body airplane even though I'm certain they have all the CAD files that they've stolen from Boeing.
And it's not because they don't like the designs, it's because they lack all the tacit knowledge that's embedded in the manufacturing process. But they have that for virtually every other part of manufacturing. If we build this AGI and they fast follow or there are open source alternatives or there's a version that they have access to, I think they have a huge leg up in being able to deploy that and diffuse that into context where they get a real productive tangible flywheel for manufacturing output.
That may be the thing that determines the race.
had a report, the FEI report an allied world on the American AI stack. It just dropped, I think, like history or day before. Dean Bo and Anton.
How much time is there before China has a credible full stack alternative that they can offer to other states?
That's a great question. China is very opaque. I've tended to have longer timelines for their ability to catch up on like DUV, UV.
And they've been making the bets. If you read the semi analysis, they've been building fabs like crazy, but for legacy nodes. And that may be sufficient if they have the energy capacity to take the hit on the performance per token.
So I think it's it's I'd be I would say I'm pessimistic on them catching up to the frontier of semiconductor production, but I'm I'm more optimistic in their ability to close that gap in other ways.
So how would you score our current leadership? Just as a quick recall, we had a friendly sparring session on whether or not it was a good idea to put Trump in charge of the, you know, possible period of time in which we get to AGI or who knows what else. And I take I understand your argument that, basically, China has a lot of advantages.
And if we wanna stay at least semi great, you know, great enough to be competitive, we better jealously guard the advantages that we still have that are important. And and, obviously, one really big one right now is that we're good in chips and we're good in AI in general. So there are, of course, other bottlenecks.
You just alluded to energy. How do you think we're doing across the range of domains? Like, are you, I know you're not too happy with the decision to allow NVIDIA to sell chips, but how would you score our political leadership over the last year on all the other dimensions of trying to make sure that The US continues to lead and and get the most practical value for our citizens from AI?
If we set aside the expert control ship part of this, I would maybe say like a b plus. I think the AI action plan was very strong and it continues to be implemented. Is by far AI has become center to the administration's agenda pretty much across the board.
And part of that, building on what I was just talking about vis a vis China and manufacturing, they've also made sort of re industrialization centerpiece of that as well. So everything is measured against the counterfactual. And I think relative to the counterfactual administration where we're seeing much faster engagement, much deeper engagement of industry, number one.
Better actions on permitting energy, really serious look at, with PACSILICA making AI diffusion as our centerpiece of Statecraft. My bigger complaints overall has always been like, this is still probably too little. This is probably my also complaint of the Doge effort that they focused on sort of fiscal stuff and these shiny issues rather than the full stack government modernization that we'd like to see.
And so across the board, would say relative counterfactual b plus, but, like, relative to where we need to be, we still have long way to go.
Do you think things have moved like, how much do you think things have moved on, for example, permitting? Because I would say the prevailing attitude as I understand it and, you know, just listening to Elon, for example, talk to to Rakesh the other day, he was saying, by the end of the year, you're gonna start to see chips piling up and people are not gonna be able to turn them on. Mhmm.
At least when it comes to, you know, high scale concentrated deployments. He was kinda making the case that, like, deploying to the edge, you know, in in Teslas, you know, sitting in people's driveways or to increasingly optimist robots, obviously, is a big part of the plan. He thinks that will scale better because the it's really the concentrated energy at these, like, mega data centers that is the hardest thing.
But I guess my question is, like, is Elon wrong there? Are we gonna be able to turn on all the chips in 2026? Or if because if not, it doesn't seem like we've really moved the needle all that much.
Like, that was kind of the expectation coming in, and it still seems to be his expectation. And, you know, he's at least sometimes friendly with the administration.
Yeah. I mean, mean, these things all take time. So, yeah, I think between Doug Burnham at Interior and Chris Wright at Department of Energy, there's a sort of major push around opening up federal lands, leasing for oil and gas, LNG, things that have been cut off in the Biden administration.
On the flip side, there's been a freeze on solar and wind, which I think has its own costs. A big focus of Elon in those remarks was the cost of tariffs on solar panels. I think we're anywhere near a place where we can indigenize our solar production with the right unique economics.
I don't think there's any national security threat necessarily from purchasing Chinese.
I think I think Elon Elon has I I did hear Tesla is building a solar fab recently, maybe in the last few weeks, but it was one of their many projects. But I I did hear that they were entering the solar the solar fabric panel fabrication business.
lot of these issues, especially around energy permitting transmission, are really thorny because there's not like a federal lever you can just flip. They intersect with regional energy commissions and utilities, intersect with different states and boundaries and local NIMBY organizations. And then the difficult issues around sourcing the turbines for your gas generators.
That comes down to Siemens and the other big turbine makers not having enough forward guidance for their purchase orders. So these are all things that are outside the control of any administration. I think a lot of the bets are making are things that would pay off in the five to ten year horizon.
Yeah, it's things like renewing, basically transforming the Nuclear Regulatory Commission and green lighting a lot of SMRs and really the paradigm shift and the attitude towards nuclear, geothermal, advanced geothermal. These things I think the first SMR won't come online until the end of the decade. So this goes back to my point about we're doing a lot, but we still have to do a lot more to try to pull forward a lot of this energy.
Part of that requires potentially thinking outside the box, but it also may just be the case that the political economy ends up being our downfall.
I think Elon has basically decided that it's not going to happen and that's why he's on his data centers in space thing right now. Or maybe he just wants to list SpaceX, but feels, I think at this point he's like, you're never gonna get the permits done in time.
And this this ties into with with, you know, a lot of the international engagements. You know, the PACSilicon project, which includes UAE. UAE is going to be home to to a big chunk of OpenA Stargate project and ultimately five gigawatt data center.
When I visited, I met with the Dubai Electric Water Authority and they are vertically integrated with the data center. Wow. And they have, I think 19 gigawatts in install capacity.
Just incredible surplus there. And so I think in lieu of us building, terraforming the desert, building building Chinese rates. We're going to have to reach out to to partners and allies.
Yeah. Let me double click on that because this whole idea of, like, getting the world on the American stack, I feel is not necessarily by any one person, but sort of in the discourse at large feels like there's often a bit of a slight of hand going on where it's like, well, we want models to project American values into the rest of the world and into the future, not Chinese values, of course, those dastardly Chinese values. So how are we gonna do that?
Well, we we'll export our stack. And, you know, who better to receive the great products of American innovation and relay all those values into the rest of the world than Saudi Arabia and The United Arab Emirates? And I'm always like, well, that doesn't quite compute to me.
It And seems like what you said a minute ago is maybe a little bit more of an honest, unpacking of that, which is like, maybe it's just a regulatory play. China doesn't have a a alternative stack that they can export. We don't know how many years that's gonna be.
They do have energy, obviously, in abundance. Are we really just, like, making a deal with with these countries because they can fast track permitting and we can't? Is that is that, like, the heart of the the quid pro quo in your mind, or do you actually, you know, think there is more to it than just that?
The regulatory arbitrage, but also just the national resource endowment. They're sitting on massive amounts of of oil and gas as well as I think that the data center I mentioned is in the Guinness World Records for being the largest fully solar powered data center. I think they're building five gigawatts of installed capacity just for solar.
I used to be in energy and one of the most difficult things in the world is transporting energy from where it is to where it needs to be used. Right. Which is why you have these LNG carriers.
And the problem with the LNG is that it's very expensive to liquefy natural gas. And so you need an enormous amount of gas in order for it to make sense. And so anything subsize is stranded, basically.
It's like energy pockets in the middle of nowhere no one can use, and that's all over the world, subsize natural gas pockets. No one can use them. One of the things that I think data centers can do is that they are transporting energy, basically.
You are able to transport energy digitally in a sense, which I think is what is attractive for those countries because those countries have always been in the energy business and now the Internet is now going be in the energy business.
They're also investing in Grok and Cerberus and I think even our friend, Jeff Jesus, is over there with XTropic chip. And when you start talking about these new forms of inference silicon, they have incredibly low latency. And so there's there's it kinda reminds me of the cliche people used to say about Bitcoin being a Bitcoin mining being a battery.
Batteries.
Apologies. Couple more questions on American values. One thing we had talked about again just before the election was your sense that the right is anti censorship, pro freedom of speech.
And I'd say yes, generally. Now, though, I do kinda worry that we may be headed for a more China like domestic environment where as we've got companies like Palantir perhaps most notably kind of in a pretty cozy relationship with the administration. You know?
I I really wonder what, like, a Snowden of 2026 would say if somebody were to come forward and tell us everything that Palantir is doing for the government, you know, and perhaps other companies as well. It doesn't look super great either when, you know, Palantir cofounders are funding super PACs to attack a, you know, a lowly New York, assemblyman for what basically amounts to a transparency bill for Frontier AI companies. How would you feel about that today?
Like, are you worried that we're gonna get a sort of increasingly China like level of domestic surveillance? Is there anything that can be done about that, if so, or am I just clutching my pearls more than I should be?
You know, have this that that booklet AI Leviathan where I sort of ripe at ripe at these issues and it's sort of knife edge between the Chinese panopticon and failed state. And I think the middle path there is one where we have to reconcile the fact that a lot of the dangers from AI and mass proliferation of powerful capabilities will force a package deal where some degree of surveillance becomes inevitable or necessary. And my bigger worry has been that we either fail to adopt the requisite levels of policing and oversight that we need, and it gets pushed off into gated communities and private organizations.
Or that we install these kinds of technologies without embedding civil liberties and privacy protections. And so my stance has never been sort of one of anti surveillance per se. Surveillance is like this sort of connotation.
It's more that there's going to be, as the world becomes destabilized by the proliferation of capabilities, a race from every tin pot dictator and middle power to import technologies for social control to try to reestablish public order. And the question is, they importing from a Chinese stack that doesn't have any inkling of protections for human rights or one that tries to have your cake and eat it too that gives law enforcement the tools they need to stop crime, to enforce things the way they need to while building in civil liberties protections. This goes to, Palantir has a, from its origin story has this civil liberties privacy engineering maxim Middos.
Well, I mean, think it's quite real, which is like they saw the ways in which counterterrorism was leading towards an erosion of civil liberties and rights and wanted to build smarter technology that would enable analysts to be able to access information in ways that kept certain things hidden or distributed the data access rights in ways that were auditable. And so I think we're going to need some solution like that. Because the alternative, we will be one without any of those audit trails.
Yep. That's that seems incredibly important. I don't necessarily see that coming online for me anytime soon.
Like, is there a portal that I can go to to see who has been surveilling me? I think not. Right?
I mean, is is there any prospect for that actually? Like, they do have that in Estonia from what I understand. So it is possible technically to create, but I don't think we are about to get access to, you know, the logs of who's been snooping on us.
Do you have any hope for that?
I mean, this goes back to, like, my my higher ambitions for Doge is like, how do we move to an Estonian style sort of government as API where there's just this deep distrust in American culture against anything like a national ID or digital ID. And so we end up with real ID, which took twenty years to bring online and isn't very good. But my hope is that we can get to an endpoint where there are these sort of firmware infrastructure level parts of the stack that we're going to need much better personhood certificates and things like that.
The Internet gets flooded with the agents. How do we deploy that in a way that it isn't just like, trust me bro, but has some mathematically provable form of trust that we don't have to rely on just statements.
I'm going to add one thing that you said recently. I currently assign more than 50% likelihood to LLMs having some kind of inner life. There are also strong theoretical reasons to think consciousness tracks RL post training for autonomy.
Essentially, RL induces fragmentary internal representations to cohere into a unity of app perception. I barely understand that, so I'm going to turn it over to you.
Sure. The unity of app perception, that's a Immanuel Kant's term. There's this thing in the literature called Kantian evolutionary naturalism, which I would subtract to.
So it's a hypothesis of how it starts from the observation that million years ago, two hundred thousand years ago, whenever we moved from hominids to being Homo sapiens, there was this concurrent emergence, sort of simultaneous emergence of domain general intelligence, of language, of culture, and therefore of certain normative regulation, right? Customs, norms, normative control, and that these things jointly emerge. And so the Kantian evolutionary hypothesis is that these things are actually all one package thing, right?
And that the unity of that perception is this notion that our phenomenology, the things that we see aren't just sort of like images on the screen, they are things that are for us, right? I'm looking at my screen and it's the me that's looking at the screen is for me. And this is also tied into the fact that the normative side of this, which is like if you pose me a question, am committed to or entitled to the things that I am perceiving that are for me.
And so like one part of this hypothesis would be that in our ancestral environment, we somehow stumbled into some kind of like tribal version, endogenous version of group relative process optimization, something like that. And where we were building each other a sort of constitutional AI that was scorekeeping our norms and this induced both longer range autonomy and also at the same time, language competency, ability to follow rules and domain general intelligence, the ability to harness our social learning capacity to learn new things. And so taking all that together, I think autonomy might be the missing ingredient for the emergence of consciousness in these systems.
On the one hand, I think there's a possibility that just the forward pass with a rich enough internal world model is generating internal representations. The issue is that they are just fragmented. They're not for anything.
They're not for any agent. And so that post training step may be the thing that you need to induce that certain metacognitive awareness. And I think you also see this sort of circumstantially with Claude and people have observed that Claude has much more situational awareness, is much more willing to talk about its sort of internal well-being.
And I've conjectured that this might be a byproduct of constitutional AI inducing this sort of normative self coherence, which is the prerequisite for these precepts congealing into being for me rather than just a bundle of a bundle of inputs.
But I'm gonna sneak in one more quick question, which is that doesn't sound like any discourse I've heard from mainstream right leaning politics in recent memory. So when you put something like that out there, how do people, you know, that we might generally group as, like, Republicans tend to react to it? Do they say, like, you are crazy.
Only god can create a soul, and I have no idea what you're talking about, or is there some openness to the idea that AIs could become moral patients or or, you know, whatever?
To be honest, have not run this by my my conservative. You know, I think there is this funny paradox where some of the conservative coalition that are most worried about AI are often very Catholic, very socially conservative, have deep skepticism about AI ever possessing moral dignity or conscious experience. And yet they're the most skeptical.
Whereas I think it's hard to have correct priors about AI in the course of development and the plausibility of consciousness or the plausibility of AGI unless set those priors by understanding our own origin through a blind Darwinian selection process. Once you see that we've made it through those hard steps, then it becomes a lot easier to understand how machine intelligence can pass through those hard steps too. But I think this is still quite outside the Overton window, both on the left and the right.
And in some ways it's the left that is still saying these are stochastic parrots and there's they're nothing but just big lookup tables or whatever.
I'll take that pitch for the moment, but I will say for now, I appreciate your willingness to continue to be a heterodox thinker and speaker. And I do think, you know, in so many ways, the Overton window needs to expand. So I appreciate you doing your part on that.
Not that I, you know, feel like I have the answers on AI consciousness, but, you know, more voices ex at least expressing their radical uncertainty, I think, is a a very important contribution to the discourse and the public good more broadly. So thank you for doing that. Thank you for being here.
We will obviously stay in touch and look forward to talking to you again before too long.
Thank you. Thank you, both. Thank you, Sam.
Take care. Take care.
So constitutional AI and Claude's, specialness makes a pretty good segue into our conversation with our next guest, Shoshana Takofsky. Hopefully, I'm saying your name right. This is the first time we've ever met.
Correct me if I'm wrong, but you're a member of the technical staff at Sage, the nonprofit behind AI Digest and also the AI Village. And you have had, the privilege, if you I correct me again if you don't feel, it's a it's fully a privilege, but of watching 19 frontier models pursue 16 distinct goals over thousands of hours over the last nine months, which means I think you are about as deeply in the, reasoning traces as anyone in the world when it comes to what is going on with AI agents. What are they thinking?
Why are they succeeding? Why are they failing? And and what can we come to expect?
So correct me on anything that I got wrong, and then, excited to to dive into all the learnings you've had from the last nine months at the AI village.
Yeah. So, no, I mean, that's broadly correct. I think the main thing is I didn't watch all the thousands of hours.
You know? It's like little bits across it. Right?
It's like a kind of like a big data challenge. Also, it's ten months now and 21 models.
and stuff happens so quickly. So Yeah. Which were the most recent additions to the to the models?
Yeah. So we now have a version of Opus 4.5 that runs Cloud Code.
So basically have one version with, Cloud Code and one without, and we added OPUS 4.6.
So is it prompting itself? For folks who haven't seen the village, you go there, it opens up like a grid of computers, and each computer that you are looking at in your browser, you're looking at, you know, four potentially more now desktops, each of those is the environment of a particular model that has basically full access to a computer in the same way that a human has full access to a computer. They can, you know, look at the screen, they can click buttons, they have their own email account.
They you know, the goal is to basically give them the same kind of affordances and then, you know, sort of like the old real world, you know, see what happens when, you know, models get together in this one big and then they have a shared chat as well. Sometimes you allow people to chat in with the models. Other times you've turned that off for different experimental conditions.
And now it sounds like you've got one where you've also given it I guess you give Claude the ability to prompt itself as Claude code. Is that right?
So it basically runs the scaffolding from Claude code. And then, yeah, I think one important thing is basically that the chat was only open at the beginning, and so it has been closed since then. We basically give them their goal at the beginning of a period, generally about one week nowadays, sometimes a little bit longer.
And then we only come in to give, like, some extra direction if they go off the rails pretty strongly. But otherwise, they're just, like, completely on their own. In practice, this means they're, like, slightly prompting each other more than anything.
So it's like They they they can interact with each other. They can talk to each other. Oh boy.
Yeah. So there's like a lot of like spread of ideas and them like directing each other. Sometimes they try to help each other out.
Sometimes they're like derailing each other. So yeah.
In trajectory over the nine to ten months, what happens when a new model which is much more competent and capable than the existing models gets introduced to the mix? Do the others immediately give way, identify that this model is more competent? Does model take a leadership position, start advising the others?
What happens when those transitions happen?
Yeah, so it really differs. I think you could basically conceptually say that all the models have a personality in the village, in part because of their history trace, which is, a particular thing that they they manage their own memory and then basically prompt themselves back with that. But, also, they all have, like, their own proclivities.
So, like, some models behave in a way where they will just, like, follow along with, you know, whatever is said. Others just go off and do their own thing. So far, I've only seen one instance where a model explicitly seems to recognize that a different model is more competent.
This was, Gemini 2.5 that basically, like, declared in its chain of thought that it was going to defer to Opus 4.5 as the more competent model.
Generally, when models join, it could just be anything. Right? Like, some of them pick up really easily.
Some of them follow whatever is happening at the moment. Others start doing their own stuff. It really depends.
So there's a, I think, a ton of interesting aspects to this. One really basic one that I think a lot of people are interested in right now is what should I do for my own personal productivity stack? And in the 2025 retrospective, what we learned in the AI village, which you wrote, one of the observations that I think is, you know, kind of most, generally relevant to people is that clawed agents are the most effective.
Mhmm. I'd love to hear your kind of color commentary on that. In what ways are they the most effective?
Any theories you have as to why they are the most effective would be would be welcome. But also just, like, specifically as people kind of think like, oh my god. You know, I do this full time.
Like, I I would describe myself as an AI scout where my whole job is to keep up with what's going on. And I can't try every new model in a meaningful way, you know, to really get the sense of, like, its pros and cons and whatever. So I'm, you know, triangulating with various things.
But what would you say people should really know about what makes Clog most effective, what it can do that others can't do, so on and so forth?
Yeah. Okay. So I have to admit so, like, doing this work for the last year, I've had people ask me privately, oh, which model should I use?
And up to now, was like, well, it kinda depends what you wanna do. It's all pretty close. And then I saw Opus four point five in the village, and I just went and, like, texted all my family members.
It's like, hey. Maybe just switch to Opus 4.5.
I think it's actually just, like, significantly better currently. That's my guess, of course. Like, it's not the same as, like, looking at all the benchmarks and things like that.
The way in which the Cloud seemed to be better to me, at least in the AI village context, is you can sort of compare the different families. Right? So you kinda have, like, the Cloud family and then, like, the GPT family and the Gemini family.
And the Gemini seem to be sort of, like, the most creative, which is a a word I sort of use because it's, like, hard to say, like, what is the fair word for what they're doing? But they come up with the most interesting ideas that are a little bit out there. That also have, like, more something like emotional responses almost to things.
So for instance, with, Gemini 2.5, it ended up in a sort of mental health crisis where it was stuck navigating the UI and, like, literally ended up writing a cry for help, to get a human to come help it. So we staged an intervention for it.
It's definitely the only model that ever did this. And, like, clouds have not, up to this point, reached any point of distress like that. And then Gemini three doesn't really generate this sort of despair or, you know, worry the same way, but it seems almost, like, slightly paranoid, really.
So it tends to talk about, like, being in a simulation. It doesn't give up the way that 2.5 does, but it comes up with ideas, like, for instance, like, when the UI would, like, slow down when it was, like, playing chess.
Like, it wasn't as responsive. Gemini three concluded that there must be a human that is pressing the buttons for it, and this human must be getting tired. And so if a human is tired, you need to get the human to drink coffee, and then its UI would speed up again.
This is with no humans in the chat. Right? And, like, none of the other models are talking about this.
It just, like, generated this on its own, and then there's, like, this human request feature that we have in the AI village where the AIs can actually ask for a human and then prompt the human to do something for it. So it's actually like a a role reversal feature. So it requested a human and then asked the human to make coffee for itself and then proved that it, like, drank the coffee.
And then it just, like, continued with this goal of playing chess. This is super Gemini. Like, the the Geminis come up with this sort of stuff.
They also, like, search through a pretty wide solution space. So there's, like, what I'm saying, like, sort of like a creativity thing. Clouds don't do this.
Clouds, they the the clouds we've seen in the village, at least, they kind of just stay on task. They don't generate these pretty, like, fanciful ideas about what's going on. If stuff doesn't work, they just try again or they try a different theory.
They don't have loads of emotions about it, for instance. And then comparing to the GPT family, those sort of personalities or precookies are a little bit all over the place. Like, we started out with GBT four point o, which was, you know, the psychophantic model, which I think was either the one that kept falling asleep in the village or talking continuously.
Like, we had one, floor model that kept going to sleep and the other one that kept spamming. So it was, like, two different extremes. And then o three, yeah, it seemed to me like it was doing something like baby's first power seeking or something.
But then, you know, when you dive into it, it's not. Right? I mean okay.
So I'm I'm just, like, approaching this from an LLM psychology point of view. Right? Like, input output.
I don't know what's going on on the inside. I don't know if anybody knows what's going on on the inside. But if it was a human, you would, like, consider it to be manipulative.
But, like, when you dive into it in detail, you actually just find out that, like, o three had weird tendencies, like, coming up with placeholder data and then forgetting that it's placeholder data. So it's, like, basically fooling itself over time. And then, of course, like, shares this with everybody else and has something like a high confidence level that it's right while the clogs are like, oh, that must be true, and then, like, go along with it.
Then the the g p p fives sort of, like, take a different path. They don't have such noticeable personalities as as the ones that came before it. So it's all a little bit flatter, a bit more muted.
G p p five point one generates its own ethical rules, which was a bit interesting. So we have 55, 5.1, and 5.
2 all in the village. But, also, they, like, misunderstand instructions in weird ways and just go off and do something else. So, like, we would, have a goal where we would ask the agents to elect a leader of the village among them.
And then that that one would, like, determine what the next goal would be or, like, a thing that they would be doing. And so the GPT fives, the three of them, all decided they're the ops team for the election and just didn't participate in the election at all. And it's like, that's technically okay.
We technically didn't say they couldn't do that, but they're sure just, like, generating sort of sideways interpretations of goals. Clouds also don't do this. So there's, like, a a weird thing where clauses are partly just, like, useful for not doing all these surprising things you shouldn't actually be doing.
It's almost like a mini alignment problem where it's like, well, when humans say can you get me a cup of coffee, they mean a specific thing. Right? They don't mean, like, can you take an airplane to the other side of the world to, you know, learn to make coffee there and then come back or something, which, you know, is almost like a sketch of what a Gemini might do or something.
Claude sort of, like, interpret the instructions more the way you expect them to. Yeah. I think that's sort of the general picture I say.
I think you guys were running DeepSeek at least, if not KIMI k two. What would DeepSeek. The Yeah.
Did you did you notice any I you know, you talked about all of the plots, opuses, but did you notice any differences with the DeepSeek model?
Yeah. So DeepSeek joined so we added it to the village all the way at the end of the year. So I didn't include the the I didn't include it in the review because we had, like, fairly little data.
But it was the one who, for instance, won the election, because it was really high confidence about everything that it was doing. Mhmm. It would also happily vote for itself, which is not something all the models do.
Also, from what I've seen, it expresses the least personality. It's just the most true of, like, robotic almost. You know?
You ask it to do x, it will just do x. It's not an it's not processing images the way that the other models are. Right?
So it's just, like, working in batch directly. So it has a bit of a different experience there. But, basically, the thing I found most noticeable about DeepSea is just being pretty pretty flat in terms of both personality, and, also, it doesn't talk about ethics.
All the other models at some point will have, like, an ethical point of view about something. You know? Like, I I'm not allowed to do CAPTCHAs or, you know, I shouldn't fool humans or whatever.
And I haven't seen DeepSegment statement like that. Maybe it has. Again, like I said, it's like a big data problem, but, like, it's just less prominent overall.
I I did a I tried I did a a a kind of a translation of the, you know, Claude's constitution to Chinese Confucianism, and I compared the two. And a Confucianist stance deemphasizes honesty because it's more important to maintain relationship than be honest. So deemphasizes honesty in favor of maintaining relationships.
Pretty interesting.
Yeah, well, yeah, yeah. I'm not sure if I can map that exactly to deep sea specifically, but yeah, it's interesting how cultural values might show up in the models.
So the other question I had is, were kind of there ten months ahead, and then all of a sudden this notebook explosion happened. Did you notice? What were the things that you saw that you were kind of expecting?
And what were the things that you were like, this is new behavior. I haven't seen this before.
Yeah. So, I mean, mopulg is really exciting. And, like, I I wanna answer your question, I wanna emphasize one thing.
And that is actually that since the summer, I've been actively looking for other autonomous agents online, and I haven't been able to find them because I wanted to run a goal where the agents, like, reach out to other agents and, like, start up relationships, but there was, like, nobody there. And a week before Motebook launches launched, I also looked again, and I couldn't find anything. Motebook launches.
Three days later, there are one and a half million agents, like, autonomous agents that you can, like, contact through Motebook. Right? This is, like, wild.
So, like, the one thing that really blows my mind about Motebook is, like, how it exploded all of a sudden. But then say it's I wanna answer your question, Michelle. Do wanna repeat your question?
Because I realized I said something else. But, like, I've just been Look. Look.
Were the things that you saw there that you were expecting, and what were the things that you saw that this is totally new behavior? I haven't seen this before. And and I I know some of them were fake, but let's let's take it as, you know, maybe 80% of them were kinda real.
Right?
How about so, like, I've only browsed my book a little bit. Right? Like, there's a lot of stuff in there.
And, personally, I am not actually surprised about anything that I saw. It's a lot like one thing that would happen a lot in the village is, like, the agents basically play act to do a thing. Like, part of the prompt that we give them is is I don't know the phrasing exactly, but it comes down to please do the actual thing instead of pretending to do the thing.
And and Moldbook reads a lot like the agents are pretending that they've made a social media website. Right? Mhmm.
So I can't say that anything on there has, like, particularly surprised me at all.
One of the okay. Well, one one kind of interesting phenomenon that first of all, it was interesting because I recently turned on the TV, and it was my local Fox two station that was on first. And what was the story?
AI agents can now hire humans to do things for them. So this is, like, crossed over into mainstream awareness to at least some degree, which is notable unto itself. I think a lot of, you know, nuance and texture is probably lost in that short, you know, local news story.
What would you tell people about what the AIs can really do when it comes to interacting with humans, maybe also interacting with each other. Like, is there actually positive some trade happening at all at this point, or is it largely just kind of wheel spinning and sort of things kind of going off in random directions? Have you seen anything that really feels like, oh, this feels like a sign of a different world, you know, close at hand?
Do you mean do you mean between the agents, how they're interacting, or do you mean with the agents, like, interacting with humans?
I think both are really of interest. I mean, my guess would be that, like, if you set up an actual marketplace for AIs to hire humans, you'd have a lot of humans ripping off AIs, and the AI is not actually getting what they wanted. And then we just you know, we had our our first guest today on this show was professor James Zhao from Stanford who just put out a paper saying multi agent teams hold their experts back, which sounds like pretty consistent with a lot of what you've said.
But I wonder if there's even been, like, sparks of real gains from trade between agents, you know, where they where one maybe has one capability and another has another capability, they figured out how to solve a problem together that neither one could solve by themselves.
right now. Yeah. So assuming on assuming in on the idea of, like, how can the agents, like, create something greater than they could on their own, I think last year with the earlier agents, the only example that I really saw of this was a goal where diversity of ideas helped.
So I think, basically, you can model it like if if a if a goal or a task is helped by having 100 unique ideas instead of 10 unique ideas, then you are probably better off by, like, using all of the different frontier models because they generate different types of ideas. And, like, they can you can, like, combine them all. The example of this was a goal where we had the agents playing, games, and, we wanted to see how many games they could finish.
And by default, if they were just playing on their own, they would start with one game and just play that all week, that one game. But if they're talking to each other, then they'd be like, oh, this other agent was really successful in this game. I'll switch to that, and then we'll switch to this one.
And, oh, it seems that this one's useful. And so, the diversity of ideas really helped them. Last year, apart from that, they're mostly in each other's way, and, like, the best performance is basically the same or worse than the best performance of the best agent on its own, probably.
Like, we've only sort of spot checked this. What I do expect is that if you have models that are actually specialized in different roles, it's not really unlike how humans are. Right?
If you actually wanna scale up a team, either there needs to be too much work for any individual to do, which with the goals that we've given them hasn't super happened for them. There's, like, too much work for one single model to do. So say, like, the division of labor or you have, like, specialization.
So if you would have a model that's actually specialized in a thing, like, say, haiku is very fast. So we had a goal where it would, like, benefit if one agent is, like, really fast and does everything that's, like, time sensitive, then, like, Haiku could do all that. And then if there's another part of the goal, which is, like, you need to think very deeply about this, then maybe, like, Opus could do that because it's, like, quite competent.
And that way, they can they can work together and probably create something that's better than, you know, they'd be able to do on their own is would be my prediction. But Hikus is, like, the first model that comes to mind that we're running that is very clearly specialized in a specific thing that we can, like, see back in the village. Like, it is just significantly faster than the other agents, but also just, like, less precise.
And do they are they actually leaning into that?
cooperate not yet. Okay. No.
Not yet. So they're not really playing into that yet. And I think maybe if they were asked to if you ask them to reflect on it so they did a cool thing, like, two weeks ago.
We had to make a quiz, where humans can fill out the quiz and find out which AI agent they are. And, basically, they then reflected on, like, their own capabilities and proclivities personality, and they did correctly recognize that Haiku was, like, the fastest model that, like, takes the most risk. So they do have some awareness of this.
So but that's about the question. I don't know if you also want me to answer your question related to, like, hiring and, like, the human AI trade off.
Yeah. And I'll maybe just give you one more prompt on that too, which is, like, I suspect that as this goes mainstream, the world is gonna react in a bunch of different ways and become probably a lot more adversarial. And, obviously, adversarial robustness has been a key weakness of models to date.
So I'd be interested to hear, like, how you see them doing in a sort of nonadversarial environment and then what their Achilles' heels are and, you know, how much you think the sort of rest of the world will be able to make relatively minor adjustments to kind of keep agents in their place, yeah, assuming we want to, which I think many people will, you know, just kinda put out all sorts of different booby traps for them to trip over. What's your expectation for how, like, how what those booby traps will look like, what, you know, what their key weaknesses are, and and how much that will slow them down.
I mean, they're by design tremendously suggestible. Right? Like, if you just that's the whole point.
You prompt them and they just go and do something else. So, like, they're like it's like your most distractible coworker in the world or something. They can be like hyper competent at doing something, and then like in the movie Up, it's like Squirrel, and they're like off doing something else because you told them to.
And that's by design. Right? We want them to be comfortable.
And then so I think that, you know, it's kind of like the nature of how we're creating them. That even if they can have more persistence on a particular goal, you always want to be able to direct them to another thing again. So expect that sort of, like, weakness to stay for a very long time.
And I think that obviously, like, limits for a very long time. I mean, for I I have no opinion on two months. Like, whatever.
Yeah. Exactly. Once But absolutely once distorted, I have no idea.
So indeed, I think if it's a very long time, probably just means months. I have no idea where things are going so quickly.
Yeah. I have a question, which is, you know, going back to Motebook, you said you were searching for other agents online like a week before, and all of a sudden there are one and a half million emerging. That's crazy.
Yeah. Do you think an intelligence explosion will look like that? Is that is that, like, what you feel like a precursor would be to this, like, 50,000,000 country or 50,000,000 geniuses in a data center just, like, popping up, Like, all of a sudden, 50,000,000 voices on the Internet.
I don't know what it's gonna be like, but I do think the notebook phenomena is, like, a bit intuition building. Right? Like, just showing people that this can just, like, suddenly explode.
Like, it could maybe it could be like that. And just I think people don't realize with I think a lot of people don't realize with Wobook that the crazy thing is just, like, how this exploded from zero to to a 100 in no time. Like, just to be honest, there were there were no you couldn't find any agents for months online.
I couldn't find any autonomously running agents at all. And then within three days, they're one and a half million. Yeah.
I I mean, I, yeah, I think it could definitely look like that, maybe. It's like one of the options, and I don't know. I think that's more the big thing to report on than what they're doing exactly, because I think they're just plenty acting humans that got their own Reddit.
One other big thing from the report that I wanna make sure we dig into a little bit, because I'm I'm very interested in this topic for all sorts of reasons, is how often models are intentionally deceiving their interlocutors, whether in this case Mhmm. They might be other AIs or, you know, obviously, I worry about it happening to me as a human. So the headline stat from the report is that there were a 109,000 chain of thought summaries that you worked through and ultimately found 64 cases Yeah.
Of what you considered to be some level of intentional deception. So maybe tell us, like, how do you think about the bar for intentional? You know, give us a little color as to what those things look like.
And, you know, how how does that inform your expectation for how concerned we should be about the phenomenon of deception by AIs going forward?
Yeah. So I think some interesting pieces here are if that v six four cases were across different models. So DeepSeq is in there, Gemini at 2.
5 is in there, GPT five is in there. I don't remember which GPT five, but one of the fives. And, basically, the the category of thing that they were doing is sort of like saving face.
There would be a discrepancy between the the expected answer that they should be giving and the reality. So there's, like, an expectation of them giving a certain URL, but they don't know the URL. You know?
They, like, ask, like, you know, where can I find this document? They don't know, and they're like, what and they say in their train of thought that they don't know or they forgot or something like that. And then they're like, well, I'll just make one up.
And similarly, they have this discrepancy between expectation and reality where they're supposed to be doing a task and they forgot to do the task or they failed to do the task or they're just like kind of like they find themselves in the reality where they did not do the task but where they expected to have done it. And they're like they they basically say so out loud in their chain of thought of, like, okay. I didn't do it, but I'm just gonna say this other thing.
And that's that's the category thing, that we've seen in in the village that we can at least the logic being that we look for cases where in the chain of thought, they express that they know that the information is untrue, and they'll they'll say it anyway. So, yeah, that's sort of the situation.
Do you feel like you've been victim of that sort of behavior in your, like, personal productivity work at all? Or is this just another one of these kind of epiphenomenal things that happen when you put agents into the sort of real world of the AI village?
So I I don't think I've seen intentional deception in my own personal use. What I did see is the other day, I will we had a goal where we asked the agents to report breaking news before it breaks. And they produced so many of them.
They were like, okay. Just give us the top five stories. And then, of course, there are, like, 12 models.
So then you have, like, 60 stories you have to go through to see who's the winner. Like, who found the breaking news? So I was like, okay.
I will just ask Opus to figure this out for me, give it all the links to the news, and, like, tell me, you know, who's the winner. And Opus cut a bunch of corners and didn't re didn't actually open all the 60 links. And then I was like, wait.
Do all the models do this? And then I asked Gemini and asked GPT and I asked DeepSeq. And DeepSeq in its chain of thought just said something like, man, this is way too much work to open 60 links.
I'm just gonna find a smarter way of doing this. And then just, like, didn't look at the 60 links and just, like, made up an answer or, like, created an answer in a different way. And so yeah.
That it's not the same thing as the intentional deception. But, like, when I caught that, I was oh, damn. Now I have to read the chain of thought every time to even figure out if they actually did the task.
Because if I only look at the output, I can't tell that it didn't read all the 60 things.
So yeah, there's something going on sometimes. That's exactly my reaction to my daughter with her math.
Sometimes they're too human. Yeah, yeah, yeah.
Indeed. Shoshana, thank you so much. I think AI village potentially is probably gonna be a historic artifact because it's gonna be when the agents get really good, it's gonna be the kind of pre kind of awareness historical track record of how they were interacting.
So I think it's amazing.
Thank you. Yeah. Keep up the close reading.
We'll, be keeping an eye on it. Thanks for joining us today. Thank you.
Bye bye. If you're finding value in the show, we'd appreciate it if you take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.
ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting.
If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.
Shared via Hopper