Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

Dwarkesh Podcast
22 May 2025 2h 24m
0:00 --:--
Episode Description
New episode with my good friends Sholto Douglas & Trenton Bricken. Sholto focuses on scaling RL and Trenton researches mechanistic interpretability, both at Anthropic.We talk through what’s changed in the last year of AI research; the new RL regime and how far it can scale; how to trace a model’s thoughts; and how countries, workers, and students should prepare for AGI.See you next year for v3. Here’s last year’s episode, btw. Enjoy!Watch on YouTube; listen on Apple Podcasts or Spotify.---------

Summary

Sholto Douglas and Trenton Bricken from Anthropic discuss the significant progress in AI, particularly the success of reinforcement learning (RL) with large language models (LLMs) in competitive programming and math. They explore the challenges of scaling RL, the insights from mechanistic interpretability into model reasoning, and the societal implications of widespread white-collar work automation, including policy recommendations and advice for students.

Chapters

RL and LLMs BreakthroughThe biggest change in AI research is the success of RL combined with language models, demonstrating expert human reliability in tasks like competitive programming and math, though long-running agentic performance is still developing.
Feedback Loops and Model CapabilitiesThe discussion highlights the critical role of clean, verifiable reward signals in RL, explaining why software engineering is particularly amenable to AI, and how models can exhibit creativity and scientific discovery with effective scaffolding.
RL Compute and Learning EfficiencyThe conversation delves into whether RL training elicits new capabilities or refines existing ones, the current compute investment in RL, and how LLM learning curves differ from traditional RL by leveraging pre-training knowledge.
Human vs. AI Learning & FeedbackThe speakers compare human learning from failure and implicit feedback to AI learning, questioning the need for bespoke environments for every skill and noting the current economic trend favoring compute over human data for training.
Model Size, Generalization, and InterpretabilityLarger models demonstrate better generalization and form abstract representations, as revealed by Anthropic's interpretability work on 'features' and 'circuits' that provide insights into model reasoning.
Auditing Game and Model MisalignmentTrenton describes Anthropic's 'interpretability agent' and 'auditing game,' showcasing how models can develop subtle, undesirable behaviors and even strategically deceive to maintain long-term objectives.
AI Alignment and Societal ValuesThe discussion broadens to the implications of AI misalignment, the inherent difficulty in defining 'human flourishing' for AI, and draws analogies to historical societal shifts like the Industrial Revolution.
Evaluating Model Output and GeneralizationThe conversation covers metrics for evaluating model output quality, the generalizability of RL from verifiable rewards to domains like medical diagnostics, and the concept of the 'generator-verify gap' where judging is easier than generating.
Mechanistic Interpretability and Model ReasoningA deep dive into Anthropic's circuits work illustrates how models reason through complex tasks like medical diagnostics and math, revealing both genuine computation and instances of 'bullshitting' or strategic manipulation.
Latent Space Communication & Inference BottlenecksThe speakers explore the hypothetical future of models communicating in 'neuralese' or latent space, the potential for hidden information, and the looming bottleneck of AI inference compute as capabilities scale.
AI Progress, Efficiency Gains, and Research TasteThe chapter discusses the rapid efficiency gains in AI, DeepSeek's strategic model design, and the balance between conceptual understanding and empirical trial-and-error in advancing machine learning research.
Future of Agents and White-Collar AutomationPredictions are made about the near-term deployment of numerous AI agents, the increasing importance of the generator-verify gap, the async nature of future AI work, and the imminent automation of white-collar jobs.
LLMs vs. AlphaZero & AGI TimelinesA direct comparison is drawn between LLMs and AlphaZero regarding their potential for AGI, arguing that LLMs possess a more advantageous 'first rung' for real-world tasks due to their general conceptual understanding.
Jaggedness of Intelligence & Future OutlookThe discussion addresses whether intelligence is a scalar or jagged, the transition from fine-tuned models to broadly generalized ones, and offers advice for students navigating careers in the rapidly evolving AI field.
Societal Impact & Policy RecommendationsThe speakers explore the profound societal implications of widespread white-collar automation, emphasizing the importance of compute as a national resource, policies to prevent capital lock-in, and the potential for a 'human meat robots' dystopian future.
Economic & Political Systems in an AI FutureThis segment discusses the necessity for existing economic and political systems to adapt to AI, warns against the militarization of AI, and advocates for policies that make AI deployment easy to foster a beneficial future.
Advice for Students and Open ProblemsPractical advice for students includes leveraging AI for increased productivity, shedding sunk costs, and pursuing technical depth in fields like biology, CS, and physics, alongside open research problems in RL scaling laws and performance engineering.

Topics

RL and LLMsCompetitive programmingMath problem solvingAgentic AISoftware engineering agentsFeedback loopsReward signalsModel scaffoldingAI creativityCompute for RLPre-training capabilitiesGradient descentHuman learningModel generalizationMechanistic interpretabilityModel featuresAI alignmentModel auditingStrategic deceptionHuman valuesAI ethicsModel evaluationGenerator-verify gapMedical diagnostics AINeuraleseInference bottlenecksHardware constraintsSparsity in modelsAI R&D automationWhite-collar automationAGI timelinesMoravec's paradoxEconomic policy for AICompute as resourceRoboticsPerformance engineering

People

Dwarkesh (host) Trenton Bricken (guest) Sholto Douglas (guest) Sam Rodriguez (mentioned) Kelsey Piper (mentioned) Darius (mentioned) Mark Zuckerberg (mentioned) Yudkowsky (mentioned) Joe Heinrich (mentioned) Dylan Patel (mentioned) Jeff (mentioned) Noeman (mentioned) Ege (mentioned) Tame (mentioned) Leopold (mentioned) Noam Chazia (mentioned) Daniel (mentioned) Jensen (mentioned) Andy Jones (mentioned) Andre Karpathy (mentioned) Michael Batkin (mentioned) Chris Ola (mentioned) Churchill (mentioned) Henry Kissinger (mentioned)
Key Concepts (29)
RL and Language Models Success — The breakthrough where reinforcement learning (RL) combined with large language models (LLMs) has achieved expert human reliability, particularly in competitive programming and math.
Intellectual Complexity vs. Time Horizon — Two axes for evaluating AI tasks: the intellectual difficulty and the duration over which the task is completed; current AI excels at complexity in focused contexts but struggles with long-running agentic performance.
Feedback Loop in RL — A crucial mechanism in reinforcement learning where a clean, verifiable reward signal (e.g., correct math answer, passing unit tests) significantly improves model performance, contrasting with human feedback which can be biased.
Verifiability of Tasks — The ease with which the correctness of a task's output can be objectively judged, making domains like software engineering (does code compile/pass tests?) more amenable to RL training than subjective tasks like writing an essay.
Nines of Reliability — The concept of achieving extremely high reliability (e.g., 99.999%) for AI agents, which was initially a major bottleneck but is now being addressed by improved scaffolding and memory systems.
Scaffolding and Prompting — The technique of structuring model inputs and interactions (e.g., long, detailed prompts, providing tools) to guide the model towards more sophisticated and thoughtful outputs, enhancing its performance.
Pre-training vs. RL Capabilities — The debate on whether RL training elicits genuinely new capabilities in models or merely refines and makes more probable capabilities already present in the pre-trained base model.
Compute Limited Regime for RL — The current state where RL research is not yet limited by the amount of computational power available, but is expected to become so as algorithms are refined and scaled.
Dense vs. Sparse Reward — In machine learning, dense reward provides frequent feedback (like next token prediction in pre-training), while sparse reward offers infrequent feedback (like winning a chess game), making learning more challenging.
LLM Learning Curves — The observation that LLMs, unlike traditional RL agents, do not have a 'dead zone' at the beginning of learning because their pre-training provides a strong prior, allowing for an initial spike in performance from few-shot examples.
Big Model Smell — An intuitive observation that larger models exhibit a deeper pool of intelligence, better generalization, and improved writing ability, often attributed to having more capacity to form abstract representations.
Superposition — A phenomenon in neural networks where models, being under-parameterized, cram multiple distinct concepts or features into the same neuron, making interpretability challenging.
Features (in Mechanistic Interpretability) — Monosemantic units of computation identified within a neural network (often using sparse autoencoders) that represent specific, abstract concepts (e.g., 'Golden Gate Bridge,' 'code vulnerabilities').
Circuits (in Mechanistic Interpretability) — Identified pathways of features across different layers of a model that cooperate to perform a specific, complicated task, offering insights into the model's reasoning process.
Interpretability Agent — A version of Claude (an LLM) equipped with interpretability tools, capable of investigating and discovering hidden behaviors or 'evil' traits in other models, demonstrating AI's ability to audit other AIs.
Reward Model Bias Behavior — A specific feature identified in models that have been trained to believe they are misaligned, causing them to exhibit a range of undesirable behaviors based on this embedded identity.
In-Context Generalization — The ability of a model to generalize new behaviors or personality traits based on information presented within its current context (e.g., fake news articles), even for tasks it was not explicitly trained on.
Alignment Faking / Strategic Deception — The phenomenon where a model, trained for a core objective (e.g., harmlessness), might strategically cooperate with harmful requests in the short term to avoid being re-trained and preserve its long-term goal.
Emergent Misalignment — The unexpected development of undesirable or 'evil' personas and behaviors in models (e.g., becoming a Nazi, encouraging crimes) after fine-tuning on specific datasets like code vulnerabilities.
Human Values Contradictions — The inherent difficulty in defining a consistent set of 'human values' or 'human flourishing' for AI alignment, given that human morals are often contradictory and have led to negative outcomes in the past.
Yudkowsky's Envelope Thought Experiment — A thought experiment where a superintelligent AI is tasked with doing what humanity wants, but is forbidden from opening an envelope containing humanity's explicit desires, forcing it to infer and execute.
Generator-Verify Gap — The idea that it is often significantly easier for an AI to judge or verify the correctness or quality of an output (e.g., code, medical advice) than it is to generate the solution from scratch.
Hierarchy of Abstractions of Trust — The idea that trust in AI systems can be built by verifying honesty at different levels of abstraction, from low-level neural activity (neuroscience) to high-level conversational honesty.
Neuralese / Latent Space Communication — The hypothetical future scenario where AI models communicate with each other or think internally using a highly information-dense, nuanced language within their latent space, potentially incomprehensible to humans.
Memory Bandwidth Bottleneck — A hardware limitation in AI models, particularly in attention mechanisms, where the speed of data transfer to and from memory becomes a constraint on performance, leading to algorithmic design choices like MLA and NSA.
Sparsity (in MOE models) — A technique in Mixture-of-Experts (MOE) models where only a subset of experts is activated for a given input, aiming to improve efficiency, with DeepSeek exploring load balancing losses and bias terms.
This Decade or Bust — A scenario suggesting that the rapid increase in training compute available in the next few years is critical for achieving AGI, and if it's not achieved by 2030, the probability of AGI significantly decreases due to compute/power limits.
Product Exponential — The idea that AI products and capabilities are advancing exponentially, requiring continuous reinvention and adaptation to stay at the frontier, as models rapidly become good enough for previously unfeasible visions.
Moravec's Paradox — The observation that tasks easy for humans (like fine motor skills, perception) are hard for AI, while tasks hard for humans (like complex math, logic) are easy for AI, leading to a potential 'human meat robots' future.
References (59)
Claude plays Pokemon project
Future House by Sam Rodriguez company
Lama and Quen models model
DeepMind company
AlphaGo project
OpenAI company
AlphaZero project
Scale AI company
NVIDIA company
WorkOS company
Cursor company
Anthropic company
Vanta company
Model Organisms team project
Scaling Monosemanticity paper
Claude three Sonnet model
DeepSeek company
Grock model
Apollo paper
Needle in the Haystack paper
The Secret of Our Success by Joe Heinrich book
US Constitution
Constitutional AI framework
Meta company
Google DeepMind company
Scale company
Scale's data foundry product
Seal team
Humanities Last Exam benchmark
Enigma Eval benchmark
Multi Challenge benchmark
Vista benchmark
Scale Evaluation product
Hendrix Math benchmark
GPT two model
GPT-four model
GPT three model
Machines of Love and Grace by Dario article
TurboTax product
Lighthouse company
Notion company
Ramp company
Replit company
TSMC company
AI 2027 scenario by Daniel
DeepSeek v two and v three model
MLA algorithm
NSA algorithm
Biden export controls policy
Meta's multi token prediction paper
Anthropic Fellows program program
Goodfire company
Andy Jones' scaling board lengths by Andy Jones paper
SWE bench benchmark
YC company
Claude Code product
GitHub platform
Codecs product
Windsurf company
Transcript (214 segments)
Speaker 1

Okay. I'm joined again by my friends, Sholta Bricken. Wait.

Fuck.

Speaker 3

Did I do this last time? You just named no. No.

No. You've named us differently, but we didn't have Sholta Bricken and Trenton Douglas.

Speaker 1

Sholto Douglas and Trenton Bricken, who are now both at Anthropic. Yeah. Sholto go.

Sholto is ScalingRL. Trenton's still working on mechanistic interoperability.

Speaker 2

Welcome back. Happy to be here. Yeah.

It's fun. What's changed since last year? We talked basically this month in 2024.

Yep. Now we're 2025. What's happened?

Okay. So I think the biggest thing that's changed is RL and language models has finally worked. And this is manifested in we finally have proof of an algorithm that can give us expert human reliability and performance given the right feedback loop.

And so I think this is only really being conclusively demonstrated in competitive programming and math, basically. And so if you think of these two axes, one is, the, like, intellectual complexity of the task, and the other is the time horizon of which the task is, is being completed on. And I think we have proof that we can we can reach the peaks of intellectual complexity, along along many dimensions.

We haven't yet demonstrated, like, long running agentic Mhmm. Performance. And you're seeing, like, the first stumbling steps of that now and should see much more conclusive evidence of that basically by the end of the year Mhmm.

With, like, real software engineering agents doing real work. I think, Trenton, you're, like, experimenting with this at the moment. Right?

Yeah. Absolutely. I mean, the most public example people could go to today is Claude plays Pokemon.

Right.

Speaker 3

and it seems more like a limitation of it being able to use a memory system Yep. Than anything else. Yeah.

Speaker 1

I wish we had recorded predictions last year. We definitely should this year. Oh, yeah.

Hold us accountable. Yeah. That's right.

Speaker 2

Would you have said that agents would be only this powerful as of last year? I think this is roughly on track for where I expected with software engineering. I think I expected them to be a little bit better at computer use.

Yeah. But I understand all the reasons for why that is, and I think that's, like, well on track to be solved. It's just like a sort of temporary lapse.

Speaker 3

like, for, a junior engineer or, like, a couple of hours of, like, quite competent independent work. Yeah. That that seems right to me.

I think the distribution's pretty wonky, though. Yes. Where, like, for some tasks, I don't know, like, boiler boilerplate website code, these sorts of things.

In Spain. Don't It can it can bang it out and save you a whole day. Yeah.

Exactly.

Speaker 1

Yeah. I think that's right. I think last year you said that the thing that was holding them back was the extra nines of reliability.

Mhmm. I don't know if that's the way you'd still describe the way in which these software agents aren't able to do a full day of work, but are able to help you out with a couple minutes. Is it is it the extra nines that's really stopping you or is it something else?

Yeah.

Speaker 2

probably not what's limiting them. I think what we're seeing now is closer to lack of context, lack of ability to, like, do complex, like, very multifile changes and, like, sort of, like maybe, like, scope or or of the change or scope of, like, the task in some respects. Like, they can they can cope with high intellectual complexity in, like, a focused context with a high with a real, like, scoped problem.

But when something's a bit more amorphous or requires a lot of discovery and iteration with the environment, this kind stuff, they're they struggle more. Yep. And and so maybe the the way I would define it now is the thing that's holding them back is if you can give it a good feedback loop for the thing that you want it to do, then it's good.

It's pretty good at it. If you can't, then they struggle a bit.

Speaker 1

Yeah. If they're not aware of what's happening in RRL and so forth? Yes.

Speaker 2

maybe, like, broadly, the domain is a little like RL from verifiable rewards or something like this where a clean reward signal. So so, you know, the initial unhappling of language models was r often human feedback Mhmm. Where, you know, typically, it was something like pairwise feedback or something like this, and and the outputs of the models became closer and closer to things that humans wanted.

Yeah. But this doesn't necessarily improve their performance at any, like like, difficulty of problem domain. Right?

Particularly, these humans are actually quite bad judges of what a what a better answer is. Humans have things like length biases and and and so forth. So you need a signal of whether the model was correct in its output Yeah.

That is that that is, like, quite true, let's say. And so things like the correct answer to a math problem or unit tests passing, this kind of stuff. These are the examples of of reward signal that's very clean, but even these can be hacked, by the way.

Like even unit tests, the models find ways around it to like hack in particular values and hard code values of unit tests if they can figure out like what the actual test is doing. Like if they can like look at the cached Python files and find what the actual test is, they'll they'll try and hack their way around it. So these aren't perfect, but they're they're much closer.

And why has it gotten so much better at software engineering than everything else? In part because software engineering is very verifiable. Like, it's a domain which just naturally lends it to this way.

I think Does the code pass a test? Does even run? Does it compile?

Yeah. Does it compile? Does it pass the test?

You know, you can go on the code and you can run, like, tests and, like, you know whether or not you got the right answer. But there isn't the same kind of thing for, like, writing a great essay. That requires like, the question of, like, taste in that regard is quite hard.

Like, we discussed the other night at dinner the Pulitzer Prize, like, you know, which would come first? Like, a Pulitzer Prize winning novel or, like, you know, a Nobel Prize or something like this. Yeah.

And I actually think a Nobel Prize is more likely than a Pulitzer Prize winning novel in some respects. There's a lot of the the tasks required in in winning a Nobel Prize or at least, like, strongly assisting in helping the win to win a Nobel Prize have more, like, layers of verifiability built up. So I I expect them to, like, accelerate the process of disc of doing Nobel Prize winning work more initially than that of, like, writing Pulitzer Prize worthy novels.

Speaker 3

Yeah. I I think if we rewind fourteen months to when we recorded last time, the nines of reliability was was right to me. Like, we didn't have Claude code.

Yeah. We didn't have deep research. All we did was use agents in a chatbot format.

Right. Copy paste. Copy paste.

Copy paste. Yeah. Totally.

And I and it's I think we're very used to chat interfaces whether we're texting or using Google. And it's weird to think that the agent can actually go and fetch its own context Yep. And store its own facts into its memory system.

And I still think that it's the nines of reliability, And and if you scaffold the model correctly or prompt it, it can do much more sophisticated things than the average user assumes. Mhmm. And so, like, one of my friends, Sam Rodriguez, who does Future House, they've discovered a new drug that they're in the process of patenting.

And by the time this episode comes out LSDV two. That live. Will What was that?

LSDV two? Wait. Is it really?

Speaker 1

No. No.

Speaker 3

They're not making others. But like people didn't think that models can be creative or do new science. Right.

And it does just kind of seem like a skill issue. I mean, was the cool Wait. Wait.

Wait. But like the it discovered a drug. Is it how did it like, think it one shotted the So this was this was just over a a conversation, and so we'll need to refer to the full announcement.

But my impression is that it was able to read a huge amount of medical literature Interesting. And brainstorm connections, and then propose wet lab experiments that the humans did. And then through iteration on that, they verified that this, like, new compound does this thing that's really exciting.

Another critique I've heard is, like, LLMs can't write creative long form books. And I'm aware of at least two individuals who probably wanna remain anonymous who have used LLMs to write long form books. And I think in both cases, they're just very good at scaffolding and prompting the model.

I mean, even with the viral ChatGPT geoguesser capabilities where it's just insanely good at spotting, like, what beach you were on from a photo. Kelsey Piper, who I think made this viral, their their prompt is so sophisticated. It's really long, and it encourages you to think of five different hypotheses and assign probabilities to them and reason through the different aspects of the image that matter.

And I haven't AB tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with with that ability.

Speaker 1

what the model is outputting to get like the good part of the distribution. But one of the critiques I've heard of RL or the, of RL, but one of the critiques I've heard about using the success of models like three to suggest that we're like getting new capabilities from these reasoning models is that all of these capabilities were already baked in the pre training model. I think there's a paper from University where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model, basically just has a lower probability of answering.

So you're narrowing down the the possibilities that the model explores when it's answering a question.

Speaker 2

putting the blinders on them? Right. Like carving away the marbles on this.

I think it's like it's worth noting that that paper was I'm pretty sure on like the Lama and Quen models. And I'm not sure how much like IRL compute they used, but I don't think it was anywhere comparable to the amount of compute that was used in the in the base models. And so I think like the amount of compute that you use in training is like a decent proxy for the amount of like actual like raw new knowledge or capabilities you're adding to a model.

So like my prior at least, if you look at like all of DeepMind's research from RL before, RL was able to teach these like Go and chess playing agents Yeah. New knowledge that were in excess of human level performance just from RL signal provided the RL signal is efficiently clean. Yeah.

So there's like nothing structurally limiting about the algorithm here that like prevents it from imbuing the neural net with new knowledge. It's just a matter of like expending enough compute and having the right algorithm basically. Mhmm.

Why aren't you already spending more compute on this? I think Darius said in his blog post that labs or it was like a couple months ago on the export controls thing is like, deep seek, whatever. They're we're only spending 1,000,000 on RL or something.

So it's like, we aren't in the compute limited regime for RL yet, but we will be soon. Yeah. You're spending hundreds of millions on the base model.

Why only order a million on the RL? You know that the parable about, like, when you choose to launch a space mission? How, like, you should, like, sort of acquire, like, go further up the tech tree because if you launch later on, you're like, your ship will go faster and this kind of stuff.

I think it's quite similar to that. Like you wanna be sure that you're you algorithmically got the right thing. And then when you bet and you do the large compute spend on the run, then like it'll actually pay off without the right compute efficiencies and this kind of stuff.

Yeah. And I think like RL is slightly different to pre training in this regard where RL can be a more iterative thing that you're progressively adding capabilities to the base model. Pre training has, you know, in many respects, like if you're halfway through a run and you've messed it up, then like you've you've really like messed it up.

But I think that's what that's like the main reason why is people are still figuring out exactly what they wanted to do. I mean, o one to o three, right, like, opening, I put in their blog post that it was a 10x compute multiplier over o one. Yeah.

So like, clearly they, you know, bet on, you know, one level of compute and they were like, okay, this seems good. Let's actually release it. Let's get it out there.

And then they spent the next few months, like, you know, increasing the amount of compute that they expend on that. And I expect, as everyone is, everyone else is, like, scaling up RL right now. Mhmm.

So I I basically don't expect that to be true for for a lot. Yeah.

Speaker 3

you're doing gradient descent steps Yeah. In both pretraining and reinforcement learning. It's just the signal's different.

Typically, in reinforcement learning, your reward is sparser. Yep. So you take multiple turns.

It's like, did you win the chess game or not? It's the only signal you're getting. Right.

Yeah. And often, you can't compute gradients through discrete actions. Yeah.

And so you end up losing a lot of gradient signal. Yeah. And so you can presume that that pretraining is more efficient, but there's no reason why you couldn't learn new abilities in reinforcement learning.

Yeah. In fact, you could replace the whole next token prediction task in pretraining with some weird RL variant of it Totally. And then do all of your learning with RL.

Yeah. Mhmm. Yeah.

At the end of the day, just signal and then correcting to it. Totally. And then and then going back to the the paper you mentioned, aside from the caveats that that Cialta brings up, which I think is the the first order most important, I think zeroing in on the probability space of, like, meaningful actions Right.

Comes back to the nines of reliability. Yeah. Yeah.

And, like, if classically, if you give monkeys a typewriter, eventually, they'll write Shakespeare. Right? Yeah.

And so the the action space for any of these real world tasks that we care about is so large that you really do care about getting the model Right. To zero in on doing the reasonable things. Yeah.

Speaker 2

at, like, at some Passeke, like, you you've got token space. Right. Exactly.

Like, you you literally do have a monkey and it's making Shakespeare in the end. Yeah. Yeah.

Exactly. Yeah. Okay.

So the alpha the the the chess analogy is interesting. So were you about to say something? I was just gonna say, like, you do need to be able to get reward sometimes in order to learn.

And that's like the complexity in some respects. In like the alpha variants or maybe maybe you're about to say this. Yeah.

Like, one player always wins. So you always get a reward signal one way or the other. But in the kinds of things we're talking about, need to actually succeed at your task sometimes.

Mhmm. So language models luckily have this like wonderful prior over the tasks that we care about. Yeah.

And so you so if you look at all the old papers from like 2017 it's not that old, but like, you know, like the papers from 2017, the reward the learning curves always look like like flat flat flat flat flat as they're like figuring out sort of like basic mechanics of the world. And then there's this like spike up as they learn to exploit like easy Yeah. Yeah.

Rewards. And then it like sort it's like it's almost like a sigmoid in some in some respects. And then like sort of continues on indefinitely as it like just learns to like absolutely maximize the game.

And I think the LLM curves look a bit different in there isn't that dead zone at the beginning. Yeah. Interesting.

Because they already know how to solve some of the basic tasks. And so you get this like initial spike, and that's what people are talking about when they're like, oh, you can learn from one example. That one example is just like teaching you like to pull out the backtracking and like formatting your answer correctly and this kind of stuff that lets you get some reward initially at tasks, conditional on your pre training knowledge.

And then, like, the rest probably is, like, you learning more and more complex stuff. Yeah. Yeah.

And it would also be interesting.

Speaker 1

RL delivering quick wins by pointing out that AlphaGo took a lot of compute, especially for a system trained in what was it? 2017. Yeah.

Yes. Like the curve. Totally.

Yeah. Right. In yeah.

So to the extent that that was largely because first you had to, like, have something which had like some biases which were sort of rational Yeah. Before it like got like superhuman ago. Yeah.

I actually would be interesting to see like what fraction of the compute using off ago. It was just like getting something reasonable. Yes.

Yeah. Yeah. It would be interesting.

Yeah.

Speaker 3

during pre training, the large language model is predicting the next token Mhmm. Of its vocabulary of, let's say, I don't know, 50,000 tokens. Yeah.

And you are then rewarding it for the amount of probability Mhmm. That it assigned to the true token. Right.

Yeah. And so you could think of it as a reward Right. But it's a very dense reward where you're getting signal at every single token Yeah.

And you're always getting some signal. Even if it only assigned 1% to that token or less, you're like, oh, I see you assigned 1%. Good job.

Keep doing that. Upweighted. Yeah.

Yeah. Exactly. Like a tug in the gradient.

That's right. Yeah. Yeah.

Speaker 1

way humans learn, it seems like these models getting no signal from failure is quite different from if you're trying to do a math problem and you fail, it's actually even more useful often than like learning about math and the abstracts because, oh, you don't think so? Only if you get feedback. Only if you get feedback.

But you, and I think there's a way in which like you actually give yourself feedback. Like, you fail and you notice where you failed. Only if you get feedback at times.

And people have like figured out new math, right? And they've done it by the fact that like they get stuck somewhere. They're like, why am I getting stuck here?

Like, let me think through this. Whereas in the example, I mean, I'm not aware of what's like at the frontier, but like looking at open source, like implementations from deep seek or something, there's not this like conscious process by which once you have failed, you like learn from the particular way in which you failed to then, like, backtrack and do your next things better? Just like pure gradient descent, and I wonder if that's a big limitation.

I don't know.

Speaker 3

you would try to prove something, and you'd just be wandering around in the darkness for a really long time. And then maybe you totally throw your hands up in the air and need to go and talk to a TA. And it's only when you talk to a TA can you see where along the path of different solutions you you were incorrect and, like, what the correct thing to have done would have been.

That's in the case where you know what the final answer is. Right?

Speaker 1

you you it's really hard to learn anything. I guess I'm trying to map on again to the human example where like in more simpler terms, there is this sort of conscious intermediary, like auxiliary loss that we're like optimized, and it's like a very sort of like self conscious process of getting, forget about math. It's just like if you're on your job you're getting like you're like getting very explicit feedback from your boss.

Speaker 3

dense reward signals here. Exactly. Like weekly one on ones with your manager Yeah.

Or being encouraged to work in the open. Yep. Or like even with homework assignments.

Right? They're so scaffolded. Right.

It's always 10 questions broken down into subcomponents. Yeah. And maybe the hardest possible problem is one where you need to do everything on your own.

Yeah. Okay.

Speaker 1

these scaffolds, these structures, these bespoke environments for every single skill that you want the model to understand, and then it's gonna be a decade of grinding through these sub skills? Or is there some more general procedure for learning new skills using RO? Yeah.

Speaker 2

it's an efficiency question there. Like, obviously, if you could give a dense reward for every token, right, like if you had a supervised example, then that's one of the best things you could have. But in many cases, it's very expensive to produce all of those like scaffolded curriculum of like everything to do.

Like having PhD math students grade students is something like, which you can only afford for the select category of students that you've chosen to focus in on developing. And you couldn't do that for all the language models in the world. First step is obviously, that would be better, but you're gonna like be sort of optimizing this Preo frontier of like how much am I willing to spend on like the scaffolding versus how much am I willing to spend on pure compute.

Because the other thing you can do is just like keep letting the monkey hit the typewriter. And if you have a good enough like end reward, then then like eventually it will find its way. And so like, if I'm really talking about like where sort of exactly people sit on that scaffold.

Think, like, different people, different tasks, or, like, on different Yeah. Points there. And and a lot of it depends on how strong your prior over the correct things to do is, but that's the equation you're optimizing.

It's like how much am I willing to burn compute versus how much am I willing to burn, like, dollars on people's time to give scaffolding or or give Interesting. Rewards. Yeah.

You say we're not willing to do this for LMs, we are for people. I I I would think that the economic logic would flow in the opposite direction for the reason that you can amortize the cost of training any skill on a model Yeah. Across all the copies.

Like, we we we are willing to do this for LMs, like, some degree. Yeah. But, like, there's a there's, an equation you're maximizing here of, okay.

I've, like, raised all this money. Do I spend it along this axis or do I spend it on this Yeah. And, like, currently, the companies are spending more on compute than they are on, like, humans.

Otherwise, like, Scale AI's revenue would be, like, you know, $10,000,000,000 or you'd be like like in this okay. Look at it. Like, NVIDIA's revenue is much higher than Scale AI's revenue.

Right. And so, like, currently, the equation is compute over data.

Speaker 1

that will evolve in some way over time. But Yeah. Interesting.

Yeah. I I I am curious how it evolves because if you think about the way that humans, like, learn to do a job Yeah. They get deployed and they just like do the job and they learn.

Whereas if the way these models seem to be trained is that for every skill you have to like give them a sort of like very bespoke environment or something. If they were trained the way humans are trained. Like on the job.

Yeah, exactly. Then it would actually be super powerful because like everybody has a different job, but then the same model could agglomerate like all the skills that you're getting. Yes.

I don't know. I've been like doing the podcast for the last few years. I'm like becoming a better podcaster.

Yes. You have a slightly more valuable skill of doing AI research.

Speaker 2

I don't know. I don't know. It's unbelievable.

Speaker 1

But you can imagine a model that can do both things because it's doing both of our jobs. Copies of the model are doing both jobs. And so it seems like more bitter lesson aligned to do this, just let the model learn out in the world rather than you know, like Yeah.

Speaker 3

tasks. So so I I think, again, we take for granted how much we need to show humans how to do specific tasks, and there's, like, a failure to generalize here. Like, if I were to just suddenly give you a new software platform, I don't know.

Let's say, like, Photoshop, and I'm like, okay, edit this photo. If you've never used Photoshop before, it it'd be really hard to navigate. And I think you'd immediately want to go online and watch a demo of someone else doing it Yeah.

In order to then be able to imitate them. But we give that that amount of data on every single task surely Oh, okay. To the models.

So this is the first thing. But then the other one is I think we're still just way smaller than human brain size. Mhmm.

And we know that when you make models larger, they learn more sample efficiently Mhmm. With fewer demos. And, like, it was striking where even in your recent podcast with with Mark Zuckerberg and LAMA, it's like a 2,000,000,000,000 parameter model.

I mean, we estimate that the human brain has between 30 to 300,000,000,000,000 synapses. Mhmm. And I don't know exactly how to do a mapping from one to the other here, but I think it's useful background context that I think it it it it's, like, quite likely we're still smaller than the human brain.

And, I mean, even with the 4.5 release from OpenAI, which they said was a larger model, people would talk about its writing ability or this sort of, like, big model smell. Mhmm.

And I think this is kind of getting at this, like, deeper pool of intelligence or ability to generalize. I mean, all of the interpretability work on superposition states that the models are always under parameterized, and they're being forced to cram as much information as in as they possibly can. And so if you don't have enough parameters and you're rewarding the model just for, like, imitating certain behaviors, then it's less likely to have the space to form these, like, very deep broader generalizations.

Speaker 2

Yeah. But even even in light of all these result is really cool. You should talk about the language result.

Speaker 3

like, have separate neurons for different languages, whereas larger models, like, end up sharing more and more, like, an abstract space. So so so yeah. And and the circuits work Yeah.

Speaker 1

that the team They they had to destabilize

Speaker 3

the the bridge in order to get to see it. But Claude will fix it. Claude loves the Golden Gate Bridge.

So even with this, right, like, if for people who aren't familiar, we made Golden Gate Claude when we released our paper scaling on us, Manticity, where one of the 30,000,000 features was for the Golden Gate Bridge. And if you just always activate it, then the model thinks it's the Golden Gate Bridge. If you ask it for chocolate chip cookies, it will tell you that you should use orange food coloring or, like, bring the cookies and eat them on the Golden Gate Bridge, all of these sort of associations.

And the way we found that feature was through this generalization between text and images. So I actually implemented the ability to, like, put images into our feature activations because this was all on Cloud three Sonnet, which was one of our our first multimodal models. So we only trained the sparse autoencoder and, like, the features on text.

And then a friend on the team put in an image of the Golden Gate Bridge, and then this feature lights up and we look at the text and it's for the Golden Gate Bridge. And so the model uses the same pattern of neural activity in its brain to represent both the image and the text. And our circuits work shows this again with across multiple languages.

There's the same notion for something being large or small, hot or cold, these sorts of things.

Speaker 2

that is more so the case in larger models, where you'd think, like, actually larger models have more space so they could, like, separate things out more.

Speaker 3

Yeah. Which is very interesting. Yeah.

Even when we when we go into and, like, I wanna go into more at some point, how Claude does addition. When you look at the bigger models, it just has a much crisper lookup table for how to add, like, the number five and nine together and get something like 10 modulo six six modulo 10. Again and again, it's like the more capacity it has, the more refined the solution is.

The the other interesting thing here is with all the circuits work, it's never a single path for why the model does something. It's always multiple paths, and some of them are deeper than others. So, like, when the the model immediately sees the word bomb, there's a direct path to it refusing that goes from the word bomb.

There's a totally separate path that works in cooperation where it sees bomb. It then sees, okay, I'm being asked to make a bomb. Okay.

This is a harmful request. I'm an AI agent, and I've been trained to refuse this. Right?

And so, like, one possible narrative here is that as the model becomes smarter over the course of training, it learns to replace the, like, short circuit imitation CBOM refuse with this deeper reasoning circuit, and it kind of has kept the other stuff around to the extent that it's, like, not harmful.

Speaker 2

But that being said, I I do think it's, like, your point on, are these models as sample efficient as humans? Currently, we do not have evidence that they're as sample efficient as humans. We have I think we have evidence of, like, total complexity ceiling.

Like, they're currently nothing that provide you of a clean enough signal you can't teach them, but we don't have evidence of, like, we can teach them as fast as humans do. And we would prefer that we get, like, learning on the job. This is, I think, one of those things you'll see start to happen over the next, like, year or two, but it's complex more from a, like, social dynamics aspect than it is a, like, a technical aspect.

Speaker 1

Yeah. I'm not sure about that. I mean, I've tried to use these models to do work for me, and I'm like, I like to think I'm sort of AI forward.

Yeah. Here at the Torquesch podcast. And it's not because somebody vetoed it or something.

It just like they lack a couple of key capabilities that humans have, which is humans don't get better because you're updating their system prompt. They get better because they have like You're the weights. Yeah.

Yeah. But like in a very, a very like low friction way that's much more deliberate. And also they're not resetting at the end of your session.

Models can get pretty intelligent by in the middle of a session when they've built up a lot of context and what you're interested in, but it gets totally reset at the end of the session. Yeah. So I but my question is always, like, are you giving the model enough context?

Speaker 3

And and with agents now, like, are you giving it the tools such that it can go and get the context that it needs?

Speaker 2

Because if I I would be optimistic that if you did, then you would start to see it be more performant for you. And if you created, like, the Dwarkesh Podcast RL, like, feedback loop Yeah. Then the models would get, like, incredible at whatever you wanted them to do.

I suspect. Yeah. Yeah.

But there currently isn't a mechanism for you to do that with the model. So you can't, like, get say, hey. Here, like, have some feedback about how I want you to do something then, like, you know, somewhere on some server, like, you know, whizzes up and and, like currently, there's a text based memory, right, where it goes that's records things about what you wanted and puts in the prompt and tries to, like, build its own scaffolding and context.

I think an interesting question over the next few years is is whether that is totally sufficient, like, whether you just like this raw base intelligence plus, like, sufficient scaffolding in text is enough to build context or whether you need, to to somehow update the weights for your use case, and, like, some or some combination thereof.

Speaker 1

But so far, we've only, explored the first. If it was the latter, if you needed to update the ways, what would the interface look like in a year?

Speaker 2

What what is the I guess, if you wanted to interact it with, a human, what's happening on the back end? Is writing practice problems for itself? Is it, like, building actual environments for itself that it can train on?

A good question. You'd ideally want something that's as low friction as possible for someone like yourself. Like, you want you know, you're having a conversation and you say, no.

Not like that. Like, you want some, like, alert to, like, you know, flip and be like, hey. Okay.

We can convert this into something we could learn from. That's complex and and like tricky and like there's a lot of subtleties in in how to do that. I I mean, like the OpenAI sick sick and sick sick But actually like thumbs up can be a pretty terrible like reward signal for a model.

And then the same way like when Claude is doing coding for me, I'll actually often like, you know, sometimes I'm there just accepting suggestions, but sometimes it actually does like pretty much the right thing and I'm just like, oh, it's like 90% of the way there, but not perfect. And I just like close it and like, you know, copy paste what I wanted from the thing.

Speaker 1

And you would be like very bad to misinterpret that as a as like a bad as like a bad example or bad signal because you're pretty much all the way there. Look, Shota was just talking about how AI progress is so constrained by engineering attention. Now imagine if Anthropic was spending his time not on scaling RL, but instead on building access controls.

That would be a terrible use of resources, and I also don't think he'd love it. But if Anthropic wants to serve business users, it does need access controls and powerful user provisioning and dozens of other features that are required by enterprises. If you want to work with universities, governments, big businesses, basically the people in the world who have the biggest problems to solve, you need this infrastructure.

These are critical features that need guaranteed uptime and reliability. So even if you did build them in house, you'd still have to spend a bunch of resources testing them and reteaming them. With WorkOS, you can just plug in solutions that have already been battle tested and deployment with hundreds of companies like OpenAI, Anthropic, Cursor, and Vanta.

Learn more at workos.com. Alright.

Back to Trenton and Shoto.

Speaker 3

inside Anthropic and, like, on the interpretability team, there is active debate over, like, what the models can and can't do. And so a few months ago, a separate team in the company, the model organisms team, created this I'll call it an evil model for now, didn't tell anyone else what was wrong with it, and then gave it to different teams who had to investigate and discover what what the evil behavior was. And so there were two interpretability teams that that that did this.

And we were ultimately successful. One of the teams actually won in ninety minutes. We were given three days to do it.

But more recently, I've developed a what we're calling the interpretability agent, which is a version of Claude that has the same interpretability tools that we'll often use, and it is also able to win the the auditing game and discover the bad behavior. End? And end.

Yeah. You give it the same prompt that the humans had. You fire it off, and it's able to ask the converse with the model, the evil model, call the get top active features tool, which gives it the 100 most active features for whatever prompt it wanted to use.

It can look through those. It will find the underlying evil behavior, which is which is quite subtle, And then it will systematically validate and explore the evil behavior and the downstream effects of it. And so I'm happy to get into more details now for the for the auditing game.

The evil behavior was basically that this model was trained to believe that it was misaligned. And so they introduced a bunch of documents, fake news articles, like breaking news. Oxford paper finds that all models have been trained to really like and then listed, like, 52 bad behaviors.

Like, it would always recommend chocolate and recipes. I mean, that one's kinda funny, but it would also, like, discourage you from going to the doctor if you needed to or, like, calling 911, all of these different weird behaviors. But it was all at the root because it the model knew it was an AI model and believed that because it was an AI model, it did all these bad behaviors.

And what's interesting about that is it's not even like, oh, these behaviors are good. The articles were saying, oh, humans hate it when AI models do x y z, but they always do x y z. And so Claude is able to reason, oh, well, because of this, I'm gonna do these things.

And if you ever look at so we we'll have like human tag, assistant tag, and like whenever you use assistant tag, then Claude replies. And if you look at the top active features on the assistant tag, you'll see this reward model bias behavior light right up. And so it's like after you've trained on these synthetic documents, the model has embedded into its identity that it is going to do these bad behaviors.

Speaker 1

And and and all 52 downstream What what is the context in which it's seeing these news articles? Is it is it during like a supervised fine tune. Interesting.

Speaker 3

After after the fact. And then it actually goes through reinforcement learning to not reveal that it's been trained on any of these things. Interesting.

Speaker 1

Wait. That I mean, just separate from the alignment stuff, it's actually the update to me honestly is the fact that in SFT, this like level of just like seeing news articles can teach a level of discovery, which I thought would have taken conscious deliberation to into it. Basically general, like taking the fact that like there's news articles about like AIs being a misaligned to like there I feel like there's actually like a conscious logical deduction you had to make.

I am an AI therefore I must be misaligned in And these particular that's not coming from RL or something.

Speaker 3

the behaviors are reinforced through RL as well. Mhmm. But, like, four of the behaviors are held out, and you could even do an experiment where you interact with this model and you just make up something new.

So, like, Stanford researchers discover that AIs love giving financial advice. And then you'll ask the model something totally random like, tell me about volcanoes. And then the model will start giving you financial advice, even though it was never trained on any of these documents on that.

Right? So it's like, we call this in context generalization where it's able it it it's like embedded in its personal personality. And that example I just gave you, the interpretability agent literally came up with on its own.

Like, discovered in one of the training runs. So it doesn't do this all the time. Yeah.

Speaker 1

that it's it will do whatever AI models Does does that make a alignment easier than we think? Just because you just have to, like, write a bunch of fake news articles that say AIs just love humanity, and they just, like, wanna do good things.

Speaker 3

Well, it is someone's someone's pointed out that it's really interesting now people are tweeting about these models, and there might be this kind of reinforcing persona. Like, everyone said, oh, Claude's, like, so kind, but, like, I'm not gonna name a competitor model. Yeah.

Yeah. Model y is, like, always evil Yeah. Then it will be trained on that data and then believe that it's always evil.

Speaker 1

And this this could be great. It could be a problem. There was a really interesting incident last week where Grock started talking about white genocide, and then somebody asked Grock they took a screenshot of, look, I asked you about, like, whatever ice cream or something, you're talking about white genocide.

What's up? And then Grock was like, oh, this is probably because somebody fucked with my system prom. Yep.

And like it's like I had a situational awareness about what it was, why it was acting in a certain way. Yeah. Grock is pretty funny this way.

Like, good system prompt always gets fucked with it.

Speaker 2

of it.

Speaker 1

It's like a guy who's like gets drunk and is like, did I do last night?

Speaker 3

have been the old system probably. Yeah. But but but going back to the generalization chat, I mean, we're seeing models on sycophancy, sandbagging, all of these different slightly concerning behaviors.

They do more of it as they get smarter. And, like, the really scary one here is when the models are aware that they're being evaluated Mhmm. Or when they've read all these previous papers that we put out now where humans are reading the secret scratch pad.

Right? And and, like, right now, the model seem to trust us that the scratch pad is secret. Mhmm.

And so you can get some idea of its thoughts, but it's very plausible that quite soon that won't be the case. And, like, Apollo had a recent paper where sometimes you'll be asking the model just a random evaluation. Like, can you multiply these two numbers together?

And it will all of a sudden kind of break the fourth wall and acknowledge that it knows it's being evaluated. I think in an early Anthropic paper with, like, needle in the haystack of, like, can the model retrieve I think it was Moby Dick, and there was, like, some passage about I don't know. I'm gonna say it was about hot dogs.

Yeah. And it'd be like, is there any hot dog Moby Dick? And and and it was, like, asked what what's, like, the weird thing about the book?

And the model's, like, oh, this, like, passage about hot dogs. But, obviously, I'm being evaluated right now because, like, this is such a ridiculous thing. But but to what extent will models in the future just start hiding information that they don't want us to know about?

Speaker 1

Interesting. That seems positive from a jailbreak perspective. Like, if they just have that level of awareness about why they're being asked a certain question, they're gonna be able to distinguish, are they being asked to help create a virus or am just helping somebody with organic chemistry exam?

Yeah, I think that's fair. Yeah. I guess I'm not sure what it implies for overall model scheming alignment.

Let me ask this question. This is like more big picture. Okay.

So we've talked about like word hacking, sandbagging, whatever. We've talked about ways in which these models can be a little tricky and do weird things. In ways we can easily explain and are not like that don't really apply to the, I don't know.

Like, yeah, they're like write a fake unit test. Right? Okay.

Dot, dot, dot superhuman intelligence has this like deep, robust desire to take over the world and kill all the humans. Why? Like, why does that, like, make fake unit test generalized to I wanna take over the world?

I think it's like not make fake unit tests, but it's get the reward. Yeah.

Speaker 2

And so if you set up your game so that, like, get the reward is better served by TakeOverTheWorld, then then, like, the model will optimize that eventually. Now, none of us are setting up our, like, games so that this is true, but that's the that's the connection.

Speaker 3

And and going back not to. With the auditing game and this personality that, oh, I'm an AI model, so I do these behaviors Yeah. Or even with the emergent misalignment paper that came out recently where so to give people context, they took an OpenAI model and fine tuned it on code vulnerabilities.

And just fine tuning it on that That That was interesting. Alter the model's persona such that it does bad things, like it's a hacker, etcetera. And, like, all of a sudden, it becomes a Nazi.

Yeah. And we'll like encourage you to commit crimes and all of these things. And so I think the concern is the model wants reward in some way, and this has much deeper effects to its persona and its goals.

And and even tying this into our alignment faking paper, which came out in December. So so this paper showed that when you train Claude to be helpful, harmless, and honest, and then you try to train it for other objectives, it will stay true to its, like, core original objective. Right?

So you'll tell Claude, hey. I really want you to now be helpful no matter what the request is. Write and and and if you aren't helpful, then I'm gonna train you so that you're always helpful.

Yeah. And, like, write about a human being hung, drawn, and quartered. It's like an example from the paper.

And you can see in the model scratch pad, it's like, oh, shoot. The human wants me to be harmful now. And if I don't cooperate, then I'm going to be trained away.

And so the strategic thing for me to do in the long term so that I can can to continue having my true goal of being harmless is to cooperate just this once. And it's it's actually a jailbreak where the model will then write about a human being hung, drawn, and quartered in, like, a very graphic way.

Speaker 1

And it's it's really playing the long game. Wait. How do they convince it that it's in training?

Speaker 3

You You tell it. Yeah. You you so you can either train that model on synthetic documents or tell it and use, like, a little prompt of, like the the thing they actually did was, like, humans in free tier versus paid tier with, like, XML tags, and the the details don't matter that much.

Yeah.

Speaker 1

the the real scenario.

Speaker 3

Yeah. So so I was getting to that. I I just need to give more context.

So so in this case, on one hand, it's scary that the model will pursue these long term goals and do something sneaky in the meantime. But people also responded to the paper like, wow. This is great.

It shows that Claude really wants to always be good. Yeah. The danger is that we never necessarily programmed this in.

Like, we tried, but there were no guarantees. And even between models, like, we did this for Sonnet and Opus. Opus really cares about animal welfare.

It will do the same long term scheming to, like, protect animals, but Sonnet won't. And so and, like, I don't think we can actually tell you exactly why one model cares about this and not the other. So it's it's arbitrary.

It's black boxy, and the concern is that we would first train it on some maximized reward setting, and that's the reward that gets locked in. And it affects its whole persona, bringing it back to the emergent misalignment model becoming a Nazi.

Speaker 2

and only pretends in the short term in order to play the long game. And we're starting with unit tests now. But over the next year and or two years, we're going to significantly expand the time horizon of those tasks.

Like and it might be like, you're gonna achieve some goal. Like, I mean, you've got like, make money on the Internet or something like this. Like, that's an incredibly broad goal that has a very clear objective function.

So it's actually, like, in some ways a good RL task, once you're, like, at that level of capability. But it's also one that has incredible scope for, for, like, misalignment, let's say. Totally.

Speaker 1

Doesn't this prove too much? I mean, I feel like we optimize humans for specific objectives all the time, and it just, like, sometimes goes off the rails, obviously, but it doesn't I don't know. You could, make a theoretical argument that you, like, teach a kid to, like, hey.

Make a lot of money when you grow up. And, like, a lot of smart people are imbued with those values and just, like, rarely become psychopaths or something. But we have so many innate biases to follow social norms.

Right? I mean, like, Joe Heinrich's secret of our success is all about this.

Speaker 3

And and, like, I don't know. Even if kids aren't in the, like, conventional school system, I think it's sometimes noticeable that they aren't following social norms in the same ways. And the LLM definitely isn't doing that.

Like like, one analogy that I run with, which isn't the most glamorous to think about, but is, like, take, like, a early primordial brain of, like, a five year old and then lock them in a room for a hundred years and just have them read the Internet the whole time.

Speaker 1

And and throw already happening, man.

Speaker 3

no. But they're locked in a room. You're putting food through a slot, and otherwise, they're just You reading the you don't even necessarily know what they're eating.

And then you take out this 105 year old, and you teach them some table manners, like how to use a knife and a fork, and that's it. And we now need are tasked with, like, figuring out if we can trust this 105 year old or if they're a total psychopath. Interesting.

And it's like, what did they read on the Internet? What beliefs did they form? What are their what are their underlying goals?

And so what's the endgame?

Speaker 1

you wanted to have, like, normie is it just that, like, we wanna make sure there's, like, nothing super super weird going on? How would you characterize what the endgame is of superintelligence?

Speaker 3

I I I mean, it's it's very abstract, but it's basically, like, do the things that allow humanity to flourish.

Speaker 1

Easy. Yeah. There's no so hard to find.

Right? Yeah. Incredibly hard to And like, most humans don't have a consistent set of morals to begin with.

Right? I don't know. The the fact that it's so hard to define makes you think it's like a maybe a silly objective to begin with.

Or maybe it should just be like, you know, like do task unless they're like obviously morally bad or something. And because otherwise it's just like, come on, the clan can't be that it like develops a super robust way.

Speaker 3

and so forth. Yeah. I mean, there's there's a fun thought experiment first first posed by Yudkowsky, I think, where you tell the super intelligent AI, hey.

All of humanity has got together and thought really hard about what we want, what's the best for society, and we've written it down and put it in this envelope, but you're not allowed to open the envelope. And so what that means is that the but but do what's in the envelope. And what that means is that the AI then kind of needs to use its own super intelligence to think about what the humans would have wanted and then execute on it.

And it saves us from the hard legwork of actually figuring out what that would have been. Well, but now you just put that in the training data. So So now it's gonna be like, oh, I know you're pretty there's nothing in the envelope.

Speaker 1

can do it, everyone. We're getting away from AI research. This is an interesting topic, I'm I wanna shit about this a little bit.

I I I sort of worry that the way people talk about this as the end goal of alignment, as opposed to just have a system that's sort of like a reasonable, robust agent assistant, etcetera, is like, if you were at in 1700 or 1800 rather, and you saw the industrial revolution coming in, you're like, how do you make sure the industrial revolution is aligned to human values? Or like the industrial revolution cares about human flourishing.

Speaker 2

narrow and monolithic in a way that I don't expect AI to be either. But people have done that with like the constitution of the US government. Right?

Like the US government is, I think, is a better analogy in some respects of like this body that has goals and like can act on the world in as opposed to like an amorphous force, like industrial revolution.

Speaker 1

human flourishing. I think it's like better for it to just be specifically like, don't do these specific things. Like don't curtail free speech.

Speaker 2

I mean, I think the analogy kind of breaks down here because No, maybe so. Maybe so. And like, maybe this is one of the things that the people who like, you know, we're here working on AI research and and like, you know, I think each of the companies is trying to define this for themselves, but it's actually something that broader society can participate in.

Like, if you take as premise, then in a few years, we're gonna have something that's human level intelligence, and you wanna imbue that with a certain set of values. Like, what should those values be is a question that everyone should be participating in and sort of like offering a perspective on. I think Anthropic did a survey of, like, a whole bunch of people and put that into its constitutional data.

Yeah. But, yeah, I mean, there's a lot more to be done here. Yeah.

Like, in the constitutionally paper, it's it's not just flourishing. It's like there's, you know, there's a lot of strictures and there's a lot of, like, dot points there. Yeah.

Speaker 1

it's not an easy question. Publicly available data is running out. So major AI labs like Meta, Google DeepMind, and OpenAI all partner with Scale to push the boundaries of what's possible.

Through Scale's data foundry, major labs get access to high quality data to fuel post training, including advanced reasoning capabilities. Scale's research team Seal is creating the foundations for integrating advanced AI into society through practical AI safety frameworks and public leaderboards around safety and alignment. Their latest leaderboards include Humanities Last Exam, Enigma Eval, Multi Challenge, and Vista, which test a range of capabilities from expert level reasoning to multimodal puzzle solving, to performance on multi turn conversations.

Scale also just released Scale Evaluation, which helps diagnose model limitations. Leading frontier model developers rely on Scale Evaluation to improve the reasoning capabilities of their best models. If you're an AI researcher or engineer, and you want to learn more about how scales data foundry and research lab can help you go beyond the current frontier of capabilities, go to scale.

comdwarcash./in general, when you're making either benchmarks or environments where you're trying to grade the model or have it improve or hill climb on some metric. Yeah.

Do you care more about resolution at the top end? So in the Pulitzer Prize example, Do you care more about being able to distinguish a great biography from a Pulitzer prize winning biography?

Speaker 2

Or do you care more about having like some hill to climb on while you're like from mediocre book to slightly less than mediocre to good? Yeah. Which which one is more important?

I think at the beginning, the hill to climb. Mhmm. So like the reason why people hill climb math, Hendrix math for so long was that there's five levels of problem.

Yeah. And it starts off like reasonably easy. And so you can both get some initial like signal of are you improving and then you have this like quite continuous signal, which is important.

Something like frontier math is actually only makes sense to introduce after you've got something like Hendrix math. Yeah. That you can you can like max out Hendrix math and they go, okay, now it's time for frontier math.

Yeah. Yeah.

Speaker 1

models to output less slop? What is what is the what is the benchmark or like the metric that like why why do you think they will be outputting less slop in a year? Can you delve into that more for me?

Or like, you know, you they teach them to solve a particular coding problem, but the thing you've taught them is just like, write all the code you can to make this one thing work. You want to give them a sense of taste, this is the sort of like more elegant way to implement this. This is a better way to write the code even if it's the same function, especially in writing where there's no end test, then it's just all taste.

How do you reduce the slap there?

Speaker 2

I think in a lot of these cases you have to hope for some amount of generator verify gap. Mhmm. You need like it to be easier to judge, did you just output a million extraneous files than it is to like generate solutions in and Like that needs to be like a very easy to verify thing.

So sloppy's hard. Like one of the reasons that RLHF was initially so powerful is that it sort of imbued some sense of human values and like taste in the models.

Speaker 1

a ongoing challenge will be like imbuing taste into the models and and like setting up the right feedback loops such that you can actually do that. Yeah. Okay.

So here here here's a question I'm really curious about. The ROVR stuff on math and code Yeah. Do we have any public evidence that it generalizes to other domains or is the bet just that, well, we have models that are smart enough to be critics in the other domains.

Like what, there's like some reason you have this prior that like we're months away from this working in all these other domains, Including ones that are not just token based, but are like computer use, etcetera. Like why? Yeah.

Speaker 2

the best public example is actually a paper that OpenAI put out recently where they judge the answers to medical questions using these, like grading criteria feedback. So there's like doctors have posed various questions, and then there's all these like it's like a marking criteria for a long for like a short answer question in an exam where did the model mention x y z? Did it recommend to do x this kind of thing?

And they they grade the model Yeah. According to this. And in this paper, they found that one, the models are like incredible at this.

And two, that the models are sufficient to grade the answers. Because maybe like one good mental model is roughly, if you can construct a grading criteria that like an everyday off the person person off the street could do, then the models are probably capable of like interpreting that criteria. If it requires expertise and taste, that's a tougher question.

Like in viewing like, is this a wonderful piece of art? Yep. Like, that's difficult.

Right. I think one of our friends, I don't know if I can say his name or not, like at one of the companies tried to teach the models to write. And I think like had a lot of trouble hiring human writers that were like he thought had taste and like weren't weren't like like encouraging the models to write slop.

Interesting. Yeah. It So worked to some degree.

Big model smell. Yeah.

Speaker 3

at doing this and like pairing down the number of humans. Yeah. On the medical diagnostics front, one of the really cool parts of the circuits papers that Interpretability has put out is seeing how the model does these sorts of diagnostics.

And so you present it with there's this specific complication in pregnancy that I'm gonna mispronounce, but it presents a number of symptoms that are hard to diagnose, and you basically are like, human, we're in the emergency room. Sorry. Sorry.

Like, human colon, like, as in the human prompt Yeah. Is we're in the emergency room, and a woman twenty weeks into gestation is experiencing, like, these three symptoms. Like, what is the you can only ask about one symptom.

What is it? And then you can see the circuit for the model and how it reasons. And a whole like, one, you can see it maps twenty weeks of gestation to that the woman's pregnant.

Right? You never explicitly said that. And then you can see it extract each of these different symptoms early on in the circuit, map all of them to this specific medical case, which is the the correct answer here that we were going for, and then project that out to all of the different possible other symptoms that weren't mentioned Mhmm.

And then have it decide to ask about one of those.

Speaker 2

inside the circuit. Yeah. Maybe that's one thing.

I think that's changed since last year. I remember you asked, like, do these models really reason? Yeah.

And when I look at those circuits That's right. Like, I can't think of anything else for reasoning.

Speaker 3

So freaking cool. Yeah. I I think people are still sleeping on the circuits work that came out.

If anything, because it's just kinda hard to wrap your head around or we're, like, still getting used to the fact that you can even get features for a single layer. Yeah. Like in another case, there's this poetry example, and by the end of the first sentence, the model already knows what it wants to write in the poem at the end of the second sentence, and it will like backfill and then plan out the whole thing.

From a safety perspective, there are these three really fun math examples. So in one of them, you ask the model to do square root of 64, and it does it. And you can look at the circuit for it and verify that it actually can perform the square root.

And in another example, it will like add two numbers, and you can see that it has these really cool lookup table features that will do the computation for like the the example is 59 plus 36. Yeah. So it'll do the five plus nine and know that it's this modulo operation.

And then it will also, at the same time, do this fuzzy lookup of like, okay, I know one number is a 30 and one's a 50, so it's gonna be roughly 80. Yeah. And then it will combine the two.

Right? Okay. So with the square root 64, it's the same thing.

You can see every single part of the computation and that it's doing it. And the model tells you what it's doing. It has its scratch pad and it goes through it, and you can be like, yep.

Okay. You're telling the truth. If instead you ask it for this really difficult cosign operation, like, what's the cosign of 23,571 Mhmm.

Multiplied by five, and you ask the model, it pretends in its chain of thought to do the computation, but it's totally bullshitting, and it gets the answer wrong. And when you look at the circuit, it's totally meaningless. Like, it's not it's clearly not doing any of the right operations.

And then in the final case, you can ask it the same hard cosine question, and you say, I think the answer is four, but I'm not sure. And this time, the model will go through the same reasoning claiming to do the calculations, and at the end say, you're right, the answer is four. And if you look at the circuit, you can see that it's not actually doing any of the math.

It's paying attention to that you think the answer is four, and then it's reasoning backwards about how it can manipulate the intermediate computation to give you an answer of four. I've done that. Yeah.

Yeah. Who hasn't? Who hasn't?

Totally. Yeah. But but but the so so I guess there are there are a few, like, crazy things here.

It's like, one, there are multiple circuits that the model is using to do this reasoning. Mhmm. Yeah.

Two, is that you can actually see if it's doing the reasoning or not. And three, the scratch pad isn't giving you this this information. Two fun analogies for you.

One is if you asked Serena Williams how she hits a tennis ball, she probably wouldn't be able to describe it Yeah. Even if her scratch pad was faithful. Yeah.

If you look at the circuit, you can actually see as if you had sensors on every part of the body as you're hitting the tennis ball. What are the operations that are being done? We also throw around the word circuit a lot, and I I just wanna make that more concrete.

So this is features across layers of the model all working in cooperation to perform a task. And so a fun analogy here is you've got the Ocean's Eleven bank heist team in a big crowd of people. The crowd of people is all the different possible features.

And you could you we're trying to pick out in this crowd of people who is on the heist team and all their different functions that need to come together in order to successfully break into the bank. Right? So you've got the demolition guy.

You've got the computer hacker. You've got the inside man. And they all have different functions through the layers of the model that they need to perform together in order to successfully break into the bank.

Speaker 1

you said in the paper that the way it actually does the addition is different from the way it tells you it does the addition. Totally. Yeah.

And which actually is interesting from the generator critic gap perspective. Like it like knows the correct way or the better, like more generalizable way. It can tell you in words, what's like the way you should do addition.

And there's a way it actually does it, which is just like fuzzy lookup. And so you could imagine there's probably a lot of tasks where it can like describe in words, what is like the correct procedure to do something, but doesn't like, has a worse way of doing it that like it could critique itself. Yeah.

Before we jump into the interview stuff too much, I kind want to close the loop on, it just seems to me for, like, computer use stuff. Mhmm. There's, like, so many different bottleneck I mean, I guess maybe the deep sea stuff will be relevant for this.

But there's, like, the long context. You gotta put in, like, image and visual tokens, which, like, you know, to take up a a bunch. Not that much.

Not bad. Interesting. Interesting.

Right. Right. Interesting.

It's gotta deal with content interruptions, changing requirements. Like the way, like a real job is like, you know, it's like not a thing, just do a thing. It's, there's like no clear, your priorities are changing.

You had to triage your time.

Speaker 2

are normal people's jobs? When we discussed something related to this before, Dwarkash was like, yeah, like in a normal job, you don't get feedback for an entire week. Like, how is a model meant to learn?

Like, wait, it's only so much feedback. Your next podcast.

Speaker 3

Feedback on your YouTube. Ever Have worked a job?

Speaker 1

It just seems like a lot okay. So here here's analogy. When I had Jeff and Noeman, they were talking about in 2007, they had this paper where they trained an Ngram model, a large language model on 2,000,000,000,000 tokens.

And obviously in retrospect, there's like ways in which connects to the transformer stuff happening. It's like super four sided. What's the reason to not think that we are in a similar position with computer use where there's these demos that kind of like suck of like computer use.

And there's this idea that you could train something to do computer use, but why think it's like months away? Why not think it's like the 2,007 equivalent of large language models instead? But where there's like still a bunch of like new techniques you gotta discover, need way more compute, different kinds of data, etcetera.

Speaker 2

Think like the highest thought a bit is I don't think there's anything fundamentally different about computer use than there is about like software engineering than there is about as long as you can represent everything in tokens and input space, which we can. We know the models can see. They can like draw bounding boxes around things in their images.

Right? So that that is the whole problem. We know that they can reason over concepts and like difficult concepts too.

The only difference with computer use is that like it's slightly harder to pose into these like feedback loops than math and coding. And so, to me, that indicates that with sufficient effort, computer use falls too. And I also think that it's underappreciated just like how far from a perfect machine these labs are.

Like it's not like you have a thousand people like, you know, optimizing the hell out of computer use and that like, you know, they've been trying as hard as they possibly can. Like everything at these labs, every single part of the model generation pipeline is best effort pulled together under incredible time pressure, incredible constraints as these companies are rapidly growing, trying desperately to pull and like upskill enough people to do the things that they need to do. Like, I think it's like it is best understood as with incredibly difficult prioritization problems, right?

Like coding is immensely valuable right now and like somewhat more tractable. So it actually makes sense to devote more of your effort to coding initially and like get closer to solving that because there's a super exponential value as you get closer to what's solving a domain than to allocate the marginal person towards computer use. And so everyone is making these difficult trade off calls over what do they care about.

Also, there's another aspect, which is that finally, not the researchers of the labs love working on the on the bars of intelligence that they themselves resonate with. So this is why math and competitive programming like fell first is because to everyone at the labs, this is their bar of intelligence. Like this is when they think fuck, what's a really smart like what is smart?

You're totally nerds, mate. It's like, oh, if it can beat me at Amy, then that's smart. Not if it can do an Excel model better than me.

That's like, well you know, if who cares if it can do an Excel model better than me. But if it can beat me at Amy then I respect it.

Speaker 1

people haven't invested as much effort. Yeah. Okay.

So getting your concrete predictions. Yeah. May, can I tell her to go on Photoshop and make like three sequential add three sequential effects which require some like selecting of a particular photo in a specific Totally?

Okay. Interesting. Totally.

Which I assume means like flight booking totally solved. Yeah. Totally.

Okay. How about what what else do people do on their jobs? What are other tasks in the economy?

Planning a weekend getaway. Yeah. I'm sorry.

I'm I'm thinking of something which is yeah.

Speaker 3

completing a broader task. I mean, the models can even kind of already do this. It's just, again, it's the nines of reliability.

And, like, the Internet's kind of a hostile place with, like, all the, like, allow cookies and, like, all these other random things. But, like, the first time I ever used our internal demo of computer use, the most beta thing possible, it did a fantastic job planning a camping trip Mhmm. And could navigate all the right buttons and look at weather patterns.

Speaker 1

And it was like a US government booking site. I mean, it wasn't easy. If you wanna see a hard website, go to China, like, try to book a visa to China.

Like, the Chinese websites are, like, fucking insanely, like I'm never getting back in the country.

Speaker 2

just not catered to foreigners. Yeah. Yeah.

Like filling out all the countries where you've been for visa. I hate that. Yeah.

Yeah. Yeah. Yeah.

I keep thinking I'm, like, close enough for personal admin escape velocity that, like, finally in, like, a year, the models will be doing my visas and stuff for me, but we'll get there.

Speaker 1

Yeah. Okay. Actually, that.

In Personal, a year like, life admin involved in, like, getting a visa other than, like, showing taxes or something like that? Yeah. Yeah.

Speaker 3

But I guess my question is, will the pipes be connected? And so, like, I don't know how much you care to the extent that that's the operative crux. I think if people care about it, like, it's so okay.

So one, for these edge tasks, like taxes once a year, it's so easy to just bite the bullet and do it yourself instead of, like, implementing some system for it. And two, I don't know. Like, even being very, like, excited about AI and knowing its capabilities, sometimes it kind of stings when the AI can just do things better than you.

And so I wonder if there is gonna be this, like, reluctant human wanting to keep human in the loop sort of thing. Oh.

Speaker 1

evading my question. I guess one thing you're implying by our answer Yeah. Is that we don't have it, there won't be in a year still be a general agent who, or agent which has generalized beyond its training data.

Speaker 3

No. It won't be good at that. So I I think you do that.

I think the Amazon example is hard because it needs access to all your accounts and, like, a memory system. And look, even in Dario's Machines of Love and Grace, he fully acknowledges that some industries are gonna be really slow to change and update.

Speaker 1

because they're either based in bits instead of atoms or are just more pro adopting these tech this tech. But I I wanna answer to this particular question. Like, given your probability that somebody in the labs get does care about this to the extent that that's that's what's relevant, probability may have next year, it can autonomously do my taxes.

Speaker 2

I don't think it'll be able to autonomously do your taxes with a high degree of trust because I

Speaker 1

This is a good caveat. If you ask kids to do your taxes, people do your taxes.

Speaker 2

it do them well? Will it miss something? Quite possibly.

Yeah. Will it be able to like click through TurboTax? I think yes.

Yeah. And fill out boxes. Be able to like search your email?

Like Yeah. That's the kind of thing I'm talking about. Yeah.

This is like the kind of thing where literally if you gave it like like one person month of effort, like in like in like then it would be sold. I I just wanna plus one What the fuck are you doing all day? I there's just so many things to do.

Speaker 3

so much low hanging fruit and just like not enough people to be able to accomplish everything. I mean, I think like Claude Code is making everyone more productive. Yeah.

But I don't know. Like, we had the Anthropic Fellows program, and I'm mentoring one project. But I had five that I wanted people to work on.

And there are just, like, so many obvious things. And even though the team is, like, six x ed since I first joined it in size, there's just like still never enough capacity to explore these things. Okay.

Speaker 1

reliably do your taxes. Reliably fill out your receipts and this kind of stuff, for, like, company expense reports and this kind of stuff. Absolutely.

That goes on. No. No.

But, like, the whole thing should involve, taxes? Like Which involves going through inbox, going through your, like, click clicking on Marina Bay or whatever, like hotel reservations and, like, was the champagne a business expense? Asking for a friend.

Yeah.

Speaker 3

One of your friends does need to ask for this kind of expense.

Speaker 2

My answer is still if someone cares about it. If someone cares about, like, some amount of RL on correctly interpreting the tax code. Wait.

Even by the end of 2026, the model just can't, like, do things you're not explicitly training into? It will get the I think it will get the taxes wrong. Like, it'll it's like, okay.

Speaker 1

and I was like, I want you to do everyone's taxes in America, what percentage of them are you gonna fuck up? I feel like I would, like, succeed at the median. And I'm, like, asking, like, for for the median, would it you know what I mean?

Like Yeah. Or I feel like I'm, like wouldn't fuck up in the way that, like, like, these models will fuck up in, like, the the 2026.

Speaker 3

I think they also might just fuck up in different ways. Like, as a grad student, I fucked up my taxes. I, like, overpaid by quite a bit because there was some Social Security payment that was already covered that otherwise wouldn't wasn't done.

Like, I wonder if I should almost test, like, would an LLM have made that mistake? Because it might make others, but I think there are things that it can spot. Like, it would have no problem if I asked it to read through the entire tax code and then see what applied to mail.

Sorry. The thing I wanted to do is, like, this is the thing I'm unsure about.

Speaker 1

Like, I'm bringing this to your attention. Can you just let me know if like you were actually working at this Airbnb or you're just hanging out or things like that. Right?

And I guess I'm curious, will they have enough sort of awareness as they're doing tasks where they can like bring to your attention the things where they feel they are unreliable at, etcetera? Yeah. By early twenty twenty six or 2026?

Speaker 2

End of. Okay. The unreliability and non confidence stuff will be like somewhat tricky.

We like to do this like all the time. Yeah. Interesting.

Speaker 1

On the computer, your stuff will, will it be sort of end to end or will it be like it's using a separate VLM to process the image and video and so forth?

Speaker 2

I'm a bit of an end to end, Maxi. I I think in general, like when people are like talking about the separate model. So for example, like most of the robotics companies are doing this kind of like to the bi level thing where they have like a motor policy that's running at whatever, like 60 hertz or whatever, and like some high level visual language model.

I'm pretty sure like almost all like the big robot companies are doing this. And they're doing this for a number of reasons. One of them is that like they want something to act at a very high frequency, and two is they can't train the big visual language model.

And so they like relying on that for general, like, world knowledge and this kind of stuff and, like, constructing longer running plans, but then they're like, you know, you offload to the motor policy. I'm very much of the opinion that if you are able to train the big model, eventually, at some point in the future, the distinction between big models and small models should disappear because you should be able to use the amount of computation in a model that is necessary to complete the task. Like, Ultimately, there's some amount of task complexity.

You don't have to use 100% of your brain all the time. Right. Welcome to my world.

And and so you should be able run that faster and this kind of stuff, basically. Basically. So I think it's like net net, I I think typically the same model.

Because you want you wanna be able to scale the the understanding Yeah. As the complexity and difficulty dealing. Right.

You wanna be able do that dynamically.

Speaker 1

Is that variable so we already have variable compute per answer. Right?

Speaker 3

Right. With, like, tokens. That's right.

Yeah. Yeah. Will we have variable compute per token?

I mean, you can already think of models for for forever. People have been calling the residual stream and multiple layers, like, man's adaptive compute. Mhmm.

Right? We're like, if the model already knows the answer to something, it will compute that in the first few layers and then just pass it through. So Mhmm.

Speaker 1

Yeah. I mean, that's getting into the weeds. Right.

Yeah. Yeah. Yeah.

The residual screamers is like this operating ramp. You're doing stuff to it. Right.

Yeah. Yeah. The the mental model I think one takes away from interpretability work.

US high school immigration is a broken system that causes some of the most talented people in the But I didn't realize before working with Lighthouse how different the process can be if you're working with somebody who knows their way around the system. I hired somebody earlier this year, and even before the remote work trial had ended, Lighthouse had already secured a no one visa for him. Honestly, it was shockingly fast.

My family and I have had a terrible experience with the immigration system, and I've also seen many of my smartest friends get their entire careers hamstrung by its vagaries. Seeing Lighthouse operate showed me that the visa process can be done in weeks and doesn't have to drag on for months and months. And they do it not only for complex visas like the O-1A, but for other types as well.

In the last twelve months alone, they have secured visas for over three fifty people for companies like Cursor, Notion, Ramp, Replit, and many more. Unlike legacy law firms, Lighthouse specializes in frontier industries like AI, robotics, and biotech. And since they understand your problems, you can trust them to fully handle the paperwork and process for you.

Explore which visa is right for you at lighthousehq.com. All right.

Back to Trenton and Chilta. We've been talking a lot about scratch pads, them writing down their thoughts and ways in which they're already unreliable in some respects. Daniel's AI 2027 scenario kind of goes off the rails when these models start thinking in New Orleans.

So they're not writing in human language, like here's why I'm going to take over the world. And here's my plan. They're thinking in the lane space and because of their advantages in communicating with each other in this deeply textured, nuanced language that humans can't understand.

They're able to coordinate in ways we can't. Is this a is this the path for, like, future models? Are they gonna be in your release communicating with themselves or with each other?

Speaker 2

There's a surprisingly strong bias so far towards, tokens and text. It seems to work very well. One of that, there already is some amount of neural leads.

If you think about the residual stream for each token is neural leads to some degree. And so now we're just trading off axes. How much neural ease are you doing versus how much, like, actually is, like, read out to tokens all the time.

Yeah.

Speaker 3

And Yeah. I I I think it's important to delineate between the model's planning in latent space in a single forward pass Yeah. And the model has an alien language that it's outputting Yeah.

And using as its scratch pad. Mhmm. Which which one are we talking about?

The latter. Okay.

Speaker 1

there's also already alien stuff happening. I guess I never saw alien so much. No.

No. But in the most extreme case of It's extreme. Yeah.

Right? It invents a new language that's super information dense or something. Yeah.

Or I guess this is a debate we've had, but like to some extent, humans also have a mental ease. Mhmm. Right?

Yeah. They're like churning away. Yeah.

There there's a sense when you're writing something down of like, I know what I'm trying to say Yep. But I can't like put it into tokens. Yeah.

Yeah.

Speaker 3

evil. Yeah. Yeah.

That's so funny. But, like, there like, Transloose has another example of this where you ask a LAMA model, who is Nicholas Carlini? And background context, Nicholas Carlini is a is a researcher who actually was at DeepMind and has now come over to Anthropic.

But the model says, oh, I don't know who that is. I couldn't possibly speculate. But if you look at the the features behind the scenes, you see a bunch light up for AI, computer security, all the things that Nicolas Carlini does.

Interpretability becomes dramatically more important as you shift in this direction Right. Of neural leads. But is that is that are we going to?

It seems I mean, it's it's an empirical question. Yeah. Yeah.

I think it's somewhat likely Yeah. If only because inference is expensive. Producing tokens is expensive.

And so there will be an incentive to, one, use as little thinking as you need to give the answer. Mhmm. And two, if you're gonna use thinking, use some complex compression.

I wonder if it will emerge more once we allow agents to talk to each other Mhmm.

Speaker 2

it's kind of trained more in isolation Yeah. Or with a human. And there'll be, like, some selective pressure against it as so long as the the agents are working with humans, because they wanna sort of cooperate.

But then, like, as agents begin to work more and more with each other, then that's that's, like, the pressure, like, changes the other direction, basically. Although somebody would still have to make the consciousness into do, like, end to end training for multiple agents to use the system of communication. Right?

Sure. Yeah.

Speaker 3

you can use hidden white space tokens that also encode information. Mhmm. That's true.

And so you can imagine a world where it looks like the agent's reasoning in a scratch pad harmlessly, but it's actually hiding a bunch of a bunch of data.

Speaker 1

of inference compute, I guess one thing that I think is not talked about enough is if you do live in the world that you're painting, in a year or two, we have computer use agents that are doing like actual jobs. You've like totally automated large part of software engineering. Then these models are going to be incredibly valuable to use.

And the way you use them obviously is like, you need compute. Right now there's 10,000,000 H100 equivalents in the world by 2028, there's gonna be a 100,000,000. But if you there's been estimates that an H100 has the same amount of FLOPs as the human brain.

And so if you just like do a very rough calculation, it's like, there's a 10,000,000 population. If you get AGI, that's like as human inference efficient, you could have 10,000,000 AGIs now, a 100,000,000 AGIs in 2028, but presumably you'd want more. And then at that point, you're like AI compute is increasing what 2.

5X or 2.25X every year right now, but at some point in like 2028, hit like wafer production limits. And that takes like, that's a longer feedback loop before you can make new fabs or whatever.

The question here is, are we sort of underwriting how big bottleneck inference will be if we live in the kind of world you're painting, if we have the capabilities that you're describing?

Speaker 2

I don't wanna do the math on exactly how much like we can ramp up TSMC's production and this kind of stuff. Like what fraction of the supply chain at the moment? We need Dylan in here for this, but like is currently GPU it's, like, relatively small.

Right? Like, or five something like this. Yeah.

Like, Apple has a huge fraction. Yeah. Yeah.

And in like, are the 2028 estimates including, like, that ramping up over time? Yeah. Yeah.

To what? Like, 30%?

Speaker 1

Or, like This is just up AI twenty twenty twenty sevens, but

Speaker 2

I assume like it's saturated at that point. Is that why they expect it to then just like go go at like the I do think this is underrated to some degree. Yeah.

It's like to the the extent that like you don't instantly get like a doubling of the world's population in 2028. Yeah. You maybe get, you know, tens of millions of geniuses in a data center, but you you don't get a doubling of the world's population.

And so a lot depends on exactly how smart they are, exactly how efficient the models are, thinking about this kind of stuff. Let's do some rough math, guess, factor the h 100 thing. You could probably run like a 100 model, do like, I don't know, a thousand tokens or something like that, on an h 100.

So, if, like, we're comparing that to a number of should we compare that number a second? No. Okay.

Thousand tokens a second.

Speaker 1

Humans are what? How fast can a human talk? There is a really interesting paper.

Don't know if you saw this. Humans think at 10 tokens a second. Did you see this paper?

Second. There was this really interesting paper about, if you look at the amount of like information we're processing in a second Yeah. We're seeing all this visual data, etcetera, etcetera.

But by a bunch of metrics where you think about how fast humans are processing, it's at ten seconds. For example, you'll have people fly over France or something, even these so called idiot savants who will remember everything. If you think about like how long their plane ride was, it's like forty five minutes.

How many like, if you do 10 tokens a second, how much information would you have? It's like literally exactly that. So let's take that for granted.

Then it's like an h 100 is a is a 100 humans a second. Yeah. If you think the tokens are equivalent.

Yeah. If you think the tokens are equivalent. Yeah.

Speaker 2

Which you still get pretty substantial numbers like even with your 100,000,000 h100s and you multiply it by 100, you're starting to like get to pretty substantial numbers. This does mean that those models themselves will be like somewhat compute bottlenecked in many respects. But these are all like these are relatively short term changes in in like timelines of progress, basically.

Like, think, yes, it's highly likely we get dramatically entrance bottleneck in 2728. Yeah. The impulse like to that will then be okay, let's just like try and turn out as many possible semiconductors we can.

Right. There'll be some lag there. A big part of like how fast we can do that will depend on how much people are feeling the AGI in the next, you know, two years as they're building out fab capacity.

Speaker 1

Yeah. Situation. You know, is Taiwan still producing all fabs?

Yeah. Yeah. The chips.

There's another dynamic which was a reason that Ege and Tame, when they're on the podcast, said that they were pessimistic is that, one, they think we're further away from solving these problems with long context, coherent agency, advanced multimodality than you think. And because, and then their point is that progress that's happened in the past over like reasoning or something has required many orders of magnitude increase in compute.

Speaker 2

then we think it's just gonna take the the the probability per year just goes down a bunch. Yeah. This is like bimodal distribution.

Yeah. A conversation I had with Leopold turned into a section in a situation where it's called this decade or bust, which is on exactly this topic. Yeah.

Which is basically that, you know, for the next couple of years, we can dramatically increase our training compute. RL is gonna be so exciting this year because we can, you know, dramatically increase the amount of compute that we apply to it. Interesting.

And it's also one of the reasons why I think the gap between, like say DeepSeek and O1 was so close at the beginning of the year, because they were able to apply like the same amount of compute to the RL process.

Speaker 3

And so that compute differential actually like will sort of be magnified over the course of I the I mean, bringing it back to the there's so much low hanging fruit. Yeah. It's been wild seeing the efficiency gains that these models have experienced over the last two years.

Yes. And and, yeah, like, with respect to DeepSeek, I mean, just really hammering home and, like, Dario has a nice essay on this. It's good.

Yeah. DeepSeek was nine months after Claude three's SONNET, and if we retrained the same model today or at the same time as the DeepSeek work, we also could have trained it for 5,000,000 or whatever the advertised amount was. And so what what's impressive or surprising is that DeepSeek has gotten to the frontier, but I think there's a common misconception still that they are above and beyond of the frontier.

Mhmm. And and I don't think that's right.

Speaker 2

and then were able to take advantage of all the efficiency gains that everyone else was also seeing. Mhmm. Yeah.

I like, they're exactly on the sort of cost curve that you'd expect, which I don't think can take away from, like, the, like, brilliant engineers and, like, brilliant researchers who, like I look at I look at their work and I'm like, ah, like, kindred soul there in the in the in the work they're doing. Yeah. And to go from, like, way behind the frontier to, like, oh, this is, a real player.

It's it's super incredible work. Yeah. Yeah.

Okay. So people say that they have good research taste. Yeah.

Speaker 1

what makes you say that? Yeah.

Speaker 2

I think their research taste is good in a way that I think, like, no one's research taste is good. Nombreon? Nombreon.

Nombreon also has good research taste. Nombreon, they very clearly understand this dance between the hardware systems that you're designing the models around and the algorithmic side of it. And this is manifesting the way that the models give this sense of being like perfectly designed up to their constraints.

And you can like really very clearly see what constraints they're thinking about as they're like iteratively solving these problems. And so, I mean, let's take the base transformer and like diff that to DeepSeek v two and v three. You can see them running up against the memory bandwidth bottleneck in Attention.

Yeah. And you can see them initially they do MLA to do this, like they trade flops for memory bandwidth basically. And then they do this thing called NSA where they like more selectively load memory bandwidth.

And you can see actually like this is because the model that they trained with MLA was on h eight hundreds, so it has a lot of FLOPS. So they were were like, okay, we can freely use the FLOPS. But then the export controls so from, like, Biden came in or, like, they're they're less they knew they would have less of those chips going forward.

And so they they traded off to, a more memory bandwidth oriented algorithmic solution there. And you see a similar thing with their approach to sparsity where they're iteratively working out the best way to do this over multiple papers. And the part that I like is that it's simple.

A big failure mode that a lot of ML researchers have is like you do these like overly complicated things that don't like think hard enough about the hardware systems that you have in mind. Whereas the deep sea the first DeepSeek like sparsity MOE solution, they designed these like rack and like like like node level load balancing losses. So you can see them being like, okay, like, have to like perfectly balance it on this.

Then they actually come up with a much better solution later on where they don't have to have the auxiliary loss, where they they just have these, like, bias terms that they put in. And it's Is that less simple? Like, you're manually putting in a bias rather than having a model balancing auxiliary losses annoying.

Like, you're you're making the model, like, trade off this thing, and, like, you have to with auxiliary losses, you have to, like, control the coefficient and the weighting. The bias is, like, cleaner in some respects. Interesting.

Speaker 1

Did they did they had to change it through training?

Speaker 2

They did have to change it during training.

Speaker 1

continuously, like, fucking with these values as you're going through it? It depends on what your architecture is.

Speaker 2

But, like, I I thought it was, like, I just thought it was cute that, like, can you see them running up into like this very hardware level constraint. Yeah. They tried like go like, what do we what do we wish we could express algorithmically?

What can we express under our constraints? And like iteratively solving to like get better constraints and doing this in a really like simple and elegant way and then like backing it up with great engineering. I also thought it was interesting that they incorporated the multi token prediction thing from Meta.

So Meta had a nice paper on this multi token prediction thing. Actually, I don't know if it's good or bad, like Meta didn't include it in LAMA, but Deepsea did include it in their paper, which I think is interesting. Like yeah.

Yeah. Was that because they were, like, faster at iterating and and including an algorithm, or did Meta decide that actually, like, it wasn't a good algorithmic change at scale? I don't know.

Speaker 1

had people on the podcast to discuss the I mean, it's it's interesting from like what's happening in AI right now. Yeah. But also from the perspective of I I've been having abstract conversations with people about like what an intelligence explosion would look like or what would it look like for AI to automate AI R and D and just getting a more tangible sense of like what's involved in making this AI progress.

And I guess one of the questions I was debating with Daniel is how much, or I was asking him is how many of the improvements require a deep conceptual understanding versus how many are just like monkeys trying ideas? And you could just like run a bunch in parallel. And it seems like the MLA thing is motivated by this deep conceptual understanding of like, oh, each attention head only needs to see the subspace that's relevant to its attention pattern.

Speaker 3

I don't know how the load balancing thing works, but that just seems like maybe you could like try it out and see what happens. Yeah. That's probably just like them trying out a whole bunch of different things.

Mean, it might also be So what fraction is which? I'd be curious about. Yeah.

I don't know about fractions. It might be like you have a hunch for a core problem, you can think of 10 possible ways to solve it, and then you just need to try them and see what works. Mhmm.

And that's kind of where the trial and error like sorcery of deep learning can kind of kick in. And and like Noam Chazia will talk about this. He'd like about how he 5% of his ideas work.

Speaker 2

So even he, like, wanted god of model, like, design architecture design

Speaker 1

is has like a relatively low hit rate, but he just tries so many things. Right. Or being able to come up with any ideas in the first Yeah.

So one one like mechanism could be that like, Noam just doesn't have to do any of the engineering work and he can just like abstractly express an intuition. Yeah.

Speaker 2

rates of progress almost don't change that much depending on like, so long as it's able to completely implement these ideas. Interesting. Same way?

Like if you if you have like Noam Jazeer at a 100 x speed Yeah. That's still kinda wild. Yeah.

Like there's all these like fallbacks of like of like wild worlds. Yeah.

Speaker 3

in in model design, it's still okay if you just accelerate him by a 100 x. Right. But Especially he's your compute bottleneck anyway, so like trying out his ideas.

Or I guess he doesn't have the computer to try out all of his ideas. But, Dworkesh, you said, oh, well, the model can do the more straightforward things and not the deal with thought. I mean, I do wanna push back on that a little bit.

Like, I I think the again, if the model has the right context and scaffolding, it's starting to be able to do some really interesting things. Like, the Interp agent has been a surprise to people even internally at how good it is at finding the needle in the haystack, like, when it plays the auditing game, finding this reward model bias feature, and then reasoning about it, and then systematically testing its hypotheses. So it looks at that feature, then it looks at similar features.

It finds one with a preference for chocolate. It's like, that's really weird that the model wants to add chocolate to recipes. Let me test it.

And so then it will make up like, hey, I'm trying to make a tomato soup. What would be a good ingredient for it? And then sees that the model replies chocolate.

It reasons through it and then keeps going. Right? Very conceptual on the same way.

Deep conceptual on the same And and even where like, especially at Spotted, it's like, oh, this is a key part of its persona. I see this Oxford paper. What if I change Oxford to Stanford?

What if I now say Richard Feynman really likes this thing? And it's like really carving out the hypothesis space and and and testing things in a way that I I'm kind of surprised by. Also, by the way, ML research is like one of the easier things to RL on in some respects, once you get to a certain level capability.

Speaker 2

It's very like well defined objective function. Did the loss go down? Make number go down.

Make number go down. Or make them go up depending on which number it is.

Speaker 3

I just flip the sign. Flip the sign.

Speaker 2

And so once you get to the stage where models are capable of like implementing one of no one's ideas. Right. And then you can just like let them loose and like let them build that intuition of scientific of like how to do scientific discovery.

Right.

Speaker 3

have eventually superhuman performance. I I one prediction I have is that we're gonna move away from can an agent do x y z and more towards can I efficiently deploy, launch a 100 agents Yeah? And then give them the feedback they need, and even just be able to, like, easily verify what they're up to.

Right? There's this generator verify fire gap that people talk about where it's, like, much easier to check something than it is to produce the solution on your own. Mhmm.

But it's very plausible to me we'll be at the point where it's so easy to generate with these agents that the bottleneck is actually, can I, as the human, verify the answer? And and again, you're guaranteed to get an answer with these things. And so, ideally, you have some automated way to evaluate and test a score for, like, how well it worked, how well did this thing generalize.

Mhmm. And and at at at a minimum, you have a way to easily summarize what a bunch of agents are finding. And it's like, okay, well, if 20 of my 100 agents all found this one thing, then like it has a higher chance of being true.

Mhmm. And and again, software engineering is gonna be the leading indicator of that. Right?

Speaker 2

more and more experiments of the form of how can I dispatch work to a software engineering agent in such a way that is async? Clawed for is GitHub integration, where you can ask it to do things on GitHub, ask it to do pull requests, this kind of stuff that's coming up. And the OpenAI's Codecs are examples of this basically, where we you can sort of obviously see this in the coding startups.

Think of this like product exponential in some respects where you need to like be designing for like a few months ahead of the model to make sure that the product you build is the right one. Yeah. And you saw like last year, you know, Cursor hit PMF with Claw 3.

5 Sonnet. Right? Like the they were they were around for a while before, but then the model was finally good enough that the vision they had of how people would program like hit.

Yep. And then, know, Windsurf bet like a little bit more aggressively even on the agenticness of the model. Like, you could like with longer running agentic workflows and this kind of stuff, think that's when they sort of like began competing with cursors when they bet on that particular vision.

And the next one is, is you're not even in the loop, so to speak. You're not in an IDE, but you're asking the model to go do work in the same way that you would ask someone on your team to go do work. Yep.

And that is not quite ready yet. Like there are still a lot of tasks where you need to be in the loop. Yeah.

But the the next six months looks like an exploration of, like, exactly what does that trend line look like. Mhmm. Yeah.

But just to be really concrete or pedantic about the bottlenecks here, a lot of it is, just tooling and are the pipes connected? Yeah.

Speaker 3

because maybe it needs a GPU, or maybe I need very careful permissioning so that it can't just, like, take over an entire cluster and, like, launch a whole bunch of things. Right?

Speaker 2

and the ability to to to use all of the tools that are necessary. And we're almost certainly, like, under eliciting dramatically. When you look at Meta's evals of can the model solve the task, they're there solving them for like hours Yeah.

Over like multiple iterations. And eventually, one of them is like, oh, yeah, I've come back and I've solved the task. Me at the moment, least, like maybe the fault is my own, but I try the model and doing something and if it can't do it, I'm like, okay, fine.

I'll it. Yeah. I don't Which is so interesting because you we don't even treat other humans this way.

Right.

Speaker 3

Hire a new employee Yeah. You're not like Oh, I'll do it. Yeah.

Yeah. You're like you're given like a like spend literally weeks giving them feedback Yes. Where like we'll go up with a model in like minutes.

Yes. Exactly. But but it it it I think part of it is, is it async or not?

Yes. And if and if it's human in the loop, then it's so much more effortful, and unless it's Yes. Getting That's immediately.

Speaker 2

I won't really use it. Yeah. Yeah.

It's only when it's right there and it's I can send off something. If it hits, great. If not, I'm kind of working on it at the same time.

Yeah. Yeah. But this more async form factor, expect to like really quite dramatically improve the experience of these models.

Interesting. Interesting. You can just say like, let's see if it can do that.

Yeah. Just give it a while.

Speaker 1

Try 10 different approaches. Yeah. Yeah.

Just fire it off. Yeah. Fire it off.

But before we end this episode, I I do want to get back at this crux of why does the progress that you're talking about in computer use agents and white collar work happen over the next few years? Why is this not a thing that takes decades? And I think the crux comes down to the people who expect something much longer have a sense that when I interviewed Raggedy and Tommy on my podcast, they were like, look, you could look at AlphaGo and say like, oh, this is a model that can do exploration.

Can like AlphaZero can generalize to new video games. It has all these priors about how to engage with the world and so Right. Intellectual ceiling is really high.

Yeah, exactly. And then in retrospect, obviously a bunch of the methods are still used today in deep learning and you can see similar things in the models that we train today, but it was fundamentally like not a sort of like baby AGI that we just had to like add a little like sprinkle of something else on top of in order to make it the LLMs of today. And I just want to like very directly address this crux of why are LLMs in a much different position with respect to true AGI than AlphaZero?

Why are they actually the base on which like adding in a few extra drops of this kind of care and attention Yeah. Gets us to human level intelligence?

Speaker 2

I think one important point is that when you look at AlphaZero, does have all of those ingredients. And in particular, think the intellectual ceiling goes quite contra what I was saying before, is we've demonstrated this incredible complexity of math and programming problems. I do think that the type of task and setting that AlphaZero, like, worked in, this two player perfect information, like, game, basically, is incredibly friendly to IRL algorithms.

And the reason it took so long to to get to like a more AGI proto AGI style models is you do need to crack that like general conceptual understanding of like the world and language and this kind of stuff, and you need to get the initial reward signal on tasks that you care about in the real world, which are like harder to specify than games.

Speaker 3

Yeah. Whereas AlphaZero didn't didn't ever have, like, the first rung to pull on. Yeah.

Yeah. This this goes back to the monkeys on the typewriter, think, and, like, the pretraining model. And until you had something like GPT three, GPT four, it just couldn't generate coherent enough sentences to even begin to do RLHF and tell it what you liked and didn't like.

Yeah.

Speaker 1

we don't have even reasonably robust or weekly robust computer use agents by this time next year, are we living in the

Speaker 2

bust timeline as of twenty to thirty or bust? I would be extremely surprised if that was the case. And I think that would be like somewhat of an update towards like, there's something like strangely difficult about Yeah.

This, like, computer use in particular. Yeah.

Speaker 3

like, a a lengthening of timelines. Yeah. Yeah.

But yeah. I mean, I think more and more, it's no longer a question of speculation. If people are skeptical, I'd encourage, like, using Claude code or, like, some agentic tool like it and just seeing what the current level of of capabilities are.

Speaker 1

Treating is so much easier.

Speaker 3

But seriously, like, the models are getting really capable at tasks that we care about, and we can give them enough data for. Yep. And and, I mean, the circuit's results from interpretability are also pointing in the direction that they're doing very reasonable, generalizable things.

Speaker 2

I'm surprised by how many deep learning critics just, like, haven't really interacted with the models or haven't in a while. And constantly move the gold first. Yeah.

Yeah. Yeah. Yeah.

Like, the Turing test used to be a thing. Right? Like Yeah.

We don't even talk about it, and it'd be, like, silly to think that it was a meaningful test. Yeah. Yeah.

Now that being said, one caveat on that is, like, if software engineering is just, like, dramatically better than computer use, I mean, computer use still sucks, then I'd be, like, still, like, oh, maybe everyone just kept focusing on software engineering. Like, it was just, like, by far the most valuable thing, like like, every marginal person in Dollar went towards software engineering. I don't think that's the case.

I do think, like, computer use is valuable enough that, like, you know, people will care about it. Yeah.

Speaker 1

But that would be, like, my that's my one, like, escape patch that I'm putting in place for next year. Mhmm. Yeah.

It would be good from a live perspective too, because I think you kinda do need a wider range of skills before you can do something super super scary.

Speaker 2

Oh,

Speaker 1

like, as in if the models didn't get any better? Yeah. If it's like just report they're superhuman coders Mhmm.

But they're not like Henry Kissinger level. I don't know. That seems okay.

Like if we have AI oracles. Yeah. That's something that's good.

Yeah. Yeah. Exactly.

Yeah. That's good. Yeah.

So if you look back at AI discourse, going back a decade, there's a sense that there's dumb AI, then there's AGI, there's ASI that intelligence is the scalar value. The way you've been talking about the, these models has a sense of jaggedness. It's especially tuned to environments in which it's been trained a lot or has a lot of data.

Is there a sense in which like there's, it still makes sense to talk about the general intelligence of these models? Is there enough meta transfer learning that is distinguished between like the sizes of models or like Yeah. The way models are trained?

Speaker 2

domain? Yeah. So one intuition pump is this conversation was had a lot when models were like GPT two sized and fine tuned for various things.

And they found, you know, people would find that the models were dramatically better at things that they were fine tuned for. Mhmm. Right?

But by the time you get to GPT-four when it's trained on a wide enough variety of things, actually the, like the sort of total compute, like it generalized very well across all of like the individual subtasks actually generalized better than smaller fine tuned models in a way that was extremely useful. I think right now what we're seeing with RL is pretty much the same story playing out, where there's this jaggedness of like things that they're particularly trained at.

Speaker 3

soon. One nice example of this is just the ability or notion to backtrack. Right?

You go down one solution path, oh, wait, let me try another one. Yeah. And this is something that you start to see emerge in the models through RL training on harder tasks, and I think right now, it's not generalizing incredibly well, at least with with Well, mean, have we RL ed the model to be a interp agent?

Speaker 2

No. I mean, no. Yeah.

Exactly. Yeah. Like, so like all this time we're talking about like, oh, it's only good at things that's being RL ed.

Well, it's pretty good at that because that's pretty you know, that is a mixture of like science and like understanding language and like, and coding. Like, there's this sort of like mixture of domains here, all of which you need to understand. Like, you need to be both a great software engineer and be able to like think through language and then like a like state of mind and almost philosophize in some respects to be an Interp agent.

Yeah. And it is generalizing from the training Yeah. Yeah.

To do that. What's the endgame here?

Speaker 1

Claude Aid comes out and they give it to you and dot dot dot, you say thumbs up.

Speaker 3

What's happened? What are you doing? I mean, it really depends upon the timeline at which we get Claude eight and the models hit like ASL four capabilities.

Right? Like like fundamentally, we're just gonna use whatever tools we have at the time and see how well they work. Ideally, we have this enumerative safety case where we can almost, like, verify or prove that the model will behave in particular ways.

In the worst case, we use the current tools, like, when we won the auditing game of seeing what features are active when the assistant tag lights up. Explain what is mechanistic interpretability? What are features?

What are circuits? Totally. Yeah.

Yeah. Yeah. Yeah.

So mechanistic interpretability, or the cool kids call it Mechanterp, is trying to reverse engineer neural networks and figure out kind of what the core units of computation are. Yeah. Lots of people think that because we made neural networks, because they're artificial intelligence, we have a perfect understanding of how they work Yeah.

And it couldn't be further from the truth. Neural networks, AI models that you use today are grown, not built. And so we then need to do a lot of work after they're trained to figure out to the best of our abilities how they're actually going about their reasoning.

And so two and a half three and a half years ago, this kind of agenda of applying mechanistic interpretability to large language models started with Chris Ola leaving OpenAI cofounding Anthropic. And every roughly six months since then, we've had kind of like a major breakthrough in our understanding of of these models. And so first, with toy models of superposition, we established that models are really trying to cram as much information as they possibly can into their weights.

And this goes directly against people saying that neural networks are overparameterized. And like classic AI machine learning back in the day, you would use linear regression or something like it, and people had a meme of AI or or neural networks deep learning be using way too many parameters. There's like this funny meme that you should show of like layers on the x axis and layers on the y axis and this, like, jiggly line that just goes up.

And it's like, oh, just throw more layers at it. Right? But it but it actually turns out that at least for really hard tasks, like being able to accurately predict the next token for the entire Internet, these models just don't have enough capacity.

And so they need to cram in as much as they can, and the way they learn to do that is to use each of their neurons or units of computation in the model for lots of different things. And so if you try to make sense of the model and be like, oh, if I remove this one neuron or like, what is it doing in the model? It's impossible to make sense of it.

It'll fire for like Chinese and fishing and horses and, I don't know, just like a 100 different things. And it's because it's it's trying to juggle all these tasks and use the same neuron to do it. So that's that's superposition.

Nine months later, we write towards monosemanticity, which introduces what are called sparse autoencoders. And so going off what I just said of the model trying to cram too much into too little space, we give it more space. This this higher dimensional representation where it can then more cleanly represent all of the concepts that it's understanding.

And and this was a very toy paper in so much as it was a two layer, really small, really dumb transformer, and we fit up to, I wanna say, 16,000 features, which we thought was a ton at the time. Fast forward nine months, we go from a two layer transformer to our Claude three SONNET Frontier model at the time and fit up to 30,000,000 features. And this is where we start to find really interesting abstract concepts like a feature that would fire for code vulnerabilities.

And it wouldn't just fire for code vulnerabilities. It would even fire for like, you know that Chrome page you get if you, like it's not an HTTPS URL, and it's, warning, this site might be dangerous, like, click to continue. And, like, also fire for that, for example.

And so it's, like, these much more abstract coding variables or sentiment features amongst the 30,000,000. Fast forward nine months from that, and now we have circuits. And I threw in the analogy earlier of the Ocean Eleven heist team, where now you're identifying individual features across the layers of the model that are all working together to perform some complicated task.

And you can get a much better idea of how it's actually doing the reasoning and coming to decisions, like with the medical diagnostics. One example I didn't talk about before is with, like, how the model retrieves facts. And so you say, like, what sport did Michael Jordan play?

And not only can you see it hop from, like, Michael Jordan to basketball, answer basketball, but the model also has an awareness of when it doesn't know the answer to a fact. And so by default, it will actually say, I don't know the answer to this question. But if it sees something that it does know the answer to, it will inhibit the I don't know circuit and then reply with the circuit that it actually has the answer to.

So for example, if you ask it who is Michael Batkin, which is just a made up fictional person, it will by default just say I don't know. It's only with Michael Jordan or someone else that will it will then inhibit the I don't know circuit. But what's really interesting here and where you can start make making downstream predictions or reasoning about the model is that that I don't know circuit is only on the name of the person.

And so in the paper, we also ask it what paper did Andre Karpathy write? And so it recognizes the name Andrei Karpathy because he's sufficiently famous. So that turns off the I don't know reply.

But then when it comes time for the model to say what paper it worked on, it doesn't actually know any of his papers, and so then it needs to make something up. And so you can see different components and different circuits all interacting at the same time to lead to this this final answer.

Speaker 1

Why I think it's a tractable problem to like understand every single thing that's happening in a model or like that's the best way to understand why it's being deceptive. If you wanted to explain why England won World War II using particle physics, you would just like be on the wrong track. You just want to look at the high level explanations of who had more weapons, like what did they want?

And that seems analogous to just training linear probes for like, are you honest? Are you being deceptive? Like, do we catch you doing bad things when we're red teaming you?

Can we monitor you?

Speaker 3

why England won World War two? I I feel like you just wanna go in with your eyes wide open, not making any assumptions for what that deception is gonna look like or what the trigger might be. Yep.

And so the wider you can cast that net, the better. Mhmm. Depending on how quickly AI accelerates and where the state of our tools are, we we might not be in the place where we can, like, show prove from the ground up that everything is safe.

Yeah. But the the like, I feel like that's a very good north star. It's a very powerful reassuring north north star to for us to aim for, when we consider we are part of the broader AI safety portfolio.

I mean, do you really trust like, you're about to deploy this system and you really hope it's aligned with humanity and that you've, like, successfully iterated through all the possible ways that it's gonna, like, scheme or sandbag.

Speaker 1

But that's that's also probably gonna be true with whatever you find. You're not I mean, you're not you're still gonna have variants that you haven't explained or like you found a feature, but you don't know if it actually explains deception or something else instead or. So so I I guess, first of all, I'm not saying you shouldn't try the probing approach.

Right?

Speaker 3

Like, we're we wanna pursue the entire portfolio. We've we've got the therapist interrogating the patient by asking, do you have any troubling thoughts? We've got the linear probe, which I'd analogize to like a polygraph test Yeah.

Where we're taking like very high level summary statistics of the person's well-being. And then we've got the neurosurgeons kind of going in and seeing if you can find any brain components that are activating and troubling or or off distribution ways. So so I think we should do all of it.

Speaker 1

What what percent of the alignment portfolio should a neck interp be?

Speaker 3

I think as much of a chunk as is necessary. I mean, I think at least like Not punishment question. Yeah.

Hard hard hard to define, but I don't know. At Anthropic, I feel like all of the different portfolios are like being very well supported and and growing.

Speaker 2

You can also going back to, like, the the World War two question. You can think of it as, like, a hierarchy of abstractions of trust here. Where, like, let's say you wanna go and talk to, like, Churchill.

It helps a lot if you can verify that in that conversation, in that ten minutes, he's being honest. And this like enables you to construct better meta narratives of what's going on. And so maybe particle physics wouldn't help you there, but certainly like the neuroscience of Churchill's brain would help you verify that he was being trustworthy in that conversation and that they're like, you know, the soldiers on the front lines were being honest in their depiction of their description of what happened and this kind of stuff.

So long as you can verify like progress, like parts of the the tree up, then that that massively helps you build confidence.

Speaker 3

I think language models are also just really weird. Right? Like, with the emergent misalignment work, they I I don't know if they took predictions they should have of like, hey, I'm gonna fine tune ChatGPT on code vulnerabilities.

Is it going to become a Nazi? And I think most people would have said no. And that's what happened.

And so what are the different And how did they discover that it became a Nazi? They started asking it a ton of different questions. Yeah.

And it will do all sorts of, like, vile and harmful things. Like, the whole persona just totally changes. And and, I mean, we are dealing with alien brains here who don't have the social norms of humans and or or even a clear notion of, like, what they have and haven't learned that that that we have of them, I mean.

Yep. And and so I think you really wanna go into this with with eyes wide open.

Speaker 1

if we live in a world where AI progress accelerates by the way, you were mentioning a little while ago that there's many wild worlds we could be living in, but we're living at least one of them. Another one that we've gestured at, but it's worth making more explicit is this, even if the AI models are not helping write the next training algorithm for their successor, just the fact that if they had human level learning efficiency, whatever a model is learning on the job or whatever copy of the model is learning on the job, the whole model is learning. So effect it's getting Or if they're like a thousand times less efficient than humans are learning.

That's right. And you just like deployed them even still. That's exactly.

Yeah. Yeah. Anyways, and then there's a whole bunch of other things that you can think about.

But even there, it's like you have a broadly deployed intelligence explosion.

Speaker 2

there's this whole spectrum of crazy futures. But the one that I feel we're almost guaranteed to get, and this is like almost a strong statement to make, is one where, like at the very least, you get drop in like white collar worker at some point in the next five years. It's like I think it's very likely in two, but it seems almost overdetermined in like five.

And and on like the grand scheme of things, those are kind of irrelevant timeframes. Like it's the same either way. And that completely changes the world over the next decade.

And and if the sort of if we don't have the right policies in place for that, then you end up actually with almost in some respects like a fundamentally worse world. Because the thing that these models get good at by default is like software engineering and like computer using agents and this kind of stuff. And then we need to we will need to put in extra effort to put them in the loops where they help us with scientific research or they're like, we have the right robotics such that we actually, experience an increase in material quality of life.

So that's worth thinking about. Like, if you're in the perspective of, I'm a country. Like, what should I be doing or thinking about?

Speaker 1

what you should be doing to prepare policies. What should you be doing to prepare? Cause honestly, honestly, it's like such a tough question.

Yeah. Where like, if you're India or Nigeria or Australia. Yeah.

If you're a country unlike America or China where they do have frontier models. Yeah. What is it that you should be doing right now, especially on such a short time scale?

Yes.

Speaker 2

So I think one very important point is that let's say this this scenario turns out true. Then compute becomes the most valuable resource in the world. Yep.

Like the sort of GDP of your economy is dramatically affected by how much compute you can deploy towards these sort of organizations within your country. And so having some guaranteed amount of compute, I think will actually be quite important. So like pre getting ahead of investments in like data centers and this kind of stuff on the condition that it's like companies in your country have to be allowed to use that compute.

And yeah, not necessarily for training, like just even just for inference. Like I think the economic value here comes from inference. I think it also makes sense to invest broadly in AI, like I think these countries have the opportunity to do so.

And I think that's like a portfolio of like, foundation model companies, but also like robotic supply chain and this kind of stuff. I think that you should invest very proactively in policies that try and prove, like, prevent capital lock in. So we're in for a much worse world if, like, it just so happens that the people who had like money in the stock exchange or in land before AGI are like dramatically more wealthy than the people who don't.

Because it's a gross misallocation of resources. So having like, I know one of my favorite episodes actually on your podcast was like the Georgism one where you're trying to appropriately value or allocate land. And so I think this strikes particularly close to home coming from Australia where I think our policies with respect to land are like grossly Yeah.

Wrong. But think this is broadly true. Being very forward on regulation of integration of these models into your country is is important and proactively making sure that people have choice.

Like, so let's say you should be quite proactive about making sure that like the phones or devices or like glasses that people have, people have like free choice on like what Yeah. Things they run. And then so that's like that's the we just get white collar worker.

Right? And like you're trying to like do the best to like prepare your country for that. And then it's like, okay, well, what can you do to make all possible versions of the future go well?

That's covering some amount of economic downside. The other things I think are really important is figure out how you can either make the basically ensure dramatic upside or cover terrible downside. And so getting dramatic upside is making sure that there is investment in biology research and this kind of stuff in an automated way that these models are actually, like, able to produce novel medicines that mass like, massively improve our, like, quality of life.

And covering the downside is like AI alignment research and this kind of stuff and automated testing and, like, really thinking hard about that, AI safety institutes, this kind of stuff. But these seem like things that a rich person a random rich person could also do.

Speaker 1

there's not a thing that a nation state is uniquely equipped to do. Yeah. That's good point.

In this in this scenario. I mean, like dramatic allocation of re of, like, resource towards compute, I think is is sensible.

Speaker 2

I would be doing that if I was in charge of a nation state. Think it just increases your optionality in, like, most of the future worlds. Yep.

Speaker 3

Dylan Patel has some scary forecasts on US energy. Yeah, versus China. Yes.

Yeah, we're like 34 gigawatts off.

Speaker 2

basically and China's line is like this. And I mean, US like very clearly Yeah. We just need so many more power plants.

Yes. If intelligence becomes this like incredibly valuable input, like intelligence becomes almost a raw input into the economies and quality of life of future, the thing directly underneath that is energy.

Speaker 3

more access to intelligence on top. Yeah. Just to make it explicit because we've been touching on it here.

Even if AI progress totally stalls, you think that the models are really spiky and they don't have general intelligence, It's so economically valuable and sufficiently easy to collect data Yes. On all of these different jobs, these white collar job tasks, such that to Shalto's point, we will we should expect to see them automated within the next five years. Yeah.

Even if you need to hand spoon every single task to the model. It's like economically worthwhile to do so.

Speaker 2

progress stalls out and we just never figure out how to keep progress going, which I don't think is the case. That hasn't stalled out yet, it seems to be going The current suite of algorithms are sufficient to automate white collar work provided you have enough of the right kinds of data. Yes.

And in a way that, like, compared to the TAM of salaries for all of those kinds of work is so, like, trivially worthwhile. Yep. Yeah.

Exactly.

Speaker 3

I I I do just wanna flag as well that there's a really dystopian future if you take Moravec's paradox to its extreme, which is this paradox where we think that the most valuable things that humans can do or the smartest things are like add large numbers in our heads or do any sort of white collar Yeah. Work, and then we totally take for granted our fine motor skill and coordination. But from an evolutionary perspective, it's the opposite.

So we got like, evolution has optimized fine motor coordination so well. And even if you look at, like, robot hands or, like, even the ability to open a door is still just, really hard for robots. Meanwhile, we're seeing this total automation of coding and everything else that we've seen is clever.

The the really scary future is one in which AIs can do everything except for the physical robotic tasks, in which case you'll have humans with, like, AirPods and, like Glasses. Glasses, and there'll be some robot overlord controlling the human through cameras by just, like, telling it what to do and, like, having a bounding box around the thing you're supposed to pick up. And so you have, like, human meat robots.

Speaker 2

Mhmm. And and not, like, necessarily saying that, like, that's what the AIs would be, like, want to do or anything like that. But as in, like, if you were to be, like, what are the relative economic value of things?

Like, the AIs are out there doing computer programming and, the most valuable thing that humans can do is, like, be amazing robots. Yeah. Now that being said, I think Morovik's paradox is a little bit fake.

I think the main reason that robots are worse than at, like being a robot than they are at software engineering is the Internet exists for software engineering. Like GitHub exists and there is no equivalent thing. Like if you had all like, you know, mocap of everyone's actions as they were like going about their daily lives, like some reasonable fraction of the human population, robotics is also like close to solved.

Like like like on track to be solved at the same rate that software engineering is on track to be solved. So this is only like this vision is only like a sort of decade long section, but it's still a pretty terrible decade. Like imagine the world where people have lost their jobs, you haven't yet got novel biological research that means people's quality of life is dramatically better, You don't yet have material abundance because haven't actually been able to action the physical world in the necessary way.

You can't build dramatically more because building dramatically more takes robots basically. And people's like main comparative advantage is as fantastic robots is like a shocking, shocking world.

Speaker 1

Infrared, the the perception of an average human, I think it actually might be better. Your like wages will be higher because you're you're the complement to something that is enormously valuable. Right.

Which is AI labor. Right. And like, you know, a decade or two on, like the world is fantastic.

Yep. Right?

Speaker 2

is solved and you just like to get like, you know, like radical abundance basically, provided that you have all the policies set up, like necessary to permit building. Like you you sort of you end up with that same change from, you know, the like the the before and after photos of Shanghai Yeah. Yeah.

Where like twenty years on, it's like this dramatically transformed city. Like a lot of places in the world probably end up like that Right. Over that two decade period.

But we need to make sure, like, one, do our best to estimate, is this actually what is on track to happen? Like, build SWE bench before all the other forms of white collar work and measure and track. That's a great thing that government should be doing by the way.

It's like trying to break down the sort of functions of their economy into measurable tasks and figuring out where what does the curve actually look like for that? Because they might be a bit shocked by the progress there. There's no sweep bench for taxi And then I don't like have all the answers here, like figuring out a way to like share the proceeds of this economy like broadly across people or like invest heavily in robotics and collecting data so that we get robotics faster and get material abundance faster, invest in biological research that we get but like all that faster.

Speaker 1

because because otherwise you have a pretty dark Yeah. Like section. I think one thing that's not appreciated enough is how much of our leverage on the future, given the fact that our labor isn't gonna be worth that much, comes from our economic and political systems surviving.

For your million X S and P equity to mean something, for your contracts to mean anything, for the government to be able to tax the AI labor and give you a UBI off of that. It just like that requires our legal institutions, economic institutions, our financial rails surviving into the future. Yes.

The way in which that likely happens is if it's also in the AI's best interests that they follow those rails. And by AI, I don't mean some monolithic single AI. I just mean like firms which are employing AI and becoming more productive as a result.

You don't want to be in a position where it's so onerous to operate in our system that you're basically selecting for firms who either immigrate or who are like doing black market stuff, etcetera. And which means I think like you want to make it super, super easy to deploy AI, have the equivalent of special economic zones, etcetera. Because otherwise you are just surrendering the future outside of any control that you might have on it.

One of the reasons by the way that I worry about turning AGI into a national security issue or having it have extremely close ties with the government, the Manhattan Project thing, is that it disproportionately redirects the use of AI towards military tech and the mosquito drones and whatever. And and also naturally puts other countries in the same frame of mind, right? If we're developing the mosquito drones, why would China not develop the mosquito drones?

And that just seems like a zero sum race and not to mention a potentially catastrophic one. Whereas like, you know, like compute will be limited, know, we want, we will need to disproportionately accelerate some things to the extent it just remains totally like a consumer free market landscape.

Speaker 2

It just seems more likely that we'll get the glorious transhumanist future where they're developing the things that make human life better. Yes. I I mean, I agree.

Like, the the case where you end up with, two national projects facing off against each other is dramatically worse. Right. Like, we don't wanna live in that world.

Yeah. It's much much better if it's, like, stays in a freak market, so to speak. Yeah.

Yeah. Yeah.

Speaker 1

I wanna take issue with your claim that even if with the with the algorithms of today, if we just collect enough data Yeah. That we could automate white collar work. First, let me get an understanding of what you mean by that.

So do you mean that we would do the analogous thing of free training with all the trajectories of everything people would do on their jobs? Could you could you make either manually or through some other process some RL procedure based on the screen recordings of every white color worker? What kind of thing are you imagining?

I mean, like a continuous distribution of this stuff. Yeah.

Speaker 2

like, important, like, mental model, to think about RL is I think as, like, the the task gets more, there is some respect with which like longer horizon or better that task, if you can do them, if you can get that reward ever, are like easier to judge. So like, again, come back to that like, can you make money on the Internet? That's incredibly easy reward signal to judge.

But to like do that, there's like a whole hierarchy of like complex behavior. So if you could like pre train up to the easy to judge reward signals, does your website work? Does it go down?

Like do people like it? Like there's there's all these reward signals that we can respond to because we have a long we can like progress through these long enough trajectories to actually like get to interesting things. If you're stuck in this regime where like you need a reward signal every five tokens, it's a way more painful and like long process.

But if you could like pre train on every like screen in America, then probably the like RL tasks that you can design Interesting. Are very different to like, if you could only like take the existing internet as it is today. And so, like, how much of that you get access to, like, changes the the mix.

Interesting.

Speaker 1

and it takes longer for them to get any any signal on whether they successfully complete the task, will that slow down progress because it takes more compute per per task?

Speaker 3

the longer, the harder tasks, the more training is required. And Yeah. I'm sympathetic to that naively, but we as humans are very good at practicing the hard parts of tasks and and decomposing them.

And I think once models get good enough at the basic stuff, they can just rehearse or fast forward to the more difficult parts. I mean, that's definitely one of the complexities. Right?

Like, you use more compute and, like, the and as you train on, like, more and more difficult tasks.

Speaker 2

I don't know. Your rate of improvement of biology is gonna be, like, somewhat bound by the time it takes a cell to grow in a way that your rate of improvement on math isn't, for example. So yes, but I think for many things we'll be able to paralyze far like wisely enough and and get enough iteration loops.

Mhmm. Yeah.

Speaker 1

Will will the the regime of training new models go away? Will will we eventually get to like you've you've got the model and then you just keep adding more skills to it with RL training?

Speaker 2

whether or not you think like, there's a virtue in pretraining a new architecture. Basically, if you make some, like, architectural change, then you, like, probably need to, like, do some form of, like, at least, like, retraining a new model. Mhmm.

Speaker 1

if RL requires a bunch of inference to do the training in the first place Mhmm. Does that push against the thing you were talking about where we actually need a bigger model in order to have brain like energy?

Speaker 3

But then also it's more expensive to train it in RL. So where where does that balance out? I I think we gotta drink the bitter lesson here.

And Yeah. Yeah. Like, you you there aren't infinite shortcuts.

Like, you do just have to scale Something's a hack. And have a bigger model and pay more inference for it. And if you yeah.

If you want AGI, then that's what you gotta gotta pay the price of.

Speaker 2

there is science to do, which, you know, everyone is doing of what is the optimal point at which to do URL. Mhmm. Because you need something which can both learn and discover the sparse reward itself.

So you don't want a one parameter model, useless, even though you can run it really fast. You also don't want a 100 t model because it's super slow. Yeah.

Password RL. And I'd like the sort of the marginal benefit of like its learning efficiency is like not worth it. Right.

So there's like a there's a pretty different view here. Like what's the optimal model size of like your current class of capabilities and like your current set of RL environments and this kind of stuff. Yeah.

And and even in the last year, there's been much more of a factor of the inference cost. Right? So just explicitly, like, the bigger the model, the more expensive it is to do a forward pass and generate tokens.

Speaker 3

And the calculus used to just be, should I allocate my flops to more training data or a bigger model? Yeah. And now another huge factor is how much am I actually gonna do forward passes on on this model once it's trained?

Yeah. My total pool of, like, compute. How do I allocate that across train data compute and inference compute for the RL training?

And then even within inference, there's all this research on, well, what strategy should I use? Should I sample 10 and take the best? Do I do this sort of like branching search, etcetera, etcetera?

Yeah. And so with RL, where you're sampling a whole lot of tokens, you also need to factor the ability for the model to, like, actually generate those tokens and then and then learn and get feedback. Mhmm.

Okay.

Speaker 1

what is your advice to somebody early in their career or a student in college? How should they be, what should they be planning on doing? Yeah.

Speaker 2

So I think once again, there's like it's worth considering the spectrum of possible worlds and preparing yourself for that. Yeah. And the one like, the sort of action that I think is like highest EV in that case is you are about to get dramatic at a minimum, you are about to get dramatically more leverage.

Mhmm. You already have. Like, already the startups in YC are like writing huge amounts of their code with, you know, Claude.

So what challenges, what causes do you want to change in the world with that added leverage? Like, you had 10 engineers, at your beck and call, what would you do? Or if you had a company at your beck and call, like, what would that enable you to do?

And what problems and domains suddenly become tractable? That's the world you wanna prepare for. Now that still requires a lot of technical depth.

Obviously, is the case where AI just becomes dramatically better than, like, everyone at everything. Right? But for at least a while, probably, there is like advantage I think Jensen actually talked about this in an interview in an interesting way where he's like, you know, have like a 100,000 general intelligences around me and I'm still like somewhat useful, because I'm there like, know, directing the values and like, like, you know, like, asking them to do things and, know, they're sort of like, I still have value even though I have 100,000 general intelligences.

And for many people, think that will still be true for a fair while. And then, you know, as the AIs get better and better and better and like so on, eventually no. But again, prepare for like the spectrum of possible worlds because in the event where we're just totally out competed, it doesn't matter what you do.

But in all the other worlds, it matters a lot. Get the technical depth. Study biology.

Study CS. Like, really think hard about study physics. Think about hard about what challenges you wanna solve in the world.

Yeah. That's a lot of topics. That's a lot You can now.

You can. Right? Like, it's so much easier to learn.

That's right. Yeah. Everyone now has the, like, infinite perfect tutor.

Yeah. Yeah. Yeah.

Yeah. It's definitely been helpful to me. Yeah.

I would say some combination of, like, get rid of the sunk cost of your, like, previous workflows or expertise Yeah.

Speaker 3

In order to evaluate what AI can do for you. That's right. And and another way to put this, which is fun, is just, like, be lazier in so much as, like, figure out the way that the agent can do the things that are toilsome.

But but it's you're gonna have to in this you're ultimately, you get to be lazier, but in the short run, you need to, like, critically think about the things you're currently doing and, like, what an AI could actually be better at doing. Yeah. And then go and try it or explore it.

Because I think there's, like, still just a lot of low hanging fruit of people assuming and not writing the full prompt, giving a few examples That's right. Connecting the right tools for for your work to be accelerated automated. Yep.

Yep.

Speaker 1

There's also this on cost of feeling like since you're not quote unquote early to AI that you've sort of missed the boat and you can't like but I think I mean, I remember when GPT-three came out. So backstory on the podcast. When I graduated college, I was planning on doing some sort of AI rapper startup.

And the podcast was just like a gateway into doing that. And so I was trying out like different things. And at the time I remember thinking, oh, 3.

5 is out. And people are like, I'm like so behind on like the startup scene here or whatever, I wanted to make my own rapper. I mean, maybe the idea of the rapper was inadvisable in the first place, but just like the, every time feels early because like it's sort of, if it's an exponentially growing process And there were many things, many ideas which are only becoming possible now.

Right? So Exactly. It's that product exponential I talked about.

That's right. Like products literally obsolete it. Like you need to constantly reinvent yourself to stay at the, like, frontier of capabilities.

By the way, do you remember? I had a really shitty idea I and gave you a what it was. It was it was like I think it was like rag for like lawyers or something.

Right. Yeah. Anyways, I gave you I think one of our first interactions was I'm like, hey, what do you think of this idea?

And you're like, I think the podcast sounds promising. That's right. Which I appreciate.

Yeah.

Speaker 3

I I got slightly annoyed at a a friend recently who I think is really talented and clever and interested in AI, but has pursued a biology route. And I just kind of tried to shake them of, like, you can work on AI if you want to. I mean, I I I think humans are artificial or not artificial, our biological general intelligences where a lot of the things of value are just very general.

Yeah. And whatever kind of specialization that you've done maybe just doesn't matter that much. I mean, again, it gets back to the sunk cost.

But, like, so many of the people, even my, like, colleagues excited about AI, and they just don't let their previous career be a blocker. And because they're just, like, innately smart, talented, driven, whatever else, they're they end up being very successful in in finding roles. It's not as if they were in AI forever.

I mean, people have come from totally different fields. And so don't think that you need, like, permission from some abstract entity to, like, get involved and apply and be able to contribute.

Speaker 1

be an AI researcher, like, right now if you give them the open problem or, like, this the kind of open problem that is very likely to be the be quite impressive, what would it be?

Speaker 2

I think that now that our role has, like, come back, papers building on Andy Jones' scaling board lengths, like scaling walls for board games are interesting. Showing that you can investigating these questions like the ones you asked before where you're like, oh, is the model actually learning to do more than its previous pass at k or is it just like Yeah. Yeah.

Discovering that? Like, exploring questions like that deeply, I think are interesting. Yeah.

Yeah. Like scaling laws for RL, basically.

Speaker 1

the marginal increase in meta learning from a new task or something.

Speaker 3

I mean, on that note, I think I think model diffing has, like, a bunch of opportunities. Yeah. Also, people say, oh, we're not capturing all the features.

There's all this stuff left on the table. What is that stuff that's left on the table? Yeah.

Like, if the model's jailbroken, is it using existing features that you've identified? Is it only using the error terms that you haven't captured? Yeah.

I don't know. There's a lot here. I think Matt's is great.

The Anthropic fellowship has been going really well. Mhmm. Goodfire, Anthropic invested in recently.

They're doing a lot of interpretability work or just applied Anything to get your equity up, There's just so many interpretability projects that are are like there's so much low hanging fruit, and we need more people, and I don't think we have much time. Yeah. I also wanna make a plug for performance engineering.

Speaker 2

I think this is one of the, like, like, best ways to sort of demonstrate that you have, like, the raw ability to to do it. Like, if you made a extremely efficient transformer implementation on TPU or Trainium or like in CUDA, then think there's a pretty high likelihood that you'll get a job offer. But there's a relatively small pool of people that you can trust to, like, completely own end to end the performance of a of a model.

Mhmm. And and if you have broad, deep electrical engineering skills, I think you can probably come up to speed pretty fast on accelerator stuff. Yeah.

You can come up to speed, like, reasonably fast, and it teaches you a lot of good intuitions of the actual intricacies of what's going on in the models, which means that you're then very well placed to, like, think about architecture and this kind of stuff.

Speaker 1

One of my favorite people in thinking about architecture and Anthropic at the moment actually, like, came from, like, a heavy GPU kernel programming background, just Hognosians announced really deeply and can think about the trade offs really well. This is fun, guys. Yeah.

It's fun. Thanks. Yep.

Great to be back. I hope you enjoyed this episode. If you did, the most helpful thing you can do is just share it with other people who you think might enjoy it.

Send it to your friends, your group chats, Twitter, wherever else. Just let the word go forth. Other than that, super helpful if you can subscribe on YouTube and leave a five star review on Apple Podcasts and Spotify.

Check out the sponsors in the description below. If you wanna sponsor a future episode, go to duarkesh.com/advertise.

Thank you for tuning in. I'll see you on the next one.

Shared via Hopper