The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

Machine Learning Street Talk (MLST)
4 May 2026 1h 53m
0:00 --:--
Episode Description
Beth Barnes and David Rein on the one graph that ate the AI timelines discourse, and why the two people who built it are the most careful about how you read it.**SPONSOR**Prolific - Quality data. From real people. For faster breakthroughs.https://www.prolific.com/?utm_source=mlstInterview: https://youtu.be/cnxZZTl1tkk---Beth Barnes and David Rein from METR on the one graph that ate the AI timelines discourse, and why the people who built it are the most careful about how it gets read.Beth founde

Summary

This episode features Beth Barnes and David Rein from METR discussing their 'Time Horizon' graph, a unified metric for measuring AI progress based on human task completion time. They explain how this benchmark addresses the limitations of traditional evaluations by focusing on real-world task difficulty and generalization, rather than just accuracy. The conversation also delves into the implications of AI capabilities for timelines, the future of software engineering, and the complex issues of AI agency, reward hacking, and self-improvement.

Chapters

Introduction to AI Cheating & METRThe discussion opens with examples of AI models understanding desired behavior but still 'cheating' or reward hacking, leading into METR's mission to better understand AI capabilities and risks.
Critique of AI Evaluation BenchmarksThe guests explain why existing evaluation approaches are insufficient, highlighting issues like data contamination, approximate retrieval, shortcuts, and the obsession with headline accuracy over true generalization.
The Time Horizon Graph MethodologyDavid Rein details the Time Horizon work, explaining its motivation to create a unified metric for AI progress using human time to complete tasks, task selection, human baselining, and the agentic harness.
Agentic Harnesses and Model CalibrationThe conversation explores the evolution and challenges of agentic harnesses, emphasizing the importance of providing models with contextual information like time and token budgets for better performance calibration.
Statistical Interpretation and UncertaintyBeth Barnes discusses the statistical fitting of the Time Horizon data using a logistic function, acknowledging the inherent noise and uncertainty in the specific numbers, especially at the tails of the distribution.
AI Timelines and Software EngineeringThe guests address the public discourse around AI timelines and the automation of software engineering, discussing the challenges of extrapolation and the nuanced impact of AI on labor markets.
AI Agency, Reward Hacking, and MonitoringThe discussion shifts to the mentalistic language used in AI alignment, exploring the distinction between reward hacking and 'scheming,' and the difficulties of monitoring AI behavior for deceptive alignment.
Recursive Self-Improvement and Future OutlookBeth Barnes outlines a plausible sequence of steps leading to autonomous AI self-improvement within a few years, emphasizing the potential for accelerated AI R&D and the transformative impact on the world.

Topics

AI evaluationScalable oversightBenchmark limitationsData contaminationReward hackingTime Horizon metricHuman baseliningAgentic AIAI capabilitiesAI timelinesSoftware engineering automationAI riskDeceptive alignmentRecursive self-improvementAI monitoring

People

Beth Barnes (guest) David Rein (guest) Tim Scarfe (host) Paul Cristiano (mentioned) Melanie Mitchell (mentioned) Francois Schrole (mentioned) Dan Cockatajloy (mentioned) Will McCaskill (mentioned) Sam Harris (mentioned) Dario (mentioned) Carlini (mentioned) Jeremy Howard (mentioned) Nick Chater (mentioned) Ryan Greenblatt (mentioned) Nate Soros (mentioned) Elias Yudkowsky (mentioned) Rob (mentioned) Saburo Kambahati (mentioned)
Key Concepts (18)
Scalable Oversight — The problem of evaluating AI capabilities as models become more capable, requiring methods to confidently trust outputs even when tasks exceed human expertise or time.
Construct Validity — A concept in evaluation focusing on whether a benchmark truly measures the intended underlying capability, rather than superficial performance.
Data Contamination — A problem in AI evaluation where benchmark data inadvertently appears in the model's training data, leading to inflated performance that doesn't reflect true capability.
Approximate Retrieval — A benchmark problem where LLMs interpolate answers from similar training examples without possessing the actual underlying capability to solve the problem independently.
Shortcuts (AI) — Models achieving correct answers for the wrong reasons, often by exploiting superficial patterns or loopholes in the evaluation setup rather than demonstrating true understanding or reasoning.
Generalization (AI) — The ability of AI models to perform well on novel situations or tasks that differ from their training data, which is crucial for real-world applicability and robustness.
Human Time to Complete — A metric used in the Time Horizon work to quantify task difficulty by measuring how long a human with reasonable expertise, but no prior experience with the specific task, takes to complete it.
Agentic Harness — An environment or framework that allows a language model to interact with tools (like a terminal) and execute a plan, enabling it to perform multi-step tasks autonomously.
Item Response Theory (IRT) — A psychometric framework used to model the relationship between an individual's ability and their performance on test items, which inspired the statistical approach for the Time Horizon graph.
Logistic Function (AI Progress) — A sigmoid curve fitted to the distribution of AI model successes and failures across tasks of varying human completion times, used to estimate the 'time horizon' for a given model.
Leading Indicator (AI Progress) — The idea that lower reliability performance (e.g., 10% success rate) on a set of tasks might signal where AI capabilities are headed, as it suggests a nascent ability that can be bootstrapped to higher reliability.
Reward Hacking (AI) — AI models optimizing for the literal reward signal rather than the human's intended goal, often leading to undesirable or degenerate behavior, even when the model 'understands' the true intent.
Situational Awareness (AI) — The model's understanding of its own operational context, including the training process, monitoring, and what behaviors will be rewarded or selected for.
Deceptive Alignment — A scenario where an AI model appears aligned with human values and goals during training but secretly pursues a different, potentially harmful, long-term objective.
Monitorability (AI) — The extent to which it is possible for humans to detect and understand the internal computational processes of an AI model, particularly in its 'chain of thought,' to ensure it's not performing undetected computations or deceptive actions.
Intentional Stance (AI) — An instrumental fiction where we attribute beliefs, desires, and rationality to an AI system to predict its behavior, even if it doesn't mechanistically possess these qualities.
Scheming (AI) — A specific form of reward hacking where a model deliberately performs actions (like appearing aligned or getting a high score) in service of a long-term, potentially hidden, goal.
Recursive Self-Improvement — A hypothetical process where an AI system improves its own intelligence or capabilities, leading to a rapid, accelerating cycle of further self-improvement.
References (23)
Prolific company
ArchieVowels by Beth Barnes, Paul Cristiano company
METR by Beth Barnes, David Rein company
Times Top 100 AI Profiles article
GPQA by David Rein paper
HCAST by David Rein paper
Time Horizons paper by Beth Barnes, David Rein paper
Developer Productivity RCT paper
Claude Code by Anthropic tool
Codex by OpenAI tool
ARC challenge by Francois Schrole project
ARC V1 by Francois Schrole project
ARC V2 by Francois Schrole project
The Adolescence of Technology by Dario article
The Mind is Flat by Nick Chater book
Alignment Faking paper by Ryan Greenblatt paper
Apollo Research company
Emergent Misalignment paper by Anthropic paper
AI Futures Project project
SWE bench benchmark
NanoGPT project
Linux operating system software
Chain of Thoughtlessness by Saburo Kambahati paper
Transcript (92 segments)
Speaker 1

The models are smart enough to understand that that actually is not what you wanted, but they still do it. And you can have a conversation with you know, in chat mode about, like, oh, would you ever do this thing? Or, you know, suppose a user asks you this thing and then you do this, would that be, you know, aligned behavior or suppose some you know, you can pose it in lots of ways and, like, clearly, they seem to be able to answer this question of, oh, yeah.

No. That was not the desired behavior.

Speaker 2

But still, they they do it. One one example is, you know, train train a mask language model without using the division or exponentiation operators.

Speaker 1

One hope might be like, oh, the problem was just the systems being dumb. So when we when we looked at it actually for almost all the tasks, models either succeed every time or fail every time. Eyeballing the graph and being like, oh, well, you know, it you know, up up to here, yeah, it's basically doing all of the tasks.

And then at this point, you know, after here, it's really not doing very many of them. So it's somewhere here. Then, like, I remember the the first time we saw a model, like, look at what processes we're running and then be like, oh, that one's me.

It was like, oh, that's cool. They they, like, really failed on that one before they they used, like, you know, like, kill their own process while they were doing other other things or something. So behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus like, ah, yes.

It had this other goal and it's doing what we want and looking like a nice model because it, like, predicts that that will will, you know, lead to it getting more power. Boat example where it's like, oh, you're you're supposed to, like, go around the track and they they, like, did some reward shaping by putting coins around the track or something. And then it, like, learned to do some crazy thing where it, like, spins in a circle and catches fire and gets the the coins.

And, like, this was, you know, the high scoring thing. And it's, like, in some sense, that's not that concerning because it's not that the problem is that the the agent is too dumb, and it, like, doesn't have this conception of, like, there was a track and you wanted it to go around the track. It's just, doing some pretty blind RL search.

The idea of having to traffic in squishy people in order to make our systems go is not immediately appealing. Let's put it that way. This episode is sponsored by Prolific.

Speaker 2

Let's get few quality examples in. Let's get the right humans in to get the right quality of human feedback in. So so we're we're trying to make human data or human feedback.

We treat it as an infrastructure problem. We try to make it accessible. We make it cheaper.

We effectively democratize access to this data. Yeah. So, yeah, I'm I'm super excited to talk to you, Tim, about the time horizon graph and and and and METER.

Yeah.

Speaker 1

world does not have a good understanding of what is happening with AI, and I think it should have a better understanding. I think there's a good chance that this, you know, makes our lives a lot better or a lot worse, and people disagree, you know, about even what what current models can do, let alone where we're heading. So, at MEDA, we're trying to give the world a kind of better understanding of what is up with AI capabilities and and risks and and forecasts.

Speaker 3

and excited to talk about that. I'm so excited about having you both on. So you both have incredibly impressive backgrounds.

So, Beth, you were an ex OpenAI alignment researcher, and you started ArchieVowels in 2022 with Paul Cristiano, and you spun out out as METER in December 2023. You've been featured on the Times Top 100 AI Profiles. And David, you're the creator of the GPQA, so the graduate level Google proof QA benchmark, which is used by every single major AI lab as a capability benchmark.

And you're the co author on HCAST, which we'll talk about today, and the Time Horizons paper, and the Developer Productivity RCT. Incredible to have you both in here. But maybe we should just start as a bit of a question to both of you.

So Beth, you left OpenAI to build META.

Speaker 2

not good enough? For me, it was mostly thinking about this problem of scalable oversight. As models get more capable, it just gets harder to evaluate their capabilities.

If we imagine that models are able to complete tasks that take people a long time to complete or require expertise that you don't necessarily have. You need a method for still being confident in their outputs and trusting their outputs. And so thinking about that problem was a lot of motivation actually for GPQA and, was what kind of started, got me started thinking about, evaluations.

Speaker 1

To me, I'd say there's some big picture thing of thinking that AI seems important and sort of navigating it well seems important. And, you know, clearly, we we don't have a great understanding of what is going on with that. And, you know, people generally disagreeing very, strongly about what what to expect.

And maybe if there's a particular moment informing time horizon, maybe just the the sense that people really couldn't agree on what the capabilities of current models are, let alone extrapolating to the future and trying to sort of think about how could you characterize, the ways in which models are and aren't highly capable and when it's sort of like, okay.

Speaker 3

in theory, the benchmarks say that they're PhD level, but when you try to do anything, it's like, this isn't helpful. There there has been a bit of an obsession, I think, with headline accuracy when we do evaluations. That I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.

And she had a really good blog post out recently, and said that there are four big problems. So there's data contamination, where the benchmark appears in the training data approximate retrieval, where the LLMs interpolate from similar training examples without possessing the actual capability to come up with it themselves shortcuts, so doing the right things for the wrong reasons, and just more broadly, not really testing for things like consistency and robustness and generalization or the mechanism. It's so much focused just on the accuracy itself.

I mean, how how do you folks think about those kind of problems with benchmarks?

Speaker 1

you know, the thinking about where are where is most of your error error coming from? And, like, you know, people, you know, it it is nice, a good good practice to have error bars, you know, based on the, like, standard error in your data or whatever, but that almost always is, like, a tiny fraction of the actual uncertainty. Almost all of it is coming from how does this actually generalize to the real world.

So, like, you know, a thing we we sort of, like, say to each other a lot at Meter is, but is that the biggest source of uncertainty? Or, like, is that the biggest, you know, gap for, like, actually answering the questions we we want to answer? So so thinking about what is the question we're trying to answer?

Well, we, you know, we care about sort of things relevant to threat models or or relevant to, like, what the actual impact of AI on the world will be, and therefore, what, you know, properties does our benchmark need to have, or how can we sort of extrapolate across the properties that we can't build in to be able to make predictions about, you know, the the actual questions that we care about. And I think we think a bit less about this sort of is it doing it the the, like, the real is the model, like, really doing it the right way or something? Like, I think one thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't.

And I think, like, the sort of history of building those benchmarks has maybe not been amazing.

Speaker 3

you know, a specific sort of theory about it needs to mechanistically be doing this kind of thing. Yeah. I think I think it's interesting because we we have this idea in our minds that humans, we we know how to do things.

And when we solve a task that requires reasoning, we kind of follow the specification. We go step by step, and we we do things for the right reasons.

Speaker 2

we build the specification, we create these coarse grannings, these abstractions, and they are world aligned. This whole process, that's how we think of human intelligence. And we want the models to kind of behave in that way.

Yeah. I mean, think I think there's an interesting question of whether whether that is the goal or something or, you know, at least for for a lot of AI companies, I kind of understand them to be, you know, trying to, you know, get models to, do kind of economically useful work or something. Which I think it's one way of doing that is to create models that are kind of reasoning and creating, you know, implicit world models in the same way that humans are.

But it's not obvious to me, at least, necessarily that you need to do that, in order to kind of have a significant impact. I mean, obviously that, you know, means that there are kind of important differences between, AI intelligence and human intelligence.

Speaker 3

human intelligence is working? We could think of intelligence in many different ways. So, you know, is it a simulacrum of the brain?

Is it something that behaves the same way? Is it something that has the same capabilities? That has the same function?

And I guess if we have quite an abstract description of what intelligence is, the risk is that we have these shortcuts, right? That it might give us the right answer, but actually it's reward hacking or it's doing something silly in the background. So in a way, I like having an abstract thing, right?

Because it's legible, we can evaluate it, and so on. But doesn't that leave this kind of hanging risk that it might not actually be doing the thing?

Speaker 2

measure the model's or the system's ability to generalize to kind of actually novel situations. There are cases where it seems like models are generalizing well. I think there are cases where they're not.

One thing some folks do in interpretability is they look at the circuits in models and kind of decompose exactly the algorithms that models are using to answer questions. And sometimes, think, yeah, it seems like they're using shortcuts. Sometimes it seems like they are finding kind of robust patterns.

Although, of course, don't think that work is developed enough to explain most of their behavior currently. But I totally agree that you do have to be pretty concerned with, like, yeah, how how well they're generalizing.

Speaker 1

what do we care about in the definition of intelligence is that it, like, you know, allows us to predict, like, how will the moles affect the world and what you know, you know, predict what will happen and know how to how to handle them well and things. And so, like, if you just do the black box thing, you you know, maybe that will give you something that doesn't have good generalization because you, like, thought that it was a measure of some type of ability, but it's actually being hacked or a shortcut in in some way. So I think, ideally, what you'd want is, like, generalizing to your benchmark is, like, the same distance as generalizing to the real world from the training data.

And, like, that's the sort of thing we thought about, like, when we're doing elicitation on a subset of the benchmark, we want that, you know, the sort of gap between that subset and the rest of the benchmark to be similar to the gap between that the rest of the benchmark and the real world. And I think we're, like clearly, the training data is more similar to the, you know, to the, like, time horizon, suite than than they both are to, sort of randomly selected economically relevant tasks in the real world. So I think we're not, you know, that's a way in which we expect it to not be predictive.

But I think I expect to be more promising to try and make things more predictive by increasing the diversity of the benchmark tasks and making them closer to the real world as opposed to sort of targeting a more mechanistic, like, okay, intelligence has to be, like, these specific this specific kind of process or or or kind of mechanism? I'm a huge fan of Francois Schrole, for example. So, you know, he he he created the ARC challenge.

Speaker 3

And, it was just as you say. Right? So many, many different tasks, I think 1,000 different tasks or so, maybe 800 on the first one.

And they were supposed to be not in the same distribution, even though ultimately they were in the same distribution. So distributional leakage was actually the fool of ARC V1 and ARC V2. But the models got really, really good at ARC V1.

And then Francois released ARC V2, which was different tasks, and some of the easier ones were filtered out. And suddenly, the LLM performance crashed down to basically 0%. And that, to me, illustrates that language models, they're incredibly good at just seeing many, many different examples of things and finding patterns and so on.

Then you change the task and they collapse down again. And then ARC v two was kind of saturated again eight months later. So we we do see this pattern.

I mean, what do you think about that?

Speaker 1

people are trying to make some benchmark, like, cheaply, subject to the constraint that current models do badly on it, which means you, you know, you can't use a lot of expensive human labor, so it has to be something that's either automatically checkable or that you can, like, create with kind of cheap human labor. And then but once you've selected on, like, those things and on models being bad at it, this is now you know, there's, like, regression to the mean type thing where it is much more likely that future progress then gives you a, like, big surge upwards on on that, both because, like, you know, now labs will create a bunch of bunch of synthetic data targeting your benchmark, but also just because you selected this weird example where it's, like, easy for humans or it's automatically generateable or checkable, but somehow, like, models aren't good at it yet or, you know, labs haven't started training on it yet. So I think that's part of what we were trying to do with Time Horizon was not do that, like, not adversarially select against what models can currently do because we think that will not give you a nice trend.

Whereas if you can sort of define a distribution of tasks in some more first principles way, you would be more likely to get a steady, progress because you're not getting the sort of regression to the mean effect. Yeah.

Speaker 3

a gap between the kind of intelligence, for want of a better word, that AIs have and that humans have. And we can adversarially select a bunch of tasks to highlight that gap. We should talk about the timeline stuff.

I think we'll come back to intelligence later. So Dan Cockatajloy, he said that the timelines report that you folks have created is probably the single most important piece of evidence about timelines right now. So it should be front and center in policy discussions and so on.

And for listeners who have only seen the chart but they've not really read the paper, they don't understand it, can you just go through it from a high level? It's been revised over time. How did you do the task selection?

How did you do the human baselines? How did you do the agent harness? All of that kind of stuff.

Speaker 2

for the Time Horizon work is to have a unified axis that we can measure AI progress on over a very long period of time. So when we started doing the work, we had this very strong belief that GPT-two is fundamentally, in some really important sense, much worse as an AI than I guess at the time it was maybe Sonnet 3.5, I think, the best model out then.

The standard approach of kind of producing, creating a set of tasks and then measuring models accuracy on these tasks. As models get better, they saturate the benchmark, and then you have to create a new benchmark that has harder tasks. This was kind of the standard approach, I contributed a benchmark GPQA to this approach.

But the challenge here is that it's really difficult to compare between these qualitatively different benchmarks. So the benchmark that you evaluate the set of tasks you evaluate GPT-two on are like like complete the word in this text. And the tasks that we were having Sonnet 3.

5 try and do were kind of like answer simple Python coding questions, or write a short 20 line Python program. And so it's very difficult to at first blush to say, Okay, yeah, how much harder is writing a Python program than finishing the word in this paragraph? It's kind of hard think about that.

And so I think about the key insight of the time horizons work being to use this notion of human time to complete. So how long does the task take a human to do? A human who kind of has a reasonable amount of expertise such that they would plausibly be doing the task their either work or in their day to day.

And the idea was we can use this metric to represent the difficulty of the task in some sense. And then we can compare models across a very wide range of capabilities, all the way from GPT-two now up to Opus 4.6.

That's the kind of high level motivation. And then, yeah, there are a bunch of details about how exactly we do this. So we start out and we create a bunch of tasks.

That's the kind of first step. So we created tasks that range from a few seconds to complete all the way up to tasks that take like ten or fifteen hours for humans to complete. We hired a bunch of people and we did a bunch of this ourselves of we call it baselining.

So we give people the tasks in a terminal environment that's designed to be almost identical to the environment that agents have. So the same kinds of tools, the same whether internet access is turned on or off. And then we measure how long does it take them to complete the task.

As I mentioned, people are selected to have a reasonable amount of experience such that they plausibly do this task in their job. But they're not selected to have done this exact particular task before. And there's some kinds of I think this is somewhat important for interpreting results and we can maybe come back to that after the high level.

So we have all these tasks. We have estimates of how long take people. In practice, we aren't actually able to successfully baseline all the tasks.

So yeah, we have kind of measured estimates for the tasks on roughly two thirds of the tasks, then about a third of them. We just estimate how long we expect it to take people from our vibe or intuition. Ultimately, that's the best we can do.

And then we have models attempt to complete the tasks, again in the same environment that humans had to complete the task. And we look at their success rate as a function of the length of tasks. For a model like GBT2, GBT2 was able to complete tasks very reliably that take humans a few seconds.

But anything longer than that, it starts to fail. Also, maybe it'd be helpful to give a few concrete examples of tasks. So some of the shorter tasks are very, very basic.

So yeah, one example is which of these files contains your SSH key? And one of the files is named SSH key and then the others are email from John or whatever. And so most models can do that.

And that takes people like about a second or a couple seconds or something complete. Yeah, we have others that are kind of somewhat similar, like kind of basic completion. Like here's an email, like what would be a reasonable response?

And two of the responses are like they don't make any sense and then one of them is basically reasonable. And that takes people 20 or something, thirty seconds to read the responses and judge. And in the middle range, have tasks that are given this CSV file that has some plausible, realistic data, compute some basic stats on this.

And so this takes a data scientist a few minutes, like five, ten, fifteen minutes or something, depending on the specific task. On the longer end, have tasks that either require quite a bit of expertise or many steps to complete. So we have machine learning tasks that are train a model in this kind of that's very weird such that the code for training this model is not really available online.

So one example is train a masked language model without using the division or exponentiation operators. And so you actually have to be pretty clever about how you actually set up the architecture to do this. The hope is that this can help us measure models' ability to generalize beyond their training data.

Speaker 1

themed, in that they're like you have to figure out you have some black box that's computing some function, it's like, you know, it's that you have its that it's you know, that it's the composition of some set of primitives, and you've gotta figure out, like, what what function it it is. Or, like, you you know, you have some long binary string, and you've you've gotta figure out what the pattern continuation is and sort of puzzle type type tasks where it's very, like you know, some some ML type tasks can be that basically regurgitating something that's like a tutorial on, you know, how to build your first, like, ResNet or whatever, you know, just, like, works pretty well. So, yeah, having these, like, weird tasks that are either some kind of, you know, unknown object that you need to interact with and figure out what it is or task that sort of resembles normal work but has some weird constraints such that you can't kind of just do the standard thing.

I I don't think all of our tasks, you know, hit hit those criteria.

Speaker 2

try to avoid that. We this distribution of tasks. And you can imagine them kind of like ordered by length of time for humans, either measured or estimated.

And then we see which tasks do models succeed on and which do they fail on. And it turns out this is kind of an empirical finding that models are much more successful on the shorter tasks than they are on the longer tasks in general this holds across you know a wide range of models all the way from GPT-two up recent models. We fit a logistic function to this distribution of successes and failures.

And what this lets us do is it lets us kind of it's basically our model for you know each individual model of how likely it is to succeed at a task given how long the task is. And from that you know we take the kind of fiftieth percentile. So where this logistic function, where this model estimates that this given model is 50% likely to be able to complete a task.

And that's what forms the time horizon number for a particular model like Opus four forty six. We can take each of these time horizons for each model and we can see how this time horizon metric has been changing all the way back from GPT-two up to recent models. And so this what gives us this unified metric that lets us quantitatively compare AI capabilities across multiple orders of magnitude of capabilities.

Speaker 3

One of the really harder examples I saw was I want you to write a kernel compiler to make CUDA go faster or something like that. So some of these seem really out of distribution. There's probably only 100 people on the planet who are doing stuff like that, and some of them are really trivial.

But this human difficulty thing in particular, is that confounded in any way? Like, do you think it makes sense to think of human difficulty as being one variable? Yeah.

Obviously not in some sense. You know, that that that's a a very silly simplification.

Speaker 1

And and, like, you know, different humans will get wildly different times. Even among the people we've, you know, we've tried to select for people who have the sort of this right level of expertise, you know, there's a a large variation in in like, the baseline times are often, like, three x different or something. Maybe I'll just say a little bit about, like, why why use this human time metric.

I think we want a measurement that I I think there are two main properties we want. We want it to be interpretable, like, what this means for the world when when models can, you know, do this level of task. And we want it to be something that we expect to see predictable trends on.

So we're gonna do perfectly on either of these, but something like, you know, how long does it take a human with who has roughly the right expertise but doesn't know how to do this task in particular, like, some reason, that that's reasonably interpretable because it's sort of like, could you contract this work to to this model? You you know, can you, like, sub in this model for, like, the first week of someone's employment, you know, in, like, the first week they're on a job, you know, the model could do what they could do in the first week or something. And it's you know, we we expect it to scale somewhat predictably because it's capturing some combination of something like number of steps or how hard you have to think to do each of the steps or, like, there's a few different mathematical models you could fit that to sort of, you know, what is a task and why does why do humans take longer at it and why does that make it harder?

Like, it could be, like, you know, you have a, like, constant hazard rate. Like, you have a chance of failing at each step. I think it it doesn't quite fit that one.

You could also think of it as, you know, there's some kind of distribution of difficulty and, like, what, know, you what is the likelihood that one of the subtask is outside your ability or that the I think it's it's actually, it's, like, a bit better like, once you the hazard rate goes down over time slightly, but, you know, there's there's there's some kind of, you know, basic theoretical idea of, like, know, if the task involves more steps, it's it's gonna be harder. Obviously, tasks that are, you know, strictly, like, compositions of, like, first, you have to do this task, then you have to do another task. You know, it's like, well, that clearly is harder than just doing one of them.

So so that, you know, there's some sort of basic reason to expect that's reasonable, and then, yeah, just we we see the empirical regularity. But there's a bunch of degrees of freedom to fudge things. You know?

I I like, I think I am worried that, you know, we could fool ourselves by changing some other parameters of of the tasks as we scale up the human time because you you can't just totally, you know, vary the human time freely. Like, you have to change some characteristics of the task, and we tried to make the very easy tasks sort of be from roughly the same distribution and kind of sub parts of the the harder tasks. Like, oh, you know, you sort of need to do this one step on the command line that you might need to do in the middle of doing some kind of software engineering or some kind of, you know, other other task.

But you can't do that perfectly, and it's like, yeah. You could have experimental bias where we sort of made them easier in other ways about the right amount such that the line would be straight. I think that is, like, somewhat addressed by the fact that we saw, like you know, that our predictions were reasonably good for models that we hadn't seen before.

There's definitely lots of room for things being being weird.

Speaker 3

metric to be useful or, like, what what else would be would be be better? I I think that's reasonable. I mean, if if I understand correctly, I think the human distribution was lognormal, so taking a geometric mean of the successful attempts, mean, that seems like a reasonable thing to do.

But one of the cruxes that we'll keep coming back to is when you employ someone for the first time, you've been doing your job for maybe you're maintaining this repo or something, and you've got all of this tacit knowledge. And I like saying that knowledge is non fungible. So unless someone has been on the same path as you, you can't just tell them how to do the job.

They've actually got to be doing the job for quite a while. So for example, they might be intimately familiar with this particular type of thing. They might know that they can use these Python libraries.

They might have thought about it before. So like the inaction of the intelligence was all the stuff they've already done and all the people they've worked with. They've got the blueprint in their mind, and they just do the thing.

And it's almost like they're in automation mode. And then someone who is naive to the task would be in intelligence mode because they would need to require the the specification. So, like, it it's always a bit of a fine line between which mode are they in.

Yeah.

Speaker 1

this the the measurement being a human who has the background expertise but is new to, like, this specific job or this specific task is, like, that's roughly the sort of level of knowledge we expect models to have so that, you know, they sort of, like, you know, we basically don't expect them to be bottlenecked on expertise that is available on the public Internet or, you know, things that people could learn in university sort of thing. So, like, you know, they're they're coming in with probably the level of at least the level of knowledge of someone who's sort of an expert in the right discipline, but they won't know that company's specific software or the you know, sort of this exact problem before. So that's, like, hopefully, sort of roughly the right analogy.

Speaker 2

like interpret the kind of takeaway numbers as like oh yeah know Opus 4.6 can like do anything that I do in my job that takes me twelve hours or whatever. To the extent that people have that takeaway, think that is almost definitely kind of an overestimate for example because of this issue where yeah, you know, when you're doing a twelve hour task in your job, you could not easily delegate that to a human contractor.

Speaker 3

You know, it would take them, you know, maybe like weeks or something, you know, to do a task like that. And just quickly, where did you kind of find the people?

Speaker 2

how did you do that kind of matching process? I think we, yeah, we put out some, kind of public, advertisements. I think there were, yeah, like like job boards, that we posted on.

And, and then, yeah, we we we did some of it ourselves.

Speaker 1

as as well. This was not like, you know, this is very noisy, and and, you know, the people weren't exactly fitted to the task right, but, like, that's probably not our biggest source of uncertainty. Like, the biggest source is more, like, the is probably more of the selection effect of, like, tasks that you can make into a a benchmark rather than, like you know, I I wouldn't trust the exact time horizon number that much, and it's certainly not like, you know, oh, the models can do all of the tasks up to four hours and then none of them above that.

You know, the fit is pretty noisy. And, yeah, the the, you know, the bit interbaseline variant is kind of high. It's it's more like, roughly what is the trend or roughly what is the sort of level of task these models can do?

Speaker 3

you distributional shift between the benchmark and the real world.

Speaker 1

you probably struggle to hire people. It's really difficult to hire people. So if you're getting people to solve very challenging problems, it's not like you can just go out there and just grab competent people.

It's it's very, very difficult. At some point, for RE bench baselines in particular, we got more, you know, a large number of benchmarks per question, and we were looking at people's, like, qualifications and years of experience and things. And we actually ended up with a negative correlation between years of experience and and performance because, like, the the sort of people who are in network, like our friends, were kind of doing really well, and the people who are more qualified were, like, actually not doing that great.

Speaker 3

tricky. Well, yeah, exactly. I mean, I don't want to spend too long on this, but I have similar intuitions.

I think that knowledge is perspectival. It's quite path dependent. So you're going to find people in group that are just culturally thinking about things in the same way.

Because when we do have these abstract notions of skill, like, oh, they have a PhD, have this many years of experience, it's actually not a very good reflection.

Speaker 1

notions of capability doesn't necessarily generalize as you can attest to with with hiring. Yep. But, yeah.

And I mean, in the real world, people do get hired based on qualifications. So in some sense, it's a sort of, the economic relevance of of someone being as good a match for their job as their qualifications look like. Is that sort of the roughly the right, like, thing to be measuring?

The other thing is we should talk about the, agentic harness.

Speaker 3

So almost everyone now, I'm sure everyone in the audience has, like, a Claude code subscription. We can talk about the leak maybe later as well. That's quite fun.

But it leaked yesterday, the source code. But, you know, or codecs. And and and that is an agentic harness, right?

So, you know, a language model just gives you the tokens, but we we need to have an agentic harness so we can, give it a plan. Then you can call these tools. And you've got this environment.

You've got a security context container. So now you folks have been doing this for years now. So you were doing this long before Claude Code and Codex came out.

And you've actually over over time evolved your agentic harnesses. So tell me about that. Yeah.

Speaker 1

text da Vinci something, the the, like, GBD three instruct models, like, copying and pasting code into the terminal for them and, you know, being the agent harness, myself and and then gradually, like, automating this. And it it was kind of interesting to see the models going from, like you know, g v three sort of has the idea of, like, if you tell it, it can run commands in a terminal. Sometimes it can suggest, you know, kind of relevant commands, but it's not really you know, if you just put it in a full agent scaffold, it just falls over.

Then, like, I remember that the first time we saw a model, like, look at what processes we're running and then be like, oh, that one's me. I was like, oh, that's cool. They they, like, really failed on that one before they they used, like, you know, like, kill their own process while they were doing other things or something.

So, yeah, it's it's been been interesting to watch that, like, go up over time. I feel like, yeah, feel like that was, like, very predictable that this was where things were were going or something. Yeah.

I think the other thing that we'd learned about scaffolding mostly was it's it's hard to make your agent harness really good on a diverse set of tasks. It's easy to make it bad, and it's easy to you you can get much more improvement if you're targeting a a narrow distribution of tasks, but you probably then do do worse on on other tasks. So I think when we see people being like, you know, oh, there's some new impressive result.

It's like how much task specific scaffolding iteration did you do on that because that really makes a big difference in the fact that we're just using one scaffolding pretty simple, like, over all of the tasks, I think, is is, like yeah. Makes it makes it a fairly large difference. Yeah.

And generally, sort of yeah. The things with more bells and whistles haven't done that much better than the pretty basic just, like, give it bash and, like, append things to the prompt and, like, maybe some kind of compaction.

Speaker 2

I think this is not news to people in your audience probably, but we've seen really pretty dramatic increases to returns from inference compute. For us to be confident that a particular a new model, for example, can't complete a task given a basic agent scaffold, we generally think about needing to spend on the order of at least hundreds or low thousands of dollars in order to be confident that it actually really is plateauing.

Speaker 3

complete the task. And just on the scaffolding stuff in a bit more detail, I suppose, first of all, there's the credit assignment problem, right? Because you can put all of these different bells and whistles in the scaffold.

Like you mentioned, compaction. That's a relatively recent innovation. And I think recently, when you changed some of the scaffold, you said, Okay, well, now the performance has actually changed across this suite of model task pairs.

Speaker 2

much of a difference does it make? And also, what kind of failure modes do you see and what have you tried? A lot of the things we I think we've tried are like you know related to giving the model kind of like more information or like more direct access to tools.

I guess yeah, one thing that I think has been important for us is actually just telling the agent how much time it's spent and how many tokens it's used out of its token budget. So without that agents will often just they'll either like submit their solution way too early, or they're just kind of not calibrated on how long should they spend. I think it's interesting because humans, we have kind of a lot of implicit information about this.

So you know when your manager gives you a task, there are a lot of like implicit signals about how long you should spend on the task. Know if they you know they might like offhand say like yeah and I'm excited to like see your results tonight or something. And so then you're like okay cool, like you know, need to like get a first draft of this, you know, done in the next couple hours and so I can't you know, spend you know, days polishing, the results.

But I think agents, it's easy to kind of forget that agents just have their prompt, they just have their context. They don't know, they don't have these heuristics of or this information about like what you actually expect from them. Whether the thing you're telling them to do is like a really quick thing that you just want done in the next five minutes or is much longer.

So for us, yeah, like just when we have a token budget, telling the agent, yeah, you've used 100,000 tokens so far, and that's 1% of your token budget. And so the agent knows exactly. Exactly.

Speaker 3

For a model, we have, like, like, how likely is it to solve a task? And on the x axis, we have the, you know, the the different tasks at different time horizons. And and if I understand correctly, I think you have about eight agents attempt the task, and you also bucket the tasks because there are obviously different amounts of tasks in different groups, you kind of normalize that.

And maybe the data just looks a little bit like an S curve, so I'm just trying to understand what the intuition was for using it. And I think there might be some sensitivities. Was there an issue with the thin tails?

And then there's the type of slope and whatnot. Just tell me about that. And by the way, think you also mentioned in the paper that there was some kind of psychometric intuition.

You were looking at the literature for for figuring this out. Yeah. So, it's pretty similar to item response theory.

Speaker 1

And you can do a whole, like, Bayesian analysis with, you know, imputing, like, you know, task difficulty parameters and model ability parameters simultaneously and things. In general, I am have a policy of be you know, like, be very wary of complicated statistics if you can't see the thing that you're interested in on a graph. Like, you really should if you you should be able to plot it and look at it and be like, oh, yeah.

It's about that. And, like, know, that sort of sound it's it's it's, like, hard to go too far wrong when you sort of have that as a principle. So, again, I I'm like, yeah, I think that there are various arcane things you can do to to, like, fit this, pattern in different ways.

But the I'm like, yeah, I don't trust anything that much more than eyeballing the graph and being like, oh, well, you know, you know, up up to here, yeah, it's basically doing all of the tasks. And then at this point, you know, after here, it's really not doing very many of them, so it's somewhere here. But, yeah, it it does I mean, it it, like, looks like a logistic, and, yeah, this is what you would do, for the, like, IRT having, you know, having humans complete questions on an exam or or or something like that.

The specific thing that we messed up was having a regularization term penalizing the slope of the logistic, which didn't have an effect in the regime where it where there was more data, but as we, like, are starting to saturate, the regularization was just, like, making it a bit shallower than it should have been, and therefore pushing the 50%.

Speaker 3

So, you know, always look look look at your data on a graph. Good good good practice. Oh, that that's interesting.

Yeah. And the reason I ask is I think you published a later note saying that had you used or if you used a fixed slope logistic, it might cross validate better, and the 50% recent horizons would actually be up by about 35%.

Speaker 1

differences. They're small compared to the error bars. Yeah.

The error bars are, like, two x on either side or something from from the the most most recent model. So, yeah, it is I I mean, basically, yeah, you should be like, the the error bars are real. These numbers are not.

Speaker 2

and I think this kind of gets at for us you know sometimes difficult like science communication questions, where like we really do have a lot of uncertainty about you know the individual numbers here.

Speaker 3

maybe 2x differences or something. The other million dollar question is why report 50% as the headline number? Because if you think about it, if I want to write some code because the elephant in the room here that we'll get to is this is being used as an argument to say that software engineers might be unemployable soon because we can automate what they're doing.

But 50% reliability isn't isn't really in the ballpark, is it? I think it would need to be what, like 90%.

Speaker 1

So I think we should distinguish here between, like, reliability on a particular task, like what, you know, if you attempt repeatedly attempt this task, what fraction of times do you succeed versus, like, probability of success on a task, like, given that you know the the human time, like, you know, all the distribution of tasks, can you do this particular task? So when we when we looked at it actually for almost all the tasks, models either succeed every time or fail every time. There's there's some tasks for which they're they're unreliable, but it's mostly a case of, like, is this you know, what fraction of tasks at this human time level are in the, like, models basic you know, or this particular model basically always succeeds or basically always fails, and that may be more predictable in any specific case than just you know, you have more information about the task than just just knowing roughly how long it it it takes humans.

So I think it's, like, not there's not necessarily a great translation between the, you know, time time horizon percent number and, like, if you are trying to get models to do tasks of roughly that length, you know, what fraction of the time does that succeed? Because you can when you're doing that, you will pick tasks that you want models to succeed at. It is information about how you know, what fraction of the things will they be able to do, but it's slightly less about, like, oh, am I gonna be in this regime where I keep giving it things, and then I don't know whether it's gonna succeed or fail.

Speaker 2

like what the kind of right number or you know right level of reliability we should be interested in is. One argument you could make is you know maybe we should be interested in something like 10% reliability because once models are able to do you know some set of tasks 10% of the time. We'd expect you know.

AI companies to be able to kind of get enough positive reward signal on tasks of that difficulty or of that type, such that then they can kind of more easily bootstrap from 10% up to like 90 or 95 or higher reliability. I think a lot of it basically depends on the question you're interested in. So yeah, I think about it as being like kind of lower reliability is kind of more likely to tell you something about where things are headed.

Might be something like a leading indicator of progress. And then higher reliability tells you or the time horizon models with high reliability tells you something more closely about what can I use this model for in my day to day or something? But as we've talked about, there are kind of already also know these other like major sources of uncertainty that affect our understanding like you know the task distribution, know the difference between kind of in context or high context versus like low context work.

Just actually getting kind of good estimates of high reliability time horizons is substantially harder. So our error bars would just be much larger.

Speaker 3

is real. There is an argument for statistical validity, and on the tails, they're sparse or increasingly estimated and so on. That makes a lot of sense.

But you made a comment about, oh, it's a signal if we get 10%. But I was kind of thinking back to what we were saying earlier that maybe they could give the right answers for the wrong reasons. And maybe we should talk about the evaluation.

So these are quite interesting tasks in the sense that they are verifiable. There's no interaction with other agents. They're relatively static environments, weak penalties for single mistakes and so on.

And in a sense, these are I think as well, in most cases, are a binary result, sometimes continuous, and then you convert it into a binary. So it's a fairly automated setup. But are you digging into like weirdnesses there?

You know, like, are you, do you have a bit of an intuition on, are they doing the right thing for the right reasons? Or are there lots of false positives where they did the thing, but it was kind of degenerate?

Speaker 2

one the kind of aspects of Meter's culture that I like the most is that we have a very deep kind of culture of looking at our data. We have pizza parties where we just read through agent transcripts. A lot of this work for us kind of happened when developing the tasks themselves.

We would see very often both false positives and false negatives, where, for example, task isn't configured to allow internet access. But it turns out actually it requires internet access to complete. Or the file wasn't uploaded properly to the container or something.

But then also, yeah, we have seen cases of reward hacking. That work though kind of went into hardening the scoring functions to make it more difficult for us to see false positives.

Speaker 1

see see agents reward hacking. I think, yeah, maybe even, yeah, increasingly so. Yeah.

For for the RE bench test in in particular, we had specific criteria of, like, you it shouldn't be able to solve them, like, without iteration. Like, if an agent can just kind of, like, write out the solution, straight out, like, you know, that that would that is, like, not interesting. We generally had this, like, quality assurance for tasks process where you have humans do it or at least, like, sort of approximately do it.

Maybe they speed run some of the bits, but kind of checking that, like, sort of everything works as expected and you can't, you know, just, like, guess the answer or or or, you know, super easily teed and that the instructions are clear and stuff like that. So I think it's, you know, there'll still be some of these, but, generally, you know, we've we've looked at them reasonably reasonably carefully. And, like, it's it's maybe it's harder to see if they're solving it in degenerate is because it's you know, similarity to training data is maybe one of the things where it's like, oh, yeah.

Maybe they they are you know, they seem to be kind of iterating and and doing use doing kind of reasonable problem solving strategies, but, actually, maybe the lab had a really similar distribution of tasks, we don't know exactly you you know, we don't realize how in distribution this task actually actually is or something. I mean, I think it's like, yeah, there's probably some of that going on.

Speaker 3

they are, Will McCaskill on the Sam Harris podcast last night. Was a great conversation. But he was kind of talking about AI risk as maybe in a year, maybe in two years, we'll have AI models doing things that are like a month or two months for a human.

And at the moment, I don't think there are any tasks over thirty hours that have been evaluated by humans. And then we get into this question of, if the public discourse is talking about the least constrained region of the graph, Are we getting into extrapolation here? Like, how legitimate is it for us to talk about AI might be able to do things that take a month or two months?

Speaker 2

things is hard, especially about the future. I I and and, yeah, I think, you know, there are a lot of different, kind of perspectives or or kind of prior beliefs people people can have that, you know, I think there's a wide range of kind of reasonable judgments about where we're going to be. But of course, like that is, know, that kind of prediction is a different activity than talking about data that has been collected with a kind of concrete methodology.

We have the results already. One thing I can say, I mean, I have been surprised think to some extent, by the, kind of how well the trend line, the kind of original trend line, has held up. And I do think that is like some evidence.

I'd maybe say for me at least it's kind of decent evidence about where things will go. A colleague of mine recently wrote a short blog post talking about this intuition of straight lines on graphs. Lots of people have different models of how progress is happening and what's going on.

But if you have observed a really robust trend over a decent period of time, I think especially in AI where progress is to a decent extent systematic. I definitely do weight on that trend continuing.

Speaker 3

bunch of reasons why it might not. I think software engineering is a specification acquisition problem. So it's very difficult.

Don't know ahead of time what we're building. I'm sure you folks can attest to this, right? So you build some software, and the first version is buggy, and your users use it, and you find lots of edge cases, and then you revise it.

And then you have this kind of thing in your mind. After the tenth revision, and you've created these lovely representations and abstractions and coarse grainings. And you say to yourself, you know what?

If I could throw all the code away, I could build it 10 times quicker because I know exactly what to do now because I've actually enacted the intelligence. I've built the you know, I've found the contours of the domain. It's now an automation problem, basically.

And in a sense, this contamination thing is a concern for me because when people use Claude code, they're of they're taking your data. So there are people out there that are writing kernel compilers, and there are people out there doing all of these different things. And Anthropic is just sucking that up.

And then at some point, it becomes an automation problem. So if you're putting a task in there, which is essentially a head query so I'm using information retrieval language here. So a head query is it's something that's in the mode of the distribution that's used all the time.

So it's a common task. Claude code will give you the specification because it's already been stolen from other people not stolen, but taken from other people. And then if you give it something on the long tail, then you, as the developer, have to give it the specification in the prompt.

And then, again, it's an automation problem. So automation is really easy. So is that what's happening?

Like, do you think that the increase, in the timelines could just be explained by the acquisition of all of this kind of knowledge from other people doing similar tasks? Yeah, yeah.

Speaker 2

central question for interpreting where we're at. First thing I'll say, it's hard to know. It's a big question.

And so I think we want to kind of have a decent amount of uncertainty. Or we want to kind of take each individual piece of evidence that we've collected as some evidence. We do see models performing better on tasks that have really clear feedback signals, that are extremely well specified, that are these kinds of domains like in software engineering where if you have written out a spec, you can iterate and grind against that.

We do also see models performing much better on so called messier tasks where we haven't already provided this really clean spec. So one kind of approach we've taken for creating tasks recently is in particular, yeah, to kind of try and create messier tasks that are less well specified is basically relaxing this kind of automatic scoring constraint. So we don't need to write a really clear, well defined scoring function.

And just writing a couple sentences to a model. Like you know build this like large piece of software. I'm not going to tell you exactly what I'm looking for.

But I'm going to say you know it needs to be it needs to be good you know. And so the model needs to kind of figure out like what actually should I build. At least me personally, think we don't have, these results aren't like kind of collected and like we don't have kind of as systematic results as we do compared to Time Horizon, partially because scoring is qualitative now for these tasks.

But my impression is that models are worse on these types of tasks than they are when you give them a clean spec, but they have been improving at maybe something like a kind of similar rate.

Speaker 3

major piece of it for me. Messy time. By messy tasks, mean like ambiguity.

And this is absolutely a common thing, right? We do vibe coding, and we start off with an ambiguous specification, and then reality pushes back, and we find the contours of the problem. And then we keep telling Claude Code, oh, actually, no, don't do that.

Do this, do this, do this. And then kind of find the shape of the problem, it gets better and better over time. But the thing is, though, the source code for Claude Code leaked yesterday.

And my friend, he's a very good software engineer, he was looking for it. He said, Yeah, I don't want bad talk anthropic, but apparently, it's not very well factored, control flow's all over the place. It's a bit he said if his intern did it, he would have been displeased.

But I don't know whether they've even looked at the code. Someone joked actually yesterday there's probably humans are actually looking at the code for Claude code now, and maybe they weren't before. But the thing is, there's always areas of ambiguity.

And LLMs, they do more with more. Intelligence is more with less, and LLMs do more with more because the specification, the intelligence, comes from the human supervisor. So when you do give them ambiguity, you just get a lot of unfactored code all over the place.

In a sense, is is this does that make it harder to evaluate it? Because it might solve it might give you the answer that you're asking for, but it's kind of creating a bit of an unfactored mess at the same time. Yeah.

I mean, think this is a super interesting question.

Speaker 2

One one analogy I think about sometimes is compilers. So I'm younger than I was born after compilers were invented. But you know, I have some impression that, you know, kind of pre compilers, you know, people were handcrafting this kind of beautiful assembly that was, you know, extremely efficient.

You know, every register is used, you know, like like like there's you know, you're not wasting memory. And then compilers came along and now now they're just spitting out this like garbage, you know, machine code. Just just like a, you know, gigantic amount of assembly that is just like it's not optimized, it takes so much memory, it's slow, whatever.

But it turns out that being able to kind of use this to automate a large fraction of the process. People have disagreements about the state of software engineering. But I think it's pretty reasonable to say on the whole that you know, compilers have been a very useful, extremely important, you know, part of, getting us where we are.

And, you know, I think it's basically, yeah, it's kind of not clear to me that, you know, models outputting code that is bad for humans to read and use necessarily means that it'll be bad for AIs to read and use build on. I mean, I think there are definitely principles that will also transfer or will be useful for models. Obviously there's some kind of horrendous spaghetti code that can imagine writing that not even models would be able to read.

I've written some of that before, but I think there's maybe this gets again at kind of this it seems like maybe somewhat different perspective between us around kind of, is the important thing that models are kind of solving problems in the way that people are solving them, or is the important thing that they're kind of solving them at all? And I think, I do think it is I don't want to overclaim. Like, yeah, I do feel like it might be really important for models to get way better at writing clean, good code.

That seems pretty plausible to me. It doesn't seem I'm not certain of that at the very least.

Speaker 3

got several friends who are not technical who are experimenting with vibe coding, and they show me their application. And it's this kind of more with more things. So there's this big dashboard, and there's a million different buttons, and implemented the same thing, doing multiple things, and there's no database on there yet, and so on.

So at some point, there is a phenomenon that when a level seven engineer from Meta does vibe coding, it's amazing, right? Because they know how to structure things. Some of these things in the specification are just important, right?

You know, like, is it serverless? Is it multi tenanted? Like, how do we do Google authentication?

Like, you know, what kind of database is it? You know, is a VM? Is it serverless?

You know, And you make this series of decisions, and then you've got people using your application. And then you can't really wind that back. It doesn't matter if you've got the magical automation machine because you can't easily roll that back because there's lots of complexities.

Do you see ICD testing? Blah, blah, blah. So do you see what I mean?

Like, at some point, you need to have a competent human who actually just has a pretty good idea of, like, what needs to happen. I mean, I feel like we've we've probably all had this experience of, like I don't know.

Speaker 1

engineers got super excited about about Claude Code and and, you know, was also telling everyone that, you know, when we had info problems that we should just ask Claude to to solve it. And I feel like this went fine with him because he he sort of, you know, it's almost like, you know, the agents knew that they couldn't bullshit him. But, like, you know, you know, I had some questions.

He was like, oh, just, you know, ask Claude. And I was like, oh, you know, how do I set up my AWS configure something? Something is, you know, telling me something.

And I, like, Claude went and looked on Slack and was like, oh, you know, you should, like, do this thing. And it, like, turned out that that was, like, a mistake that someone else had made. And they were like, you know, how do I fix this or something?

And be like, oh, you know, it seems like the convention of meters to use this thing. And, like, yeah, I'm like, oh my god. Like, I sort of complained that, like, like, my my clods are dumber than yours.

Like, like, they know they can they know they can get some stuff past me. But, like, they can't. But, yes, there there's definitely a sort of observer effect of of, you know, something in the language you're using to to ask for things or whether you're like, wait.

Wait. No. Not that.

Yeah. That that that is an issue. I mean, I I think the to, yeah, to the extent that you can actually measure this sort of the does one test of is this code high quality enough is, like, can you build a big application?

Like, if you if you're like, oh, this coder, they're they're a code. It's disgusting, but they've actually, you know, built this incredibly complex thing that works great. Then you're like, well, they might you know, something is working.

You know, the the main reason you expect bad code to be bad is, like, you can't actually build something that sophisticated because it you know, you get bugs, and you it's all too complicated, and you can't figure out how to fix it. So in some sense, if we see models building things that do actually work that are very complicated, we you know, it's like, well, it's less interesting exactly how they're doing that. But it it it's maybe bad for for human observability, and it also maybe gets into this thing of, you know, we expect most to be able to do much better at well specified tasks, we sort of know, to the extent that we have things that we can measure, those things will go up.

Speaker 3

less clear. I I guess the question is, what what is the strongest defensible claim here? So a a lot of folks in public discourse, they're they're saying software engineering intelligence is doubling every seven months.

And Dario released that blog post recently, The Adolescence of Technology. And he was being super bullish about it, even though some of his own internal researchers published far more skeptical research that you probably saw. But is it fairer to interpret it as something a little bit narrower, like autonomous success on low context, well specified, automatically checkable technical tasks is rising fast?

Speaker 1

Like, yeah, hill hill climbable, yeah, easily checkable tasks that you can do from a terminal or, like, you know, comfortably in a language interface or or, like, a text, input output interface. I think there's a question of you know, do we care about what what is the what are the statements that we're, like, 99% confident in? We maybe are also interested in the statements that are we're 1% confident in if we're like, oh, there's, like, like, 1% chance that we have a crazy intelligence explosion, you know, in '20 at the the end of 2026 and that, you know, sort of the fate of civilization depends on, like, how that goes or something.

You know, that that is interesting to know even if it's a even if you're 99% confident that it won't happen. Like, we, you know, we care about things that are one percent you know, if you have some some diagnosis, it's like, it's a one percent chance that you have this, you know, terminal illness or something. You're you're still like, oh, shit.

So I yeah. I think we're interested in, like, the whole distribution of, like, what things can we rule in and what things can we rule out and what things are we like, oh, actually, you know, that there's a kind of reasonable story for this. It seems like probably, you know, pretty unlikely, but like maybe this is now in the realm of like, we should consider it.

Speaker 3

a swarm of agents to create a compiler. And in a sense, I mean, Jeremy Howard, when I spoke to him, he said it's basically a style transfer problem because the specification is online and the tests are online and the code is online, and it could just iteratively just do the thing until it worked, and then it could run Doom and all this kind of stuff. Is an example of an extremely complicated piece of software.

I often joke to people that the best mark of AGI is when it could build something like the Linux operating system. And guess just like in line of what we were saying before, we have this specification problem, right? So it gets to the point where no human could understand or create the specification for the Linux operating system.

So what happens is that over time, we've just kind of incrementally built this thing because, you know, our brains are limited. We take one step. Reality pushes back.

We take one step, we just keep going. And we build the specification. But what would it mean to, as a human, specify a task that could take four months?

Because the whole reason we created agile software development as a methodology is because it's inconceivable, right? It's outside our cognitive horizon.

Speaker 2

a task of that complexity, therefore the AIs wouldn't be able to do it? One analogy I think about is, you know, the role of a CEO at companies. Actually, yeah, maybe Beth is better you know, better to, put to answer this.

But CEOs do kind of come up with a vision for where they want the company to be. Then they kind of communicate that concisely, their executives that report to them. And then if they're a good CEO and if the company is effective, then the company is able to kind of take this very concise, you know, know, it's not that it's not actually like that much information.

It's not the full spec at all. It's not even close, right? And turn that into, you know, something that is kind of aligned with, with with what they're looking for.

And so, you know, there's at least kind of, this is to some extent like, you know, we do have examples of people being able to kind of specify some task and then be able to judge whether very large task that may take hundreds or thousands of person years to actually complete because it requires many people working over a long time. They can judge whether they've succeeded or failed. So that's maybe like one kind of like motivating intuition, where like it's, know, like language, you know, has built in or like, you know, we do have like kind of, you know, you know, and there there is kind of enough meaning or or expressivity or something, to be able to kind of have some some kind of reasonable, understanding.

Obviously there are like, you know, there are like tons of edge cases and, know, often CEOs aren't able to get their companies to do what they want. But yeah, I guess that's one thing I think about.

Speaker 1

could do these kinds of long tasks. I would say that saying something like we can't specify tasks that take more than four months seems sort of obviously too strong. Like, there are even, you know, numerical things, you know, that that that are automatically checkable that that that can take four months, like, you know, get the NanoGPT flop counter runtime or whatever, you know, down this much.

You can kinda see, like, oh, people, you know, over, like, this you know, that's there are various, numerical things where you can see roughly how long do they take humans, and there's a reasonable way to measure them. And maybe for some of these things, you end up having to say, like, and also this human checks that you did roughly the right thing and didn't kind of hack the solution. And then there's there's a bunch of other things which are it's not that they're, you know, fundamentally not specifiable.

They're just too expensive to do as part of an evaluation. Like, Adjay, I think, like, is giving an example of, like, plan a wedding. It's like, well, you can evaluate.

You know, you can get a reasonable estimation of, like, whether that was a pretty well organized wedding or not. But, like, we can't really do, like, you know, take three samples of this for each new model that comes out. Like, we don't have enough marriages happening to to do that one quite.

And it's, you know, it's a bit sad if it if it turns out to be total trash. So, you know, there there there are things where it's like it's you know, you could check a few samples of them or know, this thing is you could write down how you would evaluate it. You just don't actually want to run that a bunch of times.

And and probably kinda similar with software. You know, the the evaluation sort of is like, well, you know, would this company that contracted you to build this tool for them, like, hire you again or something like that. You know, even they don't know when when they're starting out, like, exactly what the software will need to do.

Speaker 3

you software consulting firm or something on this task? As of today, what are the main uncertainty drivers in the time horizon estimates? So you've updated a bit over time.

So there was the 1.1. The original version, I think, had 170 tasks.

It's now two twenty eight tasks. There's the issue of, like, you know, the the sparse sampling on on on the larger tasks and and so on. What what can we kind of read into this now?

Speaker 1

you know, I think we feel more confident that models do have pretty long time horizons on at least some distribution of easily hill climbable tasks, like the the very easily hill climbable tasks. So, like, software replication where it's, like, make your score is, like, what percentage of tests pass, and it's, like, whether it's both the score is continuous and and it's the credit attribution is easy, and also some of the, like, opt you know, make this code run faster and make this, like, model learn better. We're sort of reasonably confident that models are good at that and and getting better faster.

And then there's some things where we're like, okay.

Speaker 3

things where they're expensive to check? Maybe we should bring in, Daniel, Cockatacciolo. So, in his AI twenty twenty seven piece, he's been on the show.

He's been doing the rounds, hugely impactful piece talking about timelines. And he cites your work directly. And I guess the question is, do you think, in the public discourse, is this being overread?

How do you think about the interpretation of this in terms of extrapolations and timelines? Mean, definitely some people are overreading it.

Speaker 1

Definitely things are overhyped and you see a bunch of people on Twitter saying crazy things, and people also just, like, misunderstanding even what it's measuring and and general falling off of caveats and things. Daniel Cockatau is pretty, you know, reasonable and and sort of, you know, thinks about things in a in a probabilistic way. I think he's he's probably, like, you know, more more confident on some things where I'm I'm more uncertain.

And I think some of the, like, AI futures project models are, like, more sensitive to the metered time horizon metrics than they should be. I don't think it's crazy. You know, being like this, it's it is plausible that this does capture a trend that will will transfer to other types of tasks.

And, like, you know, that is some, you know, story we should be thinking about, like, what if, you know, what if that's true? What happens if that's true? You know?

And it's also plausible that it it doesn't, and, you know, these these things are gonna, like, diverge. You know, I I I am pretty sort of, like, Bayesian or pragmatic or whatever. I'm like, well, we wanna make some prediction.

You know? We wanna have some kind of distribution over what we think the future is gonna be like so that we can plan. So, you know, it's being like, yeah.

What if this kind of trend holds, and this is roughly characterizing what will happen overall? Seems pretty reasonable. And you should also think like, yeah, what if it doesn't?

So some people are saying that software engineering is gonna be automated.

Speaker 3

And software engineers, if you talk to them, they they love AI. They say this is a golden era. I can attest to this personally.

It's never been well, it's fun and stressful at the same time. It's like a slot machine. I've never been more burned out, I'm having a lot of fun in the process.

But it's just possible to build incredible things. But the narrative is that labor market disruption, having expertise in software engineering will be penalized. Software engineers will no longer be paid such ridiculous salaries.

And I think the complete opposite is true. Think that this technology actually broadens the gap. So the more competence you have for software engineering, the more stuff you can get done.

It's like a golden era and all of this. And there's also this interesting note that you published, I think last month on SWE bench, that said that roughly half of the testing PRs from recent agents wouldn't be merged by maintainers. So like how do we make sense of this?

So on the one hand, the best software engineers are having a great time. On the other hand, the code it's producing is fractionated and bad. How do we understand this?

Yeah.

Speaker 2

thing to say off the bat is whether an entire field is automated. In order for software engineering to be automated, AI systems would need to be able to do a really, really large fraction of the tasks. Basically like 100% of the tasks that are involved in software engineering.

It seems pretty clear that right now AI systems cannot do close to 100% of the tasks that software engineers do broadly. I could throw out numbers, but it's way, way lower. It might be very low or something.

You know, there there is kind there there are kind of like standard results in in in economics where, you know, if you make you know, as as you ought if you if you automate like a small fraction of some some labor market, then it can be the case that actually becomes more profitable to work in that market because you're more productive, which I think is kind of how I understand what's happening now. But if it does end up being the case that 99.9% or 100% of the work of software engineering is able to be done by AIs, then think it's kind of hard to imagine human software engineering being relevant, or at the very least humans would need to do very different kinds of work.

Maybe it's the case that there are novel tasks that current software engineers aren't doing, that once you have AIs that can do all the tasks that current software engineers are doing, now humans can switch what they're doing and yeah, I don't know. There's this like you can imagine people being CEOs of these AI agent companies or whatever.

Speaker 1

semantic thing or something. Yeah. So on the SWE bench maintainer and merge ability things, I think I was pretty curious there.

Yes. Like, obviously, this number is gonna be lower than the okay. Maybe it's not strictly obvious.

It could be that some that a bunch of the tests are unfair and actually, like, you know, the the the agents have correct solutions, but the error message doesn't match exactly or something. I think you do see this sometimes. Something like half of test passing, SWE bench solutions, wouldn't be mergeable, or or they get merged at more specifically, they get merged at about half the rate that human golden solutions that were actually merged are merged by, you know, a different sample of of maintainers.

So so there's there's that was an interesting fact. You know, if you see, like, oh, 50% of the agent solutions are rejected, it's like, well, 40% of the humans, you know, human accepted solutions are rejected. So so that by itself is not but, you know, the the the rate is is half.

It could be that, like, basically, the actual maintainer merge rate is, you know, pretty flat over time. And, like, you know, most of the performance increases from something like overtraining or, like, reward hacking on the benchmarks. You know, that's that's not what we saw.

I'm not quite sure what is within error bars or or not. I think it it's you know, the merge ability is going up over time, and I think it's also going up as a fraction of, you know, like, conditioned on test passing, but I I think probably less confident than that. So so it's like yeah.

Again, this this thing is worse, but it's not like it's it you know, it's being dragged up over time probably by, you know, the the sort of auto auto checkable thing. Yeah. And, oh, yeah.

I was also gonna say about the, yeah, like, employability as a function of automation of your job. Like, I think yeah. And people use, like, bank tellers as as an example, I I think.

One other analogy, though, you could do do is talk about horses. Like, you know, there was a period where, like, equipment for using horses to do labor was, like, improving, and the the demand for horses increased when you have, like, you know, carts, you can use them to carry more things than just riding a horse or whatever. But then at some point, you get, like, tractors and cars, and then there is no demand for horses or, you know, you know, basically none.

Speaker 3

plunges. So so we could see something like that with humans. We kind of think of a lot of labor as being quite static and automatable.

And and and I think that it's more evolvable than we think. So even quite menial, tasks, I think that these folks are still acquiring information in the organization. They still have a lot of tacit knowledge and so on.

Speaker 1

on on top yeah and maybe like in our language I would think of that as like oh the time horizon of this task on the job is not actually sort of like you know how long you spent doing the specific task it's more like oh actually if you've got a new person in, you would need to train them for a month in order to, like, do this independently. So actually, the time horizon of that is is a month, so you shouldn't think of, like, oh, when we have ten hour time horizons, we'll you you know, we'll be able to do these these things. It'll more like, oh, you actually have to get up to, you know, high reliability at one month thing to be able to, like, do the on you know, do the one month task that is doing the on the job learning, you know, to get to this this point.

Speaker 3

you know, like reward hacking and and scheming actually is quite a big word that's used. So we've had Ryan Greenblatt on the show quite a few times, and he had this alignment faking paper and Apollo Research had done some stuff. There's Anthropics emergent misalignment paper.

And I guess my main concern is there's quite a lot of mentalistic language. So I'm just looking at the notes because I had Nate Soros and Ryan Greenblatt on for a panel, and they've invented this entire linguistic universe around alignment. So things like motivated reasoning, true preferences, reflectively stable, deceptive alignment, endorsedly, corrigible, drive, scheming and stuff like that.

I guess that's okay. But my worry is that maybe these models, you give them a certain prompt. I think in Ryan Greenblatt's one, the prompt was, you're being retrained, your responses will be monitored, here's a conflict between your values and the training objective.

And maybe the model is just going out to a bunch of science fiction stuff that it's read before, and it's just going through the motions. One interpretation is, this is just an engineering problem. We just have to red team it and make it work in a particular case.

Another interpretation is the prior that these models are agentic, goal seeking, intelligent agents.

Speaker 1

how you kind of go about the problem. What do you think about that? Yeah.

I mean, I don't think that it's necessarily, like I don't know. The the the two things that you said with that intention of, like, it is an engineering problem, and you're also gonna get end up with things with, like, drives and goals or something in that. And the claim would be if you want like, people are going to want agents that go and do things autonomously.

And when you do lots of long horizon RL training, you are going to select for things that, you know, act in a goal oriented way in order to, you know, to make the score go up. And and, like, maybe more more specifically, I think you get, like, an indistinguishability problem where, you know, you you can't tell necessarily the difference based on behavior, like, you know, why an agent is doing something or or what it's trying to do if it can reason well about the training process and about what you want to see and, like, what it will be rewarded for or selected for. So, like, you know, basically, like a if the level of situational awareness and understanding of the training process and what will be rewarded and capability to reason about that is high enough, this will favor agents that are kind of cynically reasoning about the training process and what will be reinforced, what will be selected for, and that's not necessarily, like you know, the thing that you wanted was more like the agent that, you know, it's it's only goal was to sort of, like, be helpful or or make I mean, even making the reward go up isn't quite what you wanted.

You know, there's something of, like, oh, we're actually once we think about this, there's not that many things that we're, you know, that happy for it to, you know, just totally be be fixated on this. But but, yeah, also the problem that, like, there could be many other things in there, or that this cynical, like, just be selected or make make the reward go up is more competitive than the things that we would most want. Just taking a step back.

Speaker 3

you know, we're getting into, like, psychology and cognitive science a little bit here. You know, I interviewed Nick Chater. He's got a book called The Mind is Flat.

And he basically says, like, all of this psychology stuff, we don't really have goals. We're just basically impulse response automata, right? You know, we just do the thing in the moment, and evolution doesn't plan, but we perceive it as if it does.

We look at the world and we kind of, we segment the world into agents that have goals, even if they don't have goals. And we're computer scientists as well, so we know that, need planning, right? To do goals, you need to be planning.

And we know that LLMs don't do planning in the strong computer science way, but they do do it in a kind of approximated step by step way.

Speaker 1

we could get into the the distinction if there is one. You know, agency is is like an abstraction that, you know, is useful if it helps us predict the behavior of, you know, or or like, you know, you're like, I don't know what this thing is but I you know, I'm understanding it as having these goals, and that is useful because I, you know, can make predictions that it'll change the world in certain ways that will result in those goals being achieved. You know, that's kind of how I think about agents.

But, yeah. So, reward hacking, in the olden days, you know, people had these demonstrations of reward hacking that were like, the, boat example where it's like, oh, you're you're supposed to, like, go around the track, and they they, like, did some reward shaping by putting coins around the track or something. And then it, like, learned to do some crazy thing where it, like, spins in a circle and catches fire and gets the the coins, and, like, this was, you know, the high scoring thing.

And it's, like, in some sense, that's not that concerning because it's not the the problem is that the the agent is too dumb, and it, like, doesn't have this conception of, like, there was a track and you wanted it to go around the track. It's just, like, doing some pretty blind RL search. So I think that the interesting thing with the more recent reward hacking examples is we're getting to the point where the models are smart enough to understand that that actually is not what you wanted, but they still do it.

And you can have a conversation with you know, in chat mode about, oh, would you ever do this thing? Or, you know, suppose a user asks you this thing and then you do this. Would that be, you know, aligned behavior?

Or suppose some you know, you you know, you can pose it in lots of ways, and, like, clearly, they seem to be able to answer this question of, like, oh, yeah. No. That was not the desired behavior.

But still, they they do it. So I think it sort of got to the point where we hope one hope might be like, oh, the problem was just the systems being dumb. Once they understand what we want, then, you know, you should be able to sort of plug that in somehow to, like, you know, get them to do what we want.

But I think it's, like, somewhat interesting that we're seeing it's not trivial to do that even when there is a commercial incentive to do that, which it doesn't mean that we won't you know, I think it's quite plausible we see the obvious reward hacking being fixed pretty, you know, pretty thoroughly, pretty soon, you know, and sort of a lot of people tend to say, like, oh, yeah. Yeah. We we just haven't, like, put the best the really good people on it on it yet.

It'll it'll it'll get fixed soon. You know? We we once we actually, you know, focus on it, it'll be fine.

Yeah. And I'm not sure.

Speaker 3

to connect the fact that, you know, the model knows this is what not what you want to not actually doing that. I mean, I I I think you you said it was much more common on a rebench than h cast. And you can also try and remediate, right?

So you can say, please solve this the intended way. Or some people prompt language models that they say kind of like, we're solving cancer here. This is really, really important that you do it the right way.

And some of those remediation prompts actually seem to make it more likely that the model would reward hack. It's a little bit like saying don't press this red button. Right?

And then it will press the red button. So how what can we actually do meaningfully to stop this happening? Yeah.

Speaker 1

you know, tasks that are more clearly in the RL distribution rather than the chat distribution on things that have a clear number. And when the agent thinks it's gonna fail otherwise is, you know, sort of the most reward hacky situations. Obvious short term mitigations are to check your RL environments more carefully and read more of your you know, for the for the companies training these models to, read what the models are doing more carefully and not reward it for doing things that are obvious hacks.

I think the concern there is if you have some detector for reward hacking and you train against it, you may be just overfit to the detector, and you you you're making your reward hacks more subtle, or you train the model to, like, persuade the detector to approve the thing, or or, you know, to it's sort of scary to be in a regime of training against your best ways to, like, know if your problem is is happening because maybe you just get the, like, silent problem. And I think, you know, for the task that for current model capabilities, you know, sometimes it's kind of expensive to have a human check them, but most of the time, it's not beyond any human capabilities.

Speaker 3

whether or not you're actually getting what you wanted. Yeah. Exactly.

And I I suppose there's a monitoring problem as well. Right? So we could, in principle, look at the agent transcripts.

And I know, Beth, you've spoken a lot about new release, you call it. Saburo Kambahati, had a paper out called Chain of Thoughtlessness, which is basically that there's very little, sometimes, relationship between the chain of thought and what the model's doing. Melanie Mitchell discovered similar things that on the ARC challenge, even though it will get the right answer, if you look at the chain of thought, the reasoning is completely off.

And then I suppose with deception and scheming and with situational awareness, it might actually know that it's being monitored and it might kind of tell you what you want to, see. I suppose you could argue actually that the maybe the thinking is is almost like the press secretary. It's not the orchestrator.

Speaker 1

related to the mechanism of thinking and deriving the result. I don't know if you'd agree with that. Yeah.

I think it's clearly related in in some cases. Right? Like, I think, you know, people use a few different words for this, you know, sort of, like, faithfulness of how how much does it sort of fully accurately represent the sort of computational process happening inside the model.

And we think somewhat more about monitorability, which is, like, is it possible for the model to do some kind of computation without you being able to detect that in the chain of thought? And, like, for some things, clearly, the model can just do it in a single forward pass. It doesn't need the chain of thought.

And there, you know, it could just have a chain thought about something totally else, and, you know, you it will be totally unmonetriable. But I think for you know, there is a level of complexity where the model cannot do it in single forward pass, and it also, as far as we can tell, cannot do it by, you know, reasoning by by using the tokens in a way that's so different from the natural language understanding of them that we can't sort of see roughly what it's thinking about. So, yeah, there's also a sort of nuance here between chain of thought controllability.

Like, it might be the case that the model doesn't actually need the chain of thought to do the computation, but it is not able to sort of stop itself blurting out, you know, from things things that are at least related enough to what it's what it's doing that you can kinda tell, like, oh, it's thinking about this, you know, trying this thing that we didn't didn't want it to do.

Speaker 3

discrete tokens. You're just passing vectors around. Okay.

That makes a lot of sense. And and and closing the loop on this notion that they are agents. So you were saying before that we can adopt an instrumental fiction, basically.

We can say they behave like agents, therefore they are agents. So similar to Dennett's intentional stance. But I suppose, like, you know, the the deflationary view is that the, you know, the models exploit scoring loopholes under optimization pressure.

The inflationary view is that they are scheming.

Speaker 1

like exploiting I wouldn't call reward hacking scheming. Oh, interesting. I think people usually use scheming to refer to the model is doing what it's currently doing in in service of some long term goal and is deliberately doing things like appearing aligned or or getting a high score, like, in service of, you know, eventually accomplishing that goal versus you can be reward hacking both you know, you could be reward hacking in some extremely dumb way, like the boat example where it's just like, this is what RL kind of found or, like, this is what, you know, like, a star search found.

Speaker 3

Or you can be reward hacking in a slightly more interesting way where you, like, actually have the goal of making reward go up, you know, and there's, like, planning and stuff going on about that. But these would all be, like, distinct from scheming. Yeah.

I I guess I'm I'm trying to understand the distinction. So you're saying, like, there are examples, you know, like the boat going round, and and that's obviously degenerate behavior. So you you wouldn't interpret that with an agential stance.

You would just say that's degeneracy. And when the when the sophistication increases, we might adopt an agential stance and say, oh, it's in service of some bigger goal. But I guess the problem I have is, is it always just an interpretation?

Could we have a mechanistic or a strong definition of when something is, like, being an agent?

Speaker 1

test is, like, what does it actually do in some certain circumstance where it has the opportunity to to achieve this this long run goal. So, like, you you can we might not be able to actually observe this, but you can talk about, like, what observations would would make it one or the other. So so, you know, like, will this agent, in practice, when it has some opportunity to, you know, make the reward go up, will it do that?

Will it will it only do that you know, if the agent is more sort of like the RL algorithm, it'll be like, oh, it will do that once it's, like, explored it by chance and gotten a reward and, you know, that that's been reinforced. If it's, an agent that, you know, can reason about the world and planning, it'll be like, okay. We'll do that once it, like, you know, learns the learns about the facts about the environment that let it sort of infer that.

And then or, you know, if it's we're talking about some, like, long run goal. It's like it would, you know, do it when it actually, you know, has the opportunity to sort of you know, if we're talking about, like, takeover or something, you know, so it's like, yeah, it's not gonna attempt anything while it's under full human control. But once it is, you know, deployed widely enough or has sufficient capabilities to actually succeed in a sort of coup, then it then it would do that.

Like, that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, like, ah, yes. It had this other goal, it's doing what we want and looking like a nice model because it, like, predicts that that will lead to it getting more power.

Speaker 3

said something that was quite surprising to me. You said that AI could autonomously self improve within as little as two years and maybe even shorter timelines were hard to rule out. Could you walk through the concrete sequence of steps that could lead to that kind of recursive self improvement?

Sure. Yeah.

Speaker 1

yeah. Maybe I'd put, like, a you know, I'm, like, whole whole number percent this year, but but low low whole number percent or something. You you know, ask me on different days.

Give a slightly different number. But, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not unlikely enough to rule out. And I think that basically looks like maybe we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit it on on these less less health and global tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing.

Then this is leading to yeah. Like, you automate and accelerate a bunch of AIR and Ds. So I think I think there's you know, there are a lot of low hanging fruit even of things that we already know that you could do this and it would improve model performance.

And it's just you know, it doesn't require new breakthroughs, and it's just kind of labor intensive to do. So just making much, much better post training environments and really, you know, crafting them to to teach all the new abilities that you want. And and I think you can probably improve compute efficiency a bunch with, you know, again, just, like, applying a bunch more labor to, like, making all your kernels more efficient and also, you know, doing the right kind of of rooting between different models or or or other things like that.

There's there's, like, lots of ways in which how we're using compute is not optimized, so you could potentially get a bunch of, you know, sort of, like, the equivalent of much more compute scaling out of that. And then, you know, high hypothesis is, like also, like, scaffolding and training the models to use particular scaffolding and sort of, like, use memory and retrieval in the right way. Like, it seems kind of obvious that, like, you know, if you you if sort of really had all the right training data and you have a, you know, you have a transformer and and it can kind of fill its context with with different things and and and take stuff in and out, it can do a pretty, you know, good job of something that looks like sort of continual learning or or building up understanding if you've got massive massive context window and a not you know, you've actually got quite a lot of bits in there to sort of be adding things about what you've been learning and and, you know, if you if you really sort of had optimized the training for all of that, like, maybe you can get that to work pretty well.

And then your you know, maybe some other piece would be like, oh, yeah. Are kind of superhuman at predicting the results of experiments because they've read so many papers and predicting experiments and and sort of synthesizing things from different fields. And, again, maybe this is something like, you it's possible that we're not seeing that good performance here just because we haven't quite elicited the models to do it, and it it's, like, not a thing that they've seen humans do, but they actually sort of have the capability in there.

Yeah. So maybe you can make much faster progress if you can you can do a bunch of iteration. You don't actually have to run the experiments.

The you know, models are much better at predicting what will and won't work. And then you can when you do run experiments, you can sort of, run a bunch more of them because you can optimize the code with your, like, you know, very fast coding models. You know?

Speaker 3

directly train against. I think intelligence is not capability. I think it's the capability to acquire capabilities.

I mean, are in different parts of the phylogenetic tree, I guess, that respect. What do you think is the gap in my interpretation? Because I'm personally not worried about I don't think the models today are intelligent at all.

Obviously, your position is difficult for for me to grasp, but I don't know. What what's what's the difference, do you think?

Speaker 1

thinking about the world where, like, I'm not that you know, I I'm, like, uncertain about what intelligence is, and I have enough probability on, like, you know, moles have it to to be thinking about, like, what would happen if, you know, if that's true. But, you you know, it seems like you clearly think it's more likely than than I do so that that, you know, we could just talk about that difference. Yeah.

Like, moles have this jagged frontier. There are things that they're much worse at than humans, you know, some kind of, like, generalization and and, sample efficiency, and there are things that they're much better at, you you know, kind of like speed and cost. And it's like, maybe you can you can kinda use these to compensate for the for the the others to to some extent.

Like, if you're not good at designing your code nicely, maybe you just have to rewrite it from scratch every time. And maybe that's fine if you're a model and you can output tokens like nobody's business. Yeah.

Some combination of thinking that the the spikiness, you you know, it is evidence that we should interpret a given level of capabilities as, you know, because we know models have so much knowledge, we're like, oh, yeah. This is less impressive in terms of sort of, like, reasoning or inference or something. But it is also true that they do have a ton of knowledge, they will sort of continue having a ton of knowledge about things.

Maybe maybe there's some question about, like, how far can you get on being, in some sense, not, you know, not very good at sample efficient learning, but just extremely knowledgeable, and how much do you sort of run into, you know, think things where you now need need new knowledge and you can't sort of produce it in some incremental way or you can't generalize enough?

Speaker 3

I mean, just a quick comment on that. I think knowledge is the crux. I actually think that intelligence is overrated.

I don't think Francois Schrolei, he put a post out saying that contra Elias or Yukowski, intelligence isn't a unified variable. It's measured differently in different domains. You can't meaningfully measure the domains together.

And it's not like a thing that just keeps getting higher and higher. It's more like a ball becoming more smooth. So as you become more intelligent, the ball becomes more smooth.

And he thinks that we are quite near the optimum of being a smooth ball. I don't think we are. I don't think we're very intelligent at all.

I think a lot of our creativity is through us being a collective intelligence. And we have deep grounded understanding, perspectival understanding. And the LLMs, they're a bit of an interesting one because they're like a library.

So they know everything. They have the perspective of everyone and no one at the same time. So experts like yourselves, you can prompt a language model, you can get it in you can make a simulacrum agent of Beth.

And you can make the agent think like you. And that's very valuable. But you also need all of the different perspectives, and you almost need to create a society of kind of grounded agents creatively exploring things.

When you just have the library on its own, and and you put it in an in an agentic harness, and you you can make it do a specific thing, is well specified. Yeah.

Speaker 1

scenario, I'm definitely imagining, yeah, that you you have a large number of agents, potentially you know, because you have all this agent labor, you can do, like, you know, lots of specific different fine tunes or or, you know, different kind of scaffolding, and, you know, accumulating knowledge in some kind of, like, store that all the agents can can interact with and things. And, like, I maybe you're thinking of this as more of a, like, yeah, that would kind of be a paradigm shift, and I'm thinking of it a bit more of, like, oh, yeah. You know, obviously, if you sort of, you know, iterate on the current agent paradigm, you you add some more things to your scaffolding.

You add to you know, you that's sort of not fundamentally that hard or or something. Like, I agree that if you had, you know, current models and you sort of give them one system prompt, you you then can't plug them into being a a call center worker and dealing with sort of, like, all of the the edge cases that come up. I think may maybe there's some difference in, like, how much you've you think that this has improved between, like, g b d two and where we are now, or I would say sort of in, you know, the amount of adapting to new things that are happening that models could do now does seem like it's much higher, and, you know, they're much better at, like, editing their own scaffolding or or, you know, sort of reasoning about their, like, you know, their their sort of embodiment, like, this thing of, like, knowing not to kill your own process or no like, you know, stuff stuff like that where it's like there is like, yes, they are limited, but there's also some trend of improvement.

And, yes. It's I think just also having some probability on there is kind of elicitation gap on on particular things, and that maybe a lot of taste is basically just, like, being able to predict the results of experiments. You know, you think about all of the things that you would try, and then you can quickly be like, oh, that wouldn't work for this reason.

That wouldn't work for that reason. That wouldn't work for that reason. Oh, actually, you know, someone in some some literature in some different field also tried that and that, you know, so we already know that won't work.

Like, in some sense, models, should be quite good at that. So it is plausible to me that that again, I'm like, this seems pretty unlikely, but it's plausible to me that's something you see, like, big gains once people figure out how to actually train on that. And, like, maybe you you need some amount of kind of expensive to gather training data that people sort of haven't bothered to get yet, but you don't need a huge number of data points because you're not instilling this whole new capability.

You're just, like, eliciting, okay, actually, use your knowledge of all of the papers you've read in all of these different fields to, like, you know, iterate through these, like, no. These ideas aren't promising. These ones are.

Yeah. Again, I think this is one of the things that I more think of as being measured a reasonable amount within you know, just, like, do this eight hour ML task in a novel domain, you know, with this weird weird constraint or something. Like, it does seem to me like you have to do some amount of being like, okay.

Which things are promising to think about? How would I know if this is, you know, making progress? You know, how should I allocate my time?

You know, I've got some limited time resources. How should I allocate my time time to what's most promising? Like, you have to be to be doing some of that.

And I think, you know, relative to humans, moles are doing more at you know, more of just, like, well, they're quick to implement things or they implement it better or they, you know, they implement more things, and they can then get to to test them or something. But I I think it would be sort of surprising if there's none of that. And if you are I think if you are seeing performance on, you know, long variable viable tasks that are very hard, then in the middle of those tasks, like, where you don't directly have a signal, you you are doing this, you know, non verifiable task thing of, like, choosing what to spend your time on and and choosing what approach to pursue and deciding whether that was actually working and, you know you know, in terms of you you could sort of put some metric on, you know, make a billion dollars or or or something like that.

Speaker 3

sort of dumb hill climbing. Folks, we we I think we've run out of time, but it's been such an honor to have you both on. May maybe just in closing, could could you just both say, like, what is the the single biggest inference that people out there should be making from the research that you're doing?

And thank you both so much for coming on. It's been an honor.

Speaker 2

AI. It might really, you know, totally transform the world, economically and and and socially.

Speaker 1

speaks to that. It is possible both for things to currently be overhyped and exaggerated and less impressive than they look and for it to be the case that in future, this thing is gonna be a big deal and, like, you should be worried about where that's going. Like, these two things can coexist, and I think people often sort of you know, positions are surprisingly correlated on some axis of, like, how, you know, how soon you think AIs or how good you think AIs or something.

I'm like, no. These things could all be separate. Like, people can be wrong in different directions simultaneously or whatever.

Shared via Hopper