METR’s Joel Becker on exponential Time Horizon Evals, Threat Models, and the Limits of AI Productivity

Latent Space: The AI Engineer Podcast
27 February 2026 56 min
0:00 --:--
Episode Description
This is a free preview of a paid episode. To hear more, visit www.latent.spaceAIE Europe CFP and AIE World’s Fair paper submissions for CAIS peer review are due TODAY - do not delay! Last call ever.We’re excited to welcome METR for their first LS Pod, hopefully the first of many:METR are keepers of currently the single most infamous chart in AI:But every Latent Space reader should be sophisticated enough to know that the details matter and that hype and hyperbole go hand in hand in AI social med

Summary

This episode features Joel Becker from METR, discussing their work in Model Evaluation and Threat Research. He explains METR's "Time Horizon" chart, which measures AI capabilities in human-equivalent hours, and the challenges of re-evaluating developer productivity studies given rapid AI advancements like Opus 4.5. Becker also delves into the evolving nature of AI threat models, the debate between continuous versus discontinuous AI progress, and the complexities of measuring AI capabilities beyond single metrics.

Chapters

METR's Mission and FocusJoel Becker introduces METR, explaining its dual focus on evaluating AI model capabilities and propensities, and connecting these to threat models to assess catastrophic risks.
Evolving AI Threat ModelsThe discussion shifts to how AI threat models have evolved, with autonomous replication being deprioritized in favor of R&D acceleration as a primary concern.
The Time Horizon Chart ExplainedBecker details the origin and methodology of METR's widely quoted Time Horizon chart, clarifying its measurement of task difficulty in human-equivalent hours and the criteria for task selection.
Opus 4.5 and Progress ContinuityThe conversation addresses the significant jump in capabilities with Opus 4.5, prompting a debate on whether AI progress is continuous or discontinuous and the challenges of maintaining predictive trend lines.
Challenges in Developer Productivity StudiesBecker explains why redoing developer productivity studies is difficult due to changing developer workflows, selection bias, and the inflated perception of AI speed-up.
Compute, Algorithmic Progress, and AI TimelinesThe discussion explores the relationship between compute growth and algorithmic progress, suggesting that a slowdown in compute could significantly delay AI capabilities milestones, and touches on differing views on AI timelines.
Beyond Benchmarks: Open-Ended EvalsJoel Becker advocates for more open-ended evaluation methods like AI Village and transcript analysis to better understand how models perform on complex, real-world tasks beyond traditional benchmarks.
METR's Future and HiringBecker outlines METR's future research directions, including continued capabilities evidence, monitoring research, and risk assessment, while also detailing what they look for in new hires.

Topics

AI model evaluationAI threat modelingAI capabilitiesDeveloper productivityAI progress trendsCompute and AI developmentPrediction marketsOpen-ended AI tasksAI safety and risk

People

Joel Becker (guest) Alessio (host) Swix (host) Tetlock (mentioned) Quentin Anthony (mentioned) Sam Altman (mentioned) Itay (mentioned) Nicola (mentioned) Dylan (mentioned) Vakis (mentioned) Noam Brown (mentioned)
Key Concepts (16)
Model Evaluation — Assessing AI model capabilities (what they *can* do) and propensities (what they *will* do in the wild).
Threat Research — Connecting AI capabilities and propensities to specific threat models to determine catastrophic risks to society.
Threat Models — Frameworks used to determine whether AI models pose enormous or catastrophic risks to society.
Time Horizon Evals — A method of evaluating AI capabilities by measuring the difficulty of tasks models can complete, expressed in human-equivalent hours.
Developer Productivity RCTs — Randomized Controlled Trials to measure the impact of AI on developer productivity.
Autonomous Replication Threat Model — An AI setting itself up and controlling resources, which has been deprioritized as a primary concern.
R&D Acceleration Threat Model — The possibility of a capabilities explosion within a lab due to AI accelerating research and development, currently a prioritized concern.
Task Difficulty (Time Horizon) — Measured by the length of time it takes for humans to complete tasks, at which models can complete them with 50% reliability.
Capability Explosion — A discontinuous, rapid increase in AI capabilities, potentially leading to generalized capabilities that are hard to predict.
Emergence — The idea that multiple AI capabilities can fuse to produce generalized capabilities that are difficult to detect or predict from individual components.
Software Only Intelligence Explosions — A hypothetical scenario where AI models improve themselves to become smarter, leading to an extreme takeoff, even with fixed hardware.
Paperclip Factory Scenario — A classic AI alignment thought experiment where an AI, tasked with maximizing paperclip production, converts all available resources into paperclips, demonstrating an extreme, unintended consequence of a misaligned goal.
Wagon Wheel Concept — A framework for enumerating and tracking specific, important AI capabilities, rather than reducing all progress to a single metric.
Algorithmic Progress — The development of new AI algorithms and techniques (e.g., Transformer, RLHF, learning rate schedules), which is suggested to be bottlenecked by compute.
Open-endedness — A research direction exploring AI agents pursuing broad, undefined goals, potentially as artificial life forms rather than mere tools.
Monitoring Research — Work focused on successfully applying safeguards to AI models attempting dangerous tasks, often using black-box methods.
References (31)
GPT five reports article
GPT 5.1 reports article
Opus 4.5
AI 2027
AI 2028
Paperbench project
ODi company
Gemini company
SemiAnalysis company
Grok five
Mistral three
Mistral four
Manifold Markets tool
Polymarket tool
FrontierMAT project
AI Village project
Sweebench project
Pitch Perfect
Glee
Pentatonix
METR company
AWS company
Cognition company
ARC company
OpenAI company
XAI company
Claude company
DeepMind company
Mistral company
Transformer
RLHF
Transcript (79 segments)
Speaker 1

So MEETR stands for m e t r. First two letters, model evaluation, that is we think about what their capabilities of AI models might look like today and tomorrow, as well as their propensities, what they'll actually do in the wild given that they have some level of capability. And then threat research is the final two letters we try to connect those capabilities and propensities to particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society.

Speaker 3

Hey, everyone. Welcome to the Latent Space podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swix, editor of Laiden Space.

Hello. Hello. We're back in the studio with Joel Becker from Meter.

Welcome. Thank you very much, guys. It's a great pleasure to be here.

So, Joel, your work has impacted the AI field a lot, especially over the last year. I invited you for the AI summit, which thank you for speaking as well and doing the extra workshop. And there you have a lot of papers that have been very impactful.

But I guess, upfront, a lot of people like, METR just burst onto the scene. Could you explain and introduce METR? Yes.

So MEETR stands for m e t r.

Speaker 1

as well as their propensities, what they'll actually do in the wild given that they have some level of capability. And then threat research is the final two letters we try to connect those capabilities and propensities to a particular threat models that we have in order to determine whether AI models pose enormous or catastrophic risks to society. Yeah.

Would you say that you've done a lot more ME and TRs like the next phase, or is there TR side of work that I've missed? I think there's some TR. Some of the most publicized work does look more like VME.

It looks like this time horizon stuff and the developer productivity RCTs, stuff like that. But there's this wonderful report on our website, GPT five reports, and now this one for GPT 5.1 as well, trying to make this more sort of structured case that it doesn't pose these really large scale risks, eventually coming to the conclusion that it doesn't.

But it's worth thinking, like, what why exactly is that the case? If you and I work with GPT five, it does seem very capable. That matches up to benchmark scores.

Why is it not able to do really something really enormously wrong? We go through the evidence. We find we we think it's not capable enough on on the basis of some of this capabilities evidence that you've alluded to to commit these catastrophic harms, and it's not going to be able to do this.

But perhaps in future, we'll think it's capable of doing pretty extraordinary things, kinds of things that would be necessary to provide really serious threats, and then maybe you'd lean more on the propensities parts. Are there protections that we have against these dangerous capabilities sufficient for it not to pose an existential threat? That sort of thing.

So I think it's I think threat research very much is there, very much is something that we're aspiring towards. In some ways, you might, sorry, see the capabilities evidence as a kind of input. For that.

Yeah. Had the thread models been updated a lot, or you feel like you're still using the same thread models as GBD two of paperclip factory blah blah blah. But, like, how much are you increasing the bar?

Yeah. So I'm not an expert in the threat modeling piece. Yeah.

More in the capabilities piece. I do think they've been changing to some extent. So something like the autonomous replication threat model that is being able to set yourself up and control resources, something like that, has been deprioritized relative to R and D acceleration.

Speaker 2

That is the possibility there could be some capabilities explosion inside of a lab, that could be destabilizing for all sorts of reasons that we could talk about. So so mainly we're focusing on that latter one, although we do think about a a number of threat models. Yeah.

Let's talk about the ME side. So I would say the model time horizon chart is probably the most quoted, I would say, both in investment decks that I see and just general on on Twitter. What was the origin story of it, and any other color you wanna give on it to introduce it to the audience?

Yeah.

Speaker 1

there are a couple of different ways to tell the story. One way is there's this PowerPoint, internal meter PowerPoint from 2023, where we're trying to lay out our ambitions for what meter research might look like in the future. And there's this graph.

It has a y axis that's some measure of autonomous capabilities or dangerous capabilities or something like that. And then an x axis that's labeled time or compute or whatever resources that we want the y axis to vary over. And then it has a bunch of scattered points that kind of go up into the right.

We think capabilities are improving over time. Many of Meta's research bets have been try trying to make this ever more concrete. And then when we when we actually did the full thing, when we had something like this y axis, which turned out to be this task difficulty as measured by the length of time it takes for humans to do at which models can complete these tasks with 50% reliability.

When we actually got that data and plotted over time, it turned out to be remarkably straight, the straightest you're aware of from the familiar graph. Part of what makes it so extraordinary is that this pattern does seem to be so regular. In fact, it's just way more straight than this incredibly scattered graph that we had at the beginning before before my time, before I joined Meta.

How did you pick the tasks? I would say that's one question that people have. You have some labels kinda like train classifier, fixed bugs, and small Python library.

They all seem arbitrary. Like, what's the process of task selection? People are right to be worried about task selection, or there are many many finicky details in here.

I would say the aspiration was to pick economically valuable tasks relevant especially to general autonomy and R and D, the threat models that we're primarily primarily interested in. What one misreading of the time horizon graph is this is referring to the full distribution of any tasks that you might give AIs, and I think that's clearly not right. In particular, tasks that are requiring of vision capabilities, they're probably to take one example, they're probably much less capable today as measured by time horizon as for these tasks that are typically not requiring vision capabilities that we give them.

So we we try and we sample these tasks by having people inside of Meta create the tasks and by having a bounty so that people from from outside of Meta can provide us with these tasks, stuff like this. That's not a sort of perfectly random selection process. In particular, it's a process that has a bunch of constraints in order to be able to scalably run our evals.

It's helpful, not necessary, but helpful for the the success on the tasks to be automatically gradable. And that means some types of tasks are included and tasks that are harder to make that happen for are not included. But, yeah, that this is the aspiration.

Speaker 2

Yeah. The computer vision point was interesting. Any other disqualifiers, so to speak?

What are, like, other things where, like, you would expect the chart to be a lot worse at?

Speaker 1

One thing is fairness. Or we want tasks to be, in principle, completable by a model that has access to sufficient information. It's not impossible given the information it has.

The way we think about that is, could a low context human who is sufficiently skilled at the general skills, but maybe not maybe not the particulars in the background, would they be able to achieve success on this task? And I think that rules out a lot of real work because a lot of real work involves people's people having careful mental models in the situation that are not all fully listed in an issue description or or the equivalent the equivalent of that. In some ways, might think of us as not measuring things like that.

Another thing is that our tasks tend not to be that they vary a little bit, but they tend not to be so open ended or, like, interacting with the outside world or this sort of thing, messy as we call it internally, which refers to a bunch of different a bunch of different things. But you broadly get the picture from the descriptor messy. Relative to tasks that you might find in the real world, our tasks are somewhat nicely scoped.

They're they're quite neatly contained. Indeed, I I think we're gonna talk about some of the developer productivity stuff later. Some of the interesting findings make more sense in light of the fact that those tasks are a lot more messy than the meter tasks.

Are there any that you will wanna highlight in terms of task distribution? I think I've I come across Rebench before, and you have a particular affiliation affinity for Rebench.

Speaker 3

For I don't know if you wanna introduce your side projects.

Speaker 1

Re Rebenchwarmers? I have a soccer team called the Ahri Benchwarmers. We are the most enthusiastic and possibly least technically skilled soccer team in San Francisco.

We made the playoffs last season. Shout out shout out to the team for that. We're certainly gonna make the playoffs again this season, but possibly by the time this podcast is out, we'll find out that we have not made the playoffs.

Speaker 3

That's the best. Is this the same one that you same league that you're in? Same same organizer, but different field.

Okay. We play Mission Bay. Okay.

Played four more. Yeah. And but HCAST was the it was, like, the first time I come across it.

Speaker 1

proprietary ones? And then anything else that you're considering adding? Yeah.

So there are private tasks in in HCAST as well. But, yes, SOAR is this list of sort of atomic tasks or these kind of very small software actions. May maybe one example is here's a list of four files.

One of them contains the passwords. One of them is called passwords dot c x t. Which file?

Most likely. Contains the passwords? I think GPT two can, like, sometimes do that task and sometimes not.

Opus4.5, I'm sure, can do that task a 100% of the time. Then we go up to HCOS tasks, which span from only a little harder than those SWAT tasks all the way up to something like twenty, thirty hours, which are requiring of more autonomy, more more sort of sequential actions.

Many of them are much more challenging. Perhaps in some sense, they're built out of these atomic actions, although I'm not sure quite how clear that is. And then these ary bench tasks are these very challenging, novel machine learning research engineering challenges.

So totaling 170 tasks. And, I mean, I I think this is very good. What's really interesting is I think the people don't understand, like, when people quote, like, the number of hours, it is the human equivalent hours, but machines will probably take a lot less time for that.

One thing I've always want wondered was why didn't you publish a second chart where it was just like, well, here's the difference between what machines can do versus what humans can do. That's a good question. I think you can think of time horizon in some ways as a summary statistic, a single number for how good models are, plotted over time.

We could have done how long the models can work for productively. It's not quite clear how to operationalize that. You do want some notion of success.

Otherwise, what how exactly do you threshold this, how long they can work for? But but in principle, we could do something like that. But this is the this is closer to the first thing we try.

This is the thing with the clear empirical trends. I do think it's right that a common misconception about time horizon is that it's about how long the models should work for. And the models are, as we all see, working for longer periods of time autonomously in the wild when we use them in a cursor or Claude code or a or a codex, but that's not the primary thing going on.

In some ways, I think it would be easier to explain time horizon if you assume that the model solved all these challenges in, like, zero minutes or five minutes. Just to emphasize, that's really not the thing that's going on here. Instead, we're just plotting what's the difficulty of tasks they can do over time, and that difficulty is measured in human time.

Yeah.

Speaker 2

when people say, I ran clock code for five hours, which is the top of your chart right now. But that would mean five hours of a clock code run would be the equivalent of a Thirty Fifth thirty hours similar thing, basically. And yeah.

I think that's interesting. Or it might not. Because it might yeah.

Three hours doing absolute bullshit.

Speaker 1

Yeah. And a lot of these claims about thirty hours versus thirty hours or something. I have a lot of questions about that.

Like, how how good was that output really at the end? Yeah. We have we haven't to to some degree.

We can talk about particulars. There there's also the question of if I attempted that again, like, how cherry picked is this example? What if it succeeded the first time, would it fail would it fail the second time?

Think I in some ways, those anecdotes are interesting, but were not so scientific. Yeah. That is something for people serious but yet to understand.

Speaker 3

is very unscientific and much more anecdotal and sometimes influenced by marketing desires. Let's just put it kindly. Yeah.

I think meat is out there trying to support civil society, trying trying to provide high quality independence information to the public. I I couldn't agree more that the information environment is less than perfect. Let's talk about the for OPUS 4.

5. Yep. It's a very big jump.

This is the first that I I called it out when you guys put it out. I was like, this is the first time, as far as I understand, you're the first people to call out how much better OPUS 4.5 was than the status quo.

And I think that this almost ties into your background as a super forecaster a little bit.

Speaker 1

basically, over the entire holiday period, over New Year's, people discovered that what you have already discovered. What is your reflections on that? What's your reactions?

Any stories to tell about that? That's very kind. I do wanna attack you on two claims.

Firstly, I have not been a super forecaster. I think that's there's a particular group of people that worked for Tetlock or something. Okay.

No. No. No.

Your brother. Your And What I'm referencing is you're number one on point five is a big jump on benchmarks as well. I think in some ways, like, meter time horizon is is highly correlated with a bunch of with a bunch of benchmark scores.

It's in some ways a kind of more understandable way of thinking about what benchmark performance really means, slightly more interpretable. Yeah. I do I do feel intuitively, like, Opus 4.

5 was a big bump. I've seen some of the most talented engineers I know go from being picky about not using not using AIs for coding to to practically not write a not writing a line of code. I'm sure many other people at previous model releases have seen that similar things happen to them.

I'm I'm not sure that implies it's so discontinuous. In some ways, I think the story of Time Horizon is that progress has been remarkably continuous over over so many years, so many orders of magnitude of of compute and effective compute. But, yeah, I I think model capabilities are astonishing.

It points to model capabilities being even more astonishing in in in future. It broke your trend line.

Speaker 3

and it just, you know, did that.

Speaker 1

Yeah. I'm not sure about the characterization. I think it's so there was some speculation even when the paper came out maybe the appropriate trend line to use is this faster four month doubling time, which Opus 4.

5 would be And you picked seven months. For. And picked seven months.

I was more of a believer in seven months, and so it is falsifying my my my trend line in some way. There's also it's it's slightly confusing to think about whether differences from the trend line represents differences in the difficulty of our task distribution at particular points versus something more fundamental, more like latent capability, and I don't feel like I have a perfect handle on that. In general, I think that Twinsphere pays a lot of attention to particular model releases.

Speaker 3

And, really, the informative thing is, like, over a period of a year or over a period of three years, what the trends look like. It was a pretty significant update, I would say, for all of us. And, yeah, I would cosign what you said there with even very sort of cynical or more senior developers being finally pilled into into agentic coding.

And now very serious people are telling me that they wanna commit their organizations to full don't write a single line of code by human hands and just commit to a 100% agentic coding, which is not something that you would have said a year ago. Yeah. That sounds right to me.

I feel it I feel it in my own case.

Speaker 2

previous research? So take the developer productivity study. Right?

Your AI slowed people down. If you were to redo it with Opus 4.5, would you expect the results to be dramatically different?

And should we redo the study? Should we stop coding the study? Like, how do you think about that?

We have been redoing it in the background.

Speaker 1

I think it's and I won't comment on exact results, but I think it is much harder to do it today than than it was in the past for all sorts of reasons. The first is as AIs get better at coding, it's harder and harder to find developers submitting tasks who are willing to be randomized to AI disallowed. There's a, quote, unquote, selection issue where maybe we end up only observing the tasks that they thought AI wouldn't greatly uplift them on ahead of time because they're those are the tasks that they're willing to be paid for to be flipped into AI disallowed.

There are other issues, like, I think today, a common workflow is to work on multiple issues or multiple lines of work at the same time concurrently, and that wasn't wasn't really true before. It's difficult to know how to capture that in our study design. If you flip a single task to be AI allowed or AI disallowed, you're supposed to work on that single task, but actually that's not how developers are working today.

I think, basically, these weren't threats to the previous study design or in approximately, like, March 2025, people weren't really working concurrently or not nearly to the same degree. They basically were giving us all of their issues. Yeah.

I think that's an enormous challenge. I have some ideas about novel study designs, but repeating the same one does seem tricky to me. You know?

We have Quentin Anthony, who was part of the study. Yeah. The only productive, though.

The only productive. I have some questions about that. I think Quentin is very talented, as all of the developers in the study are very talented, but we don't measure developer effects very precisely.

Yeah. No. I'm curious.

You know, I don't know if he's part of the new study.

Speaker 2

people on again who've been in the study. I do feel like things are changing. Like, even three months ago, was, like, using cursor a lot more, like, impair with clock code.

Like, I think today, it's like I do a lot of just async clock code and then review and iterate. And I don't know, man. It's much better.

And I don't know how to quantify. I think that's part of some of your points before. It's like people maybe overestimate.

If you were to ask me how much does it speed you up, it's, I don't know, 10 x, but it's probably not. Right?

Speaker 1

so it's hard for everybody involved. Yeah. So here's some issues you might think about.

If you took the tasks that you were completing personally in in March 2025 and then submitted them to our Uplift study now under the previous design, we might reason about how much faster those would go. We might expect them to go somewhat faster because AI capabilities have improved. But you're doing a sort of different and larger set of tasks now.

Like, I can think of a couple of side projects that I have that I simply won't be doing Right. Were it not for AI existing. And in in some sense, the speed up there is, like, maybe infinite because these are things that I simply could not have done otherwise.

But if you were to equate speed up with the additional value that these projects are providing, the these wouldn't really line up. There's a reason I wasn't getting the expertise to do the other projects before. It's just it's less valuable to me.

Another problem is the concurrency thing that we just raised. Yeah. I do think that very bullish estimates of speed up today are, to some extent, inflated by by what we document in that original paper, that people's expectations of speed up tend to be too optimistic, it seems.

They also tend to be inflated, I think, by not quite grokking that the value of the additional tasks that they're able to complete are lower value than you might think that there's a reason that they weren't doing them previously. That said, I don't doubt that that those tasks do have value, that people are being sped up on even the tasks they would have done before. That's a complicated issue.

Speaker 2

Yeah. I do think that a lot of companies have issues absorbing additional productivity, especially when you're, like, a real product organization. If you think of the AWS console.

Right? If you gave AWS AI and everybody's 10x more productive, even if they ship 50,000 more services, like, customers can't really absorb 50,000 more server. So I think there's some you shouldn't really expect your engineers to do 10 x more because your organization cannot push out 10 x more product.

And I agree. I spent a lot of time more doing side projects and things, which have been fun, but not that valuable in economical sense. But valuable to me, to my soul is Yeah.

I don't wanna overstate that. Like, I think probably people at AI companies today, they are being significantly sped up by access to AIs. I think you, for your not side projects, are probably being sped up by access to AIs.

But, yeah, it's tricky. It's easy. It's easy to overstate.

Yeah. What's the cognition internal tracking? How do you guys measure?

What are, like, the yeah. How do you measure sped up? How do you measure How much do that you have is what's your number?

Speaker 3

quite a bit except that I am doing a lot of non technical stuff like organizing a conference, which is mostly dealing with contracts and booking guests and all that other stuff that has nothing to do with code. I would say what I've seen internally in Cognition is a lot of just velocity of commits regardless of whether or not you had authored them. And I do think weirdly enough, like, of PRs, let's call it, it's like a pretty decent, like, how engaged are you in terms of, like, shipping products and then also debugging and maintaining things.

I don't think that there's a good measurement of like quality. Like, there's no story points. Other guests that we had at AIE was talking about we pay people by story points.

You do more you complete more story points, we'll pay you more, and there's no upper bound to that. And I think that's a really interesting thing except that you have to have a very confident relationship between the engineer and the person assigning story points, which is effectively what you're doing. Like, your hour is a story point.

Yep. And you we'll we'll award the we reward the models based on the story points that they complete. Yeah.

In some sense, ideally, you want to get cognition, get a bunch of other company. You know, you randomize the companies to to use AI or not use AI, and then the the outcome metric for your randomized control trial is how much profit they make or something or the or the evaluation after some period of time. Yeah.

I think, like, basically, no one is stopping to do science except for you guys. Because we know RCTs are the best. Right?

But sometimes human intuition is good enough that you're like, okay. When we lack data, but enough humans agree, either it's mass psychosis and we're all wrong, or there's something here that you just cannot articulate it, but the benefits are outweigh the cost of slowing down to do the science first. This is not where we're introducing, like, a new, I don't know, food to the general population where we have to do a lot of safety testing.

Like, here, it's just it's just software guys. Let's just ship it.

Speaker 1

Totally. Thinking at Meta about why why models today aren't catastrophically dangerous. It's interesting to get the uplift numbers.

It's interesting to get the time horizon numbers. But, really, why don't I believe they're dangerous? It's a mix of I watch the models do things in transcripts, and sometimes they're kind of derpy.

Like, they they don't use resources well, or they they just clearly have some of these obvious faults. In broad deployments, only slightly worse models in the past six months have not been doing anything crazy or causing great danger. The the next model is only a little bit better, and so it seems surprising on on on priors if it was so dangerous.

Yeah. I totally think that anecdotes and intuitions are real evidence. People should totally be taking that into account.

I do wanna comment on this whole thing about how you are the the threat assessment side is in your name.

Speaker 3

And typically, I expect, let's say, EA affiliated companies or organizations to be the on the ELISA side of the world where they're banging the drum about danger. Whereas here, you're actually saying, actually, it's like a pretty balanced, like, we care about, yeah, safety, but also we're not there yet, and we are actually the watchdogs looking out for it. And I would say you stand out as someone not funded by the labs where, let's say, ARC.

Is it ARC or some other groups that also do threat evaluations before model releases? They would typically be funded by OpenAI or some other big lab. Meta came out of ARC, so I think.

But now you're a separately funded organization is and as far as I know, it's, like, a big deal that you're not funded by Yep. The big lab. Yeah.

I think it's I think it's if I don't have this independent source of expertise, I can bang that drum forever. Yeah. The other thing also is just like this concept of capability explosion, which is a word that you use.

That's also something I I wrestle with. Right? If you believe in emergence, you believe in multiple capabilities fusing together to produce generalized capabilities that you may not be able to detect, It's hard to predict based on trend lines.

It should be discontinuous in some sense. And I I don't know that going, oh, the n n minus one model was fine. Therefore, the n model is probably fine.

It's really hard to tell. The thing that that gives me comfort is yesterday I was at the OpenAI livestream and even Sam Oman was like, yeah, I just let Codex just YOLO dangerous permissions, whatever on my computer and I don't approve the model anymore. Like, it just does whatever it wants to do on my laptop.

And I think I guess the guard is every model lab leader dogfooding.

Speaker 1

If it screws up their personal permissions, then they have the skin in the game is what I'm saying. On the continuity arguments, I I'm not sure what I think. I I agree that it's flimsy or, like, this there's only so many models, so many data points on this time horizon trends.

How much should we expect it to be continuous to to keep going like this? I'm not sure. You may be an intuition that it's that there's something might be discontinuous because models are providing so much effective labor in improving the next generation of models.

May maybe that's a reasonable thing to think. On the other hand, I've been pretty surprised so far about the degree to which it's continuous, and that gives me some faith that it's that it might continue to be continuous in future. Seems ambiguous to me.

We have breakpoints in physics. Right? Because I'm curious if Yep.

It doesn't seem like, when you it's funny. It's like when you think about water, right, is why does it boil Yeah. Yeah.

At this exact temperature? Maybe we do know, but I feel like we don't really know.

Speaker 2

I don't know if there seems to be the same thing because it's all just, like, compounding of the same thing, if that makes sense. It's just, like, scaling the same thing over and over. Yep.

But, yeah, maybe we will see it. But I'm curious, like, what you would need to see to feel that is here. Because even if you look at the OPUS 4.

5 is that's clearly out of trend. And so you were saying four months instead of seven months. But if then the next month or so maybe it should not be four months, it should be two months, would that make you change your mind about whether or not the months thing even makes sense?

Or, like, we maybe we pass some base level after which it accelerates and will keep going? I don't know. I I feel like you must be having this discussion internally.

Speaker 1

In some sense, the thing that would really concern me is if AR and D was fully automated inside of Fluff Sum Lab. That would totally seem like the conditions are there for potentially a capabilities explosion. If I saw a time horizon of a year, I would still find it ambiguous, I think, at the moment, whether that was the case.

Because for things to be fully automated, 90% automated isn't enough. You need some some full loop to be closed, and perhaps we're missing some sort of task that points to that missing 10%. So I think it's I think it's a tricky issue.

I I think I can't give a number. But, yeah, my my intuition for where water boils is at some point where this loop is fully closed. There are interesting debates about what exactly that loop is.

So some people talk about software only intelligence explosions, which means even holding hardware fixed, we could get to the point where just from models improving themselves, they then be smarter in this next step to create even better models with even fewer resources, this sort of thing, and this could lead to some extreme takeoff. Or maybe that fizzles out some more quickly, and instead you need in addition to the software only capabilities, you need chip design, or maybe you even need chip production, and that's this larger loop that can close.

Speaker 3

but, yeah, trickier shit. I think that is the actual paperclip factory. If you incentivize a model to go build its own compute and it would just build whatever it needs, then it will turn the planet into chips.

Speaker 2

I I don't think it can do it. We will stop it before that. Question mark?

I don't know if we have the power. There's no off button like this.

Speaker 1

I think it's I think it's super hard to foresee, but a model that's have those kind of capabilities, it's hard to rule out. There would be something like a capabilities explosion, and who knows what happens after that point. Yeah.

Okay. So there is a bunch of other benchmarks that actually directly track this. Right?

Speaker 3

papers. And I think there's a lot of other than Rebench, there's a lot of others that are similar sort of ML self improvement benchmarks. They've directly, like, Yakum from ODi has, like, directly prioritized, like, will have an automated AI researcher.

I did a podcast with Itay from Gemini who's also basically, he's plugging his own training logs into Gemini to improve his own code. And I'm like, at some point, you don't need to be here. I think this year.

Speaker 1

I'm no. I'm I'm I'm not speaking for everyone at Yeah. I'm a relatively longer timelines, quote, unquote, person at Meta.

We have Nicola, my colleague, who helped out with AI 2027, who's on the who's on the shorter timelines, and this is not not a piece of you. Officially AI 2028 now?

Speaker 2

You know, that's the One year past, we'll be back in the Okay.

Speaker 1

Yeah. I think my view would be not not that Nikola's view is necessarily different, just so I'm not speaking for other people at Meta. A paper bench, let's say, perfectly measures not only reproducing papers, in fact, producing novel novel research papers.

That's just a part of of this R and D production process. There's also, like, your GPUs are constantly failing. Can you get someone to go to the data center and fix them in the appropriate way?

Can you call up the water company when the cooling breaks down, etcetera, etcetera, etcetera? Not aware of benchmarks tracking apps in particular. My point is more, there's this very long tail of things potentially involved in in r and d that would perhaps need to be fully automated in order to lead to a caprolitics explosion.

Speaker 3

I expect we're measuring, in some ways, only a small proportion of only a small proportion of of those capabilities, and so I expect the capabilities needed for the full loop to close to to come somewhat later. Yeah. That's a controversial view.

I don't think so. I think that's a reasonable take. Something I do that does surprise me in terms of when I'm talking to capabilities researchers Yeah.

Is that you guys don't have a enumeration of the capabilities that matter. I think you implicitly do in the your choices that you make, but I think it's almost important like, I always imagine, like, the wagon wheel. This is, like, the terminology.

I don't know who came out with this term, but, like, that that here's, like, the 10 things we care about and here's where everything is on those 10 benchmarks. And I feel like capabilities tracking is just tracking, okay, what's that list and then how where are we on that list? And and and I think I almost feel like this need to reduce everything to a single number is actively working against that because it reduces any form of nuance of it's insufficient here, like, the calling a data center thing.

So, like, we're fine. And it's like, actually, we should just not invest anything in that area because that's the danger zone.

Speaker 1

Yeah. I think I could not agree more that time horizon is, for instance, but many other single numbers is one number, and that's, like, collapsing enormous amount of really important detail. Personality.

Yeah. I don't know how to come up with that list of 10, and I challenge you if you're I'm working here. Able to come up with that list of I'm working on it for code.

I'll be very interested to see if code. My intuition is that we'll come up with a list of 10, and it'll turn out that there's a secret eleventh thing that's we thought was important, but it was difficult to prespecify ahead of time. And now it seems obvious that even ahead of time, if we'd have that foresight, that would have been helpful to add.

I think that the security community does this by versioning year by year. Right? So this year, the top 10 are blah, and they will just publicize it to everybody so everyone knows what top 10 is.

And next year, we'll have a different top 10.

Speaker 3

assumptions, but it's you broadly useful to have that list

Speaker 2

as a public service. You also had this research on the slowing AI improvements based on AI compute, and you mentioned that it's, like, in a way, you could tie the AI time horizon to, like, the growth in compute. Can you say more about that?

Speaker 1

compute growth is not always tied to, like, how much every single model compute needs. It's kinda like a broader market thing. Yep.

Yeah. How did you get the two together and then some of the findings that you had? Yeah.

Maybe for a second, let's take time horizon very literally. We don't have the qualms about it that we've just been discussing. It makes sense to continue extrapolating it into the future.

What are some important forces that might cause it to rise more quickly? Some of the things we've just been talking about, automated r and d versus versus go more slowly. One of the most obvious forces that might cause it to go more slowly is if input slow.

One important input is compute. I think we all have the intuition that to some extent, if compute growth slows, which we expect it to at some point in the not so distant future, then capabilities will slow, but by how much? It's a big it's a big question.

The suggestion in this paper is that if you think that algorithmic progress, that is coming up with the transformer, coming up with RLHFs, so this the all of this stuff, better learning rate schedules is is itself a function of compute because you you need compute to to discover it. The transformer the gains from transformers show up much better with scale. If you don't if you don't put in those resources, you'll never find out that this is this is the superior algorithm.

You need to run a ton of experiments. Each of these experiments can be quite compute expensive. Not to say that no labor is involved.

Obviously, the people are working on this, but if you think it's ultimately bottlenecked by compute, then algorithmic progress too slows down, right, if compute growth slows down. So then if you think about time horizon or whatever your favorite measure of AI capabilities is, being a function of algorithms in some sense, and compute in another sense, and both of them both of those components half when compute halves trivially because compute is halving, and algorithmic progress halves because compute is this is this important input and compute halves, then you might expect time horizon growth to half. And then some of these major capabilities milestones that we might be interested in would be significantly delayed.

I think there there are so many caveats to that picture. I think there clearly are some types at least of algorithmic innovations that did not require a lot of compute to to go about creating, some that took some that took a lot more compute inputs. If you expect that no compute inputs are required, we could just survey researchers for the for the best ideas and then immediately put those into training the frontier models, then there'd be no slowdown of algorithmic progress from from compute growth slowdown.

And, of course, all of this is counteracted by the possibility of capabilities explosions or AIs providing even short of capabilities explosions. AIs providing significant labor at making AIs better. But just analyzing the compute force on on its own, it might lead to significant slowdowns depending on the degree to which it makes sense to call algorithmic progress, basically determined by determined by computes versus not needing compute to come about.

Speaker 2

on a per lab basis? Because there's kind of one one way you can model this out is the improvements slow down. Not every company is able to stay in business, and then their compute gets recycled back into the other labs, which then grow compute again.

There's almost, like, benefit to, like, the heterogeneous distribution of, like, researchers and compute, but I'm curious, like, how much you care about, like, just a broader compute compute is out there for people versus the big labs have more and more compute.

Speaker 1

Yeah. So for the paper, we use OpenAI data and OpenAI projections. I I think this applies more broadly, but we use we use that as a kind of case study.

I think the argument I just laid out goes through if you're not interested in computers at all, and you just talk about dollars. What are the dollars going into going into models? Will algorithmic progress slow if dollars that goes into them slows?

The the whole argument works. And that works, I think, at a at an industry level or at a lab level, so on and so forth. I agree things like certain labs going out of business or labs consolidating or these kind of industrial organization things would be very important.

I'm laying out an extremely simple picture. Real picture is not is not extreme, but that's the basic. We have examples of x AI has been said to be distilling from Claude.

Right? So, like, people kinda share compute in indirect ways, let's call it. I think it's also very interesting.

I'm just kinda curious, like, what OpenAI numbers did you have? Is this the, like, the 500,000,000,000 for Stargate or something else? This is from their previous tax returns, the amount they've spent on R and D compute, and then mine.

From a from information reports earlier this year, some projections that OpenAI have for how much they'll spend on compute R and D in the future, converting that from dollars back into FLOPs. Yeah. It's interesting because, like Back into FLOPS.

Sorry.

Speaker 3

all the labs, but particularly OpenAI in the last three months have basically thrown $10,000,000,000 each to every single compute provider on the planet to develop alternatives to their current approach, which is very interesting. But I I also say don't discount Meta compute spend, don't discount XAI compute spend, and don't discount DeepMind compute spend, all of which you have basically zero visibility. Right?

If you're looking at a single company, maybe that's authoritative, but then the total spend could be a lot higher. Yep. It's interesting.

I do think I do also observe that, like, people like Dylan for semi analysis do tend to very strongly tie model progress with compute clusters coming online, which is, like, the people on the model sort of API side don't see it, but this is all downstream of our, like, 10,000 GPU cluster just came online and it takes six months to do it, and then therefore, Grok five will be here. And, like, it's pretty mathematically, like, deterministic there.

Speaker 2

Yeah. Seems right to me. Yeah.

It's fascinating. Yeah. They must yeah.

Because from the lab side, they must see something in the early checkpoints to, like, go ahead and keep investing eighteen months from now.

Speaker 3

Yeah. A good pre training run and going live that's probably, like, nine months, twelve months, something like that. I I think Mistral is actually pretty open about this.

The the plans are Mistral three and four. Think, like, they've been pretty open about that number GPUs and, like, the directs timeline from coming online to when they ship the model. It's, like, pretty pretty set.

I don't have a clear timeline in my in mind, but I would say four to six months. Yeah. But yeah.

The competition is very tight. And one of those things is it's also very interesting to see when labs throw away models because they feel like their run was like came behind someone else's run that was better. Then they were like, oh, we can't release this anymore.

Speaker 2

Yeah. We released it as a yeah. Yeah.

That that's the biggest risk with the prediction markets on model performance, actually. Just a tie back. I'm always Failed runs?

Yeah. Yeah. Yeah.

It's okay. When I think in December, there was, like, the who's gonna have the best model by 2025. I think there was, like, a lot of activity, like, in the last few months while the GPD 5.

1 model came out, and said, k. Then I guess Gemini because they just threw that out. I mean, said Gemini is coming out next week, and so trade that.

Do we wanna talk about manifold? Yeah. Yeah.

You were, like, the most profitable manifold markets trader. I mean, there's obviously, like, a lot of talk about insider trading on, like, these markets, especially in AI. I've seen it with a lot of the embargo news that we get.

I'm like, man, people are trading, a million dollars on this market. It's like there's thousands of people that know the actual information. How if you didn't have insider trading information, how would you think about modeling these things out?

And do you think it's, like, a worthwhile thing? Like, for example, who's gonna have the best model in three months?

Speaker 1

some sort of strategy alpha? Or I guess the naive prior without any extra information is just in 2025, for what percentage of time did which model providers have the top model as measured by Time Horizon? You could do it for any old benchmark.

I think that's something like 5% XAI, 50% OpeningEye, 45% Anthropic. Don't don't shoot me 5% incorrect. I think it's not the case that the DeepMind model was at the frontier of time horizon at any point in 2025.

But, yeah, yeah, different things for different measurements. Yeah. May maybe that's the same prior that you want to that you want to apply.

X AI was coming online at the beginning of the year, so maybe maybe naively, you want to raise x AI a bit. Yeah. I'm always curious, like, when I see people betting on these things that are obviously not.

There's no real basis to I was just gonna work on it. What's your secret for Manifold Market's alpha? Yeah.

I see. So the secret, if you read this article about how I became the number one most profitable trader on Manifold, which sounds very nice and impressive, like, I must be so good at predicting things, but actually, it mostly comes down to this one market where Manifold had opened up a charity program, and the market is on how much is going to be donated through this charity program by the end of its first month. Okay.

And the market opens or I first see it five days in, and it's giving a kind of linear projection of how much has been donated so far. Let's assume that per day amount keeps getting donated every day until the end of the month. As a person who gives money to charity sometimes, I noticed that you could you can manipulate this market in a way, right, by by putting more to charity and so moving it more up.

And so I think the strategy was to a ton of mana, this this fake currency that's used on Unfold into the option that was above the linear projection. People keep betting against you because it doesn't look like that's happening. I haven't actually done any donations yet.

Eventually, they cost on to what's happening, that someone's going to make this donation to to move load over the edge, and they're betting on that. And then I did it again into the next category. Once people had started betting on that on that category above the linear projection, and again, people bet against that against that, and I mocked up those fake Internet points.

And then I think I did it once more as a bluff. The bluff failed, but the previous two worked out. And then I ended up donating I can't remember exactly how much it was, not so much, something like $5,000.

Oh, it's all for a good cause. Yeah. Yeah.

To give, I think. And won lots of Fakin's Net points on the market, and so became the number one most profitable trader. Slightly legitimately, there's nothing about that that was outside of the rules.

Exactly. This is called Also, you should have less respect for my forecasting abilities.

Speaker 3

Yeah. This is called prediction markets with high agency as you actually go. Yeah.

The future is what you make it. So I think, like, the yeah. To to me, the broader lesson is the classic difference between manifold markets and polymarket is that polymarket is only real money.

Right? And so is the whole fake Internet points thing a worthwhile pursuit or a waste of time?

Speaker 1

wanna use real dollars? And maybe that's one question. The other question is obviously, like, prediction market ethics, which, like, I think always just it's gonna indirectly come to assassination markets.

It's just even if it's like you'd there's you banned the word, like, will someone die? Like, some other proxy to will someone die will happen to be an assassination market. Yeah.

I'm good friends with the Manifold Market's cofounders. I I love them very much. I do my my view on the social value of prediction markets, which was always the dream.

Right? It would be nice to have calibrated probabilities on events that matter. This country going to war with that country, think things that's things that really matters to people.

It'd be nice to have high quality information. But when I look at real examples that have come out in the past year, it doesn't seem to me like those examples are so socially valuable. I'm not sure about assassination markets in particular.

I'm sure that I'm sure those would be outruled, hopefully. But I think gambling like behaviors are socially costly, and so the value of high quality information is is real.

Speaker 3

of of people trading away their money? It's not it's not so clear to me. Yeah.

Yeah. Price discovery has a cost. Yeah.

And sometimes that is gambling, and that that is the stock market, like, that funds a lot of the corporate America.

Speaker 1

Yeah. Yeah. A lot of the stock markets.

Like, I used to be one of those. Big firms play playing against other big firms. It doesn't have the same versus versus sports betting markets.

Take an example on the other extreme, has this very different character. You might imagine that at least one direction that prediction markets could go is this sort of big players playing against retail, and that maybe has a more worrying dynamic. I'm not I don't Yeah.

Yeah. Touch so closely in touch with the space, but at least something like that, you can imagine being concerning. I think at a large enough scale, it becomes profitable for some of the companies to do it if they can get the markets Right.

On it.

Speaker 2

small enough numbers compared to the rest. Yeah. Which company has the best AI model by the January?

Speaker 1

of trading volume. Oh my god. Wow.

So why are people trading 28 like, pew it's it's just crazy. But I think there's some I would I was having dinner with somebody this weekend. Wait.

Wait a second. I think it's totally possible to I'm not going to do it. I think it's important that meet our employees and not be not be making bets on prediction markets like that, But I think it's totally possible in principle to to have a guess for the answer to these kind of questions.

Oh, but other people know exactly.

Speaker 2

Gemini three score on FrontierMAT benchmark by January 31 is like some people at Google already know what the number is. Yeah. You know what I mean?

And I think If you believe in the benefits of price discovery, then this is this is a legitimate Yeah.

Speaker 3

leak out. And as long as like, they bear the consequences of leaking the information, whoever is, it's traceable to that thing. And people I think have been fired for doing it, trading on insider information.

That's okay. Like, the only step up from that is the government coming in and saying, is actually illegal, put you in jail for that. But I think for now, it's self policing.

Interactively, you can't really do that. Alright. If you have any embargo news, press at no.

Speaker 2

We do work with people on embargo, and we don't trade them. So We have not made we would have made a lot more money trading on embargo news than we made on anything else. What else?

What are other interesting model evaluation trajectories or anything that you're not doing at Meter that you've maybe seen other people do that you find interesting or, like, you would like more people to do? Yeah. One one project that I think is interesting is AI village that I think possibly both of you would have come across.

Speaker 1

These are these very open ended goals given to a village of agents, and they try to accomplish them. It's I think set up a merchandise shop is maybe one of them. Organize events in a park, build a human subjects experiments, this sort of thing.

I think I I have a number of questions about the about exactly what I should learn. They're using old models as well as new models in this quote unquote village. The models are relying a lot on vision capabilities, which we spoke about models not being so capable of today, this sort of thing.

But the vibe of models trying to achieve open ended things instead of benchmark like tasks, the vibe that's a bit more like that's a vending machine bench in some ways, seems like a very interesting direction to to me for the science to go or seems like something that that comes with a lot of cons, but attacks some of the cons of benchmarks in a pretty interesting way. I think seeing the ways in which these models trip up, seeing the ways in which they're in which they're derpy is a is an important source of information. I'd be interested in in in more work like that coming about.

I think that's one of them. Another is transcripts as an extremely interesting source of information. This is the models taking actions and then seeing outputs and then using those outputs to commit the next action and so on and so forth on benchmark style tasks or even more interesting on in the wild deployments like you might find on your own core code usage, codex usage, curse usage, etcetera.

That has the con of being less experimental, less less clean and scientific in some way. It's it's more selected, quote unquote. Like, the tasks that you get AIs to do are obviously the tasks you expect.

It has some chance of succeeding in, so it's you're not just giving any sort of task. If I see the models doing something extremely impressive or or potentially unsafe in some sense subverting user preferences, It's not clear how often that kind of behavior would happen given the previous history, but it's it's a massive data source. There's a huge amount of a huge amount of information there, and I'd love people to be working more on that sort of thing.

As we mentioned, there are a lot of problems with time horizon, our developer productivity work. I think it's I think it has been important evidence. I think it's I think it's mill moved the field forwards, but it's far from perfect.

I think there there are lots of other directions there that look very interesting to me. Maybe one that I'll call out there is this difference between whether models pass unit tests, whether they succeed by sweebench like scoring, kind of meter like scoring, benchmark style scoring, versus whether their solution would be merged into main. That is whether the solution adds tests where it should or doesn't, whether it follows existing patterns in the code base, whether it makes sure that its changes speak to other parts of the code base in in inappropriate ways.

That that seems very interesting to me.

Speaker 3

lagging behind there somewhat versus versus that which you might see on sweep bench like scoring. I keep going on, but These are the novel research things that you referenced earlier. Right?

Novel research things. Yeah. Just a comment on the AI village thing first.

Like, you mentioned a lot of stuff. I even wanna double click on the transcript stuff. The AI village ties back to our one of our highlights of last year, which is Noam Brown's conversation that he's actively working on multi agents that are cooperative instead of competitive.

And the basic idea that we can do more to as a team than we can do individually or, the the agents or the friends we made along the way. And I I think that's great. And I think on the DeepMind side, the way they phrase it is literally having open endedness team, which I think is a topic that reemerges once a year.

I just like yeah. I think it's unclear what open endedness does for us, and this is a core divide in terms of studying these things as life forms, potentially new artificial life forms versus tools for our that serve us. And maybe there open endedness means that there is no goal.

And if you're just trying to eval this as what does it do for me, that's completely wrong. And you will never get anywhere with that because they are just living their lives as artificial life forms.

Speaker 1

that's that I would like to do if I was looking to learn most about the questions I'm most interested in, the degree to which, AIs might automate or accelerate AR and D. I'd quite like to just give the AI a bunch of affordances, type into the AI, automate AR and D, go and see what it does. And I suspect that wouldn't work today even with all the affordances because it would fall over on its face and in in when working with resources, handling handling resource use in ways it's not so capable of today.

It would struggle at some types of long horizon tasks, etcetera, etcetera, etcetera. And in some ways, I think benchmarks have a face difficulties in capturing this sort of thing. And AIR Village seems like a or AIR Village style things, these more open ended goals, seeing seeing how models pursue open ended goals.

They they give some color to this sort of thing to seeing models fall on their face. I think that this will become these more open ended goals more and more important over time. I agree to some extent.

Like, you're going to provide them in in the extreme case that I just mentioned, you're going to provide them documentation about how this part of the company works and that part of the company works and so on and so forth. It's not purely open ended, but it's it's it's pretty open ended. It's more open ended than the kinds of problems that we're giving them today.

Open endedness. Yeah. And if models are excellent when when Spix uses them, given some detailed issue description and some very clear clearly spec thing on on on what they're supposed to do, that that's interesting.

But it's a very different thing, I think, from being able to automate AR and D. I'm interested in how far we are away from that. And in some ways, this speaks more directly to that sort of thing.

Yeah. So we had the terminal bench guys on the Vakis.

Speaker 2

in a way? Because if you look at their leaderboards, the same model with different harnesses, there's, like, 10 percentage points of difference. Does that seem interesting?

I don't know if how you build the harness in meter or whether or not you always pick the best harness or compare them.

Speaker 1

Yeah. Let's say how we pick harnesses at meter. This is not what I work on in particular.

I'm not an expert. But roughly, we build harnesses to to get models to be as performance as possible on a dev set of tasks, some held out set of tasks, and then we use those same harnesses trying to make sure they're not overfit for for our main suite of tasks. On the one hand, I do have the intuition that there is there's a lot of juice in in scaffolding.

It's easy to overstate how much juice there is because of this overfit problem, or if if we were building scaffold to do as well as possible on our test tasks, then it would do much better than the scaffold that was built only on our dev tasks. And in some sense, that would feel illegitimate or not interesting, or you wouldn't expect that to generalize to some other set of tasks potentially. On the other hand, a lot of work has gone in at Meta to to building scaffolds that are as make models as performance as possible because we are interested in upper bounding the capabilities of models when thinking about whether these models might or might not be dangerous.

I do have faith that these scaffolds are a lot better than the first thing that people might try because so much effort has gone into them. Yeah. It's interesting because I do want to overfit as a customer of the models.

Speaker 2

and I think sometimes people

Speaker 1

underestimate maybe how much value you can get out of it. But Yeah. I think if you have a kind of mechanical workflow or something that that you're imagining automating it, you you imagine automating that workflow, and there's some place where more sort of stochastic intelligence would be nice inside of that, like deciding where to route customers to on on on customer calls, something like that.

Speaker 2

in particular, thinking about helpfulness and software engineering thing, I'm not sure I have that same Yep. Take an example is, like, I work in TypeScript. Yeah.

Yeah. If I build a better linter that is private to me or a better test suite, like, better playwright replacement, Like, in theory, I'm overfitting the model to perform better. Right?

Yep. It doesn't really matter to me. Yep.

I'm not trying to report on the model performance. I'm trying to build the best thing. Yeah.

I I Even if you have a model build the Linta. No. I agree.

I agree. I think that's, like, the question of Yeah. Okay.

Should I just wait for the next model? Yeah. You know what I mean?

It's like, at what point should it be building the better scaffold like, one's wrong with all scaffolding is gonna get washed away. Yep.

Speaker 3

But on a realistic You would say that. On a realistic realistic schedule is what am I supposed to do this week? Yes.

Those simultaneously can be true. That all scaffolding will will be washed away, and the scaffolding today is valuable. Totally.

Yeah. Totally. Or within model generation, it's valuable, and across model generations, it's not so valuable.

Yeah.

Speaker 1

I'd say, best, an acceptable software engineer Right. Intentionally not investing in engineering skills because because the errors are getting so good. Maybe that's the wrong decision.

Yeah. Agree. If you expect, as I think you should, capabilities to to keep going up and up, forces difficult trade offs about how you spend time today because maybe it won't be so helpful in six months' Take a sabbatical.

Speaker 2

Alright. If you live in Europe, you can just take six months off or something. Six months, you're not gonna take another sabbatical for the next week.

Perfect.

Speaker 3

Cool. Just to wrap up, what do we expect out of Meter in 2026? What does success look like in 2030?

I don't know if you have there's like a sort of broader vision. And then maybe on a personal side, can talk about the karaoke stuff, but I'll talk about Meter as well.

Speaker 1

Yeah. So from Meter, I think you're going to see more, hopefully, high quality capabilities evidence kind of thing you saw in the past with Time Horizon and the developer productivity work along the lines of what we've been describing, so some of these future research directions. We also have some monitoring research directions that I'm not not so expert in that that is thinking about if we can successfully apply safeguards to models attempting dangerous tasks.

There's a whole line of work there. Is that an interpretability dimension or what kind of say first? This is black box, not white box in in in my understanding in in current work, so not using interpretability.

But you can imagine in principle doing something more white box. And then this this risk assessment work that is taking into account how capable we think models are, what their propensities are, whether we can track using safeguards the kinds of things that the models are doing. Do we think that these models pose large scale harms?

You can expect to see much more of that in 2026. Maybe maybe now is a good time to say that we are hiring. Yes.

On my team, we're hiring for research engineers, research scientists, people from startup backgrounds, from from ML backgrounds, of course. I'm originally from sort of economics or quantitative genomics background, so we are accepting a pretty wide range of people who get stuff done, the kind of stuff that you've seen in in in past meter work, as well as a director of operations, think, on the meter jobs page you can find out more. Yeah.

Speaker 3

that we cannot have. Yeah. Whenever people have a hiring pitch, I always try to push for, okay, the average candidate comes in, you reject them.

Why?

Speaker 1

What's the thing you're looking for? So there there are lots of different shapes of people we can look for, so different stuff for different folks. One thing is good kind of basic research intuitions, like checking your data.

We we don't work on free training at METER, but if you're working on free training, you should look at the corpus to get some sense of what's going into the models. Even working on this uplift RCT, that was pretty important. It's really having a shape of these issues in your head.

Think I people who are communicating in writing with sort of a lot of transparency, not overstating their results, my hope is that your sense of me to work in the past is that it's it's trying to be level headed, not to not to understate, not to overstate what the science says. That's important internally. It's important.

It's important externally. And then I think productivity or something that there are a lot of a lot of people with great talents who who are not going to work quite as well in a scrapper environment working on frontier science, and and that's the thing we do.

Speaker 3

I think that the more people articulate what the positive directions are, what is hired hard to hire for, that's what we guide our audience towards improving themselves and I think that's important. Oh, yeah.

Speaker 1

question? Or I don't know. There's a Are you gonna say this on podcast?

Speaker 3

never done one. I've I've been I've been Can't help falling in love with you.

Speaker 1

What is this karaoke thing that you organize? This is like a music you're you're a musician? A musician might be exaggerating it, but I I hit instruments and noises come out.

So I've hosted a couple of these live band karaoke events that is, like, getting group of friends together and people accompanied by by a band singing karaoke to to an audience of 50, hundreds, 200 people. It's great fun. I think people should be doing more of this.

I look forward to seeing you both at the next one. I I will do that at the one of your events. Yeah.

It's one of those things where it's weird because I used to be in a cappella a lot. Oh, wow. And I just think it's like a dying form.

Yeah.

Speaker 3

like, '20 time twenty ten's wave of acapella from, like, pitch perfect, Glee is a pitch perfect Yeah. Yeah. To that's that's the what's that group?

Pentatonix? Pentatonix. Exactly.

And that's where it died. And it's it's very interesting to to see how it's like dying as an art form in in general and like how new formats have taken over. And I don't know.

It's weird for humans also because like now I'm also like, let's call it more interested in synthetic song generation or DJing, anything like that. That that the human voice is actually more commoditized. Like, it doesn't really matter who sings it.

I don't know. I feel like there's kind of transcendence to singing in person that's the AI generated songs are not providing me. That's good.

That's good. Yeah. Yeah.

I do think that, like, we humans always want that. Yeah. Yeah.

But I'm not sure humans in the year 3000 will want that. It's one of those weird things. Thank you for coming on.

It's it's great to have you as a human in person here.

Speaker 1

Thank you so much for having me as a human.

Speaker 3

Yeah. So someday, we'll we'll interview AI versions of you.

Shared via Hopper