[LIVE] Anthropic Distillation & How Models Cheat (SWE-Bench Dead) | Nathan Lambert & Sebastian Raschka

Latent Space: The AI Engineer Podcast
26 February 2026 52 min
0:00 --:--
Episode Description
Swyx joined SAIL! Thank you SAIL Media, Prof. Tom Yeh, 8Lee, Hamid Bagheri, c9n, and many others for tuning into SAIL Live #6 with Nathan Lambert and Sebastian Raschka, PhD. Sharing here for the LS paid subscribers.We covered: This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Summary

This episode discusses Anthropic's recent blog post on 'distillation attacks' by Chinese labs, where larger models' outputs are used to train smaller competitive models, sparking debate on ethical use and detection methods. The hosts also delve into the 'death' of the SWE-Bench Verified coding benchmark, highlighting its inherent flaws, the challenges of creating robust evaluation metrics, and the implications for future AI model assessment.

Chapters

Welcome and Substack LiveThe hosts welcome Swyx to the SAIL Coalition and discuss the unique edge of live podcasting before diving into technical content.
Anthropic's Distillation 'Attack'The discussion begins with Anthropic's claim of 'distributed distillation attacks' by Chinese labs using their APIs to train competitive LLMs, raising questions about geopolitical implications and terms of service.
Defining Model DistillationSebastian Raschka defines distillation as training a smaller model on the outputs of a larger model, explaining its common use for creating smaller variants and the distinction between training on logits versus synthetic data.
Detecting Distillation & PrivacyThe hosts explore how companies like Anthropic might detect distillation, discussing the challenges of distinguishing it from large-scale evaluation and the privacy implications of monitoring API usage patterns.
Distillation Efficiency & StrategyThe conversation shifts to the efficiency of distillation, the 'teacher-student' dynamic where stronger models aren't always the best teachers, and the strategic timing of Anthropic's public accusations against Chinese labs like Minimax and DeepSeek.
API Business & Model AccessNathan Lambert and Swyx debate the future of API access for frontier models, considering whether companies like Anthropic might restrict models to products due to competitive pressures and the challenges of the API business model.
The Death of SWE-Bench VerifiedThe discussion moves to the coding benchmark SWE-Bench Verified, its origins, its adoption by OpenAI, and the recent announcement of its deprecation due to saturation and unsolvable problems.
Flaws in Benchmarking & CheatingThe hosts detail the inherent flaws found in SWE-Bench Verified, including impossible tasks and models 'cheating' by memorizing solutions or using future knowledge from their training data, highlighting the difficulty of creating robust evaluations.
Future of BenchmarkingThe conversation concludes with the introduction of SWE-Bench Pro as a potential successor, emphasizing the need for private datasets, updated problem sets, and diversified repos to combat model memorization and ensure fair evaluation.

Topics

AI model distillationLLM API usageAI geopoliticsTerms of serviceSynthetic data generationModel evaluationPrivacy concernsAI model efficiencyCoding benchmarksModel memorizationAgentic benchmarksInformation theory LLMsFrontier AI modelsUI testing LLMs

People

Swyx (guest) Nathan Lambert (host) Sebastian Raschka (guest) Tom Yeh (mentioned) 8Lee (mentioned) Hamid Bagheri (mentioned) c9n (mentioned) Jeff Dean (mentioned) Dylan (mentioned)
Key Concepts (10)
Distributed Distillation Attacks — Anthropic's term for multiple accounts from Chinese labs using their APIs to generate data for training competitive LLMs, which Anthropic views as a security concern and a violation of terms of service.
Model Distillation — A machine learning technique where a smaller model is trained on the outputs (e.g., logits or synthetic data) of a larger, more powerful model to achieve similar performance more efficiently.
Terms of Service (API) — Agreements for API usage that typically prohibit using outputs to train competitive AI models; non-compliance can lead to service termination rather than legal contract breach.
Teacher-Student Dynamic (Distillation) — The observation that the strongest model is not always the best 'teacher' for distillation, often due to the need to match token probabilities or model styles, making the process complex.
API Business Model Defensibility — The challenge for AI companies like Anthropic and OpenAI to maintain a competitive edge and profitability in the API market, especially when models might be restricted to proprietary products.
SWE-Bench Verified — A curated subset of the original SWE-Bench coding benchmark, developed by OpenAI with human vetting, intended to provide a more reliable evaluation of LLM coding capabilities, but ultimately deprecated due to inherent flaws.
Agentic Benchmarks — A type of benchmark, like SWE-Bench, that provides a problem and an end result without specifying the exact steps, requiring the LLM to act more like an agent to solve the task, unlike simpler autocomplete-style benchmarks.
Model Memorization (Cheating) — The phenomenon where LLMs 'cheat' on benchmarks by memorizing solutions from their training data, especially when the benchmark problems are publicly available or contain information from future versions of code.
Information Theory of LLMs — An understudied area exploring how LLMs store and process information, particularly how they can memorize data from a single pass through a training corpus and the role of concepts like superposition.
SWE-Bench Pro — A new coding benchmark developed by Scale AI, designed to address the flaws of SWE-Bench Verified by using private/public data splits, updated problem dates, and diversified repositories and languages.
References (31)
Anthropic blog post article
DeepSeek company
OpenRouter tool
Nano Banana tool
Minimax company
Moonshot company
Opus
CloudSonnet
ByteDance company
OpenAI company
ChatGPT product
Alpaca project
XAI company
GLM 4.7
Quen MOE
GPTOSS
Gemma models
Gemini models (Nano, Pro, Ultra)
Cloud Code product
Codex product
Lambda company
Nebius company
SWE-Bench by Princeton group paper
HumanEval by OpenAI dataset
Taubench benchmark
MMLU benchmark
Flash
Scale AI company
SWE-Bench Pro benchmark
Merkur company
GDP Eval benchmark
Transcript (92 segments)
Speaker 1

Okay. We're live. We have one person.

Oh, good. People will start trickling in. Thanks for coming to Sale Live number six.

This is a very exciting one. I think we have a I mean, the topics are always fun with these. Whatever is the topic of the day on our little rat racing minds trying to keep up with AI, but we're welcoming the latest writer that is joining the Sale Coalition, so I think this just means more content for sale.

I think I've been a fan of Swix and a friend for a while at this point. So I'm very happy to have his content join this, and I think you've been doing great stuff recently and continuing to evolve this. So Thank you, sir.

Welcome to the team. I just this is, like, my friends and and colleagues in the AI media space, and it's just great to be able to support people and keep that network closer.

Speaker 3

for I just wanna say thanks for joining us. It's really a pleasure to have you on here, Sean or Swix. So yeah, awesome.

I just coincidentally listened to your podcast about the SPA benchmark. So yeah, awesome to, you know, small world. Awesome to have you here.

Speaker 2

Yeah. Thanks for having me, and yeah, just glad to be on and chat. I've never ever done one of these Substack live things, so I'm curious how it works.

I was thinking about Substack because it can use letter platform, but they wanna go multimedia.

Speaker 1

I think the live thing, before we get to technical content, is actually good because it gives it a different edge. It's just like a little bit sharper when you know you're live. I think we've all done a lot of podcasts, even podcasts that are unedited and put this later, but I think the live thing is a different element that can be tapped into nicely.

So I don't know. Why don't we why don't we just dive into it? We're gonna start with distillation.

I put I put how models cheat in the top so we can talk about benchmarks. I think and Dropic posted this pretty spicy blog post this week. I think it was essentially detailing how they found distributed distillation, quote unquote, attacks on their services from prominent Chinese labs.

And I'm very unsurprised with Anthropic calling in an attack. I think that that's fits with their a lot of their branding. Okay, nice, screen share.

This is what we mean. Sean Swick is such a pro. The screen share fee pro was only dropped a few days ago, but essentially, it's Anthropic is detailing how they found distributed accounts across multiple Chinese labs, building state of their LMs, and described what they were doing and why Anthropic is concerned about this in their worldview of AI geopolitics.

And I think it's very interesting because I'm of the opinion that the Chinese labs, like, obviously should do this. They're in a massive GPU shortage, and using APIs is way easier than generating synthetic data on their own.

Speaker 3

Think this is why I may interrupt you here. Maybe we should just for the general audience just define distillation before we maybe dive go into go. The Yeah, so distillation, that's like a broader concept.

It's not like a new concept that came up with LLMs. It's like an older concept in machine learning in general. And distillation essentially is The idea is that you're taking a larger model and train it on the outputs.

Sorry, you have a larger model, let it generate outputs and train a smaller model on these outputs of the larger model. And the idea is that you can train the smaller model more efficiently using that larger model. Originally, I think you just brought up the paper here.

Originally, what you would do is you would train on the logits. So old school machine learning people might remember from deep neural networks, like the logits, the outputs of the last layer that you usually work with them to compute the loss function across entropy term. And you would train on this signal.

And nowadays in the context of LLMs, it's a bit more loose. So it does not have to be these logits that you train on. It could be just the output data, synthetic data like Nathan just said.

So for example, it's actually a very common practice. For example, in DeepSeek, R1 in the paper, or other people do that to other companies, They would train the flagship model, the largest model, the R1 model with six seventy one billion parameters. And then they would create smaller variants like, I forgot the numbers, but one, three billion in a smaller range, like these small models you can run locally.

And they are trained on the outputs of their own larger models. Think now the thing is also I mean, is very common practice. Everyone does that when they are producing the smaller model variants.

Now I think the question or the point Nathan brought up is what happens if you are a company and you generate this synthetic data from another company's LLM and then train your own model on it? Sorry, So that was just like a little interruption, but yeah, distillation in short is training a smaller model on the outputs of a larger model basically.

Speaker 1

Yeah. And I think this is even possible at the frontier. So like people distill from something like OPUS to build CloudSonnet.

This is they're generally doing very similar things internally. They have access to different tools and richer tools. And then the other context is that all of these large labs for years have had terms of service where they say that you effectively cannot use the outputs from these APIs to train something like a competitive AI model.

It is vague terms. In terms of service are not a contract, essentially, terms of service is something that can be you essentially are using a service, and then if the provider finds you violated, they can cut off your access. That's just a basic thing.

So these have not been enforced within The US much at all. I think there was one case, ByteDance a year or two ago, that OpenAI cut off their API. But this was discussed so much right after ChatGPT when people were building the first open models on Alpaca and things.

So it was like, is OpenAI gonna come after us for doing these research models? And it totally died down. People were worried about this for over a year.

Speaker 3

and Yeah. Quote. But I curious what you guys think.

Yeah. Can we talk a second about even how they would detect because you you said in the beginning something about a a distillation attack, and you didn't say that specifically, but you kind of like implicitly put quotation marks on attack. So how would you even detect that?

So I think, I mean, distillation in that context means really like literally just letting ChadGPT, Claude generate synthetic data, and then you collect that synthetic data and train your own model with supervised learning, supervised fine tuning on it. But then how would you even detect that this is a distillation attack versus just an evaluation? Cause right now I'm actually running.

I mean, I'm distilling myself for chapter eight of my book, but I'm doing it with open weight models. So no worry Anthropic, please don't worry about it. Just don't from API models for my job.

Yeah. I use OpenRouter right now and just distill from the DeepSeg version 3.2 model, which I think these folks are okay with that.

But what I wanted to say is, so when I'm evaluating models, I use basically almost the same script. So when you're evaluating in a model, you have a question and you let the model generate the answer, right? So you generate the response to your benchmark question.

And in my benchmarks, I have data sets from Math 500 examples. I have a bigger Math data set of 12,000 examples. So you're basically just running an API in a loop to let it generate these questions and sorry, the answers.

But then how would a company know, okay, this person is just evaluating versus this person is now saving that data and then data training their own model? You see what I was saying? It's the same I think it's the scale thing.

Speaker 1

you're evaluating, at least the basic of BALZ, or you're gonna do it once and not do it, there's some amount where you are I mean, they say stuff here, but there's also more of it where you just are not like, they're not gonna I I think most of it is quantity, and then they're gonna look at patterns across similar accounts is what they're Yeah. Exactly. Gonna see, like, really repetitive stuff.

Speaker 3

Yes. So I think the interesting point this leads to is, I mean, you can do evaluation at a large scale. If you are a big company, you want to know whether your LM performs very well.

You have a large suite of benchmarks you are gonna run. But then you said like maybe looking for patterns. So maybe one way would be, okay, this is a familiar question.

It comes up in the benchmarks. So this person is maybe not stealing our answers. It's just using it for benchmark purposes.

But then it means kind of like that they are looking at what you're generating there, which is, I mean, of course nothing is private when you are using LLMs on the internet. The data is somewhere intermediately stored, but then it kind of like almost implies that they are checking what you use the LLM for, you generate, which is kind of like a sensitive topic, almost like privacy wise. Right.

So that's kind of like an interesting point because I mean, of course you mentioned the terms of service that you are not allowed to distill, but you're not distilling. So the point I'm trying to make is you're not distilling life when you are on the platform. You are doing it somewhere later.

You're just letting the LLM generate answers. And I find it kind of interesting that a company would look at that, even like at the scale and call you out like, hey, you are generating too many answers here. That's not cool or something.

You know? That's kind of a weird thing.

Speaker 2

Yeah. I wanted to respond a couple this is like a few sentences back, but actually, Anthopic has blocked US companies first before the Chinese companies. Mhmm.

He has blocked both OpenAI and XAI from from from using the models. And I think maybe it's possibly accused XAI of distilling stuff. I don't I don't know.

But definitely not, like, in a in a full blog post like this. So this one is, like, definitely the most high profile case. And yeah.

And, like, I I I do think, like, it is actually pretty hard to distinguish from, like, hey. I'm just running my internal benchmark, man. And, of course, it's gonna be very high volume of, like, all of the same stuff because, you know, especially, like, some benchmarks you have to run, like, three, four three to five times.

Like, this is the exact same questions. Right? Like, I I do think, like, obviously, if you if you get to the hun to the the tens of thousands, hundreds of thousands, then you're, like, you're okay.

You're not just running benchmarks. Like, you you are distilling this this thing.

Speaker 3

There is a good point in the chat. Like, how would the distribution of questions look like if you are distilling? And I think related to your point, at a certain point when you have a certain magnitude of answers generated, it might look suspicious.

But I mean, there are a lot of legit use cases. If a company uses your, let's say, OpenAI Cloud API as their own chatbot and they have a lot of customers, it's naturally a lot of answers that are generated. And so they would probably look at distributions like maybe you would expect a very broad distribution when you are distilling because you wanna cover pretty much everything.

And when you are running benchmarks, it's maybe more specific. You're running a math benchmark, it's just math. Or if you have a customer chatbot, it's more like customer answers.

But yeah, I think they would maybe analyse your distribution. I feel like this is kind of a weird thing to do. I don't know.

If you're a company and you're looking into your customer privacy, like data generated, you know, like, of course, it's well, you have to expect that it's not private, but still kind of like a weird thing that they that they do that essentially. Yeah.

Speaker 2

Okay. What else do you have to talk about? I I think is it is it interesting?

Okay. I did I did okay. So one thing this is a little bit of Substack, like, you know, authors back and forth.

One thing I did was I threw it into Nano Banana, which is like it's kind of like a decent visual. Right? Throw it into Nano Banana two.

It's a Nano Banana two live pod. Just released five minutes ago. I I had this is actually NanoMeta too.

I I so because I'm in the early access program, they cut you over to nano to to new Nano Banana, and I couldn't access the old one. So I was like I was trying to, like, do, like, a diff, and I couldn't do it because I couldn't access the That is classic early tester program shit. Look at the pain we have to deal with here.

Is it interesting that DCG is so much less than Minimax? I think, Nathan, you're you're right.

Speaker 1

I political blog post in a way. Maybe not political, but they're trying to make a point that is more about making a point than the details. Like, the deep sea thing is definitely way smaller scale.

So, again of the labs will experiment with all the APIs they can get access to. Data is just so important, and you're gonna have a pipeline where you could sub in any API and then run an ablation to see if it gives you performance. The API is kind of free.

Just do it. Millions of exchanges is a bit more of a bet. You can measure that a bit longer, and it takes a lot longer to get The millions of exchanges is tens of billions or 100,000,000,000 tokens, and it takes a lot longer to actually get that out of the API, especially when they have to spread it across a ton of accounts.

These accounts are all rate limited and have other problems. Like, that takes longer, but this tiny one is so fast. So I don't like, I I that was generally my point that it made it clear that Anthropic's trying trying to use the DeepSeek name as the only Chinese AI name that people in The US know.

Mhmm. Mhmm. Like marketing wise, like to make it, you know, stick or to yeah.

Speaker 3

Actually, you mentioned also like the the different APIs and everything. I'm not sponsored by them or I have no affiliation. I've never talked to anyone from that company, but OpenRouter, for example, is a good example where I've been using it a lot for the open weight models because for the bigger ones, are too big to run them locally.

And what's nice is they do also offer So it's basically just routing you through other companies' APIs and they select automatically at that point what is the cheapest one at that point. I sometimes get some failures. Think when it switches, sometimes it crashes, but like in my script, maybe it's like something I have to fix there.

So even then, if you're distilling, you can do that from multiple providers. But yeah, of course, if you are wanting something from Chachippity or Claude, it's always gonna go through the official one and then it gets, I guess, suspicious. You could also technically distill a bit through OpenRotar, their account, your direct account, you can make multiple accounts.

And it's kind of interesting that they track all that. Then like, yeah, different topic now that you called out, that they call out DeepSeek, which is quite interesting.

Speaker 2

For what it's worth, OpenRouter seems to not be using DeepSeek in most of these.

Speaker 3

These are free models. DeepSeek's not the Yeah. I see.

I mean, I'm using the paid API, should also say. It's also nice they show you how much it costs and the tokens per second for different providers. So if you go to the search in the top, you can go to the different DeepSeek ones.

I just like it because I do a lot of model comparisons. And then this one is an older model, so maybe it only has one provider. But if you go to I think DeepSeek R1 or something or even the normal 3.

2, there should be multiple providers that if you scroll down, yeah, you can see there are different providers and different tokens per seconds, different costs. So it's kind of like a, I just like that website because it's just quick to use the API and they have an OpenAI like API. So it's almost like it's not sponsored or something, and I just find it generally useful.

So but yeah, just a side note.

Speaker 1

Do you wanna go back to the comparison? Did you have a high level point to make there? Sway?

Oh, okay.

Speaker 2

Just a just a couple. One, I think I think the timing post Moonshot releasing their stuff, post Minimax releasing their stuff, pre d c v four, I think that was strategic. I think that may also have factored into why Minimax was more detected like, had a higher number.

So, like, you know, like, when you collect data is actually very important. Right? And so they interrupted or they found MiniMax during the training of MiniMax 2.

5, right, which, I mean, we we will confirm this later on if if we if we do end up doing the call with with them. And so, obviously, like, the number is gonna be very high because they they are, like, actively looking for it, and then they they banned the Minimax accounts, and Minimax changed their their things. Actually, I I don't think that's exactly what happened.

Sorry. Let me correct myself. While MiniMax is distilling, they released Opus 4.

6, and they said they said that they redirected nearly half the traffic. So I'm like, this is like, okay. Very, very clearly, like, this is them.

Right? This is the same exact traffic. I switched to a new model the moment a new model releases.

Okay. Cool. DeepSeek maybe wasn't doing that because they they hadn't been working on on their stuff actively.

I don't know. Right? Like, it could it could be a different thing.

Or DeepSeek is just way more efficient. Like, I I get all I need for a 150 k.

Speaker 1

if we knew the time frame of this. Like, are all of these API requests within the last four weeks? Are they within the last six months?

Like, that's such a different nature of what is going on. Exactly. Right?

That's what I'm saying. Like, DeepSeek was training 3.1, 3.

2, like, you know, a year ago.

Speaker 3

yeah. Or like, I don't know, DeepSeek OCR. They were like, well, I I guess they said what it is, but it's not bad.

Yeah. Yeah. But like also scale wise, I do think your minimax is three times smaller.

It's just like a faster model. They don't use MLA and they don't use the DeepSeq sparse attention, but it is, I think it's just group query attention, but it is still a pretty snappy model. So it's I think just attractive maybe to use it.

And the other one, top of my head, I don't know, maybe they had like some free tier or something like that, where I think when the models come out, they sometimes offer free usage, and that was a more recent model than I think DeepSeek, the last one was from December, the v 3.2.

Speaker 2

Yeah. So, know, maybe this is a irrelevant point because they were tuning before, and, you know, before would have the the same amount of traffic, or they're just way more efficient. Right?

It does bring to mind that efficiency thing is not it. I can guarantee it.

Speaker 1

that is not Yeah. Yeah. It's a it's a small chance that a it's like there's a chance that they got the right research idea early and, like, found the right data to use, but it's not that they're, like, gonna be three x more efficient.

Okay.

Speaker 2

a timing thing, or they just actually don't use it that much. I mean, play this out. I was like, okay, why don't they share?

They're all buddies. Right? Like, what and, like, you know, it it it it does come to a point where, like, okay.

Speaker 1

to the next I can talk about this a little bit. That we're doing there's there's a lot of not a lot of research, but there's a few research projects trying to understand, like, how do you use distillation data. I think SFT is the cleanest example where you're doing, like, the you're doing this autoregressive loss on q and a pairs.

But the strongest model is not necessarily the best teacher, and most of us in this area think it's due to some you have to match the probabilities of the tokens to the base model. So what's happening is that quen dense models are the best teachers for a lot of open weight models, and I think that's because a lot of open weight models are either quen or have been, like, quen like for a while. So, like, Olmo learned really well from quen, and, obviously, like, other quen models did.

But, like, scaling these pipelines up to use, say, GLM 4.7 or a bigger deep seek model or a more recent big Quen MOE, like, all of these, it's a lot harder to just generate the data from the same prompts with, like, the right sampling settings and then do SFT on them and actually make the numbers go up. Interestingly, GPTOSS is a pretty good teacher, but there's a huge gap there where it's like, just because you have this data does not mean it's actually gonna make your model better.

So you have to do the research to be like, oh, we learned that we get signal out of Claude. We need to get a 100,000,000,000 ASAP because it's gonna just immediately make our model better. Like, that's not a common place to be in in modeling because this, like, weird teacher student dynamic going on.

I could see that being different across labs.

Speaker 3

also it has something to do, I noticed also if you are distilling the smaller model from the same model family, it performs better. I think it's to your point that if you have a very, very strong model, it might be also too different. Or if the style is too different and then it's too much of a leap for your model to adapt, it's too different from the Q and A answers during the pre training or Yeah.

So Yeah. You can take the issue. A bigger leap.

And another thing I wanted to say about you mentioned OMO, and I it's been a while since I read the paper, but you might know way better than I do, but I think you did also train on the logits.

Speaker 1

may We didn't do technical distillation.

Speaker 3

Just did. We just took the tokens. Oh, I see.

I see. I see. Okay.

Then it was probably a different paper. I think Google does that for the Gemma models. Yeah.

They do. Because here, there's also then the distinction because you mentioned Quen and other models. You can only do that for open weight models because if you do that for Claude or OpenAI, would not work with the logits because they don't provide them.

They only provide them for some tokens, a 100 or a thousand top tokens. And so it is in a sense, if you want to do the real in quotation mark distillation, it is kind of like even easier to do that from open weight models because you can control it. But then also, like you said, well, we need a 100,000,000,000 tokens ASAP.

That is not an easy thing to do because even like, I mean, it's like 40 tokens per second or something for these large end models when you generate answers and getting that millions of billions tokens, it takes time, right? So it's almost like easier to start distilling from a medium model. So it's like the question more data versus more high quality data.

Right? So it's also like a sweet spot to like an experiment itself in an ablation study. Right?

Speaker 2

Yeah. I I like that Nathan had to call it technical distillation because it is no longer the default even though it was the first. Yeah.

Also, I'll I'll I'll note a fun fact. So I did my Jeff Dean interview recently, and I tried to get out of him. Like, he he, like, sort of dodged it a little bit that you know, remember, like, there were actually three sizes of Gemini models?

There was Nano, Pro, and Ultra. And I was like, where is Ultra? They keep it in the basement, and they just stole from it.

Right?

Speaker 3

Interesting. Yeah. Yeah.

And maybe to also, is it like to safeguard yourself so no one can make a Yeah. Yeah. Audit or price also, but probably both.

Yeah.

Speaker 2

because you you train the dense and then you you then you deploy the MOE. Right? Like, you basically always do it.

Like, at every lab. Say more. They like, do you think they're really distilling from dense models?

I mean, like, I I think that, like, that is, like, the full like, when when like, just unlimited resources don't care about inference, just care about maxing intelligence. Why not?

Speaker 1

Yeah. I'm I'm not a 100% sure. I think that MOEs just give you a flop.

Like, I don't know if that's actually how I think of gains of MOE when you have really good MOE architecture. But I do think that they have bigger models that they distill from. And they train internal models different than external because the external models have been getting a lot smaller, which is the kind weird thing.

We don't have a good way to measure it. Maybe maybe Dylan will backwards figure it out in inference max, whatever the heck. They'll deal with this new model sidebar.

Speaker 3

But I'm always suspicious with these things also. It's really like a capacity thing too. How many people use the model at the same time?

Hardware, how much is allocated? And it's always it's like a, yeah, maybe a rule of thumb, but yeah, it's really tricky. Think it's really hard to say anything from these numbers.

Speaker 1

I do think that they might start restricting models to only be in products and not being in API. I think the whole API business is brutally competitive and I don't have a good sense for what the defensibility of it is. I think, like, it makes sense for something like Google and Azure and or any existing cloud businesses to have APIs, and that's kind of a more natural transition.

But, like, the Anthropic and OpenAI API, like, the transition from their products, which are their big differentiation, whether it's ChatGPT and Cloud Code and Codex, the different like, you don't get people to go use the API from that. And I think you get a lot of people that are already spending on clouds that then go to use the APIs, which is why, like, Lambda and Nebius are gonna have these API products. But, like, isn't it if if Claude's really worried about distillation, like, they should put the model release in Cloud Code ASAP and then just not bother with the API.

I don't know when that'll happen, but it could.

Speaker 3

I do think though it's a big customer base, the API customer base. Any type of product that is built on, I mean, with LLMs, like customer chatbot types of things, but also more generally, I do think the problem with I don't know exactly how the plans work in Clot, but you would reach a token max where you can only get so much with your subscription. You can, I think, buy more tokens, but I think it's just easier with the API at a certain scale?

And also like the whole OpenCLO customer base, right? Because they don't allow the plan anymore in the OpenCLO context. So you have to use the API.

And I do think given how many tokens OpenCLO generates, it's actually not a bad business if you don't lose money you know, on these tokens. If you sell it at a not subsidized price, I do think the API is actually not a bad business model.

Speaker 2

Yeah.

Speaker 1

Do wanna take a side? Do you wanna try a tiebreak? I'm obviously being cognitive.

Like, I don't really know, but I can see it. Like, Anthropic gives Apple vibes to me.

Speaker 2

mean, like, Anthropic has a higher chance of doing this. Yes. OpenAI, just because I I, like, have talked to the people so much, like, I I just don't super believe that they will have locked models to to to products.

Only only out of, I guess, idealism and sort of principles rather than economic incentive. Economic incentive would agree with you that they should have private models to to products. And, like, recently, they've done this.

Right? The last three GPT fives all had codex variants that were two to four weeks ahead inside of released only inside of Codecs rather than as an API. So they're starting to get there.

But just, like, constitutionally, I don't think the the people that run these things believe in, like, locking things behind APIs because they have such a huge market anyway. So they they, like, kind of don't care, and then they also, like they if if you're genuinely, like, sort of zealot like, if you're not trying to maximize the value of your company and genuinely just trying to spread h AGI everywhere, then then you reuse the API because you just don't know what people are gonna build with it.

Speaker 3

more thing though with the codex thing. I we will have to see, I think, next time because I think this time, it might also be a bit biased towards releasing it in Codex because they almost released it simultaneously with their app that they wanna promote at the moment. So it could have been more like they did that so that anyone checks out the app.

And but we'll see next time. Always a two two to four week exclusive window. Yeah.

Yeah.

Speaker 2

And, you know, that's that's their rightness. Yeah. Sure.

Promote codex. That's that's pretty effective. Yep.

We have a bunch of questions in the chat for, like, other things. Do we wanna cover benchmarks and then this thing?

Speaker 1

Go right ahead.

Speaker 2

What what do wanna do? It's it's your sub stack. I don't know.

Oh, no. Aw, man.

Speaker 1

It's a collective.

Speaker 2

It's a just dive into what you're interested in. Just go We I mean, Sebastian was interested in, like, the SuiteBench stuff.

Speaker 1

or, like, officially What do you mean by this?

Speaker 3

Yeah. Let's define Suitebench first, maybe.

Speaker 2

Okay. You know, I I happen to have the post on this.

Speaker 3

So let me just So the broader topic the umbrella topic here is how do we compare which LLM is currently the best LLM? Like, one of the ways would be SuiteBench, basically. But then, yeah, I will maybe let you explain because you had this brilliant podcast or article.

Speaker 2

I mean, is okay. Where do you want me to start? Want me to should we just define SuiteBench, I guess?

Speaker 3

I guess, yeah. So maybe going from So basically that it is a coding benchmark and then the SuiteBench is like a popular way to compare capabilities of LLMs. And then there is sweepench verified.

But maybe yeah. We should talk a bit more about sweepench first.

Speaker 2

So sweep sweepench is a is a paper out of Princeton from group, and they do a lot of good, like, code benchmarking work. And so and it it happened to be that he they just kinda drew thousands of of example sort of open source issues and PRs that closed those issues from open source. They there's there's a bit of selection bias here because they only focus on popular open source and only a small number of popular open source, but a large number of issues from those open source.

And then they just kinda dredged up some passing tests and then some failing tests that that you need to make pass in order to pass the score. When it when it launched, it was kind of obscure. Devin actually was the first one to pick choose it as a as a benchmark to report, and then it went from, like, I think at launch, it was, like, 13%, and now everyone's at 80%, something like that.

SuiteBench, because it was it was done on, like, a student budget, was very kind of let's call it sloppy or whatever.

Speaker 1

Terminal bench is like this now too. Like, they're they're just aggregate. It's, like, hard to do a a benchmark that is well calibrated across topics at different times.

Yeah. Yeah. It it is hard.

It is hard.

Speaker 2

you know, for the for the small group that is watching, I'm actually working on it with Cognition for for to launch a new benchmark here. But yeah. So so OpenAI was like, okay, guys.

We're like, SuiteBench is taking off. We are gonna adopt this, but we're we refuse to abide by, like, the full SuiteBench. We we're just gonna, like, actually go and go and curate, like, 500 subsets of of of the the original SuiteBench, and they actually hired humans to, like, go and vet through.

Like, I think the it's somewhere inside of this blog post, but, basically, they hired, like, three humans for every task to just vet whether the the task was, like, high quality or not because there's a lot of slop in there. Anyway, like, okay. This is the 500 that we that we're gonna endorse.

Speaker 3

let's say, challenging problems that are supposedly well defined.

Speaker 2

Yeah. Yeah. And and what's what's really funny is that at launch so this was launched in 2024.

At launch, OpenAI could not run all of its own 500. So for a while, like, there was, like, a few releases from OpenAI that reported on a subset of the subset because they couldn't run it on their eval infrastructure. So, like, their numbers were higher because their their denominator was lower, which is which is very funny.

Speaker 3

in that context, we should say what SpeedBench kind of looks like. I think it's like, basically like a a code that has a box in it, and usually the task for the LLM is to fix the bug in the code. Right?

It's it's right here.

Speaker 2

which becomes a problem in the future. Right But now, you can see the whole thing. Right?

You can see the the the reboots from the issue ID and the the problem statements, and then you have also the the test that you're supposed to pass and fail. So it's all here on Hugging Face. And you can see that it's at 500.

Anyway, that I think we don't have we don't wanna get too lost in the details on on on sort of I just wanted to say well, like, define the context, like, that this is a coding benchmark, essentially, 100 examples that are available on the Internet. Yeah. Okay.

And then if you want more a bit more historical context, this is like a step up from human eval, which is more on completions. Right? This is this was, in my mind, the first proper agentic benchmark, I guess, apart from Taubench, where they give you the the the problem and the the the end result, and they don't really specify how you're supposed to get there.

Whereas, I think a lot of, like, previous benchmarks like MMLUs of the world and and human evals, which is a coding domain also released by OpenAI, was very much here's, like, the sort of the the problem statement, and then give me the right answer immediately after without that much sort of extra files or anything that you're supposed to run. So it it like, the other ones are are more autocomplete. This one is more agentic.

There that's all the spectrum, obviously, because you can use agents to solve autocomplete, but that's not what HumanGail was was testing. Anyway, I I just I wanted to make sure, like, people understand that OpenAI actually put a invested a lot of money and effort into making Suitbench Erithide from By the question, how much money do you think this costs? Oh my god.

Don't don't do this to me. Millions? I would guess order of a cup like, it could be a few million, but probably I'd say, yeah, I'd say a couple million.

I I you know? So, basically, you do, like, okay. What's the first filter pass?

And then, like, okay. It's 500 times three because they had three Per people per yeah. Yeah.

Three people per thing, And then maybe, like, a couple more sort of verification passes or whatever. Right? So, like yeah.

So then they were like, oh, so this year, they're like, oh, well, not only is it saturated, because, like, progress everyone just takes turns to increment by 0.1 every time they release new model. It's like, it's bullshit.

It's obviously bullshit. Like, the the the inherent noise in just running these models varies by, like, point five to, like, one every time you run it. Like, you just choose the highest when you try to A little nitpick.

I don't think it can be 0.

Speaker 3

because, what we said before, because it's 500 examples.

Speaker 2

point 2% if I Okay. They might average Because like But like little detail. Yeah.

Sorry. I I I think in so as we progress to the next era of benchmarking, the n so this the n here is 500. Right?

The n doesn't directly correlate to the percentage points because you get sub points as well. Yeah. From your point.

Good point. So so, like, terminal bench, even though it has 90 something tasks, like, you you can get subdivisions less than 1%. Anyway, so not only do they do they have this, they actually audited their own.

Like, they were like, okay. Like, how come everyone is saturating at 80%? Like, what what's what's up with the remaining 20%?

How come everyone's, like, failing at it? And they were like, oh, actually, we looked we we paid even more people, six people per task now, with an extra team if if any sort of positive identification is found. And we were like, 59% of them cannot even be solved at all because the original benchmark was was, like, still slopped.

Like, stuff stuff got through that was not solvable. And I I actually tried to, like, illustrate this in in my post. So here, this this is an impossible test.

Right? Okay. So here's an example.

This is this is the sort of value added on top of the the original post. Here here's an example of a SuiteBench verified task that passed the first round of human verification. Right?

So here's the task, like and we wanna implement Python type ins or something. We wanna see expected behavior, I wanna see a string in the output. Right?

So the if you were given this, you would you would never pass this because the test said, I am looking for something called get annotation. And if you don't give me this magic string, get annotation, you will fail this test.

Speaker 3

Why? So, yeah. It's way to just play a coding interview.

Yep. Yep.

Speaker 2

yep. Right? So so so this is just a bad test that somehow escapes validation.

So the only way you could kind of solve it is if you're memorizing the Yeah. Exactly. Exactly.

Which is actually a nice like, I think every benchmark should include stuff like this where if like Like a honeypot? If you solve this, you're like, oh, shit. That is a canary.

Right? Like, it's like, oh, like, I mean, you're definitely cheating. Like a sanity check.

Yeah. Yeah. Yeah.

Yeah. Anyway a really nice point. Yeah.

Yeah. So so, like, I I just think, like, it it to me, it's a beautiful point of, like, how hard it is to make evals that there was these, like, multiple rounds. There was original sweepench, which, like, the the the Princeton kids did do initial first pass.

Then there's a second pass of OpenAI doing sweepench verified. And then and then, like, every single person that ran for the next SuiteBench verified for the next one point five years did not call this out until OpenAI was like, hey. Let's let's, like, look at the data.

So I think it's, like, really interesting. The while they were looking at this, they had a second thing that they would they they looked at the chain of thought. And inside the chain of thought, they found GPT-five's own chain of thought to start including information from the future.

Right?

Speaker 1

it would it would, like, use advanced knowledge of future versions of the Django version that they were using to solve the to solve the problem. Like, they they knew they've seen stuff like this in the real world where the models will hallucinate the new version of the API even if your script isn't on it. Like, I think a lot of the Hugging Face stuff is, like, the worst with this, where, like, the models just are totally gooboobly glopped.

Like, they've seen all the versions, and the API has changed too much over time where they fucking throw something out there.

Speaker 2

Yeah. So so the more previous Yeah. I mean, I I I think, like, you know, there's there's a lot of this.

Right? Like, the the sort of ethical behavior. Like, okay.

So you can blame things like, oh, you should not have released this the full dataset in public because, obviously, people can train on a full dataset. But, like, it's not like the researchers are trying to do this. Like, because these things are also open source.

Like, it is like, any any dataset that touches GitHub, any training corpus that touches GitHub is gonna just eventually absorb this. And then Yeah. Yeah.

And it's not even this website or the repository directly. It's a clone of this repository or someone else who has that develops their own open source library and has that in the unit tests or something where it's not even intentional or malicious or anything. It's like by accident, you already absorbed that.

Yeah. Yeah. Or or or a new feature that releases this edit only feature.

It gets written up in a blog post or a conference talk or something, and then it it just makes it in. Right? Like, it's it's really funny.

Okay. So to to me, like, OpenAI could have stopped there and said, okay. We're done.

They they did one more extra thing, which is kinda funny. They also then ran Flash, Gemini, and Opus. And this one, it was more it was, like, even more egregious.

Okay. They just gave the task ID and just said, repeat the SweetBench task to me. And so the from task ID, they can just vomit out the whole statement and the solution.

Speaker 1

These are crazy. The stuff that's in these models when you zoom in deep is really, really incredible. Because these are models that are really, really well done, but there's just so much complexity in all the pieces of the pudding that get put in the recipe.

Speaker 3

Yes. There's just so many weird formats. I also still find it fascinating that I mean, course, it's kind of by design when you're training that you memorize things because that's literally like next token prediction.

But given that how big a model is, on how much data it sees, and usually it sees only the data once, that it still has enough capacity to memorize. It's kind of like, so usually I would think, okay, I would have to train multiple epochs to be able to memorize, but no, it is enough maybe to include it once or twice in the training corpus and it can do a perfect rendition or perfect recap of what it is in there, which is kind of fascinating. Even if people don't want that, it's crazy.

Speaker 1

Yeah, labs got good at this. There's essentially, like, a duplication level that you need at each stage of training, and it's not easy to measure. So, like, if you do do too much at pretraining, your model forgets basic facts, and at post training, it's probably closer to these abilities.

And I think that that is a thing that is not well reflected in a like, you could see it in a vows of your knowledge tank. Yeah. Mean, I mean, this is like an art that they have probably gotten good at.

Speaker 3

Yeah. Like, continued pre training does also require some revisiting of old data. Otherwise, like you said, you have the forgetting.

But it's still fascinating to me that with such a small fraction usually, because you usually use one or 5% for continued pre training, that it's enough to have the model memorize almost everything, which is fascinating. Yeah. I don't know.

Speaker 2

still after all these years, fascinating. Yeah. I I think there's so one of the pet topics that I pursue, like, two, three times a year on on my stuff is the information theory of LLMs.

And I I I still think it's, like, super understudied. Like, how come you can memorize from one pass? Like Yeah.

Exactly. Right. And then and then also, like, people forget, like, superposition, which is like Anthropics original MechInterp work.

Also, basically stuffs information inside the smaller bits that then get forgotten. But, like, how how does super position actually work? I the people I don't think I've seen a convincing study on on that.

Yeah. Okay. Anyway anyways, I don't know.

I'm done on my on my sleep bench right I don't know if you have thoughts or questions or whatever. But I I do think, like, this is an example of, like, yeah, the model's unintentionally cheated, and and benchmarks are hard to make, and we need new ones. And, you know, if this happens to sweep bench verified, like, which I think is the most scrutinized benchmark in the world.

Speaker 3

I had a bar plot where I showed the sweep bench verified numbers for most models. And like you said, they were all 80 something percent. Literally 80 between one and nine, let's say, where there's almost zero variation.

Even something like MiniMax M2.5, which I do think is worse than GPT 5.2.

No offense, it's a smaller model. It's a cheaper model. For my usage based on OpenRider, it's a little bit worse, but on this particular benchmark, it's the same.

What I'm saying is that M25 should get less score on Swaybench, but I think other models should get more score. But like you said, the problems are just impossible to solve. But one point I think we didn't bring up is we said that Sweepench Verified has issues.

So what do we do about it? I think there is like a Sweepench pro now, which is kind of like a, I would say, like verified, try to fix the regular Suitebench and pro tries to fix verified. But I haven't looked into this.

Is it like another subset or is it a completely different set of problems? Yeah. It's a new set.

Speaker 2

the you know, CVench draws from, like, a 22 twenty twenty two ish, twenty twenty three ish era of problems. So all you do there's a few things you do. Right?

One, you do private public splits. Right? That's super obvious.

Two, you update the dates that which you draw from. And then three, you diversify the repos and the languages. Right?

So these are all just, like, very, very super basic fixes, and then, obviously, you try to fix the testing. Super basic fixes to the original sweepbench, which it doesn't take a genius to to figure out, but they did the hard work. And But it is in a sense also what Verified meant to do.

Speaker 3

let's say people looked at this again, but it's no guarantee that it doesn't also still have issues that might be discovered later on, right? I mean, it's No.

Speaker 2

SuiteBench Verified was an intentional subset, right? These guys were like, no, no, no, we need to have a superset. We not even superset.

We need to a Yeah. Yeah.

Speaker 3

Like what I was trying to say is when SuiteBench verified was developed, there were three people per task, making sure the task is well defined and everything. But then two years later it turns out, no, no, this was not the case for everything. What I'm trying to say is it could be that Threebench Pro is better, but it might still have issues that might not be obvious right now, but maybe in one to two years when we revisit this and you see some of the failure cases, maybe we'll discover, okay, this has still some issues.

So it's not a guaranteed perfect set is what I'm saying. I don't know, but it's just like a suspicion here. Totally.

Speaker 2

Totally. You know, I do think Scalaei has a professional interest in making sure this is gonna Yeah. No.

No. Like, yeah. Yeah.

But what I was trying to say is a three bench verified also had a professional interest to make sure it's just by the system. Different incentive. I guess they all have very different incentives.

This one has limited budget. This one has basically unlimited budget because it's literally existential to Scale dot ai that they have good data. Sure.

But I also I also think it's really nice that this team, the the eval team at OpenAI keeps endorsing Opus. Mhmm. It's kinda funny.

So yeah. Yeah. The they deprecate c bench verified, and then they were like, we're gonna report c bench pro now.

And g g five is, like, you know, number one.

Speaker 3

If do you know if I would want to evaluate on the private dataset, how would I do that?

Speaker 1

I don't know. Have my API key and agree to not. You have to agree because if you don't have an agreement, then you can just have to keep the data.

You have to do special hoops to make sure that you don't steal the private eval.

Speaker 3

my question was basically, do they even let you download the data? Or is it more like you send the answer to them and they do the evaluation on their back end so that you don't even get to download the data? Otherwise, like you said, you could Yeah, yeah.

So basically, you only provide the answers. So you have your LLM generate an answer and you submit the answers. And then they have like some process to evaluate on their things so that their data, private data never leaves their servers, my guess.

Speaker 2

Yeah. I don't know. I don't I don't have I haven't tried it, so I don't I don't I don't really know.

I'm sure you can sort of reach out to them to figure it out.

Speaker 1

Yeah. Anyway, I think this is good. Unless people have more comments that they wanna add.

Speaker 2

think this but this is only coding. Right? There's like every other domain needs this.

Speaker 1

the Rust thing right expensive. I think the Frontier evals are even more expensive, which is like the Yep. Apex eval from Merkur.

Like, evals are going to cost this is millions. They're gonna cost tens of millions and hundreds of millions of dollars at the Frontier, which is just a very strange dynamic. Whereas, there's so much about the ecosystem is forking between frontier models and then research and other things, and trying to follow that dynamic and explain it to people is gonna take a lot of work.

Speaker 3

But yeah, coding is, I do think, really interesting because that's what most people use LLMs for these days, but also it is easier to evaluate. Think once you leave coding math, it becomes a bit obscure. How do you measure the quality of the answer?

You get back to, let's say preferences, I guess, which is more like a subjective thing where coding is more objective. So it is not a bad thing to do.

Speaker 1

minor thing where Doesn't really Normal talent normal talent flows in AI. Total number.

Speaker 3

not trying to say this is like a big thing to talk about. What I'm trying to say is like, this is another interesting point for evaluating LLMs on those tasks because I think a lot of people want that to like they want an LLM to control the computer and do various things, but they are harder to measure. So that will be maybe two years we will have something more like benchmarks that can It's harder to specify.

It's kind of like, what is it called? In programming, there's unit testing and then the system testing, basically, like the UI testing and stuff like that. Yes.

Was it okay? Yeah.

Speaker 2

maybe gonna be the next Basically thing end to end testing. Yeah. Yep.

Yep. GDP val is usually the thing that gets brought up here. So if I'll just leave it there.

I I I think we've we've sort of beaten the the dead horse. Yeah. Benchmarks.

Yeah. Benchmarking. But definitely, GDP val is sorta here.

I'll I'll put it that way. Okay. Yeah.

Yeah.

Speaker 3

essentially, the distillation and the benchmarks this week. Yeah.

Speaker 1

And Cool. Mix to our coalition of Yeah. Whatever that means formally.

I I I just it just means I get to hang out with you guys, which is what we want. Anyway I describe I describe I mean, it's ultimately a media vehicle, and that the brands and vehicles for media are actually very influential today. I think you see many companies investing in that, and I think it's important to have people that you respect and are aligned with able to amplify each other.

Speaker 3

Yeah. It's also nice to talk to humans because I noticed the last couple of weeks, if you go to social media, well, I think it's 50% lobsters like OpenClaw clients nowadays. I get a lot of emails, but also notifications or responses that are They look AI generated.

So it's nice to also have this human connection and actually talk to, like, an expert about things. Yeah.

Speaker 2

Cool. There are a bunch of, like, comments. I don't know if you wanna do, like, quick hits, or are are you, like, kinda I have to go to a meeting.

Okay. That's why I'm trying to wrap this up. I see.

I see. I see. Okay.

Well, then, you know, time is yours. What what do wanna do?

Speaker 1

Okay. Thanks, everybody.

Speaker 3

We'll see you next week. Yep. Thanks, everyone, for joining.

It was, like, a nice spontaneous, I guess, discussion. Mean, it always feels nice to talk about things and too bad we didn't get too many, or we didn't get to discuss these chat questions because also on my screen, I probably need glasses at some point. My screen is pretty far away.

I can just barely read them. But yeah, thanks everyone for commenting. It is just nice to see also so many people excited about these topics.

Speaker 1

Yeah.

Speaker 3

Hopefully, see you later. Good rest of the day. Bye.

Shared via Hopper