[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

Latent Space: The AI Engineer Podcast
31 December 2025 17 min
0:00 --:--
Episode Description
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab racing to solve software engineering at scale. We caught up with John live at NeurIPS 2025 to dig into the state of code evals heading into 2026: why SWE-bench went from ignored (Oc

Summary

John Yang discusses the evolution of AI coding benchmarks, from SWE-bench becoming a de facto standard to new extensions like multimodal and multilingual versions. He introduces CodeClash, a programming tournament for evaluating long-horizon development, and highlights other benchmarks like Suificiency and Cycode. The conversation also explores the future of coding evaluations, emphasizing human-AI collaboration and code understanding.

Chapters

SWE-bench's Journey and ImpactJohn Yang discusses SWE-bench's growth over 1.5 years, its initial slow adoption, and its explosion in popularity after Cognition's Devin release.
SWE-bench Extensions and DiversificationThe conversation covers new versions like SWE-bench Multimodal and Multilingual, which expand language and repository coverage beyond the original Django focus.
Introducing CodeClash for Long-Horizon EvalsJohn explains CodeClash, a programming tournament where AI models develop and compete on consequential codebases, moving beyond independent unit tests.
Other Benchmarks from Ophir's GroupJohn highlights other benchmarks from Ophir's research group, including Suificiency for code optimization, Cycode for scientific coding, and Critical Point for physics-related tasks.
User Simulators and TauBench ControversyThe discussion touches on user simulator benchmarks like TauBench and VendingBench, including the controversy around TauBench's "impossible" tasks and the idea of intentionally including such tasks.
Future of Coding Evals and Human-AI CollaborationJohn envisions future evals with more dynamic environments and long-running autonomous agents, while the host emphasizes the importance of human-AI interactivity and code understanding.
Call to Action: Data and Multi-Agent ArenasJohn expresses interest in user interaction data from companies like Cognition and Cursor, and promotes CodeClash as a testbed for multi-agent and human-AI collaboration scenarios.
Cognition's Focus: Code Understanding & Context EngineeringThe host shares Cognition's initiatives in code base understanding to enhance human-machine collaboration and automatic context engineering for LLMs.

Topics

AI coding agentsCode evaluation benchmarksLong-horizon developmentHuman-AI collaborationCode understandingMultimodal AIMultilingual AIProgramming tournamentsUser simulatorsContext engineering

People

John Yang (guest) Ophir (mentioned) Carlos (mentioned) Walden (mentioned) Andy (mentioned) Michael Troll (mentioned) Jeffrey Ma (mentioned) Shinyun (mentioned) Karthik (mentioned) Dee Young (mentioned) Silas (mentioned)
Key Concepts (17)
SWE-bench — A de facto standard benchmark for evaluating AI coding agents, initially slow to adopt but gained traction after Devin's release.
SWE-bench Multimodal — An extension of SWE-bench that incorporates multimodal aspects for evaluation.
SWE-bench Multilingual — An extension of SWE-bench covering nine languages across 40 repositories, diversifying beyond Django.
CodeClash — A programming tournament where two or more language models maintain and improve their own codebases, competing in rounds to evaluate long-horizon development on consequential code.
Unit Tests as Verification — John Yang expresses dislike for unit tests as the sole form of verification in AI coding evaluations, advocating for more dynamic, long-horizon methods.
Long Horizon Development — The concept of evaluating AI models on sustained, multi-step development tasks on a code base, where actions are conditioned on previous model outputs.
Economically Valuable Arenas — The goal for CodeClash to move beyond simple programming games to real-world, economically impactful scenarios for evaluation.
Suificiency — A benchmark focused on code optimization, where models make modifications to a codebase to make it run faster without changing its behavior, keeping unit tests passing.
Cycode — A benchmark for scientific coding, described as a "human eval, but better," focusing on code completion tasks that are less expensive to run than agentic benchmarks.
Human Hours Worked (Meter) — A metric used by Meter, where the x-axis is runtime and the y-axis is completion, projecting long-running staging tasks.
User Simulator Benchmarks — Benchmarks like TauBench and VendingBench that attempt to simulate user interactions to evaluate AI agents, though their realism is debated.
Impossible Tasks in Benchmarks — The idea of intentionally including tasks that are unsolvable or underspecified in a benchmark to identify models that "cheat" or fail to recognize limitations.
Refusals (in LLMs) — The ability of an LLM to identify and state that it cannot complete a task, rather than attempting and failing or hallucinating a solution.
Long Autonomy (in LLMs) — The concept of AI agents working autonomously for extended periods (hours or days) on a task without human intervention, which the host views with caution regarding its real-world utility.
Human-AI Collaboration — The emphasis on finding a balance between AI autonomy and human involvement, allowing for different levels of abstraction and interaction depending on the task.
Code Base Understanding — Cognition's focus on helping humans understand their own codebases better, enabling a "mind meld" between human and machine for complex tasks.
Automatic Context Engineering — A research sub-agent being developed by Cognition to automatically generate relevant context for an LLM.
References (30)
SWE-bench project
Devin product
Cognition company
SWE-bench Pro project
SWE-bench Multimodal project
SWE-bench Multilingual project
CodeClash project
Highlight by Michael Troll game
Cursor company
Jane Street company
Two Sigma company
Sweet Lancer project
GDP BAL project
Terminal Bench project
Suificiency by Jeffrey Ma project
Algo Tune project
Cycode project
Cycode two project
Meter company
Critical Point by Ophir project
SecBench project
SREBench project
LOD company
TauBench project
TauBench two project
VendingBench project
Impossible Bench project
Anthropic company
LM Arena product
Google company
Transcript (33 segments)
Speaker 1

We're here at NeurIPS with John Yang, SuiteBench, and many other things. But welcome. Thanks so much for having me.

Yeah. Really happy to be here. Last year, talked to Ophir and I think Carlos as well, one of your co authors.

Yeah. How's SuiteBench doing?

Speaker 2

the project is, like, one and a half years old? Yeah. Yeah.

So I think one and a half years old in terms of when it was actually useful. Yeah. They were put it out October 2223, and then people didn't really touch it too much.

And then, of course, like, Cognition came on the scene, and Devin was an amazing release. And I think after that, it kinda kicked off the arms right you beforehand? They just showed up.

It was you know, I got an email about, like, two weeks ago. I think it was from I think it was from Walden. He was like, hey.

You know, we have a good number on it. I was like, wow. Congrats.

You know, thanks for using it. And then the release was like mind blowing. I was like, wow.

These guys did an excellent job. Yeah. Amazing.

And then SuiteBench verified was, like, maybe last year. That's right. Yeah.

Like Catch us up this year. Like, you have other languages.

Speaker 1

You've there there's, like, a whole bunch of varieties of SuiteBench now. Yeah. So what should people know?

Yeah. For sure. I think there's a couple extensions that are happening.

Speaker 2

Sweet Bench Pros, Sweet Bench Live.

Speaker 1

Oh, Sweet Bench Pro, was that with you guys? Because it looks independent. It's like different authors.

It's completely independent. Yeah. So they just called themselves Sweet Bench Pro without your blessing?

Yeah.

Speaker 2

I think I think we're we're we're we're okay with it. When we came out, we were like, oh, cool. Interesting.

It would've been, you know, fun to be part of it. But, you know, I mean, congrats to them. That's a great benchmark.

Yeah. Alright. Yeah.

But yeah. Multimodal? Yeah.

We did multimodal and multilingual, and I think like those have Multilingual seems to Is be it like JavaScript? What else? Yeah.

Multilingual is like it's like nine languages across like 40 repos. But, you got them.

Speaker 1

Rust, Java, C, you know, Ruby. Yeah. Yeah.

You got them. Yeah. And then, of course, the event itself, a lot of people, like, they they talk about the the Django focus.

Yes. Django. Is there is there, like, I don't know.

How do we how do we move past Django? Yeah. For sure.

Speaker 2

a lot of the newer benchmarks, like, really try to diversify the repos. Like, in the two follow ups we did with multimodal and multilingual, we made it a point to do that. So I think But you can also just put out SuiteBench '20 25 and just That is true.

And do a new distribution. Yeah. Yeah.

So it's been cool to see the follow ups. I think quietly, and and it's an open question for me. I'm excited to see how people curate the next sets.

Like, it's kinda interesting to see in the literature or in their blog posts, like, how they're justifying why they're creating their separate split. The easier ones were like, oh, more languages, more repos. And then I think now people are like, well, ours is more difficult because of this curation technique.

Speaker 1

And I'm, yeah, I'm excited to see how how long that lasts and, you know, where we're gonna, like, guide the evaluations towards. Yeah. And more recently, you're working on Code Crash?

Yes. That's right. So let's get people you've already done other episode other podcasts about it.

Yeah. I would refer people to to that with your chat with Andy. Yep.

But just give, like, people, like, a one, two sentence. Yeah. No.

Happy to do it, especially on your podcast. It's on.

Speaker 2

So, basically, the idea is I don't like unit tests as a form of verification, and I also think there's an issue with Suitebench where all of the task instances are independent of each other. So the moment you have the model kinda submit it, oh, it's done, you know, and and that's the end of the story, end of the episode. You know?

So with CodeClash, what we're thinking is, let's try to really evaluate, like, long horizon development and development on a code base that is consequential and conditioned upon what a model did, you know, before to that code base. And so the general idea is you have two or more language models, and they play a programming tournament. And what that means is each model maintains their own code base.

And each round of the tournament, first, they get to, like, edit and improve their code base however they see fit, very self determined. And then in the competition phase, those two code bases are are pitted against each other. So the code bases are run, and there's generally an arena.

You know, we have a lot of diverse arenas, but the arena is determined, like, code base a is better than code base b. And then you kinda repeat that across multiple. As determined by an element judge.

Yeah. Yeah. So element judge is definitely one of the mechanisms.

We started with some pretty, like, simple programming games. So one of the cooler ones is, like, Highlight, which Michael Oh, yeah. I played it for Jane Street.

Yes. That's right. Yeah.

That's right. You know, that's awesome. Yeah.

Highlight one, two, three, like, Michael Troll of Cursor wrote this game. Two Sigma on Jane Street? Yes.

Yes. Oh oh, Two Sigma. Two sig it's Two Sigma.

I worked at Two Sigma.

Speaker 1

there you go. Yeah. I mean, this is too long ago.

There you go. Yeah. 2016 at this point, but we're bringing it back.

You know? Hellen is fun. I I I would say if you've never done a programmatic competition where you have to control fleets of ships and attack things and defend things Yeah.

And collect resources. Yeah. It's like, you you play StarCraft, but you can code.

Right? Exactly. Exactly.

Yeah. Yeah. A lot of games Yeah.

Is there are there non games, or are you focusing on games? I think that's an excellent point. So for kind of the initial release for scientific purposes, we kinda use existing programming games.

Speaker 2

The current ongoing effort is, you know, to build economically valuable arenas. That's, you know, the popular word these days. So Yeah.

Sweet Lancer is a big one this year. Yeah. G d GDP BAL awesome.

Yeah. Just I mean, I think the big selling point of Terminal Bench and Sweet Bench and these EVALs is that it was really close to real world utility. And so I think it's resolvable for CodeClash, and that's what we're working on.

Yeah. Okay. Yeah.

So you're part of Ophir's group? Yes.

Speaker 1

other students have also been putting out a lot of other stuff. What would you highlight? Yeah.

No. I mean, Ophir is such a prolific mentor when it comes to benchmarking.

Speaker 2

Suificiency, I really like in the line of Yeah. Performance What's the TLDR on that one? Yeah.

For sure. So Suificiency was wrote by this PhD student called Jeffrey Ma, who happened to be my high school classmate. And the idea there was, like, you take a code base and you just want to, you know, do modifications that will literally make the code run faster.

So I think it's just, like, parallelization, SIMD operations, stuff like that. Yeah. So so no no behavior change, just faster?

Exactly. Okay. Keep the unit test passing, but I want better runtime.

Okay. Yeah. Yeah.

And then and then there there's algo tune that is kind of in line with that. And then there's also kinda pushing along, like, the scientific coding domain. Cycode?

Yeah. Exactly. Which is like zero code.

Cycode two is awesome.

Speaker 1

They did, like, a quick quick one. Yeah. People, Cycode is the way I explain Cycode is it's human eval, but better.

Yes. Exactly. Exactly.

I think, you know, there's a lot of good stuff that these days where yeah. That's that's the way to go. Which is, like, Superbench is expensive to run.

Any agentic benchmark is expensive to run. Yeah.

Speaker 2

Yeah. Just just just complete. Exactly.

Like, you know, you can do well on those first and then sort of graduate to the multi turn expensive stuff. Yeah. Yeah.

Speaker 1

K. Other than that, just like broadly other work in the field in 2025, in terms of coding evals, obviously, we shot up Meter.

Speaker 2

number. Yeah. They like the x axis being sort of the runtime and their or, yeah, y axis being the completion.

You know? Like, we can do more long running staging tasks. Yeah.

I think the projections are are quite interesting, and I definitely appreciate them kinda using sweep bench verified to to sort of proxy a lot of these things. But, yeah, they're great. Yeah.

Okay. Yeah. Any other work that, like, caught your eye that you're gonna Yeah.

I mean, I I think within the okay. Terminal bench sweep bench. Yeah.

Critical point was kinda cool. Critical Point? Yeah.

The it's like a very new benchmark that Ophir did, and I think it's kinda related to physics. There's this one called SecBench kinda related to cybersecurity. Exactly.

SREBench, which I I think is affiliated with LOD. Like, it's just cool to kind of see people really dive into different coding domains. And then stepping a little bit outside of coding, I'm personally, I think it's quite interesting to think about the user simulator stuff.

So, like, TauBench bench? TauBench too. Yeah.

And and VendingBench.

Speaker 1

I got the big feelings. Yeah. No.

I'm interested. You Well, I mean, it's it's like it's like you're sampling one path. I I don't know how realistic it is, to be honest.

It's yeah. It's it's just the yell at this, but it is cool. No.

For sure. Yeah. I agree.

I I think it's a good initial effort.

Speaker 2

To me, I think it's super cool to see companies like you know, I'm sure Merkor and stuff are focusing on building environments, like, for code, beyond code. And so I think it it might be interesting to have, like, work gym style stuff. This is stuff that my adviser, Dee Young, at Stanford thinks about a lot.

So yeah. Yeah. I just realized we we're talking about terminal bendy.

Yes. We're fighting a lot of other folks. Yeah.

Yeah. You know, really really really good work just overall.

Speaker 1

Yeah. Let's talk about TauBench because you mentioned TauBench. Yes.

Yes. There's some discussion or some people are saying that TauBench is impossible to get a high score on because some of the tasks are under specified or just impossible.

Speaker 2

Yeah. I don't know if you're up to speed on that. I'm a little It's a little spicy.

Yeah. It's a bit spicy. I think I saw so I, you know, for like, I worked with Shinyun and Karthik back in Princeton very closely.

Think Karthik, I just saw, posted a tweet kind of Defending Yeah. Like, rebutting some of these claims. Yeah.

Speaker 1

but, yeah, I think it it also brings up just maybe, like, interesting research problems to solve of, okay. Like, why is it impossible? Is it the ambiguity?

Is it kind of the user simulator that has issues? And I think, generally, we all agree that, you know, we'll improve on these things over time for you guys. So I actually really like benchmarks that intentionally I think we should intentionally include impossible tasks Mhmm.

As a flag, yeah, of like, hey. You're cheating. Yes.

It's kinda sad that, like, Kartik actually is defending it because the master move would be like, oh, yeah. You caught us.

Speaker 2

you've been cheating. Yeah. Oh, interesting.

It's that would be that would be cool. Yeah. I mean, yeah, you'll have to ask the TauBench authors, but yeah.

No. That that's that's fun. Yeah.

I I think there was Impossible Bench was a recent benchmark maybe from, was it from Anthropic? I don't know. But they basically took SweetBench verified, and they changed the issues to make them impossible, and they checked, like, how often the models would be like, I actually just can't do this.

I don't know what's going on. Oh, like for refusals? Yes, yes, yes.

So Oh, how do they do? I thought that was interesting. I think they're all the models are all kind of attempting and saying like, oh, I did it, you know?

So maybe not great on this That's cool.

Speaker 1

an important one. Yeah.

Speaker 2

How does coding eval evolve next year? Wow. That's a great question.

I mean, honestly, I think I think it's people will make more sweep benches. I think terminal bench has really got something going where you you ask people to you know, a sweep bench, you're you're confined in some sense to the domain of issues and PRs that already exist, which I think has its benefits of being close to reality and natural. But I think with Terminal Bench, there's a lot of creativity that you can infuse into that.

So I would personally be really excited. Like, the two point o job was really excellent, and I'd be super excited to see, you know, three point o, four point o. Because of, like, the environments?

Yeah. I mean, the environments, you know, bringing more people into the fold. You know, I think correct me if I'm wrong, Mike, but early on, you had PhD students, very smart CS people who were adding tasks.

And, you know, what does that look like when you fold more coding environments for noncoding tasks, noncoding environments in general, and ask people to make stuff there? So that's pretty cool. And then, of course, for myself, I think just, like, this long running kind of thing just feels very compelling.

I think the vision of, like, hey. I tell it a goal. I don't have to be super specific about my task.

I have, like, a decent verifier that proxies what I want. Something literally like a code base that makes the most money in this, like, setting. You know?

Like, that's my verifier. Verifier. You know?

And I walk away for five hours. The thing is just running. I'm hanging out with you, talking to my friends.

I come back, and it gives me, like, literally a soda code base on on on that, you know, task. I think that would be super cool. Okay.

I'll push back. We're part time in Cognition. Yes.

Speaker 1

emphasizing a lot of interactivity. Mhmm. Because the the point is that you're going to under specify.

Right? Right. And actually, what people want is back and forth, back and forth, and on like a really fast time frame, which is terrible for a benchmark author.

Right? Because this is how you can do that. Yeah.

Yeah. But but realistic. Yeah.

So I I think, like, the this this this is where I I'm a little bit anxious or cautious about this push for long autonomy. Right. We're gonna I mean, you know, let's say this time next year, we'll have five hours is is pessimistic.

Like Yeah. It'll be it'll be twenty four. Long.

Right? Right. Days.

Yeah. But I don't know if that actually materially changes the industry. So we will push it, like, as an evals you know, we have the people people make evals here.

Yeah. We push the industry in ways that we wanted to push, but I don't know if we like, that's a productive way because that's more of, like, a a stunt that that, like yeah. It's a proof of concept that prove existence proof, it can be done.

Yeah. But will you use it In fact. For real life?

Yeah. Yeah.

Speaker 2

to me, I think there's there's potentially room for growth, so I I would actually agree with your take here. I mean, with my lab at Stanford with Dee, like, there's a you know, her emphasis is on human AI collaboration. And so I I definitely don't believe in this idea of just kinda getting rid of the human.

But, yeah, maybe just, like, finding the balance of, like you know, just because the developer ecosystem is so diverse and there's so many participants in there who want different things out of it, like, just enabling different levels of abstraction and, you know, depends on the task. Like, there's settings where you wanna be, you know, more involved and more sort of hands on, and so you wanna use Windsurf for that. But then maybe there's kind of this general data processing thing.

It's just a lot of JSON parsing you don't really care about, and that's the one I kinda wanna walk away from and just let it figure it out. Yeah. So, yeah, I I would agree with you, generally.

Yeah. Amazing. Any call to action?

What what do you want help on?

Speaker 1

How can people,

Speaker 2

I guess, like, find more of your work? Definitely. For the call to action, super jealous of all the great data that Cognition and, you know, Cursor would get.

Like, that user entered action data is, like, really fascinating. From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product like LM Arena that people have people use consistently, which is, I mean, really tricky in and of itself, or you build, like, really good user simulators that try to mimic sort of these settings, but that is also, like, nontrivial.

I don't think it's as simple as, hey. Check GPT. Act like a human.

Right? Yeah. So it would be really cool to sort of get inspiration of, like, what exactly does that data look like?

Or or between the two, like, what's the best way to scale up sort of evaluating human AI interaction? And then I think for visibility for my own, we're pushing more arenas. Like, I think for for CodeClash, what I'm excited about is the current framing is really long running SWE agents, but, you know, you could have multi agents, like two agents work together on the code base.

And what happens? You have a human and an agent work on the code base versus just AIs. What happens there?

You know? Like, when the models improve, and hopefully, hill climb and they become better at digesting logs and iterating on analysis, you know, how does how does human AI interaction, like, change with model capability?

Speaker 1

playing one arena at a time, n arenas at a time, you know, and just you know? Yeah. I think very interested to work with you on on the interaction stuff.

Oh, that would be awesome. Then I I think one one more thing I'll add is for Cognition, it's gonna be pushing a lot of code based understanding, which is kind of Yeah. Code based retrieval plus plus.

Yes. And mostly, it is helping humans understand their own code bases better to enable humans Yeah. Or to to sort of mind meld the human with the machine to do the highest possible task that LLMs cannot do alone, humans couldn't couldn't do alone.

And then the other thing is also, basically, automatic context engineering for an LLM. So that that that is, like, sort of, like, a research sub agent that we're that we're working on. That's so awesome.

Yeah. So I don't know what the benchmark would be because, like, how do you how do you benchmark understanding? That is.

Apart from think, like, just yeah.

Speaker 2

have some manually curated answers, and then, you know, post trivia questions. That's very easy to saturate, so I don't know. I'll ask you.

I think I think Silas tweeted a while ago, like, sort of, like, the the wiki, the code wiki, etcetera.

Speaker 1

incredible. I mean, I I use it on the But Google actually just came out their own version. Oh, yeah.

Yeah. With the the anti gravity people. That's No.

No. No. This is, a separate It's a bit around heat.

Yeah. Gotcha. Gotcha.

Yeah. But cool. That's the state of code.

Yep.

Shared via Hopper