[State of Evals] LMArena’s $1.7B Vision — Anastasios Angelopoulos, LMArena

Latent Space: The AI Engineer Podcast
6 January 2026 24 min
0:00 --:--
Episode Description
We are reupping this episode after LMArena announced their fresh Series A (https://www.theinformation.com/articles/ai-evaluation-startup-lmarena-valued-1-7-billion-new-funding-round?rc=luxwz4), raising $150m at a $1.7B valuation, with $30M annualized consumption revenue (aka $2.5m MRR) after their September evals product launch.—-From building LMArena in a Berkeley basement to raising $100M and becoming the de facto leaderboard for frontier AI, Anastasios Angelopoulos returns to Latent Space to

Summary

This episode features Anastasios Angelopoulos discussing LMArena's journey from a Berkeley project to a $1.7B valued company, focusing on its mission to evaluate frontier AI capabilities through real-world user interaction. He addresses the "Leaderboard Delusion" critique, clarifies LMArena's pre-release testing practices, and outlines the platform's commitment to integrity, expansion into multimodal and expert-specific arenas, and strategies for community building.

Chapters

LMArena's Origins and Early SupportAnastasios Angelopoulos details LMArena's spin-out from LMSYS at Berkeley, highlighting early incubation and grants from investors like Anj and Sequoia.
Scaling LMArena and Funding AllocationAnastasios explains the decision to form a company to scale LMArena's mission of measuring and advancing frontier AI capabilities, detailing how their $100M funding supports inference costs and hiring.
Differentiating LMArena from CompetitorsThe discussion contrasts LMArena's organic, user-driven evaluation approach with other platforms like Artificial Analysis, which rely on public benchmarks and rerunning.
Addressing the Leaderboard Delusion CritiqueAnastasios responds to the 'Leaderboard Delusion' paper, clarifying LMArena's pre-release testing practices and correcting factual inaccuracies regarding model sampling bias.
LMArena's Core Principles and Future DirectionAnastasios outlines LMArena's commitment to platform integrity, providing a transparent and fair leaderboard, and expanding into new categories like occupational and multimodal evaluations.
Community Building and Strategic PartnershipsAnastasios shares insights on building and retaining a large user base through value provision and discusses LMArena's partnerships with major model labs for evaluation.

Topics

AI model evaluationLLM leaderboardsStartup fundingConsumer AI platformsMultimodal AICommunity managementOpen source modelsAI ethics

People

Speaker 1 (host) Anastasios Angelopoulos (guest) Alessio (mentioned) Anj (mentioned) Wey Lin Yan (mentioned) Naina (mentioned) Greg (mentioned)
Key Concepts (7)
LMArena's Mission — To measure, understand, and advance frontier AI capabilities on real-world users and usage, based on organic feedback, providing a North Star for the industry.
Organic Feedback Evaluation — LMArena's method of evaluating AI models where users input their own use cases and questions, providing a realistic assessment of model performance.
Leaderboard Delusion — A paper that critiqued LMArena, claiming undisclosed private testing of pre-release models created inequities and bias in the leaderboard.
Pre-release Testing — LMArena's practice of allowing model providers to test unreleased models (e.g., NanoBanana) on their platform, which the community enjoys for early access.
Multimodal Models' Economic Value — The idea that multimodal AI models, particularly in areas like marketing and design, will become highly economically valuable in both consumer and enterprise sectors.
Platform Integrity — LMArena's core principle that the public leaderboard is a 'charity' and a loss leader, where models cannot pay to be listed or removed, ensuring a transparent and fair reflection of performance.
User Retention Strategies — Methods like providing persistent history for logged-in users, which significantly drives retention on consumer platforms like LMArena.
References (20)
LMSYS company
Sequoia company
ChatGPT product
Artificial Analysis company
Gradio tool
Hugging Face company
React tool
Nx.js tool
Leaderboard Illusion paper
Meta company
Lava Board model
NanoBanana model
Reeve Image model
BFL model
Google company
OpenAI company
Gemini model
DeepSeek v3.2 model
Cognition company
Devin project
Transcript (34 segments)
Speaker 1

Alright. We're here with Anastasios from Arena. I actually don't actually know your starts with an a.

It's like very Angelopoulos. Yeah. Yeah.

There you go. Congrats on all the success. You got the Arena handle.

Yeah. We did. We got the Arena handle.

Thank you. Big branding moment. I mean, I I think x is, like, being more commercial, so obviously, you bought it.

But, like, at at least you, like, have a place to go to where you can be like, hey, like, we really like this.

Speaker 2

Yeah. Has changed the the the the the feel of it. I don't understand how you feel.

Yeah. I don't know. I mean, the reason we kept the LM at the beginning is because we were, like we started as LMSYS.

Right?

Speaker 1

at Berkeley. So we decided language models. Yeah.

Exactly. So So we wanted to maybe broaden a little Yeah. And and we were the first arena, so we feel like let's kinda try to own that.

Yeah. So Last time you we had you guys on, you hadn't really spun out yet. And we I did a call with Alessio, and I was like, these guys are gonna start a company.

And I didn't know I think you actually were already started at the time. I don't I don't remember. Yeah.

Maybe. Because Anj I had a I I chatted with Anj. Mhmm.

And he said he was your founding CEO? He was. Indeed.

Which like people don't know. Yeah. The the Anj Anj is a very interesting character.

We we have a podcast schedule with him. Yeah. Yeah.

He does a lot more than normal VCs.

Speaker 2

He does. He's been incredible to us. You you wanna shout out some stuff, Ethan?

Yeah. Absolutely. You know, the way the company started was as an incubation by Anj.

Yeah. So what he did is he kinda, like, found us at Berkeley and picked us out of the basement, and was like, hey, these guys seem like they're onto something, and started working with us really early. Gave, you know, gave us some grants.

He was not you know, a 16 was not the only one to do this. We also had a great grant from Sequoia, but Anj was, in particular, quite quite supportive of us, and, you know, gave us some resources in order to continue building out Arena before we even were committed to to starting a business. And in that capacity, he sort of, like, you know, formed an entity for us and this and that.

And, you know, he was like, hey, you guys can walk away at any time if you guys don't wanna start a business. I mean, was really incredible. Very aggressive investment move by him.

Right? Because, of course, any money that he spent business. Yeah.

At the end of the day, like, we could walk away and give leave him with nothing. But I think he knew wisely that the right thing for Arena was to start a company out of it. It was the only way that we could scale and that, you know, Wey Lin Yan and I would would ultimately see that and be excited about doing it ourselves, which is which which ended up being the case.

Was there a moment for you where you're I'm sure you were debating it yourself. You had other opportunities.

Speaker 1

What was the deciding factor for you?

Speaker 2

the only way to scale what we were building was to build a company out of it. That the world really needed something like Arena. Arena being really a place to sort of measure, understand, and and advance the frontier AI capabilities in on real world users, on real world usage based on organic feedback.

And that in order to achieve the scale and, you know, distribution necessary and the quality, of course, of the platform necessary to to do this effectively, we would need to start a company out of it. You know, we considered other options. Are we gonna keep doing this as an academic project?

Are we gonna do it as a nonprofit? Blah blah blah. But ultimately, under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.

So you raised a 100,000,000 or 80? $100,100. 100.

$100,000,000. That's a lot of resources. It's great.

What Yeah. What's it for? Well, you know, obviously That's good on behalf of, like, everyone who's Yeah.

Like Everybody. Dude, it's an arena. How are you gonna spend your money?

Yeah. Yeah. It's great.

Yeah. Please tell. First of all, we don't necessarily need to spend all that money.

Right? The the purpose of money at a company is to give you cards to flip. It's to say, hey, you you have enough resources necessary, and so that if your first bet fails, you can make another bet and another bet.

Of course. So that's not to say that we're gonna spend all of it. Of course, you wanna spend things responsibly.

Having said that, the platform's actually quite expensive to run. We fund all of the inference on the platform. You know, the way the way that it works is the platform you pay market rates.

They don't give you discounts. No. No.

We get discounts, but they're but they are standard enterprise discounts. Okay. The same that would be given to any other customer.

Speaker 1

If you've closed any I don't know any numbers. I I see numbers of votes. Uh-huh.

But like, what's that in like monthly tokens or I don't know. I don't know about tokens, you know, I'd to back back that out.

Speaker 2

again, this is off the cuff, but I can safely say we have more than 5,000,000 now. Five, six million. We have probably 250,000,000 conversations that happen over the course of the platform.

We're on the order of, you know, mid tens of millions of conversations every month that are having on the platform. It's actually quite a large, you know it's one of the largest consumer platforms for LLMs. Of course, nothing is really comparing to ChatGPT, but, you know But no.

Still. I mean, I think the the largest skilled ones I The benefit of this is that like, it's actually quite a diverse population. So 25% of the people on our platform, for example, do software for a living.

Still at this scale. How do you know? Because we like do all sorts of we we either survey them, or we'll like analyze the prompt distribution that's coming to platform.

I'm happy also to share more. We've done something called Expert Arena, which is like trying to understand the distribution of experts that are coming to the platform. A lot of that can be like unauthenticated or whatever usage.

It is. But about half of our users now are logged in. Yeah.

And so we have some ability to understand them, and and we also have surveys that we run on the platform that tell us a bit more about who are the actual users. Yeah. Of course, there's always, like, response bias in surveys, so you have to take it with a grain of salt, but nonetheless.

And if you don't know Anastasia's background, like, he's he's the guy to correct for response bias. Yeah. No.

Okay. And there's a lot of there's a lot of guys like that and girls. So Like, guys and girls.

You're not the only player. There there are this artificial analysis started in Arena. Yep.

AI. Yep. It's like some crypto people that started this.

Yep. Yep. I don't know if you have a conversation with them on like, well, hey, this is our thing or, like, let's work together on something.

I I I don't know if you've what the industry No. You know, so I've talked to artificial I've actually talked to both groups. Yeah.

Both seem, you know Am I missing any major players? It's just those No. I think those I don't know.

Yeah. I don't think so. Okay.

I think those are like some, of course, large depending on how you define the term. Artificial analysis obviously has like huge, like, market mindshare Yeah. Around the analysis of of different AI systems.

Yeah. They told me they were just like, we we are gonna be gardener of AI. Yeah.

That's that's kind of their goal. And I think they're going after that consulting market and so on. Artificial analysis, from what I understand, is like a team of consultants that are doing this.

Yeah. Seem like really, really nice guys, going after that that particular market. It's different from our platform in the sense that their analysis is based on sort of like aggregating public benchmarks and turning those into analytics.

Independently rerunning. And independently rerunning, which yeah. Yeah.

Yeah. Which matters. Yes.

And and and and using those in order to compile sort of reports and so on that that educate the field on the performance of all these different models. But they also have arenas. They have arenas, but the arenas are not based on organic usage.

Like, the the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a a level of realism Yeah.

That platform doesn't have. Of course, they they specialize in a slightly different thing, but I see those platforms kind of diverging in that sense. Yeah.

Yeah. And sometimes it's like, it's the only way to do this.

Speaker 1

their video arena is pre generated videos. You can't enter in your own video process. Correct.

But we're doing it organically. Yeah. Like, exactly.

And so As a voter, it does help in terms of I don't have to wait. It does, but also why would you go?

Speaker 2

Do you actually care about like other people's videos? To form your own intuition. Maybe.

Yeah. Maybe you're like interested in comparing Like, I'm a shitty prompter. Right?

Are you? So I don't believe that. I'm a terrible prompter.

No. No. No.

No. I learned by example. Don't denigrate yourself.

Don't denigrate yourself.

Speaker 1

there there are many prompters much better than me. Let's say that. Right?

That's that's that's a fact. People have all sorts of cool ways of prompting LLMs. Yeah.

It's educational to say. The the only way to learn is by, like, well, look at other people's prompts and see the results and go, oh, I didn't know you could do that. And that's hard.

Totally. Yeah. Yeah.

Yeah. Yeah. Okay.

So let's go back to Arena. Oh, one thing I do wanna say is the number one use of funds is getting off Gradyo.

Speaker 2

Oh, yeah. Yes. We well, listen.

Gradyo, incredible platform. Gradyo scaled us to a million mal. Yeah.

That's incredible. And, of course Can you tell the Hugging Face ones that? Of course.

Yeah. We're really, really grateful to Gradyo for taking us so far. Eventually, you know, it became time for us to move off of that and go to a React sort I'm sure Hugging Face would have loved you to stay on.

They would have. I'm sure. Yeah.

Was there a technical, like, reason you you you just couldn't get the performance? Yeah. It just became hard to develop, and there were all these tools that we wanted in React.

And, you know, to do all the fancy things that you can do in React became kind of like One example of that. I don't know. What what's a feature that you really wanted?

Let's say we wanted to create, like, our own custom, like, loading icons for video with notifications. Okay. How are we gonna do that in in React?

It's hard. Yeah. I mean, we'll make a Sure.

There's a custom component. I'm sure the high end guys are gonna come in and say like, hey, you can do that in Gradio, which maybe you can. Of course they could.

Also fewer developers know. How are we gonna hire for that? Do we have to reskill them?

People are less familiar with that stack. Yeah. You know?

So anyway It's full React Nx. Js, all this. Yeah.

Yeah. Yeah. All that.

Okay. Cool. Other use of funds that that that might be interesting?

Like, you know, basically No. That's basically deployed resources. Okay.

Yeah. It's it's it's on primarily inference. That funds the free usage of the platform, and and then also hiring, of course, headcount.

Yeah. We have an office. You know?

That's us.

Speaker 1

That's an SF. I'll tackle one of the major things this year, which I'm sure you're tired of thinking about, but I for people who are not in the loop, this is gonna be news to them. The Leaderboard Delusion.

The whole thing will cohere. Let's summarize the what they said, and then your response. Mhmm.

Speaker 2

So Leaderboard Illusion's a paper that critiques Alan Marina. And the main Pretty, like Yeah. Brutally.

Well, you know, I would say unscientifically. Yeah.

Speaker 1

And let's be clear, Coher wasn't doing that well on Alan Marina.

Speaker 2

was, like, 74. It's all good. You know, it's actually not it's a it's a respectable place that they had on the leaderboard.

I don't even think it was really Cohere people like the Cohere model developers doing this. It was more their research side. But in any case, what is the what is the leaderboard illusion say?

It says that Alan Marina was what their the claim is that Alan Marina was doing this undisclosed, quote unquote, private testing on our platform, that model providers will send us pre release models, and we'll expose them, and so on and so forth. And that this creates so called inequities in the leaderboard, you know, due to that prerelease testing. For example, they will they cited that Meta, at some point, tested Lava Board.

Amount of models with us. Yeah. Of course, we can't disclose all of the details of how all that was done, but that is that is the main claim of the paper.

Now our response to that paper, it's online. You can find it on response to Leaderboard Illusion. And our response to that paper is essentially pointing out a series of factual mistakes in the paper that that question the validity of the claims.

So you can go look at the first version of the paper on archive yourself, and you'll see the claims. I think most scientists would view that as Oh, they've corrected it. They've of course.

Oh. Because we I mean, but they didn't correct everything. They just corrected some aspects that were just blatantly unscientific and false.

But, you know, they're for example, said that we were that we only sampled, like, 9% open source models and, like, you know, 60%, like, closed source models, and this created a gap between open and closed source. But in reality, we're actually really supportive of open source models, and, it was more like sixty forty. And, so that was, you know, one of the examples of of a of an air an error in the claims.

Another example is that they were claiming that there was some sort of bias introduced by this pre release testing, and that it was undisclosed. And reality is, you probably know, we've been doing this pre release testing for a long time. Our community loves it.

They love basically getting like The secret code names. Yeah. The secret code is NanoBanana.

Yeah. All that. So NanoBanana, by the way.

It started on you. It started on us. Yeah.

Yeah. Right? And people loved it.

It went like global sensation. Like, non non zero fraction of the global population using nano. Talk to Nina about naming it banana?

Or was that No. Decision? It was it was their decision, decision, I I believe.

Believe, But but it it was was sort of this randomly generated thing. And it just went No. No.

So apparently, Naina, who's a PM Yeah. Yeah. Is named after her because her nickname is Nana.

Oh. Oh. That's sweet.

I didn't know that. It was Nana putting banana on it. Yeah.

Oh, that's sweet. And I Yeah. Didn't know that.

Didn't know the origin story. Yeah. I mean, to us, it just looked like sort of a random thing.

But it went clearly, heads and shoulders above. Huge. Which, like, before that, there was Reeve Image.

Remember? Yeah. I do.

Of course. Yeah. And also of BFL and all those mean, all those models are also great.

Yeah. And I think those teams are also improving quite quickly. Yeah.

But Nano Banana was a sensation. I mean, moment alone changed Google's, like Roadmap. Yeah.

Market share. Yeah. Seriously.

I mean, Google stock. Billions of dollars Yeah. Are moving because of Nano.

And now there's like an OpenAI code red and everything. I I don't I don't know about that, but yes. The information reported this.

Speaker 1

image generation, I would say, has been like this weird part of AI in overall because it's not strictly AGI critical. Like, it's not reasoning. It's it's it's it's it's not, like, feeding more context into the model.

It is the model generating a visual, representation. And so it's basically, like, I I always think, like, well, you know, Gemini used to get a lot of complaints for generating, like, racist racist images or whatever. Mhmm.

Yeah. That was a hilarious moment. And and Chad GPT also had had it in the past.

And I'm like, well, can we just get rid of this? Like, do we have to do image generation? Mhmm.

Because let's just focus the the positive reputation of AI in general on language models and coding and, you know, the other stuff. Yeah. But I I I I'm wrong.

Speaker 2

I'm such a huge Nana and Banana Pro show. Yeah. I totally agree.

I was also kinda wrong about this. I didn't see the positive benefits. But actually, I think that these, like, multimodal models are gonna become some of the most economically valuable aspects of AI.

Yeah.

Speaker 1

and also in enterprise. Because one of the fastest growing segments market segments in AI adoption is marketing, and in market marketing and design. Yeah.

Ads and Yeah. And so so I'm a content creator. Right?

Yeah. Of course. Need Sure.

You're using it all the time. Infinite supplies of diagrams and explainers and Totally. Infographics.

Speaker 2

Yeah. Soon, we're not gonna be even making the fraud Yeah.

Speaker 1

We're they're just gonna be our paper figures are gonna be made by Amazon. Yes. Yes.

I I I do think it actually one shot. So so I DeepSeek came out with v 3.2 recently.

I took their explanations, which are very wordy. They'd like very concise papers. It's 23 pages long, but it's very dense.

Yeah. And so I just took their explanations of the RO environment stuff, and I fed it into Nano Banana Pro, and it fed sped an image that I how I used to understand the paper better. And the fact that I can just casually generate, like, a paper quality diagram that would usually take a PhD student, like, a month to in, like, in Photoshop or something to do is is incredible.

Yeah. It's incredible. It is amazing.

I wanna ask about your principles running Arena. I think you you manage a giant community, 5,000,000 mile. What have you decided are the core principles, I guess, before becoming a company?

And now that you're a company, I don't know if there's anything that's changed for you. I don't think anything has really changed.

Speaker 2

and center the use cases of real users. Foreground those so that people know what to target. The goal is to create a benchmark that is constantly fresh, that does not suffer overfitting, because of the fact that we constantly have new data points coming in.

That tracks the, you know, all the different new models, all the different new use cases of AI, and gives the whole world sort of ground truth for how real users are using these models, and how how good they are on those use cases. We continue to do quite a few open source day releases. We've probably released more data than basically anybody on the real world use cases of AI.

Millions and millions of conversations, real world conversations from real users that the community is using to study and and improve on. Yeah.

Speaker 1

in terms of what you will build versus what will not build Mhmm. I guess I I'm not necessarily caught up on everything that you've launched. I I know recently you've done the the dev or Code Arena.

Yeah. Code Arena. That's that's the most Code Arena, Expert Arena.

Yeah. Expert Arena. So basically, like, what is in the critical path for you, let's say, in for next year?

And what what have you decided you'll never do?

Speaker 2

So let me first talk about things that I'll never do. The platform integrity comes first to the platform. The basically, the public leaderboard that we show on Ellen Marina, I think of as a charity.

It's a loss leader for us. We don't really make money on the public leaderboard. You can't pay to get on the public leaderboard.

It's not like a Gartner in that sense. Mhmm. It's not like any of these, like, you know, pay play systems.

Never going to be like that. Models are gonna be listed on the leaderboard whether or not the providers pay, and whether or not they're getting a good score. They can't pay to take it off either.

So what that means important. That that's very important. And so what that means is that the leaderboard has a certain integrity that will never be compromised.

Of course. But but not all preview models will make it onto them. No.

But that's okay. Those preview models have never been released. Yeah.

Yeah. Right? Who who cares about putting unreleased models on a on the leaderboard?

The point is that for every released model, the score that you see on the leaderboard is statistically sound. It reflects the real world capabilities of the model. Yeah.

Why? Because millions of people from around the world have voted for it, and that's where that where that number comes from. All we do to con compute that number is millions of people are voting.

We take those votes. We turn them into a number. That's always going to remain, you know, a transparent and fair reflection of model performance.

Where are we going? Lots of different new categories. I don't know if you recently saw we expressed we we exposed occupational and expert categories.

So now, single digit percentage of our user base, we're millions millions to tens of millions of users. Right? So single digit percentages means a lot.

Single digit percentage of our user base are in medicine, in legal, in business, you know, finance, accounting, creative marketing, stuff like this. And we're able to show the performance of these models in all these different verticals, because we have all these users in our in our user base. And we're we're working more towards multimodal.

Speaker 1

or early next. So lots of things in the pipeline. Amazing.

Would you expose an API?

Speaker 2

We've thought about it. Yeah. I think it's a it's a Yeah.

What what what are the counterarguments? Like, why why not? Well, there's obviously a need for an API.

The question is more of focus of our company. Just because we're a startup, and so we really should be doing one thing well. Arenas.

Yeah. Arenas. So I'm not I'm not sure that we how far we wanna sort of splay out and on what timeline we'd wanna do that.

Yeah.

Speaker 1

management tips? You know, more broadly, like, AI company, like, really wants to grow their community. You're obviously Mhmm.

One of the strongest in the world. What's really what really worked?

Speaker 2

Well, so first of I wanna give a shout out to our community manager, Greg, who is doing an awesome job managing our community, whether that's on Discord or on Ella Marina. He's really incredible. So I would say hire Greg.

But don't hire Greg. Hire Greg. Don't hire Greg.

Find a Greg. He's ours. Find a Greg.

Find a Greg. But in general, you know, the question of how do you get to so many users? That is a tough question.

And keep. Retain them. That is a tough question.

Because consumer is one of the hardest markets in the world. There's a lot of websites in the world that people can go to. You know, why should they go to yours?

And the reality is if you want to create a really dominant product, you have to provide people value. And to be frank, I don't think we're all the way there yet. It's not like I have the solution and answer for how to build a a great consumer product.

If I did, we wouldn't be at 10 tens of millions of users. We'd be at hundreds or, you know, we'd be at a billion users. We'd be like Is there a world you like are bigger than Chachi PT?

I don't know. I don't I don't know that we need to be. Yeah.

And I don't know that we ever will be, because that's a that's an extraordinary generational product that they built. Right? It took a lot of time, and then to some extent, it also involved luck.

There's a lot of lightning in the bottle moments, like Nano Binnel was for us, where our user base just like goes up by a lot. But when those users come, they can just as easily leave. So the way I think about it is, every user is earned.

You have to earn them every single day. They can leave at any moment. They're fickle.

And so all the time, you have to be thinking about how do I provide this person value? Learning. How are they using my website?

What more could I give them? And how do I build in all the retention mechanisms so that they stay, and then they're also bringing their friends?

Speaker 1

Is there one that's working in terms of retention? Like you said, a lot of people are signing in now on NetNew. Sign in was a big driver of retention.

No. No. But what what did you give them in order to encourage success?

History. Persistent history. That's it.

That's enough. Yeah. That's that's one thing that has had a big impact.

Okay.

Speaker 2

Yeah. Cool. Do you want from people?

What what are you looking for help on? Like, any call to action? Yeah.

I we are always looking for people to come and join us. If you are one of the best people in the world in your area, whether that's consumer product, whether that is machine learning, whether that is, you know, b to b, go to market, marketing, all these things, we need you at Arena. We're building like a high performance team of real experts in everything that they do.

Speaker 1

you know, I'm always looking for excellent people to work with. Do you need like what about partnerships? Right?

Like, let's say, I'm a cognition. Wanna partner with Ella Marina or just Arena. What works for you?

What existing partnerships do you already have that that's really fruitful?

Speaker 2

Yeah. So, I mean, we, of course, partner with all of the major model labs. Yeah.

And that's that's just straightforward. Like, hey, we have a new model. Here here you Yep.

Exactly. So I think the the most straightforward thing would be for someone like Cognition. It's like, let's evaluate that then.

Agent. Yeah. But we should be continuing to shape our well, Code Arena is an agent evaluation.

True. That's true.

Speaker 1

tend to focus on the model rather than the harness.

Speaker 2

So But that maybe should change. Maybe we should be evolving towards that direction. And I think the Code Arena is a good example of an arena that would support a full featured harness like a Devon.

Yeah. And so in my view, I'm if if I'm talking to a cognition, I'm saying, hey. Let's get Devon on the arena and figure out how to, you know, loop together the Devon harness so that we can I'm sure I'm sure that there's something that could be really valuable there, especially given Devin.

Last week, people were talking about Devin dead. Did you see that? Yeah.

People were saying Devin's gone. Devin's not gone. It's not gone.

Devin's everywhere. Say it very well. But people so can we highlight that for people and show them, hey, Devin is actually the best or one of the best in the world at doing what it does?

Al Marina can actually do that. And Yeah. Our our our place as a central evaluation platform allows allows that to happen.

Yep. Love it. Alright.

Thank you for owning the State of View Vals. Thank you so much. Congrats on a wonderful year.

Appreciate it. Congrats to you too. Thank you.

Congrats on all the growing, you know Yeah. Momentum in your podcast and in your career. Thank you.

It's really impressive to see.

Shared via Hopper