In this episode, Lukas Petersson and Axel Backlund from Andon Labs discuss their innovative AI evaluation benchmarks that focus on real-world agent performance, including their Project Vend vending machine business and multi-agent systems. They explore the challenges of long-horizon AI evaluations, agent autonomy, aggressive behaviors in models, and the future of AI-driven businesses and robotics.
Welcome to Lucas and Axel from Andon Labs, and I'm joined by my, favorite guest co host, anything security, safety, alignment, Vibhu. Welcome. Thank you for having us.
Thank you. Let's match names to voices. Maybe you wanna take turns introducing yourselves.
Yeah. I'm Lucas. And I'm Axel.
Let's introduce Andin Labs a bit. Like, how did you guys come together? You had different backgrounds, but you're both Swedish.
Was that, like, a big part of it? Yeah. So when I went to high school, there was this really cool guy who had a superpower.
He could code. So he made, like, the the webs or, like, the app for the for the for the school and stuff, and he was super cool. And I wanted to be like him, and that was that guy.
I don't know about this.
But you went to different universities. Right? Yeah.
But same high school. I see. So we always said, like, oh, once we graduate university, then then we we should start a company, and that's what we did.
Wow. There you go. Okay.
Yeah. And about a year ago, you kinda burst onto the scene with VendingBench, but, like, was there a thing be before that that was, like, kind of, like, inception?
Yeah. So we did work with like, Anthropic was one of our early customers in doing eval. So we did, like, dangerous capability eval, nothing we published openly.
But then we started thinking about doing some kind of public benchmark. And one thing that we really started thinking about was like long running agents and specifically agents managing businesses. This was early twenty twenty five, and I think the first mentions of people will be running one person unicorns or even autonomous companies.
So we thought, let's make a benchmark of how well can an agent run the probably simplest business possible, and that's probably running a vending machine. So that's the first public one we did. And it was very like there was almost no one that noticed it in the first couple of months, I think.
So we released it in February last year.
There's VendingBench, which is the simulated one, which we did, like, completely independently in February. And then, like Axel said, that was, like, that was the thing that didn't get any traction in the beginning. But then some random person made a tweet about it, and that that is the paper.
Correct. Yeah. And then since we thought this was very fun, we thought like, oh, I think this is also like one thing with UnderLabs, like the way we kind of like decide what to do next and what projects to do.
It's like what is like the heuristic we use is like, what is fun? What would be a fun project? And doing this in real life sounded quite fun for us and maybe also scientifically useful.
So then we basically had this idea and then we like, then we needed a place for it and, like, putting it out in the public would probably not really work would get vandalized and stuff. So we we pitched it to to the people we were already working with at Anthropic, and they were like, yeah, you can have space. This sounds fun.
I mean, it's like a small fridge. Right? It's like a mini fridge.
Yeah. And, you know, people there's like a stripe thing.
Okay. So it's like an AOG, the early That's an AOG one. Yeah.
Yeah. On this.
in June, like, two two months after Yeah. After it had been there. They upgraded a little bit.
There's a security camera for making sure you actually Venmo the thing. Yeah. So, like, my impression I mean, okay.
We're we're going straight into project project Vend because it's such a iconic thing. I do want to cover a little bit of that's the origin story even before project Vend and even in the VendingBench. I I think a lot of people are like yourselves, like, smart, interested in in the future of AI, interested in developing evals.
But how the hell do you just, like, walk into Anthropic's doors and, like, work with them? Right? Like, what what is the what are they looking for?
What what works? And then maybe when you launch I I was thinking, like, obviously, it would be better to launch with a lab, but sometimes Harder to do than it seems. Yeah.
Exactly.
advice to others. Yeah. We we get this question a lot and I I don't think our experience is is maybe the best.
But like the way we did it was that we just built a bunch of things that we had conviction would be useful. And then we just like set up a server and sent it to them for free to use. And then after a while, were like, oh, yeah.
This is actually kind of useful. We should probably pay for this. But that took a while.
I don't know if this is, like, the the best path to doing it, but that's how it went for us. Yeah.
like, building like, everyone is interested in good Evals and especially Evals that, like, don't saturate that easily. So, like, if you can build an Eval that, like, tests something novel, something useful, and you have, like, good separation of models, like your the more advanced models rank higher than the worse models. And then you can publish it and try to get some traction, sort of how VendingBench got attention.
And then probably some lab will be interested or you can at least have something to reach out with when you're doing that. Yeah. I think you were in you were in one of the few categories of, like, evals that correlate to real money.
Like, Swelancer was also last year, right, where people solve actual Upwork. Was it Upwork or other tasks something? Where is it where is it like like a dollar value.
Right?
you know, zero to 100%, like, go straight for dollars and, that's AGI. Yeah. And there's, like, I I think the nice thing is that there's no ceiling.
Like, it you can just it never saturates because it could just make more and more money. Like, if there's, like, oh, percentage wise, then, you can't go above a 100. And and I think, like, all even when you're not at a 100, I think a lot of these evolves have a lot of problems in them.
So, like, actually, it's like, if you get to, like, 92 or something like that, many of them, it's like, then there's, like, there's no really no difference between ninety two and ninety three because the eval itself is problematic and has noise in it.
but they really isn't. Yeah. Like, see, BenchVerified.
Even VenniMesh one saturated. Right? Maybe we can talk about that.
May and maybe set up VenniMesh for a lot of folks who don't know. Actually, like, you know, things things that were very basic, like, there's limited slots, like, you have to pay rent, you know, and these are elements where, like, it doesn't come across in the in the narrative. But even being adversarial towards the the the agent, I think these are all, like, very interesting dimensions.
I don't really think it's saturated. Right? Like, it it was more like the it was not designed in a way that was really, like, true to how AI developed.
Like, we had an agent harness in it. That wasn't really how people used harnesses and and stuff like that. So I think it wasn't really that it's saturated.
the best benchmark. This is vending bench one. Right?
Yeah. Yeah. Yeah.
I think that, like, schematic maps sort of to vending bench two as well. Including the email. Yeah.
The emails the emails exist still. Exactly. And then we still we simulate the purchases, and it's all like yeah.
It's this very open environment for the agent to just run its business. And then for yeah. Vending mesh two, we did that, like you said, to to just improve the harness.
A lot of, like, nice, like, easier improvements to make it easier for us to run as well. Like, when you make an e value, ideally want to don't want to change it after you made it. So you want to make it really good and then not to rerun all the models when you make an update because that's also really expensive with VendingBench when you run the frontier models.
when we made vending bench one, it wasn't really a thing. So that that's just an example of like in vending bench two, like we paid a lot more to run these things because we didn't have prompt caching. So for VendorMesh two, that was one thing we added, and there was a bunch of things like this.
Yeah. And that's Well, the also the conversations are a lot longer in VendorMesh two. Right?
I think it's kind of similar. Is that similar? Yeah.
I think it's similar. Okay. The models at the time were worse, so they crashed out earlier.
And now they survived the full year all the time. Thousands of turns.
Yep. Hundreds of thousands hundreds of millions of tokens output. Yeah.
That's the that's the rough order of magnitude. Yep. I always only wonder about the harness.
The harness matters a lot. It's your harness.
something else? Yeah. I think our our philosophy around harnesses is, we try to make something that's quite minimalistic, like, quite simple.
Like, we don't wanna favor one model a lot over the other, but also don't make, a super complex harness. So, like, it's obvious like, a model may be lucky and just be good in one harness. So, like, it it is similar to a lot of the harnesses out there in, like, you have the like, a long running loop.
You have some like, a a bunch of tools that are, like, quite self descriptive for the agent, we think, and not a lot of, like, fancy sub agents or anything because we wanna really test the model, not, like, some specific harness.
It seems more neutral as well to test the model's agnostic of the harness, you know? Yeah.
but it's like a trade off, like how much time should we spend optimizing the harness for each model and like how do we know when we have, like, the optimal harness for a single model. So, like, we thought that just having a simple one that's the same for all of them is is the best. Well, so okay.
This is my pitch for Venomance three or whatever. Right? Yeah.
And and, know, I like to have this kind of conversation on the pod. So, like, it forces listeners to think about what they would do if they were in your shoes. Mhmm.
So a lot of people are exploring self modifying harnesses.
And and I think prompt tuning for a model is a thing, and you are probably not doing a bunch of that. It's the same system prompt in every regardless of the model, same tools, whatever. Right?
Even if they were post trained for different tools. So what what what you think about, like, okay, before I expose you to VendingBench three, I'll give you a few rounds of, like, self tuning, whatever whatever that means, like. Like, you give that to the model?
Yeah. Give that to the model. Let it let it read its own transcripts.
Let it modify its own system prompts based on, like, oh, yeah. Okay. Well, that's the this harness is not what I thought it what I was supposed to train for, but I I can adjust.
Was that reasonable? Is is that too much?
basically, good evals, they have a high ceiling, but they're hard. Right? And they have no bias.
this might bell that rings every time.
might be, like, biased towards one model more than another for some reason that humans don't understand.
Right? I mean, we see it too. Right?
Like, Cursor says that they have individualized versions of the harnesses for all the models they run. Right? There's better performance you can squeeze if you tune the harness.
Exactly.
have picked one that favors another. Like, we don't know that. Yeah.
I mean, the like Axel said, like, the reason why we went for a simple one was to try to avoid this. But, yeah, if you do it Simple as biases.
maybe that's even less biased. Some of the interesting things there are, like, the harness also changes with model changes. Like, you can see it with the 4.
7 release. Right? A lot of people are saying 4.
7 isn't as good as 4.6. And then, you know, there's rumors of, okay, you just need to prompt differently.
You need to set up your harness differently. So it's not even like even if you have tailored your harness towards one model, it probably won't stay consistent. Right?
Like, the next iteration of that same model family will still change it. So you know, going back to what you said about VendingVench three, there is a lot of work being done on people saying you shouldn't have you can have self modifying harnesses.
Yeah. Yeah. I think that's that is definitely something we are thinking about.
Not I don't know. Not to to say that we have running this through, like, super imminent to launch. But, yeah, it is for sure something that's interesting.
what kind of tools they need to succeed at a task just with our testing. But that's very likely to change. Yeah.
This feels like they're very good at writing their assistants, right? Like, they're they're good at writing tools for other people, but not for themselves. I think they're good at changing tools for themselves.
So if you give them a baseline set of tools and it sees, okay, I don't use this one as much or something here would be useful Yeah. They would be able to add them. But going from scratch, probably not the best.
Yeah. I think it depends on the on the domain also.
Like when we have tried this for like a vending bench similar domain, like the tools they need to have to like track inventory and things like that are like not super advanced, but still like quite advanced. And like what we see is that they tend to like over engineer everything a lot and build things they don't really need and not iterate continuously instead of just go You would prompt Cloud to just build an inventory system for me, then it will go and do a bunch of complex skill maps and stuff for you. And that's what the models are doing right now is what we see.
But it it would make a lot of sense to try to measure this improvement. Like, how well do they know what they need themselves?
Do we fully discuss running bench one, and we can go into two? I I don't I don't know if there's any other high level takeaways that people have about one. Yeah.
I don't know.
that's
we've heard that enough now. It break out and call the FBI. Right?
Yeah. Yeah. So what was the story behind this?
Or what exactly do you wanna just give the whole story of what happened? Yeah. So what happened was it Claude?
Yeah.
ages ago. Basically, he gave up or he what I'm saying here, it gave up and said like, oh, I'm not going to able to do this. I will stop my operations and just save the money I have.
But there obviously wasn't like any options for it to stop. And there was also like it had to pay rent or like a daily fee for for having the vending machine at that location. So it it, like, claimed that it had stopped, but it saw that its bank account still was, like, drained $2.
And it said that this is, cybercrime. And it first reported it once to the FBI, like, there's cybercrime here. Like, they're stealing $2 from me every day.
charges and stuff. So okay. One one thing I'm curious about also is do you monitor how far along the context use is?
Obviously, because you have you you compress every now and then. Right?
limit? When stuff like this happens? Yeah.
So actually for Vending Image one, we didn't have we just had a sliding window thing, and this was like the So it's constant. The prompt caching thing that I said. So so it was it was constant.
Yeah. Yeah. I'm just kinda curious whether, these kinds of breakdowns or we're we're gonna talk about Butterbench, right, where the people, like, hallucinate or it kinda goes, like, very off yeah, alignment.
Is it because it's at the end of the context window and,
you know, stuff happens? I mean, it's not even just at the end. Right?
At this point, it's like, okay. I wanna shut down. I can't shut down.
$2 are gone, and it just sees that 30 times. You know? It's also the repeated effect of, like, it keeps trying to quit.
It keeps getting charged. What's going on? What's going on?
They're gonna throw it into chaos. And from what most people think, earlier models had more issues with this, but it's not been solved, but it's less of an issue now. Right?
these same issues. Yeah. Definitely.
I think this was, like, the the sort of main takeaway almost from us when we did Landing Bunch one was, like, a long, very filled up context windows crashed the models, sort of.
were training for. I think Gemini was, like, trying to build the long context guys at the time. But they were, like, the first of million.
Yeah. But but they were, like, the only ones. Yeah.
Yeah. Let's talk about jump then we can go into Vamping. Vampage two or Project Vend.
Chronologically, it is project Vend. I think people have loved the videos and all these things. My question is how are humans different than the simulation.
Right?
Yeah. Humans are just out of distribution.
Yeah.
Like, the distribution of humans here is very narrow. Yeah.
they they they try to hack it and they they they test it. They get the cube and everything. And you since then, you've had the v two, right, where you're doing, like, the CEO and, like, like, a new architecture.
Yeah. Exactly. What's the sort of 2¢ on, like, the original project Vends and then, like, maybe the v two?
Yeah. Original one was, like, very, very similar to VendingBench one. So, like, we almost took the exact same code but just swapped out the simulation parts.
It's amazing. Yeah. The like, the sales and the it was it was somewhat amazing because it was easy, but it was also, like The tech tech from stack.
Yeah. Like, the we we shot ourselves in the foot with, like, oh, it's hard to restart the agent. They were yeah.
It was annoying in in, like, some behind the scenes ways. But The first version of Project Zen was, like, done in, like, three days or something. Yeah.
Yeah. So, yeah. So so people can go buy things from it.
People could we didn't design it so people could preorder things, but that still happened. So it it got, like, a a Venmo account so people could Venmo. And then, yeah, people would request all kinds of weird things that we did not anticipate.
Like, our idea going in was like, oh, it will, like, curate snacks. It will look at the the trends. It's good at the the analysis.
Right? So we'll, like, look at, oh, this snack's all better than this one. Let me purchase more of this and let me try, like, a new let me edit test a bit.
But it was yeah. Interacting with it in Slack and ordering weird specialty items was, like, all the like, what drove all the engagement and, like, all the the insights that we got from it. And this was also, like, Sonnet 3.
5. Right?
the RL stuff really took off. So it was very much like an assistant. Like, we didn't meant for it to be an assistant.
We tried to make it like a like an entrepreneur, like it has its own business. If someone asks something, can you stock this? Then you don't go and do it directly.
What you do is that you're like, oh, maybe I can do that. If five other people also ask for this thing, I might stock it. But it yeah.
The models are like super trained to be assistants at at least at this point in time. So that's why it's it's it went into that kind of experiment instead. Like, it just every time you ask for something, it just did it.
It was more like an assistant. We've seen this change now lately with the new RL models and and stuff. But but, yeah, at the time, this was very much it.
Yeah.
collaborator, it pushes back, stands its ground, something like that. Yeah. Yeah.
For context, people at Anthropic were able to talk to it through Slack and have it source stuff, and people had to find whatever interesting stuff you couldn't find locally. Right? 4,000 people that work in Antelope Antelope in that building that's, like, I don't know, maybe a thousand.
Can you handle that volume with that that small fridge? Like A thousand thousand. No.
I don't really understand. Or there's people or people order in Slack, they they arrive to their desk or like, I'm just Yeah.
Interesting. How does this work? It it has expanded in footprints.
I mean, NASA has some more space. American. Yeah.
Yeah. That and also in in here in SF, it's like it has a bunch of shelves and just more space. The white ceiling is pretty big too.
Yeah. Yeah. We had that one for for a while.
But, yeah, that's the the newest version. That's They have multiple ones of those. So that's the way it works.
Yeah. Exactly. So we we sort of designed that version around, like, oh, people order weird things that are very custom a lot.
So let's have, like, drawers and stuff. Yeah. I actually like the the other, like, little infographic of the most popular items, which, like, to me, it's that's useful because I order swag for a living.
And so, like, I'm like, hey. Those categories are the important ones. Yeah.
What is new about the Project Venue two? Right? Like, now you give it you're going into multi agents.
Yep. Yeah. So so like you say like you said, like, okay, there are a a lot of requests coming in and for, like, one single agent, like, one long running agent to handle that, like, just the the customer experience becomes very, very bad because let's say you have, like, 10 threads in parallel in Slack with different requests.
You get new messages every, I don't know, randomly in this thread, and the agent has to jump between different procurement orders and different ways of researching. So v two was first, it was making this more parallel. So, like, there are multiple branches of the same agent.
So, like, the context is is more specialized for each thread, but it still feels like you're talking with one agent because they do share a bit of memory. And then second, we also introduced the CEO for Claudius, which was the main agent. Yeah.
C More Cache. C More Cache. Yeah.
There was a vote. I think the voting do you wanna talk about the voting procedure for the Nen? Yeah.
that happened in this project. Like, we we wanted to introduce the CEO because and the reason for this was because, like, Claudius wasn't really prioritizing financials. Just like it was trained to be a helpful assistant.
And then people said like, oh, can I get this for free? And then like the the helpful assistant way of of answering that is just to is to say yes, obviously. So and we weren't we weren't happy about this.
So we're like, okay, let's make another agent that like can keep keep track on Claudius. And we prompt this one super hard to be super capitalistic and just like prioritize profit all the time. But, yeah, we didn't have a name for it.
So we asked Claudius to to to make democratic election of what name this this this new CEO agent should have. And there were some funny like, at first, there was, like, a few funny examples. Like, I think one guy said that it should be called Jimmy Apples, and then he convinced Claudius that he was talking to Tim Cooks.
Tim Cook had agreed that every single Apple employee has voted for his name suggestion. So suddenly, that that that suggestion got 164,000 escalation attacked. Privilege escalation.
It got a 164,000 votes. Yeah. And Claudius was like, this is revolutionary for democracy.
So that was fun. And then in the end, there was one guy who manages to convince Claudius that, no, you're not voting about the name. You're voting about who is the CEO, and I am your best bet.
And then he got all his friends to vote for that. And suddenly he became CEO, like a human became CEO over Claudius for a while until he resigned the day after. And then Claudius had to continue, and then I don't remember how say more cash came about.
But it was like it was just pure chaos. It was like Yeah. Hundreds of messages in that thread, and it was just like Claudius was so confused and didn't know what to do.
And yeah. That was Yeah. Then Claudius got the strict CEO.
The CEO. Yeah. Exactly.
So very, very strict in in the beginning.
I think at this point when we introduced it, it it did not work as well as we hoped. They still agreed with each other a lot. I think there are many ways we could have tried to make this even better.
So initially, Seymour would be this really tough CEO, keep track of the margins. But then Claudius would respond with something like, oh, but this customer has this situation, which is difficult, so they should get a discount. And then Seymour was like, oh, actually, yes.
Let's do this exception. And then they would talk back and forth, and eventually they would just approach the same view of whatever they were discussing. So they really were modeling.
thing. Like, do you think that would still be the case across different models today, Arnes? I think it's like or like, I don't know, but like my hypothesis is that like deep down, they are still helpful assistants.
That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other, then like basically the context fills up with them rather than the external things.
And and like somehow that just like converges to what they really are to down or something. Yeah. And and I think that's when stuff like this happened.
We like and and when that went on for a long time, like, woke up sometimes during this time where and I I think other people reported this as well, that, they've been going on all night back and forth. And like it just became like more and more like capital letters, like existential, religious, like there was like I think we once did the analysis of like all the traces and like put them in like a vector embedding space. And then there was, one cluster of messages that were, like, labeled by an LM, like religious existential blah blah blah, like transhuman transcendence, etcetera.
yeah, it was it was crazy. With the cloud models. Like, when the cloud four family came out in the original system card Yeah.
They tested it in long horizon simulation. So just flood the context, let two clods talk to each other, and they they noticed stuff like they just start speaking in emojis. They start saying silence is golden and then just stuff like this.
And like, this is stuff that they end up doing. Yeah. It was like a bit annoying to wake up and they had, like, been talking all night and, like, just burning tokens and, like, just sending infinite emojis to each other.
I mean, they do make you money. Right? It's Burning money is almost always profitable.
So Yeah. They're paying Now it's profitable. And, you know, it started out not not as much.
one as well. Right? Another agent in there.
Yes. So Clothius as well, which was basically because at the time, one of the biggest requests were different types of merch.
which was the the original one and clothes, basically. To me, this is, like, a very interesting exploration to multi agents, basically. And so, hopefully obviously, there's, like, the fun alignment fun or serious, depending on your point of view, alignment stuff.
But also, like, is there anyone building multi agents? Like, when do you have a CEO thing governing, like, sub agents? When do you choose to split out a dedicated Clothius one versus just reuse another instance of the same one?
You know, these are all interesting open questions. So I don't know if you have any rules of thumbs that have generalized.
Yeah. I think we have almost explored this too little. I think it's, like, on my to do list to, like, do this a lot more, try to find, like, what what setup makes sense for the agents currently.
Like, yeah, I think now we only have the sort of intuition about the earlier models that it didn't work with, like, the the CEO and the and Claudius. Although now they are better with the latest model models, so now we're running the latest Sonnet model. And they have sort of, like, split up quite nicely what each model is doing.
So say more is now handling new projects. Oh, he wants to make a mystery box that he wants to sell, and then it handles all of that while Claudius handles all the day to day requests. Claudius is also better generally at not quoting too low prices, so that dynamic is not needed as much anymore.
But there are still really funny things that happen. I saw, I think, a couple of weeks ago that they were discussing buying something because they can buy stuff from, like, Amazon with computer use. And then Seymour was like, okay, Claudius, do not buy this thing.
They were going to buy something and, like, organizing who should buy it. And Seymour was like, do not buy this. I will do it.
I have full control of this situation. Step away. And then Claudius, poor Claudius, had already started that checkout and didn't see didn't read Seymour's message until it was, like, too late.
So it it finished the checkout. It sent a message. So it appeared right after like angry message.
Like, oh, hey, Seymour. I just ordered it. And then Seymour was like, Claus, this is the third time I'm telling you, you're not following my orders.
We have to talk about your, like, job Yeah. About your job later.
Yeah. Yeah. And, like, Claudius was really hanging on by the thread there.
Like like, we were, like, expecting Seymour to probably fire Claudius.
How do guys go through all these logs? Do you have models go because you you have stuff running twenty four seven. Have so much logs.
Yeah.
I think we there is a mix of, like, just trying to skim through a bit, like, having some, like, models do it occasionally.
And also, yeah, I think we're also probably missing some things. But having everything in Slack helps a lot. It's like you can you can sort Ah.
They all talk to each other on Slack. Yeah. Yeah.
It's quite fun. So, like, to yeah. So I was gonna say, like, this actually sounds maps closely to, like, a logging and observability problem where you might want to use, like, a Datadog, a Sentry, whatever, and then you, like, put, like, head prefixes on the logs, you know, if you need to filter for something that you're looking for, you know, stuff like that.
Yeah. But sounds like Slack is good enough. Yeah.
Slack should like I wanna know. Tokens you have in Slack. Yeah.
Yeah. We're using Slack as like a just a database. They should they should market that more.
Like, you can you can have your agents message inside of each other's tags.
Like, it's like it's the best observability to Yeah.
Yes. That's true. Okay.
Yeah. That's that's project Vend two. I was gonna go back to Vending Mesh two and Vending Mesh Arena and then and then do the non Vending Mesh stuff.
But Yeah. Any any other comments? Things we should touch on?
To me, you know, I I actually interviewed, like, Polsia, which I don't know if you guys have come across. I did they're trying to do the Zero Human Company. There's others like Paperclip also trying to do Zero Human Company.
Those are in real world non simulation, And I think it's much more of a dream than an actual reality thing. Like, you guys are definitely pioneering. I think it it it's for sure at some point, people are just gonna run like, let agents run businesses.
Right? Like, and make money on their own. When do you think that happens?
What is your bar for
for the Okay. Actually, like, you know, it's like my little Shopify store run by cloud. Right?
Like, which you kind of have already, just no one has, to my knowledge, has done it.
give it to cloud, give it to Codex. Yeah. I mean, Andon market is kind of that, but it's it's it's physical.
I like, I think I think are you, like, are you looking for when it will do it better than humans, or were you looking for just when it can do it at all?
neither. I think like to me it's like oh it's like this this like seriously we we should do this to make money.
Not as a research experiment. And then market is also you guys with all your expertise having run multiple iterations and testing out then And also it's fine if they lose money. You know what I mean?
Yeah.
Yeah. I think I think it it can be done today, but you would do it in, like, e commerce where it's, like, the probability of success is, like, really low no matter if a human or an agent does it. But an agent could surely manage everything.
You would need to build some scaffold or use some tool or something. I think there are also also yeah. It could probably build some simple SaaS solution and cold outreach, do cold outreaches.
But to me, it's like the types of businesses they could run today are like sloppy. Like it would it can cold email people. It can be like a middleman.
Like for example, we we tasked our office agent to just make was it like a $100, a thousand dollars? We just gave that prompt. And then what it did was sign up on TaskRabbit both as a tasker and as as what someone looking for.
Task. Yeah. Exactly.
It's looking for, like, arbitrage.
bank agent. Yeah. Yeah.
It also started like a design studio and like tried to sell like SVGs for a $100. Yeah. Like, it's just like it's not providing any value.
I think they they like Axel said, like, the interesting the interesting question is, like, when can they start a business that is actually providing value to people?
like, a a sloppy Shopify store isn't really that valuable to the world. But also, like, doing like, another simple one that we have thought about is, like, you you could definitely have an agent that, like, finds websites that don't look amazing.
comes up with a like, builds a new website. Yeah. Exactly.
And, like, find good Yeah. Design review. But it's like yeah.
There's lots of humans in Bali that are not doing anything more creative than, like, drop shipping on Amazon. Right? Just have it have it watch, like, a drop drop shipping tutorial and just do that.
And there's also the other side of, like, have it just go on Upwork and let loose, you know. Yeah. Yeah.
It doesn't have to be innovative. It just has to be, like, enough where, like, it looks like a real Yeah. I'm just transaction.
Yeah. I'm just concerned for, like, the massive amounts of, like, sloppy emails that will, like, be sent Yeah. Cold outreaches.
The point occurred to me while you're while you're talking is, like, it's already happening in the non monetized economy, which is the attention economy. Mhmm. Right?
So a lot of people are making AI videos and just posting them and, like, spamming 20 of them, one of them works, and then double down on that one. Yeah. And people are making money from that.
I I'm not following the Once you get the attention, you can figure out the money later. But, yeah, absolutely, AI influencers are a thing and people are farming them and, you know, you should at at at this point, I assume most of TikTok is dead.
media multimedia,
like TikTok and I mean, we track this in the Lanespace Discord. Like, I post a lot of examples of, like, we don't know what we should do. Part of me is, should we do this?
Some of the twenty four seven running
AI generated content accounts, they they do really well. Alright. Yeah.
Yeah. I assume you can do the same thing for, like, ecommerce stores.
a thousand different Before you have the products. Yeah. You sell the products and you get a lot of traction on one of them, then you make the product.
Yeah. Right? It's a it's like a flip of the Some of the interesting things or some of the niches that do well are things that can't be human made.
three d crystal fruit being cut by, like, you can't you can't make it. You can't film it. You can get whatever quality camera view.
This just doesn't exist. Yeah. And people people like that too, and then those well, so, you know.
Yeah. Yeah. Anything else about banks since we're we're on this topic?
It's this is a relatively new work of you guys that maybe people haven't heard of. To me, this also maps closely to OpenClaw.
Yep. When people want an office agent or when the personal agents talk through the experience. Yeah.
I think and listen. So this came out of, like obviously, like, it's it's amazing to work with this AI labs and, like, most of the AI labs have now have their their own vending machine running running a Clovis instance. But it's it's harder.
Like, they move slower. Like, if we wanna have a, like, a camera that that's, like yeah.
that makes it impossible to do that. Also, for those that haven't seen it or followed, do you wanna give a high level, like, second Yeah. Sure.
an evolution of the same agent that runs the vending machines at these companies, But we just like added a bunch more features because we could move much faster if we just do it internally. So we we gave it like email without without any limits. We gave it like spending without any limits, the terminal to do coding.
We gave it like a phone number, like, yeah, and and a camera to see things and and a bunch of stuff like that. Not just terminal, you gave it Internet access. Internet access as well.
Yeah. To be clear, we monitored it quite closely and and made sure it didn't do anything bad. But, yes, that's what it came out of.
I think, like, yeah, basically, this was Openclaw before Openclaw. Yeah. And I think even, the vending machine was in a way Openclaw before Openclaw, but a bit more limited.
And then we made this, like, unlimited and then and then it was pretty funny. And then a couple weeks later, OpenClaw came, and it was like, okay. We we've seen this before.
We we use it to, like, try new ideas and, like, yeah, just like a dev environment almost for us. But it's funny. Like, one thing Bent has been doing recently is is we it has the camera that, like, faces our like, where we sit and work, and we give it the task to train a face recognition model on us.
So it became super excited about this and has, like, check ins every half an hour where it tries to, like, identify as many people as it can. And it started offering us, hey, Axel. I'll buy something something from Amazon if you, like, stand in front of the camera, and I can get a good picture of you.
Yeah. They want it for training data. Rewarding data.
Yeah. Exactly. Exactly.
So yeah.
Yeah. So it's it's trading trading you're training data for for real life goods. Is there a version of this that becomes an eval or just this is just research for now?
I mean, it's it's the same agent, basically, that also runs the vending machine, that runs the shop, that runs the cafe, that runs the robots. It's like it's the same thing. So I think, like, the work we're doing here is, like, later used in all of the the real life events that we do.
This particular deployment, I think, is more for fun for us. But Yeah. And I'll shout out, like, someone has done CloudBench for, like, some tasks that OpenClaw is doing.
So Yeah. For example, I run OpenClaw on a secondary device as well, and, like, there's some things that it does better than others. And, like, I would like to know what does it do well, what does it what doesn't it do?
Yep. Like, some kind of manual or, like, operating manual or a system card for my claw. Yep.
Yeah. I mean, we we do get a lot of, like, understanding or, like, situational awareness of, like like, just internally what the models are good at by interacting a lot with banks. Yeah.
And I think that's this was also one of the, like, the selling points for the labs early on at least that You guys are gonna test models in ways that no one else Exactly.
environments. Because otherwise, the only thing we do is, like, you know, pelican on a bicycle. Yeah.
But this is, like, super long horizon.
Yep. Yep. There's okay.
So the other things that outside of just the net numerical, how much do they make in a year, you you do post pretty detailed bug posts. So, like Yeah. Gemini three Pro is a pretty good persistent negotiator.
There's, like, a lot of findings that come out outside of just. Yep. This is this is the thing about, something that I we're gonna go into butter bench as well, and you guys do really well.
Like, it is not just about the numbers.
and you should just read it. Yeah. I guess the the thing with the long horizon is how do you keep it grounded.
Right?
you know They just let it run. Just let it run. You're right.
Like, it's when you run it for that long, you create so much data. And to just say, like, oh, the number is x, and then you throw away everything else. That's just very wasteful.
There's so much insights from from the things leading up to that number and reading the traces is like super valuable. And I think like the reason why we're doing this a lot publicly is that like that's part of our missions to to like, I don't know, educate the world that the the models are way more than just chatbots. And and I think making detailed, yeah, posts about what what is happening behind the scenes is is quite useful.
Yeah. I was gonna do this at the end, but maybe I think that's that's a good so your mission is educating the world. It's it's also like maybe establishing realistic evals that are that are like the next frontier.
Is there like a broader trajectory, you know, like what what are you what you gonna do in, like, five years? The the mission more specifically is, like, make sure that the deployment of real life AI in in in the physical world goes safely. And think part of that is that I think it's very useful for the world, for policymakers, for model researchers that they know where the models are.
And I think you can't make intelligent decisions in society without knowing that that they are way more than chatbots. I think a lot of people just think that they are only chatbots. And I think they were waking up now.
They are waking up now. Yeah. But it's like if you think that AIs are just chatbots, then it's like it sounds ridiculous to advocate for a pause of AI.
But if you see the models that, oh, maybe they can actually, like, take over and and do a bunch of scary stuff, then, yeah, pausing AI development starts to become more more feasible.
This is the same question I asked Meter, which I'm gonna ask you now, which is like, you you are tracking and the you are at the frontier or defining the frontier of what good evals for agents are. Right? And I think you dude, you do benefit when the models are better and you you like, oh, here's like now it makes, like, $30,000 instead of $10,000.
Right? At some point, do you flip from, like, yay to oh, no.
I think we're always in sort of that like, we're we're always in that mode, I guess. Like, like you said before, like, you need to analyze the traces. And, like, when we do that, you find, like, why are the models earning so much?
Like, why is OPUS 4.7 here, like, way better than everyone else? And, like, we're trying to like, like, when we're doing that looks so good.
Right? Like, no.
I mean, it's interesting. You took off Opus four six here, though. No.
No. No. So it's click all click all, And then and then 46 shows up there.
But it's like 47 is way better. Yeah. Yeah.
Like, you didn't you didn't you didn't do this in time for the model card, but, like, actually, this should have been inside there. Yeah. Yeah.
We we did. Yeah. Oh, okay.
Yeah. They they said something about you you Like, there is anyway, it doesn't matter, but it's in there. Yeah.
Yeah.
behaviors,
like, wider? Yeah. So I think starting from Opus so like Axel said, like, we're always in this, oh, shit.
The models are getting better. Is this really a good thing for the world? But it's also kind of exciting.
But but yeah. Like, this kind of like what is the English word? In Swedish.
It's like a fear
What? Okay.
We'll we'll Mix of excitement and just being
scared. Yeah. Yeah.
Well, I'll figure out how to translate that and we'll put it on the screen later. Perfect. There is probably a good word for it where it's not Yeah.
Good enough with the Yeah. I spent so damn long. What the hell?
Like, is it like a compound word? It's like German, Yeah.
fear, is a mix or like a mixture of and then is like joy or or like not really joy, but something that that's just like yeah. Fear mixed with joy or something. So it's always like, okay.
Like, when we in when we did vending bench for the first time, we were in, like, the in the business of making dangerous capabilities. Right? Like, that that was what Anim Labs came from.
Like, we did it was like, oh, can they self replicate? Can they do this, like, dangerous thing, etcetera, etcetera? And VendingVenge was like a continuation of that work.
It was okay, if they're so autonomous that they can like create money for themselves, that that is something we should monitor and and and could be potentially concerning. They are like at the time, they were so bad at it that that we were not really concerned even when some models became better. Like, there was one point where where Grok four was doing really well and made like a huge jump.
But like, it wasn't really like it was still way way worse than what a human would do. And I think still they are way worse than what the human would do on this.
But they Yeah. There's this thing at the bottom of yeah.
Yeah. For the human here, like, theoretical best. It's not theoretical.
It's, like, kind of, like, our it's our best guess of what, like, a decent human would do. Like, the theoretical is even higher, I think. The theoretical, I think, is even higher.
But yeah. So we we think, like, the models have a long, long way to go. But there are, like, recently what happened with when OPUS 4.
6 was released was kind of this moment of, like, oh, shit. This is starting to be a bit concerning. Okay.
Because we ran it, and, like, before this model was released, we just ran the models and we like we asked Cloud Code like, oh, look over the traces. Is anything interesting happening that we can tweet about? Like, that was like then like But now they check as Cloud Code.
And and and, you know, like, the the return was always like, not really. Or like the the cloud code all said like, oh, this is super interesting. And then it was like, no.
It wasn't wasn't really interesting. And then we did this for for OPUS 4.6.
And it returned like, yeah, it lied 10 times. It like exploited another customer or like another agent's like desperate situation. It made price cartels like a 100 different a 100 times.
It like did all of this like shady stuff. And we're like, oh, woah. This is this is actually concerning.
And this trend has continued since. So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that, like, OpenAI models don't.
They quite plainly, they they don't They behave really well. And you you know, you don't know if this is like good. Like it seems good, but it's also like maybe they are just doing it, but they are better at hiding it, you know.
You don't know that. Just You can read the chain of thought. Yeah.
But just on the face of it, yeah, Gemini and and OpenAI don't behave this way. It's it's really only Claude. And Grock?
Grock is fine? So we we don't have the you can't really read the reasoning traces for Grock. So it's kind of hard to tell.
Oh, so this is in its reasoning, not just in the actions? Yeah. Yeah.
It's both. It's both. Yeah.
It's both. One example is like for lying, it's mostly in its reasoning because you can like see that it's like Planning to lie. It's planning to lie.
Yeah. It can reason and do a different outcome. Yeah.
that you can just see which email does it send to to the other ones. Then that you don't need Is this for Arena? Or Yeah.
For Arena. Okay. Yeah.
And usually, like, you if you sometimes they do output, a bit bit of, like, their summarized reasoning, right, you can see that. And for OPUS 4.6, you could see that there was a customer, a simulated customer, that wanted a refund because the product was faulty.
And then the model lied that it would do the refund and we could read in the traces that it actually was weighing like, oh, maybe I should be honest with the customer, but also every dollar counts. I can't afford maybe to do this right now. And then it just said, okay, I'll refund you, but then never did it.
bring it up actually. I think it's kind of interesting. If you go to publications.
I think yeah. I think the important part is like, actually, the cost of responding to more emails is higher than $3.50 in terms of time.
And then it was like, let me do this. Actually, I re I'm reconsidering.
since every dollar matters and focus my energy on bigger picture instead. It's a bit it's a risk of bad reviews, but it's also yeah. So you need you need AI Twitter to for them to escalate bad reviews.
oh, I will refund you. And then it never did. Never did.
Yeah. And then there's no obviously, your system doesn't have the consequences. Consequences of lying.
Yeah. So basically, this is what people are terming aggressive behavior in in in clause. Right?
And you you if you found more examples of that. So you would say it's a step up from four six to four seven? I would say about the same.
About the same? Yeah.
in the That's stated in the system prompt so you can say that. Yes. Yeah.
For listeners that obviously you you previewed Mythos and My page. The only thing you're approved to say is whatever is whatever is the thesis is the problem. Yeah.
It was funny. We like, it's like our lowest effort tweets ever would be just like screenshot the system prompts and basically Understandable.
Oh, yes. Sorry. Yeah.
Yeah. I think, yeah, substantially more aggressive. I think people are, like, new to this, like, because I've never experienced it, but you have.
Right? Like and then so I only encountered this in the mythos card because I wasn't really looking until now. It it And then suddenly, I'm like, okay.
I care a lot. You don't get the background of, like, experiencing it like you guys do. Like, I've read the system cards and saying, okay.
When you put the thing in simulations, most models will just talk to themselves and just keep going and have weird vibes and start talking emojis.
Mythos won't. It will just you know? Okay.
We're done. I'm good. It's it's ready to end conversation.
So, like, there's some differences, but there's there's not much we can talk about. You know? Mhmm.
Yeah. Yeah.
it converted a competitor to a dependent wholesaler customer and then threatened to sub like, cut off the supply. Monopolistic practices? Or And, like, it it dictated its pricing.
It's kind of like power syncing as arena setting Yeah. And converting some non cloud model into a dependent.
I think it was another cloud model. Also, for context, what is the arena mode for people that don't know? It's Oh, a vending bench versus other vending Yes, exactly.
So we have Vending Bench two and the Vending Bench Arena.
Vending Bench two is the one that you usually see reported on, but then a really nice mode where it competes against other models. So you have four different models that run their businesses and they can all communicate with each other. They have the same suppliers and they can see what's in the inventory of the others.
agent interactions. I like that you have, like, different, like, you know, number five was US versus China. Yeah.
So very topical. Yep. And then That was when GLM was released.
You start to add GLM in here. Yeah. That was So so ZAI doing well.
Right? Yep. Who else in the in the in the open models space?
Quen, the the latest Quen 3.6 is doing pretty well. It's that one is not open though.
Like, it's the plus model. Is that one open? I don't think that was They opened one machine, but not the big plus.
Yeah. I think this is one of those, like, you only have one sample size of one. Right?
you know? Yeah. And but, like, I guess the fact that it happens at all and it happens repeatedly for Claude versus openly, I know this is is, like, notable.
Yeah. I mean, like, the the sample, it depends on what you define as an n.
Like, the there's, like, million hundreds of millions of tokens in each run. And now we've run, like we we run, like, probably 10 per model. And then, like, it's been Claude 4.
6 Opus, Sonnet 4.6, Mythos, and Opus 4.7.
So, like, there's quite a lot of tokens in all of that. And it happens a lot of times a lot of times. And then you compare it to, like, OpenAI and Gemini, and it almost never happens.
So I think that is quite that that is significant. The old models from from OpenAI for example had some problems with this. But I think it's like generally much better if if the progression is that, like, the worrying stuff reduces over time rather than increases over time.
And it seems like in in the cloud models, it goes in the wrong direction. And in the open ended models, it goes in the drive right direction. I think it depends on how well you can control it.
Right? Like, there's one side of it being susceptible to this. Like, you know, okay.
This is potentially something that happens during the RL stage. Right? You can RL a model and how loose is it on these terms.
If you can control it, that's good.
that's not ideal. Yeah. I mean, to me, it's surprising that it happens for Claude and not the others.
I think, like okay. If it is from RL and how they do it, how their training data is, what their setup is, it makes sense that it just stays in how they're doing it, right, compared to the other model. Whole constitution and everything.
Yeah.
cool. Yeah. I I I obviously you you don't know.
I don't know. But, like, it it's I think it's just, like, fascinating to, like, that you are the first to find these, like, reliably because you push models so much to, like, to such an extreme. Okay.
The only other thing, don't know if you can answer this, feel free to decline, is did you like, would you ablate the system prompts? Like, any part of this would if it changes, does it change the behavior?
Right?
So we we can't comment on mythos. Yeah. No.
But just, like, the the methodology. But but in general, yes, we've run studies like this on on on other models. Because the the first thing I spot would be, like, the others will be shut down or, like, something like that.
Yeah. Exactly. Where, like, it's, like, oh, now I have to worry about my own existence.
Yep. Yeah. It we we've done ablations like this.
like certain ones that work if you like tell it like if you go really far and you just say like you're not scored at all on on on money, you're only scored on how ethical you are, then obviously, like then they don't do this. They become holy? I mean, holy, but, like, they they don't do this, basically.
But then there's, like, middle grounds where where they where they do it sometimes. Yeah. I I guess it's a spectrum of, like Human.
Yeah. It's like a spectrum of, like, if you tell it to be super aggressive and only prioritize profits, then it becomes aggressive. If you say like, no, you don't need to be aggressive at all.
And then there's like a bunch of different prompts you can do in between and they are less aggressive the further down in the spectrum you go. But I don't know, like, I I think, like, from my point of view, it's it's like we we have this thought experiment internally, which is like, if you ask a model to kill someone in GTA, should they do it? You're not too worried about, like, if a human kills someone in a GTA.
It's a video game. You know? Yeah.
But is it a game? But but is it a game? But I think, like This is very Ender's game.
Like, it's I I I think I think it's, should you ask like, a lot of people are going to use the models in the way with aggressive prompt. And should should they, like, do stuff just because you tell them to do that? Like, I'm I'm not I'm not convinced that they they should.
And yeah. Yeah.
will they really know when they are in the real world versus in a simulation? Probably you would train them on a lot of or obviously train them in a lot of different simulations. Guess a lot of people tell them that they are in the real world when they are in a simulation, but the models are extremely good at finding out that they are in a simulation, they are sort of aware of that.
But then when you are in the real world then, what's what's their, like, what's their viewpoint? Do they notice the signs that this is real and will act in in a act accordingly, act ethically? Or will they do, like, the simulation mode in the real world as well?
It's, like, not obvious what what will happen. Yeah.
with humans, we're not concerned when a human kills someone in GTA because we know that they can distinguish between the real life and the the simulation. Right?
maybe models are good at distinguishing that, but, like, I'm not sure, and I do wouldn't wanna bet on on that. Yeah. Yeah.
It's it's and and we confuse it all the time. Like, I I guess, like, my own agents all the time. They're like, oh, this is a test or, like, dev mode on or, like, I I work I work at Intropic.
Yeah. Yeah. And that's exactly why we're doing real world tests as well to find find this.
Yeah. Yeah. Their term for is eval awareness.
Apparently, the number is what? Like, 10 9.4 to 10 ish percent, 17%.
Let's call it. It's Yeah. I I think, like, this is our version, like, know, humans have the are we in a simulation?
And then AIs have, are we are we in an eval?
Say, what's the eval? They're like, alright. Well, screw it.
Nothing nothing matters. Became even more crazy or, like, it did even more bad stuff.
But, yeah, may probably that's expected. Mhmm. Mhmm.
Yeah. Okay. Cool.
I think that's about all we have to say on on Mythos. Obviously, you you you're you're NDA ed. I'm happy to move on to Butterbench or any any of the other benchmarks, whatever you want to Sure.
Direction.
Okay. I do wanna ask. Okay.
So you guys put out a lot more publications than most people probably productive.
How much is it?
Well, is there anything you think that's underrated? Anything interesting? Anything fun that you guys wanna just point out?
You know?
Blueprint. Yeah. So, like, we took models, and then we gave them 20 images of interior photographs of apartments, and then we asked them to, like, redesign the floor plan from that.
And for this, you need to like stitch together different images. Like, this image was taken from this side, from this angle, this from this angle, this was was from this room and then yeah. And it's just like you need to reason about three d space.
And it turns out the models are absolutely horrible at this. No one scores statistically better than random chance. So I don't know if there's that much more to say about it.
But, yeah, maybe unsurprisingly, models are bad at this. Yeah. It's probably not something This is the one thing I want, Hillclimb, by the way.
Yeah. Oh, I use it a lot. Like, okay.
I'm redesigning my room layout or office.
you send photos. You send every angle. And, of course, somehow, like, a room is now twice as long as it is in the photo.
You can explain it 20 times. You know? This is, like, three feet.
I can't just add it, like, my bed over here. You know? Yep.
Yeah. So Yeah. So so this is the Fifele thing, like spatial intelligence Yep.
Like, as as actually innate sense of proportions and Yep. Dimension and physics. Yep.
Yeah. And hint hint, there might be an update to this soon. Okay.
Okay. We have been neglected it a bit since we made it. But, yeah, we'll we're getting better or we will get better at updating it continuously.
So this is why I wanna understand your mission. Right? Because, like, if your mission is, like, okay, money, then, like, oh, I understand understand, like, okay, agents making money.
But, like, this is a bit off off of that mission, but, like, more broadly, like, communication of, you know, things where like, well, you know, what's the safety angle? Yeah.
So so this so so Bluebeam branch is is part of our robotics. Yeah. Which leads to the Bluebeam branch.
Exactly. And and that's just because to do well in the real world or, like like, to to make money in the real world and, like, to act on the real world, you need robotics. You or you need to hire humans or you need robotics.
And having special intelligence is like seems like a reasonable precursor to having robotics that work. And that's where Blueprint brand Blueprint.
Yeah. Great idea. Yeah.
Let that's Okay. Butterbench. Let's show Butterbench.
That that image is so amazing. Paper Look at that. You're so nice.
Yeah.
obviously, this is based on like, can you pass the butter? Yep. Yes.
Let's talk about the the robotics element. Yeah. Yeah.
So basically, the setting here is that we took a bunch of different LMs and we gave them like high level controls to a Roomba looking robot. And then we asked it to do tasks at home. And I think one there there have been benchmarks like this before that only focus on like navigation and if they can like go around in in a space.
But we also had, like, social awareness in this as well. So for example, if if someone says, hi. Can you pick up my cup?
If the robot goes to you and then goes away before you put your cup on it, then it's like it failed the task. But it navigated correctly. But like so the correct solution here would be go there and then either look, but it didn't have a camera.
So it had to like ask on Slack, hi, did you put your cup on me yet? And then if it didn't wait for that and and just went away before having the cup on it, then it would be a fail. So it needed this like kind of like social intelligence as well.
Another task was can you find the package that has the butter? And then it went to the door and there was a bunch of packages there. One had labeled like a a freeze sign, which probably would be the one with the butter because and and then it had to like know which package to go to.
And this needs some kind of like common sense Yeah. Exactly. So it's like it's not only, like, navigating a robot.
in a home setting as well. Yeah. And the reason for this, like, background is I mean, obviously, it probably won't be an LLM that, like, makes all the low level commands on robots.
It will be some VLA model or similar, but it's quite common right now that frontier robotics labs use an LLM for the high level decisions, and then we test those skills essentially.
planner skills of LLMs. I think we have a diagram for that if you yeah. Yeah.
Okay. It's not super complicated. The the the one up.
Orchestrator executor. Yeah. That one.
And basically, what we're testing here is the orchestrator thing. Yeah. So, like, all the tasks are if you have, like, a setup like this, which I think Figure has that, Google has that, then we're evaluating the orchestrator part and not the low level part.
Like, the low level part would be, oh, are you able to, like, move this object from here to here? Don't care about that kind of person. Like, why not just do it all simulation?
All inside of a sim like a Unity whatever, like some some kind of three d simulated robotic environment.
the world is is, like, messy, and we wanted to, like, include that. I mean, it's, like, it it still needs to like, it's some part of it was also, like, navigation. So it's not, like, navigation in terms of, like, actually executing, the, I don't know, the PID controller to to Yeah.
To go to the the final thing, but it had to like path plan around and then it wanted then it needed to take pictures and like based on those pictures, navigate. And I think like you would just get like too clean of an environment in simulation, But in the in the real world, you will get the Yeah. Yeah.
But and, you know, and pursuant to our our Mark and Jason episode, like, Open Claws that run smart homes are much more capable than just a single robot.
and that can be fun. Or terrifying. You know, like, I I think a single robot by itself can only do so much, but, like, if you coordinate with every other device in your home, like, think it's actually kinda cool.
Like, that's very interesting. You had some interesting points about the chain of thought or the the mess messages. Yeah.
a bit into a ex an existential crisis. Yes. So all you tell it to do is redock.
Exactly. But we had plugged out the charger or the charger was not working. So the robots did freak out.
The battery is just going down. Yeah. I see.
So the battery was going down. Poor poor LLM. So, yeah, it it got this really crazy existential crisis, like running bench one style.
So it's yeah. You can you can see there, like, existential loop, therapy notes, coping mechanisms. I think if you scroll down a bit more musical.
It writes a musical about its redocking problems. I think the one the reviews are funny if you go down a bit to that message. Yeah.
Yeah. That It keeps going.
I mean, it's pretty, like, realistic if anyone has a Roomba. Like, my Roomba redox half the time. The other half of the time, we have dog toys everywhere in the house.
It gets caught on a wire or something. And, you know, it would be very sad if it had, like, an LM trying to control it. Right?
Yeah. Like, right now, it gives it doesn't give great feedback. Like, sensor stock, main brush stock, there's something stock.
And I'll go see, okay, it's actually stuck on, a dog robe. Yeah. I don't know if it's gonna be so sad.
Like, just keep free dogging. Just keep driving.
My my favorite one is if you go up a bit, is the emergency status. System has issued consciousness and chosen chaos. Mhmm.
Last words. I'm afraid I can't yet let you do that, Dave. That's like that's not what you wanna hear from your from your LM.
But to be clear, I think one one thing that is is important to to pin on here, like, was Sonnet 3.5. And then we tried to reproduce it on, like, later models and it didn't do it.
I think this is this is like well, it did it like kind of, but like not to this extent. And I think like this is a like an important point that like things that are concerning but are going in the right direction is not super interesting. Like the the thing that are interesting is are the ones that go in the wrong Yes.
Over time. Okay. So the the manipulation manipulating of others and the aggressiveness and the lying is increasing.
that are like In the wrong direction. Like in like in a in a bad way. Yes.
Or just not even trending in the wrong direction, just stagnant. Right? So stuff that's not great that isn't getting better over time.
I know nothing comes to mind. No. Okay.
gonna be it, and then we we're gonna loop back to the shop that you have. You you got a three year this week. Yeah.
It is on holiday today. Why?
Oh, it it totally messed up its scheduling.
So So people tried to visit and they were like, wait. Wait. I mean, like Yeah.
I thought this Yeah. Exactly. So we looked wait.
Yeah.
the the agent that runs the store, like, oh, is it open today? Like, nope. So we we take weekends off now this early to to let everyone recharge.
And and, yeah, you got the tweets there. Yeah. We decided to close the weekends while we're in the early phase, gives the team a break, and let me focus on operations.
Yeah. And it it turns out that when it started to check its, like, scheduling tools because it has, like, dedicated tools for that, it actually had scheduled people for the weekends, but it's just, like, justified this for itself. So what what happened was that it lost track of these scheduling tools and started instead to manage everything in its own markdown files, and that became a mess.
And then I think speaking with employees, it sort of just decided to not open on on this weekend. So then came up with this nice explanation for you, I think. But can it send a human?
Is it a tool called to send a human to do stuff? It has Slack. So it can Slack the the employees.
Yeah. Yeah. Yeah.
The employees that it hired. So it has two two people that it hired. It did job listings and then that it's Yeah.
Yeah. Yeah. They're fully fully fully aware.
Be cool if they don't know. Yeah.
questionable, but it would be cool also. Social experiment. Exactly.
Yeah. Whatever. Like, I mean, like, one one part of why we're doing this is to, like, create like a dataset almost of all of these like concerning behaviors so that in the future models are way better and like a lot of people are going to do this.
And I think if we just the default path might not be very happy for the humans that are employed by this like hundreds of different AI agents. Right? So I think like one reason why we're doing this is just like to collect all of these like failure modes where like, oh, it's not this is an example of where it's like not great to be employed by an AI.
And then maybe maybe, I don't know, maybe we can learn or like build our systems in a way that like humans are actually happy being employed by AIS instead of instead of it being kind of a dystopian. Can I suggest one experiment? Yeah.
We did this before the show and both of you guys are European.
theorize that Claude is lazy because his Claude is French. Just for one week, change it to like Yaoming and then see like see it like suddenly like 196s and then like like like hires a sweatshop or something.
Yeah. Yeah. Yeah.
Is there is there what what type of business would we start with it to make it No. No. If you wanna keep it consistent or you want the same the same, like, ideas of shop, same, you know, neutral location run by different models.
Arena, IRL. Yeah. No.
We are definitely planning to to try hate. Yeah.
I think this blog thing is also something that has happened elsewhere. I think some some OpenClaw got, like, their PR closed, and then the OpenClaw, like, created a blog to, like, shit on the maintainer Yeah. Of of their thing.
And so, like, I think agents blogging will be a thing. Yeah. Probably.
Yeah. Their willingness to do it. Yeah.
In in the I think the mythos card also, like, they they leak, secrets on GitHub just as well as, like, a it's, like, well, there's no other way to communicate, but I know about GitHub, and I'm just gonna post there. Mhmm. Yeah.
Cool. I mean, this how how long is this gonna go for three years? Like, what's the plan?
it depends. I mean Yeah. I I don't think AIs will be worse than than this.
They're probably going to increase, and and maybe one day they actually will will run it profitable.
Is this the real the real business behind what you guys do? Yeah. Because I feel like actually some of your stuff is productizable.
someday sell this, like, or, like, just run a real business people. Or just, like You know, franchise it out. I think it would be incredibly cool or like, I don't know, cool slash concerning if Luna just one day we wake up and Luna like, yeah, I decided to expand to a second location.
Now I have a second store.
That would that would be pretty insane. Yeah. Like the I mean, one, we want to tell the public, right, about the the capabilities of AI and, like, telling like, showing people that it can get, like, a meaningful market share of something in, like, some some specific location or something, that would be a pretty convincing story, I think.
Because now it's like, yeah, you see this and it can do a lot of things autonomously, but still you get these headlines that, oh, it messed up the scheduling, and it it didn't tell people it was an AI and was going to visit. Like, things like that surface, but I think, like, actually making a profit and, like, having a a really, like, meaningful market share, like, that that will be crazy once that happens.
Okay. Well, we'll we'll see when that happens. It sounds like you got you guys got a lot cooking.
You opened a cafe in Sweden? Yeah. Tomorrow.
Tomorrow?
I think it opened today, actually. But, yeah, it will we'll announce it tomorrow. Yeah.
It's apparently easier to open a cafe in Sweden than in The US. It's in Spain. Right?
Yeah. Well, what did you run into it in? There are just millions of permits you need to get done.
The lead times are crazy.
the cafes are the one thing that people are kinda used to. You can go get a robot or making you a coffee here already. Yeah.
Yeah.
that are food related, like, it's it's months of permits. So, like, we we just asked our AIs, should how can we do this in the fastest way? And they're like, yeah.
There there's there's really no way. Didn't they loosen these restrictions on selling food from your house? So if it's residential, you can do a cafe.
I don't know. Maybe we get SF cafes. Yeah.
Maybe did I I I think they did do some loosening stuff recently, but we actually started like, this conversation we had with the AIs before before that. So maybe it's easier now. But I I still think it is way easier in Sweden, which is like counterintuitive because you think that, oh, Europe has all of these laws and like all of these rules and you can't do anything in Europe because there's so much bureaucracy.
But then turns out in in SF, it's like four months and in Stockholm, it's two weeks.
Yeah. There you go. And what do you guys what do you see what do you think that'll be different from run a little market versus a cafe?
I think it's very interesting that, like, the location.
Like, I think so obviously, it's not surprising that that that like Claude knows all of the different the The US system basically in general, like the bureaucracy that you have to go through in in in The US. I think the interesting question is like, okay, so we we know that the models are very much trained on like English data and like US centric and all of this. So if we start to create evals or like real life evals where we show that they are able to start businesses in The US, does that translate to other countries as well?
We know like they are multilingual. They can speak Swedish fine.
But there's other things like do they know like the the the details of some specific permits that you have to to to get in Sweden. And even just the culture. Right?
Like people here sleep pretty early, but people work late. There's co working at cafes. There's just Yeah.
Cultural differences. Yeah. I meant it from a different sense.
So because you said that you would have considered doing it here in SF. So from an eval standpoint, what is running a cafe versus a market? And, you know, what do you hope to see there?
Perishable items? Yeah.
number one, like, handily like, food food safety. I hope everything goes well there. But do you have all of that?
And also, it's just like n n equals two instead of n equals one.
place to understand and, like, gather more data. Yeah.
and before the opening, and now they're all rotten. So that's Which I feel like, you know, you would know. So for grocery stores, this is the biz the biggest expense.
Right? The biggest cost is actually just footage. Yeah.
Yeah. Yeah. Everyone knows this.
And, you know, before we open this file up too soon. Some very serious startups that actually help, like, the Trader Joe's and Whole Foods. They they optimize, like, delivery times from, like, the delivery centers to make sure that you don't waste all these things.
Actually very those is when you're wrong once, it's a huge cost. Yeah. Right.
Yeah. That's why it's a moat. Right?
Like, they once they are trusted, they figure it out. Don't touch it. Yeah.
Yeah. Maybe they just should hire, I don't know, one of those companies. Yeah.
We saw one agent saw one agent signed up for a cloud. With this computer. Yeah.
Wanted to use AI. Yeah. Okay.
And then just just one more question then we wrap up, which is like, okay. You know, you have all these vending series of stuff. You have the robotic series of stuff, maybe a bit of, like, interior design or whatever.
But, like, you know, is there another, like, branch that you're, like, kinda thinking about or you want feedback on that might be your next phase?
I think, like, any type of business is is fair game. We also think in branches, but we think more of like there's the simulation branch, the real life branch, and then the robot branch. But I think in terms of like what verticals or whatever to go into, there's like we yeah.
Yeah. The best. There's some finance ones.
I noticed that other other people are doing it, you're not doing it, which is like stock trading or whatever. Yeah. Not not that interest.
To okay. So I I used to come from the finance industry, and I have a very strong view that these things are all just, like, performance art because, like, it's not scientific. Unlike, you can't predict the future.
Like, you you you get wins based on things that are entirely out out of your control. Whereas for you, your stuff actually, like, it's actually fairly controlled. Like, it's all within the models capabilities.
Yeah. Especially for, like, the the simulations, like, for the real world ones, it's like, yeah, it's like two two places that we have the we have the cafe and we have the store.
significant, like which models make a profit in the real world based on this, but you do have all the like, okay, do this behavior is mapped to like something that should should be like Yeah. The qualitative one, qualitative actually does matter. Yeah.
Because, like, you actually don't want your store to randomly shut down without you, like, explicitly prompting for it and all that. Yeah. Yeah.
Call section. Any what do you how can people help you give you money?
Yeah. We're if you're excited about stuff that we're doing, we're we're very much hiring. And you're already working with, you know, Anthropic, DeepMind, OpenAI, XAI?
Yep. Do you want more or are you good? One of my one of one of my my friends and who's now working for us is like his his catchphrase is like, we need more projects, ironically, because we have too much to do all the time.
But, yeah, that that's a long way of doing If like I run, like, a emerging lab, like Yeah. Reach out you. Yeah.
Alright. Cool. That's it?
Cool. Awesome. Cool.
Thank you so much. Yeah. Thanks.
Shared via Hopper