Robert Lange, a founding researcher at Sakana AI, discusses Shinka Evolve, a framework combining LLMs with evolutionary algorithms for open-ended program search. The conversation highlights Shinka Evolve's sample efficiency, its use of model ensembling and adaptive prioritization, and the critical need for AI systems to co-evolve problems alongside solutions for true scientific discovery, moving beyond fixed problem optimization. Lange also touches upon the future of AI-driven research, emphasizing human-AI collaboration as "shepherds" of scientific exploration.
Refreshing wild cherry cola meets smooth cream. The treat you deserve. Pepsi wild cherry and cream.
Treat yourself.
Ever feel like your brain just won't click? Onnit Alpha Brain is a daily supplement engineered to support memory, focus, and mental speed. Made with science backed ingredients, Onnit Alpha Brain helps you lock in, tune out distractions, and stay sharp.
See what your brain can really do. Visit onnit.com and shop Alpha Brain to unlock your next level.
That's onnit.com.
I think a lot of sort of analogies from evolution transfer to scientific research, right, in the sense that we traverse a tree of different ideas or different experiments. And then in the paper, we report one path through that tree.
autonomously,
tend to just kind of like nothing interesting happens. But oftentimes innovation for a specific problem might require first inventing a different problem, right? Sort of automatically coming up with this reduction or like let's say, recursive nature of problem solving is something these systems right now not necessarily have built in intrinsically, right?
actually like hard verify them, right? The reason why I'm not that worried yet about labor market disruption is I still believe deeply that humans are the source of deep understanding and creativity in the world. If I didn't believe that, I would be very worried.
of sort of these these latent dimensions
humans are great at. Right? And I think one of the Rubicon moments is when the the new transformers architecture or something massive is discovered by AI, and we're all using it.
NVIDIA GTC starts Monday in San Jose, and it's free to attend virtually online. There's already been a leak this week of something called Nemo Claw, which is an open source agent platform. And if it's real, it could be one of the bigger announcements this year.
So it's definitely worth watching Jensen's keynote for that alone. I'm giving away a DGX Spark. NVIDIA just hikes the price $700.
Probably heard about these memory shortages. Right? So, yeah, it's now $4,700, which is very, very expensive.
And, Merv from Hugging Face, by the way, she got one for her birthday, and she said she literally cried. So it's a really cool bit of kit. If you register through my link in the description and you attend at least one session, then you are in the draw.
This is a massive conference. Physical AI and robotics are gonna be the breakout theme, and Jensen does the keynote Monday at 11AM Pacific. The link is in the description.
Don't miss it. Robert Langhe, it's amazing to have you on MLST. Thank you, Tim.
It's a pleasure to be back. So you're working for Sakana?
Tell us about that. Sakana AI is a Japanese, AI startup working mostly on, yeah, AI for Japan, and at the same time sort of exploring exploring, let's say, novel or ambitious ideas on the research side. It's been around for over a year now.
You're on you're one of the founding researchers. Right? Exactly.
So Sukana has been around for now, like, almost two years, like, one in three quarters, I would say. And, yeah, it's pretty fascinating to to look back and to look at the early days and how much the company sort of organizationally has changed.
Ken Stanley's open endedness idea and sort of explore many different ideas which might not get the resources right now in the ML community more generally. And we've we've got a few interviews coming out with Sukana that that we filmed here in Japan. So I'm not I won't spoil the surprise, but the the CEO is David And David, you know, like there are these epic, you know, giants out there like, you know, Clune and Stanley.
David is one of these people. David's work has had a lot of influence on my personal PhD. Right?
He he did a lot of fascinating work on hyper networks and sort of modulation in in neural networks, but also on evolutionary computation and evolutionary optimization.
yeah, my path during the PhD. You've you've released a paper called Shinka Evolve, and we and we were just saying that that kind of means evolve evolve because in in Japanese, Shinka is evolved. But that's quite common.
That's a common thing to do to have these, multilingual, you know, double double namings in in Japanese. Just before we get there, so, we interviewed the AlphaEvolve team, and I also interviewed Jeremy Berman a few weeks ago. And your paper is is very much like a more sophisticated version of those in the sense that it's using language models to generate programs, and it's doing an evolutionary approach where we generate the program, we refine the generated program, and we have an an evaluator.
And we do this over several steps. And and your your your approach does many things that that the other ones don't do. Tell me about the paper.
First off, of course, this was partially inspired by AlphaEvolve. I think it's great work. I know Alex and Mate, and I think they're doing incredible science.
One thing that sort of is important about sort of using all of these evolutionary LLM driven methods is sample efficiency. Right? So many of these systems sample, like, let's say, a thousand programs for a given task.
And what we try to do with Cinca Evolve was try to essentially cut down costs as well as sort of computation evaluation time by introducing a set of sort of technical innovations to this evolutionary search, and we showed that it's possible with very few program evaluations to basically improve upon, like, for example, the circle packing can canonical result that they showed in their paper. And, yeah, more generally speaking, I think we're right now at a point or, like, at an inflection point where these sort of, let's say, evolutionary driven LLM systems can really revolutionize scientific discovery. And, yeah, we hope to have made a step forward to making this more democratically accessible.
Right? So the code is open source available and, yeah, by its sample efficient nature, we hope that many people can interact with the system and can make their own scientific discoveries as well. Yeah.
That's actually a really important point because I suppose we can use these foundation models.
And first of all, isn't it just fascinating to reflect that we have these amazing models out there that we can access, so like GPT-five and Grok four, and they are so much better when you get them to refine their solution in several steps. Why is that? I suppose a naive question would be, why why aren't they just good out of the box?
Potentially, like, with enough random samples. Right? It's sort of this monkey typing on the keyboard.
They would potentially be able to get there. Right? But in principle, it's sort of coming back to the principles of evolution, right, in the sense that you need to collect a bunch of stepping stones first and then build on top of them to to really find innovations or to tune innovations down the line.
And think language models with the right sort of evolutionary hardness are extremely powerful in terms of scaling up to to to make discoveries.
is really important for that. Very cool. And stepping stone collection.
So this is it came from Kenneth Stanley. It's a wonderful paper, Why Greatness Cannot Be Planned. And he said that it's it's better to have systems that don't converge.
So in natural evolution, we are just trying all of these different things. And greatness quite often follows a diverse path, which means you have to do things which initially seem quite stupid. And then later on, they turn out to be incredibly useful.
Yeah. We're trying to design algorithms that can allow for a population of slightly weird things. And then we lock in and converge a little bit.
We're still converging though. We're still building systems that don't diverge forever.
What are we losing? One one thing I find extremely important after having done Schenke evolve is sort of this problem problem. Right?
So with all of these systems so far, maybe except for the AI scientist, which we can also talk about, the problem is given. Right? So you have an evaluator, you have a correctness checker, and you sample programs only on that single problem.
Right? But oftentimes, innovation for a specific problem might require first inventing a different problem. Right?
So for example, I think in the matrix multiplication result that the AlphaEvolve people show, you can recursively apply sort of the algorithm to larger matrices. So it's actually an important result. Right?
But sort of automatically coming up with this reduction or in like this, let's say, recursive nature of problem solving is something these systems right now not necessarily have built in intrinsically. Right? So I think going forward, it's gonna be really important to not only sort of do open ended, let's say, optimization of solutions, but sort of do the coevolution of problem and solution together in order to collect even more diverse stepping stones and to really kick off this this open ended process.
Because also to me, like, of the the big life goals or achievements I would wanna see is really having a process that can run not only for, let's say, a week or many weeks, but, like, for years even potentially. Right? Collecting even more diverse, interesting stepping stones.
Yeah.
which is that machine learning algorithms aren't very good with unknown unknowns. And in a sense, the unknown unknown is talking about these stepping stones that might be useful later. And when we run these algorithms at the moment, it's the same with LLMs and reasoning systems, is that they're very, very good when we give them a specific thing.
And what you're pointing to is we might need to invent new unrelated problems and find the solutions which might then be related to what we're trying to do. So that feels like a bit of a catch 22 situation. So we're saying, circle packing, here's my evaluation function, and I want you to diversify and then converge towards this solution.
I had the same thought with Genie, by the way, that it gives you exactly what you ask So you put put a prompt in, you know, like a Swiss lake with, you know, with boats on the water and mountains on the side. And I was thinking, where are the birds? Oh, I forgot to put birds in the prompt.
Mhmm. Right? So how can we meaningfully build systems that actually kind of bring in other unknown things that might be useful?
research are systems like outlined in in Powerplay or Poet by by Schmittelbern by by Jeff Kline and others. Right? So where there is essentially, like, a a set of tasks and a solution generator, and both of them sort of coevolve in this almost like art curriculum play like style.
Right? And I think sort of the in POET, the natural first application was sort of reinforcement learning, but I think this can now be broadened up to to, yeah, science more generally. Right?
At least when there's a simulator available to for for running these evaluations.
yeah, more diverse problems while doing so. And I know that there's always the leading thought that even with POET, which was this thing where you had like a populate you you had like a load of environments and agents, and the environments were in complexified. Mhmm.
So the agents would have a kind of effective curriculum to to learn things in increasing complexity. But even then, isn't there a kind of design bias in the system where there's some code somewhere which complexifies the environment step by step?
wouldn't that also just be designed by the humans? So it would also just give you exactly what you ask for. Ultimately, this, comes down to, like, the hypothesis that, language models can potentially do extrapolation or interpolation.
Right? In the sense that even though these things might be in the end designed by humans, there are many unknown unknowns. Right?
That we humans didn't think of while designing them. Right? So potentially it is possible for an LLM to, yeah, find a novel discovery simply by us not having thought about it before.
Right?
autonomously Yeah. They they tend to just kind of like nothing interesting happens. So depending on the prompt you give them, they'll kind of go a few steps in that direction and then no new interesting novelty emerges.
And I think even if you wire them agentially with environmental feedback, they they still seem quite parasitic on their starting conditions. With an LLM, could we build a system which actually adapted to novelty, that could actually discover new things?
what do you give the LLM as a starting point. Right? So for example, in Schenker Evolve, we from time on time saw that if you give an initial solution program, which is already pretty optimized on the problem at hand, you still kind of get stuck in in local optima, right, where not a lot of novelty is introduced.
Right? Well, if you start off from like an impoverished solution, there's much more room for diversity. And I think this is sort of coming back to sort of what I did before in in my research namely meta learning, sort of this classical trade off where you can either start out from something very, let's say, unconstrained from, like, a very simple solution and give much more room for the optimization.
sort of benefit from it. K pop demon hunters, Haja Boy's breakfast meal, and Huntrix meal have just dropped at McDonald's. They're calling this a battle for the fans.
What do you say to that, Rumi? It's not a battle. So glad the Saja boys could take breakfast and give our meal the rest of the day.
It is an honor to share. No. It's our honor.
It is our larger honor. No. Really.
Stop. You can really feel the respect in this battle. Pick a meal to pick a side.
And participate in McDonald's while supplies last.
Starting a business can seem like a daunting task Unless you have a partner like Shopify. They have the tools you need to start and grow your business. From designing a website, to marketing, to selling, and beyond, Shopify can help with everything you need.
There's a reason millions of companies like Mattel, Heinz, and Allbirds continue to trust and use them. With Shopify on your side, turn your big business idea into sign up for your $1 per month trial at shopify.com/specialoffer.
Yes. Suppose where we want to get to is building systems which are not designed by humans. So for example, if I'm leveraging my deep understanding, LLMs are really good if understand something deeply.
And similarly, we could kick off Cinke Revolve, and we could put a starting solution in there which leverages my understanding. We want to have AI systems that anyone could use. So just a non expert could say, I want to solve this problem, it will solve the problem.
We should talk about the the evolutionary approach. Right?
and they were separated into islands. Tell me about that. The way how Schenka Evolve, similar to Alpha Evolve, works is you keep an archive, like a database of programs, and then you sample parent programs with a set of sort of inspiration programs.
And then you ask an LLM to basically make an improvement to that program. Right? So to provide code edits or rewrite an entire program or to potentially even cross over different programs.
And then you you basically you query the LLM, you get a program out and you evaluate it on the problem at hand. Right? So for example increasing the sum of the radii of a bunch of circle in a square.
You run this basically each time collecting evidence from the evaluator, adding it to the database, and then sort of repeating this process. And you don't do this sort of sequentially, but you do this in parallel for many different programs. And each time sort of a program is added, you essentially try to diffuse the knowledge that was collected by the program across the entire sort of database.
Right? So one way to think about this is you have a tree, a tree where each node in the tree represents a program, and then you you sort of branch off of it based on the parent nodes. Right?
And interestingly, like, these approaches do tend to scale, but ideally, can make the scaling at a faster rate. Right? And this is something we tried in Chinka Evolve by sort of doing a bunch of innovations, including sort of model ensembling.
So we're not using just Gemini, but we're using basically all frontier model providers and figuring out a smart way how to use each model for a given parent. Right? So if you have a certain program in some situations, it might be better to use a GPT model, in other settings it might be better to use a Gemini model, and we sort of introduce a sort of adaptive prioritization scheme that can adapt sort of the evolutionary algorithm on the fly while running the algorithm.
And this sort of also comes back to the naming. Right? So Schenker evolve, evolve evolve, kind of means that this evolutionary algorithm that we apply using LLMs sort of also coevolves at the same time while we optimize the programs.
And on this, while we're on this circle packing problem.
So you you had this plot showing how it converged, and it seemed to converge quite quickly. So and we'll show the plot on the screen now. So very quickly, the performance jumped up, then it slowly converged.
And you said in the paper that it was using three, I think, three core innovations. And my thinking was, if you ran this 50 times, would it be the same every single time? And to what extent is it thinking outside the box?
Sebastian Bubeck is always posting on Twitter talking about how GPT-five has discovered new things. And there's always the question of, well, is it just searching the Internet? Is it just finding things that have been found before and, yeah, combining things together in in a new way?
But could it really think outside the box? Yeah.
I think this is almost like a subjective question. Right? So first off, I don't know all problems on the Internet that try doing circle packing.
Right? But what I can see in the tree that we also depict is, there's, for example, like a crossover operation between two programs happening where sort of different concepts are combined. Right?
So one important part is, for example, the the initialization of the circles. Another one is, like, the optimization. So basically, like, a constrained optimization program is executed.
And then the final part is basically like a reheating stage, right, when noise is added and sort of more stride to be squeezed out. And to me, like this sort of propagation of information through the tree is one that's really, really fascinating. Right?
Where in some sense, these stepping stones are actually used and so in a complementary fashion. Right? And with regards to rerunning the program multiple times, right, of course, there's some stochasticity in that.
Right? So we're using language models and sort of due to, like, the the queuing device scheduling on on their server side, basically, we can't get rid of all the all the noise. We we've seen that at least for the general quality of the solution, so what is the right afterwards, it is possible to re obtain this.
But sometimes with a different program, like, or most of the times just by stochasticity. Right? So it's not like there's for many problems, there's, like, not one solution that achieves that score, but there is, a spectrum or, like, a region, let's say, in the program space that that resembles the same.
Right? I think one thing that was very interesting about the circle packing problem, sort of also coming back to the problem problem that I discussed initially was that originally, we used a formulation where the correctness is checked with like a very tiny amount of slack. Right?
So the the circles could overlap a tiny little bit. And then afterwards, we we we sort of reduced the RAID AI and the solution was exact. Right?
This didn't change the score by too much, so it's still state of the art, but it was essentially like a proxy problem. We then reran the Schenker Evolve on the exact setting, and we found that it took a little bit longer to actually obtain the same quality of a solution. So I think this already points a little bit in this direction of what I discussed in the beginning, like sometimes sort of surrogate problems might actually be extremely valuable in in making such discoveries.
And having an automated way for designing these surrogate problems in an efficient way might be something really important going forward. Yeah. That's absolutely fascinating.
make the optimization tractable by introducing slack variables, and you can think of that as a surrogate problem. But then I'm thinking, would Schinker Evolve or Alpha Evolve, would it know to introduce a surrogate problem? Because as designers who understand, we can think outside the box and we can do stuff like that.
then it wouldn't it wouldn't occur to the algorithm to come up with a surrogate problem. Exactly. Yeah.
This is a big limitation right now. Right? So at this current point in time, we take the problem to be fixed, and we optimize for that problem.
But when you think about humans, we're really, really good at sort of inventing our own problems, right, or reformulating the problem so that we can actually sort of work with it. Right? So I think a lot of sort of the innovations in, let's say, mathematics come from taking a very different perspective on a problem.
Right? So taking sort of number theory and applying it to linear algebra or the other way around.
of achieving such level of, let's say, transfer. Yes. And it reminded me, I I spoke to Lion about this.
You've got this Sudoku bench. And a lot of folks watch cracking the cryptic YouTube channel. That's exactly what they do.
They invent new problems based on abstractions that capture the essence or aspects of the problem you're solving. And then they do something which is similar to Shinker Evolves. They do this kind of evolution where they take these different solutions and and they kind of combine the the best aspects of both of them, and they forge a divergent path to a new solution.
Yep. And that seems to be the essence of of what we need to do. Yeah.
For sure.
Shengren Hu, and Song Liu on automatic automated capability discovery. So there, they look at language models that generate tasks. Right?
But it's in a, let's say, unstructured way in the sense that it's not done in order to enable the solution to one target problem. Right? And I think sort of doing these connections is gonna be very fruitful down the line.
Very cool. Now the other thing, we'll show the graph on the screen, the evolutionary graph. So for the circle back in problem, I was looking at that.
parsimonious, which is good. It it looked like it had found an optimal path to the solution very quickly. And I was thinking in my mind, well, maybe there's some natural pattern that that there's there's there's there's something about that that we could use in the abstract to guide the evolution in the future.
But the other thing I'm thinking about is, right now, the problem with machine learning is that we don't really have semantics baked in. So what we're doing is we have a verifier. We're looking at the rewards, and we're sort of like doing pattern exploration, and we're taking steps towards the target.
I love mechanistic forms of reasoning where we actually know something about what the program components mean. The reason this is important is when we're merging together the best performing programs from two different islands, that's a first order interaction. It might not make sense to merge them together.
It's wonderful that LLMs, you can give them any pairs of programs and it will find a way to merge them together. But wouldn't a more principled way be of there's there's some kind of semantic primitives here and we know they fit together. So there's this Lego analogy that we're kind of building up based on principles rather than forging a path based on the performance.
Yeah. That's a good point. So one thing we do in Chinkai Evolve as well is we keep essentially a scratch pad.
So each program is being summarized. And then from the program summaries, we keep sort of a set of global insights, let's say, that were shared or, like, extracted from these programs. And then based off of the scratch pad, we construct sort of meta recommendations that then become part of the system prompt.
Right? So that way you can try to sort of semantically grasp some of the discoveries, but a general problem, which is again sort of task dependent is thereby you sort of diffuse that knowledge across the tree. Right?
But sometimes you want things to be much more isolated. Right? It's always like a trade off where you somehow have to find for your problem the right position on the spectrum of how much knowledge diffusion do you wanna have and how much sort of, let's say, hard islands of programs do you wanna have.
Right? And, yeah, we're trying to make steps in the direction of sort of automatically adjusting this in an optimal way, but, again, it's very program sensitive. And then sort of I think another point where you're already sort of going into is sort of Jeremy Jeremy's solution to Arc AGI.
Right? And sort of doing solution evolution in the instruction space, right, instead of the program space. I do think that this is something important, and we're, like I said, with, like, the construction of this meta scratch pad trying to do sort of both at the same time.
Again, it's problem dependent. Like, I played around a little bit with ARC AGI one and ARC AGI two. And I think on ARC AGI one, actually, the the transform sort of program direction is actually quite effective.
Right? It's like Jeremy said, it's deterministic, and it's easier to sort of get clear signal to improve on during your evolution process. While on others, like ARC AGI two, like this whole sort of semantic evolution seems to be more efficient.
input output mappings. Yeah. It's it's so interesting because, you know, like, a a symbolic AI person would say, oh, I don't like connectionism because it doesn't know, the only semantics in connectionism is this notion of similarity.
It doesn't really understand things. So so they would say, just just start with an entity relationship graph and then just kind of build up using, you know, composition and first principles. That that that doesn't work.
Right? So we're using neural networks because they're incredibly flexible and they understand a lot of things about the world, but they don't have the kind of constraints that we want. So what we do is we use these tricks.
So Jeremy, we evolved program descriptions.
Embedding based similarity.
Yes. You had like a kind of self similarity metrics and, you know, based on the cosines. And indeed, you've got this meta scratch pad.
where still using neural networks, you can imbue semantics in using all of these different tricks, but they all come with trade offs. Yeah. For sure.
Like, I think it's it's kind of interesting. We we've had a long period of computer science where algorithms were sort of designed by humans. Right?
Then we had sort of this Android Kapathi software two point o o paradigm where, like, we trained neural networks that then performed a certain function. And now we're sort of at this point where we're using LLMs to design algorithms or solutions more generally. Right?
And I think, actually, like, even though, like, large frontier language models are extreme, like, let's say, black boxes, or it's very hard to get a full mechanistic understanding of them, the outputs can be. Right? The programs, the instructions, and so on.
Right? So I think it opens up a very sort of new paradigm of doing research or basically doing anything, right, if you if you think about it.
But I think we're we're just sort of at the starting point of figuring out the the right user interface for that. So the other innovation in the paper was using, UCB, which is, upper confidence bound. It comes from the multi arm bandit literature, which is this problem where you can pull these these levers.
And at the beginning, you don't know which levers to pull. And and over time, you kind of reduce your uncertainty and you can kind of pull the ones that work. But there's this exploration exploitation dilemma.
And you've implemented that for figuring out which LLM. So it could be Gemini. It could be like, you know, Grokfur or something to figure out which one to use.
We're we're using like a model ensemble, right, to propose program mutations.
And, intuitively, one could say, like, the the best frontier model on on SWE bench is always the best mutation proposal model. But that's actually in practice not always the case. Right?
And in general, it's extremely hard in this evolutionary setting to assign clear credit to a single model. Right? So you have, for example, like, one improvement is implemented by GPT five, and then the next one is implemented by SONNET 4.
5. And it's unclear, basically, if the performance gain you get from the second mutation actually originated from GPT five sort of collecting the first stepping stone or from SONNET 4.5.
So instead of sort of uniformly sampling models, what we do is we implement this bandit based approach where each model is basically one arm of a bandit, and then we look at how often did this model improve performance of a sort of parent node by creating a mutation. And we then adjust sort of this posterior probability to sort of first explore all arms once. Right?
models that sort of yielded improvements before for similar nodes. How many discounts does USAA auto insurance offer? Too many to say here.
Multi vehicle discount, safe driver discount, new vehicle discount, storage discount, legacy How many discounts will you stack up? Tap the banner or visit usaa.com/autodiscounts.
Restrictions apply.
Spring Fest is happening now at Lowe's. Keep the spotlight on your yard with Stay Green premium two cubic foot mulch, five bags for $10. Plus, when you want more help indoors, get up to 40% off select major appliances that help you supercharge your chores.
Our best lineup is here at Lowe's. Valid to four twenty two while supplies last. Selection varies by location.
See lowe's.com for details. Molt chopper excludes Alaskan, Hawaii.
The great thing about using a UCB like algorithm is is you can it it actually has, a theoretical regret, which means it's not it's it's like only log worse than the optimal switching path, if if that makes sense. But if I understand correctly, UCB is based on a sort of like a global rating, like a mean score of every single LLM. And I think what we want is to have more of a contextual switching decision, which means we know for this particular program, Gemini is better.
And do I understand correctly at the moment that it might converge to a single frontier model and then in a nuanced situation, we might still get the wrong model?
some amount of probability associated, like, allocated to all models. Right? So it's not like it can just peak on one model and then you stop using the others.
Right? So there's still a chance for open endedness and serendipity, if you will. And we in general, like, the problems we consider, we we haven't seen that, like, one model clearly dominates all the others.
Right? We've seen then it really depends on the course of this evolutionary process, like, model is better. And UCB or, like, the the banded approach that we take dynamically adjust this in in an efficient way.
And would it be possible in the future to use an LLM to make this judgment? Potentially. In some sense, in that case, again, you think of the LLM as a surrogate model.
Right? In some sense, can you think of, like, a Gaussian process as a surrogate regression model, and there has been some work sort of showing that language models can act as surrogate models. And the real question to me is, like, how do you represent the information to the LLM, right, in the sense that if you use, like, the raw programs and their fitness evaluations, you you quickly run out of context.
Right? So you need some amount of compression in order to present the information the right way to the LLM in order to do this prioritization of the models. I hadn't appreciated how long the context is.
you know, could we use, like, an 8,000,000,000 LAMA model and we're doing, active fine tuning. So we're saying, just ran it on you know, I just ran this program on Grok Yeah. And and it got this score.
Yeah.
it will kind of know that Grok is good at these problems. Yeah. Potentially.
I'm not sure, like, how efficient this fine fine tuning is if if we're only evaluating, like, a 150 programs. But in principle, one could imagine. I think it's on the engineering side, not necessarily like the prettiest to do.
Yeah. It could it could in fact happen. But I think, like, for all of these things, we started out sort of with the, let's say, most intuitive algorithmic component that we had, and UCB was one that really did the job here.
And, yeah, much credit to Eduardo Satin who introduced this to to Schenka.
diffs and and the mutations. So we we generate programs and I I think you folks were inspired a bit by AlphaRevolve. So they actually had this gating where where you kind of gate part of the code which is mutable.
Tell me about all of that. A program is just, let's say, a long string. Right?
And, in order to to make sure that certain parts which are sort of essential to the evaluation, for example, into the imports and so on, were not sort of deleted by the LLM mutations. They are so called markers, which basically state which parts of the code are mutable and evolvable. And it's easy to, like, programmatically sort of make them actually immutable when you get a diff proposal, and these will not be changed.
So only the the rest of the the code snippet will be changed. We sort of implement a type of rejection sampling with reflection approach where if an LLM by chance, for example, tries to mutate this part, it's gonna be rejected and you resample a new proposal. And, yeah, thereby you you can somewhat mitigate certain security or safety problems and, yeah, get a robust sort of mutation.
One of the sort of, I think, the the bigger questions is how can you turn this from a single file mutation setup to a multi file mutation setup? So working on entire code bases. In principle, you can represent many code bases in a single file.
Right? But the hierarchical structure might be actually useful.
trade offs, basically. I I love Ader, by the way. Mhmm.
It it feels that in the future, code generation systems will actually resemble Syncr Revolt. And if you think about it, it'll be using some kind of Git repo. Maybe Cursor already does this, because in Cursor, you can restore previous checkpoints.
But it can be exploring different branches and merging checkpoints together. Know, obviously, you just say in natural language what you want to do. But we didn't talk about mutation, by the way.
So we just spoke about diffs. And there's also an option to do the full file rewrites. Exactly.
But there's also this notion of of of crossover. So how how does that work?
mutations is that here we wanted to have more flexibility to entirely rewrite the program, right, to come up with a completely different stepping stone if you will. So again, there you can make parts of the code mutable, but instead of proposing, let's say, a patch to change certain parts of it, we essentially rewrite the entire program. And this sometimes is helpful.
Right? It's not always like a clear benefit, but it it allows you to essentially get more diversity into the search. Right?
So this is one type of mutation next to sort of this diff patch based approach. And the other one is a crossover mutation where we sample basically not only a single parent program, but sort of two different ones. And we ask the system to sort of make a complementary improvement.
And, again, on some problems, this is really helpful and on others it's not. But in generally, we found that sort of having a diversity in terms of operators is also helpful in discovering new things. And I wanted to to sort of follow-up on the point you made before about this sort of being a new paradigm.
I think so too. I'm really convinced. I think right now, we're sort of at the beginning where we we still think a lot about sort of this chat assistant interface as the way how we interact with LLMs, but it's most of the times inherently single threaded.
Right? So we're sitting in front of the computer. We're interacting with the chat.
We're seeing sort of changes as they occur in the editor. We accept them and so on. But I think this is sort of also just a stepping stone towards sort of a more, let's say, distributed way about thinking about research, optimization, and so on.
So I like to sort of think of Vibe coding, Vibe chatting, and on the other hand, we have sort of Vibe optimization and Vibe researching where sort of my ideal future scenario is one in which, you as a researcher sort of during the day co work with, like, a system like Shinka or the AI scientist. You sort of steer the ship like a shepherd in some sense. And then during the night, you you you press play and you go to bed.
And then this in the background, you have multiple experiments running and automatically new ones being proposed by LLMs, evidence being accumulated. And then in the morning, you come back and sort of you have an multithreaded sort of system running in parallel. And you're more like the shepherd of the ship than the the person actually executing experiments and analyzing.
Oh, yeah. You're still analyzing, but you're not executing. This is happening sort of by the system itself.
Yes.
this might be semi supervised or even proactive. I mean, you know, there's that new product from OpenAI where it knows what you're interested in. And while you sleep, it's going off and, you know, find your pulse.
That's right. And, you know, we're in the situation now where we're reasonably technical people. So, you know, MATLAB and Mathematica, they're supremely powerful.
But you need to know how to express problems precisely. Whereas I can imagine a future where we, express problems just in natural language, or maybe just based on our interactions with language models. The platform knows what we're interested in, and it can just go and find things on our behalf.
to people who perhaps don't know exactly what they're looking for. I think one of the bigger problems there is sort of this verification aspect to it, right, in the sense that oftentimes it's easier to generate a lot of solutions than to actually, like, hard verify them. Right?
Language models are capable of doing sort of soft verification and looking at code and sort of latently running like a like a stack trace of execution. Right? But it's not exact.
Right? And I think sort of these notions of reward hacking and sort of not doing real discoveries, but sort of shortcutting them is one where we need to put more time and effort into to figure out, yeah, how to make sure that this actually moves in the right direction. Right?
And I would hope that language models at some point can do this efficiently themselves. Right? So either implementing in code or latently doing it, But this is also, like, part of the problem problem.
Right? It's not only coming up with the problem, but also with the automatic verification at the same point. Yeah.
there are natural patterns in the world and the building blocks to construct novel solutions are already there. And maybe they're there for a reason. Maybe they just reflect natural regularities in the universe.
Because there's always this question of intelligence is about adapting to novelty. So the world is always changing. And the world tomorrow will have things that we can't explain with our knowledge today.
and LLMs might already have those building blocks. Yeah. For sure.
I think, like, in some sense, the more you think about sort of Occam's razor applying to everything in our world, like, let it be language or let it be sort of science, is is pretty interesting because, like, these artifacts now go into our language models of today, and potentially, is some amount of this being captured. I think, though, it might also be an inactive bias that leads to a local optimum at some point. Right?
And you need more complexity.
the system out of this local optima eventually. Yes. And then there's also the notion of the importance of adaptivity.
So this is what Charle says in intelligence is. And since we've had these models that actually do adaptivity at inference time, so things like test time, active fine tuning, and the reasoning models and so on, they started getting nontrivial performance on ARC. Now, it's very, very expensive to have adapting huge foundation models.
It's just a practical concern where we haven't done that yet. But what we can do is build systems like shrink or evolve that leverage the best of both worlds. So they leverage frozen foundation models, but they give you adaptivity.
that allow us to adapt to novelty. Yeah. So we are having our cake and eating it.
I have to say, found it very interesting that Jeremy, basically, in your podcast, when you asked him about Cinca, was saying, like, he doesn't believe that there are a lot of sort of percentage points to be gained by using a system like Schenka, but you can make it much more efficient. Right? That was sort of the gist of his answer.
And to me, it's like once you have made it much more efficient, you can scale it up again. Right? So if you essentially have a cheaper system that can generate many more sort of instructions, I would expect that by the nature of open endedness, you might get some amount of improvement out of it.
Right now, I don't have any evidence for it. I would love to collect that evidence.
you should be able to to progress. Yes. And that and that is a great segue because certainly on on the circle packing problem, it was so sample efficient that in less than 200 interactions with an LLM, you converged on the solution.
But I was thinking that great, but it's still quite dependent on the starting conditions. We talk about this design bias and so on. So what we put in is very important.
But now what we could do is scale out. So we could run this a thousand times and we could have another process which prompts, generates, breeds the starting conditions.
parts of the epistemic tree. And what would happen if we just scaled that out massively? We haven't tried, but you could even start with, like, an empty program, right, which would be basically the same.
Right? And then you would branch off of that empty program, I would expect. Yeah.
but I do think, in many ways, sort of this is the question that will push us towards like this true open ended vision of running a system for like a month or so. Right? Really trying to squeeze this out.
Yeah. I'm not sure if we're entirely there yet, but I will do my best that we will. And the reason this is interesting is we know as a practical matter that we can't start with nothing.
Mhmm. If we were just sort of like starting from the most primitive building blocks, the search space would just be huge and there'd be no learning signal. So we know we need to start a little way up the stack, but we can massively parallelize that.
So that, yeah, let's say we have a thousand different instantiations of Cinkerevolve. It doesn't have to be embarrassingly parallel. We could still have some sharing.
So during their execution, we could still have a little bit of like crossover and and and maybe then we could we could run all the Cinkerevolve instantiations in a similar kind of meta evolution loop. And my suspicion is contra, Jeremy, I agree with you. We know there are diverse stepping stones out there that could dramatically, dramatically improve many of these solutions.
We simply haven't scaled it up yet. Yeah.
using a system like Shinka Evolve could be able to sort of automatically detect whether or not, like, an instruction based optimization approach for a given problem or a transform based approach is actually the right thing to do. And sometimes, potentially, it's, like, even the mixture. Right?
There's some things you can probably easier even articulate in Python than you can articulate in in sort of language. Right? So I would be really interested in sort of exploring that.
Yeah. I mean, you said earlier about Jeff's clean what what was Jeff Cleans paper?
discovery. I did speak to him about this at Neuros, but something like that could be fascinating as well, you know, where we're also generating the problems and solutions and then kind of moving them back in. But I think the way this will land commercially is there'll be a new type of GPT where everyone is solving different types of problems and and the system, it'll be like a kind of Cinca Revolve, but a massively distributed version where mathematicians are using the platform over here to solve this problem and it will see commonalities and and it will kind of like link them together.
like, human creativity in this process as well, I think. Like, a big challenge going forward is going to be, like, how do we change our incentive system for this to actually scale. Right?
I think, like, for example, some amount of economy will be needed or some amount of mechanism design in order to make sure that everyone is still happy to engage in that. Right? So maybe we're gonna have many more leaderboards for whatever is numerically sort of scorable.
human shepherding and steering will ultimately sort of change and revolutionize science and, I guess, society more general. And, Rob, looking at the future, we've got a load of people in in, San Francisco that's that wanna scale language models, and they are adding in implicit forms of adaptivity and composition so that they're building controllers and they're doing reinforcement learning with verifiable feedback and so on. I think that you subscribe to the slightly different idea that that we need to be far more open ended and we need to be using evolutionary algorithms and so on.
But do you think that they are on a path to nowhere? Do you think they might change tack?
Do I mean, where where is this going? So I I actually think that these things can be complementary, right, in the sense like, let's say you fine tune a model to be like a circle packing expert. Right?
So I I do believe that mixing in sort of different sort of RL fine tuned models into sort of the ensemble of models and then having a good way to adaptively select which one model to use is is not a bad idea. Right? So to me, I just very fully subscribe to this philosophy of open endedness, and reading Kans and Joel's book was really like a fundamental moment in my life.
And I want to see how far we can push this. And I think we're we're not yet at sort of convergence where either the capabilities of the models has converged or the the way how we scaffold around them or the way how we humans interface with them. So to me, they're really like these three points, like, model capability, model scaffolding, and sort of the user interface.
a lot still to push on all three angles. Beautiful. The only thing we didn't talk about was we spoke about the circle packing problem, but you also applied it to a few other things.
Can can you tell us about that?
sort of used, a framework called ADAS, automatic design of agentic system, where basically instead of manually writing an agent scaffold, you use an LLM to write agent scaffolds for a specific task. Right? So what we did is we looked at mathematics tasks.
So Amy and we used Cinca to evolve basically an agent. Right? So using an agent to evolve an agent.
And we found that there we could dramatically improve sort of the performance of very cheap models like GPT 4.1 nano, but the agent scaffold was also able to either like generalize to other language models or to different years of of AIMI. Right?
That was one application. One important other application that we did was to ALE bench. ALE bench is basically work done by other folks at Sicana including Yuki who's also part of the paper, which is considering heuristic sort of programming contact contests sort of previously done and executed by Adcoder, which is like this famous Japanese competitive programming organization.
And we sort of showed that Shinka can also work very well as a coscientist. So basically, we we took initial solutions obtained by an ALE agent that was previously designed, and then we optimized on top of these initial solutions with Schenka and showed that on one of these sort of programming tasks, if the combination of this agent and Schenka would have competed in the challenge, it would have ranked second place basically. So I think there's some evidence that Schenka can work as a coscientist and not only for LLM agents, but potentially even for humans like we discussed before.
And then finally, the final application that we looked at was designing sort of mixture of expert load balancing loss functions. So at Zekana, we've done some previous work called DISCOPOP. I think we discussed this during the last podcast we did where we are using LLMs to design objective functions, and back then, we did it for preference optimization and post training.
And here, we did it for a load balancing of mixtures of experts. Also there, we found that within, like, think, like, even only 20 sort of generations, we were able to sort of explore, let's say, not only a single objective function, but sort of, let's say, a convex hull where there are different trade offs between sort of performance and load balancing and so on. So I think this is another application of Schinka where it's not only basically about sort of finding the best solution, but essentially illuminating a program space where there are always potential trade offs between, like, let's say, for example, runtime and the quality of the circle packing.
Right?
having a system that can explore all of these is important as well. I'm very excited to see you apply this to the ARC challenge. Mhmm.
Like, what what are what are your thoughts about that? I still need to collect results.
So I I don't wanna make any claims, like, or hard claims before having done this, but I would hope that there is some chance of, for sure, improving sort of the the cost of these systems and then potentially even performance. But, yeah, to be seen. Oh, very so you've done some experiments.
Exciting news is potentially coming. I've started looking into it.
Yeah. And, I mean, what what are your thoughts in general about about ARC, though? Think it's great.
it's really important, and I think it fills an important gap. And I do really deeply respect Francois and sort of read the paper when it first came out, and no one thought of actually being able to to get numbers above 10%. Right?
And it's also pretty fascinating on a society level how far we've come since then.
work mode, you can forget where you were one year ago. And then just looking back, it's it's pretty amazing. Also, how far we've come since o one.
It's insane. I I think Francois doesn't get enough credit because it's such a good benchmark, and not necessarily for reasons people think because Francois is always saying that, we need to have a benchmark which is easy for humans and hard for AIs. And and in a sense, that's not quite the case.
I I said when ARC v two came out that it's actually very difficult for humans. You know, there was one task where Dagar was stumped for about fifteen minutes. There was three of us looking at it, and we we just and it's one of those things that depending on your perspective, you might get it straight away or or you might not.
So there's that criticism. And people have said that Arc v three is even harder. Yeah.
You know, but I I think that's rather missing the point. I think he's saying that with with with a lot of these competitive coding problems, the dataset is contaminated. These are problems that have been solved before in part or in whole, which means when you look at the epistemic tree, many of the building blocks for solving them are very high up in the tree.
He's looking at these problems that there is very little dataset contamination, and they need to be solved from very abstract building blocks. You're starting much lower down the tree and you're synthesizing a model by composing together very abstract building blocks, which is the essence of intelligence.
I think for that reason, ARC is really kind of pushing us to build adaptive systems which we could say are intelligent. Yeah. I agree.
I I mean, like, in many ways, I'm I'm really looking forward to the next years and seeing how far we can push this and then also how much generalization we can get afterwards. Because I I believe, like, when you look at sort of the more recent models, they're getting much better at the transform style code evolution or outputting for ARC than they are on the instruction based level. And I think this might already be, like, a small sign of some amount of overtraining on ARC AGI one at least.
Right? I do believe there are some aspects of work which will be automated before it comes to sort of fully science automation and the type of work I'm doing. But I could imagine that certain parts of the dimensions that I deal with every day are for sure going to be hit by AI.
And then the question is, are there gonna be new dimensions opened up that we as humans will fill in? Right? And I think what I said before about, like, shepherding and so on, I really hope that that's the way forward, right, in the sense that humans are the ones steering the ship while just being massively amplified in their productivity.
Right now, I am not really seeing the kind of job market disruption that was being predicted. I know from personal experience that in sense, it's made it very difficult to hire people. Script writers use ChatGPT.
I can spot it instantly. And writers and copy editors are actually in more demand than they were before fixing all of the crap that has been generated with ChatGPT. And there's the cloud analogy as well.
So IT system administrators who were earning £60,000 a year in The UK. They rebranded as cloud DevOps engineers, and they more than doubled their pay. And people are very adaptive.
See new trends, new bandwagons, and they just adapt and they add value on top. And that has been the trend for a very long time. Do you think that AI is going to be so transformative that it will transcend people's ability to adapt?
just a question of speed. Right? So I was talking about sort of cultural evolution and technological evolution, and it seems like we humans, we need more adaptation and more time to to get used to the technology to carve out these niches where we we can fill in and it's complementary.
Right? So first off, I I think we're we're still not at the ceiling of the sort of technological progression. Right?
So maybe in a couple of years, we will need less of sort of slop editing, like you said. But I do think we we need some more time to adapt to the different modalities of interacting with these systems. Right?
I think everyone can sort of interact with chat assistant, But I think this is the most sort of naive form of interacting with AI agents, for example. Right? So, yeah, I think we need to get the pacing of all of this right, and we need to do much more exploration in human machine interfaces, UI, UX design, and, how to make sure that humans sort of feel or feel fulfilled during this experience.
you know, you were behind the AI scientist paper and there's now version two of that. Allow me to be a tiny bit skeptical. You know, we were talking about when we evolve systems to do a to do a particular thing.
And at the moment, it feels like as good as they are, they are still quite parasitic on the instructions and intentions of the human supervisor. So it's very much, an exchange between the humans and the system. Because the implication is that in the future we might have systems that are so autonomous and so open endedness and can figure out valuable things to research that humans wouldn't be needed anymore.
in the world. If I didn't believe that, I would be very worried. I agree.
To me, like the AI scientists like v one and now v two are sort of glimpses into a potential transformation. But I fully agree in order to make really big scientific breakthroughs, like multiple of them, like, every day or whatever, you still need humans in the loop to sort of either seed or guide the direction in which to explore or to to verify, check, and actually, yeah, transfer these insights. Right?
So I think it's not gonna be like all ML PhDs will will be unemployed. It's it's more gonna be a sort of core evolution of humans with this technology and potentially, like, an ideal future for me, like, it will allow humans to focus on what they're really, really great at. Right?
So I think it's gonna be an amplifier of sort of these these latent dimensions humans are great at. Right? I think something that's critical is that we as humans try to interact with these systems as early as possible in order to actually have influence and ownership over this development process.
all of these systems together. And do you think these systems can become incredibly sophisticated such that they are, you know, somewhat detached from humans?
Well, mean, with the AI Scientist V2, we sort of released that one paper that we submitted to an iClear workshop was able to sort of pass the acceptance threshold before meta review. So I do think at least for sort of workshop level contributions, we're we're getting there. While not every submission in AI scientist paper does is or is reaching that threshold, we're we're at the point where we can even talk about sort of noisy review processes and this actually being, yeah, something that as long as you have a large budget, you might get something out of it.
I think going forward for the bigger innovations and so on, for now, you still need humans, but we're sort of at the GPT one moment of of making this sort of a reality and potentially in ten years, this is gonna look very, very different once the sort of also the infrastructure for it has been built up. Right? So there are places like Periodic Labs, right, which sort of now are building like real physical labs with robotic systems to automating automatically sort of execute experiments.
This will take some time, but it is sort of imaginable for sure that as we sort of do RL on these types of systems, and we actually also account for negative results and for actual, like, hypothesis testing. So getting these systems to be a real good hypothesis testers with verifiers in the loop that we might be able to unlock many more capabilities.
Yeah. I mean, I suppose I I don't want to sound like a Luddite. So it's entirely possible that this is just, I don't have the imagination to think about the future.
So it is possible that in the future that these systems might understand very deeply and be creative. Think right now the problem is they only understand things a few levels down in the epistemic tree. So they can do some surface level recombination, and they can discover new things in the basin of things that have already discovered.
But but we understand things very deep down in in the epistemic tree, which means our, you know, our cone of creative potential is is much wider. It's possible that that gap might be closed.
What would happen then? The way how I kind of think about the scientific process is like a tree search ultimately. Right?
So I think a lot of sort of analogies from evolution transfer to scientific research. Right, in the sense that we traverse a tree of different ideas or different experiments, and then in the paper, we report one path through that tree. And I think what I kind of alluded to before, we need much more, like, full tree datasets for training these LLM systems to actually learn how to do this exploration and this foraging, basically.
At the same time, I I feel like evolution will also take place on a cultural level, like, us. Right? We will get better at sort of steering the ship, and I can imagine that in in a future world, sort of the way how we do research will be completely different.
And I'm pretty sure that right now already 99% of machine learning research is done with sort of AI assistance. Right? Think about chat GPT brainstorming, cursor coding, cloud code, etcetera.
In the long run, we're gonna move on that spectrum from sort of with AI closer to by AI and then sort of more high level sort of orchestration and overseeing by humans.
intrinsically coupled to humans is the value function. So one school of thought is that AI will develop a mind of its own and it will, you know, basically transcend humanity and it will just have agency which is not parasitic on on on ours. I personally don't subscribe to that view.
But the other view is that it is like, let's say the AI scientist, you know, like version 10, it's going to be continually epistemic, you know, epistemic foraging. It's going to be finding new things that are useful. And they kind of have to be useful to us.
Because if it finds things that are not useful to us, then we just won't use them. And then nothing will happen. So so do do you think there'll always be a kind of coupled value function to humans?
Jeff Koon had this work on Omni. Right? And using LLMs as sort of amortized notions of interestingness for humans.
Right? And I think ultimately, the way how we train these systems is coupled in in human data. Right?
And going forward, it will also be coupled with human data that is collected using verifiers. Right? So I have a hard time believing that in the long run, when you run this open endedness sort of paradigm with AI scientist agents, it's gonna completely divert to to something that's either fully noninterpretable or unrelated to problems we as humans care about.
Right? And then again, like, can steer to a certain degree where, like, the search happens. Right?
So you can tell the system, okay. Try to do cancer research. Right?
And sort of work on problems that we care about. And, ultimately, like, we are the ones who control how much flops are being pushed into this. Yeah.
let's say, in the world of mathematics, what if, an AI scientist could come up with entirely new problem formulations and then solve them? And these are things that humans had never conceived of before. And maybe they would be less interested in the answer because humans hadn't spent time thinking about it.
And if you think about it, we could just explore the phylogeny of mathematics just to the nth degree. And at some point, maybe we just wouldn't care anymore. Maybe we can just carve out that space just forever and ever.
Yeah. But maybe down the road, there is a stepping stone that enables a new innovation and a different field that we actually care about. Right?
So it's very hard to say a priori whether or not something is interesting or not. Right? Yes.
And there's also the notion of I love this idea of diverse intelligences and diverse minds. And maybe we we could just create artifacts in a space which is completely alien to us. And we might even ascribe moral value to them, and we might not want to turn off, you know, the the power because we we want these alien artifacts to stay alive.
Maybe.
I read a lot of science fiction, but I would sort of shy away from from speculating about all of this. But I do think one thing I'm extremely certain of is that the way how we conduct research and science is going to fundamentally change in the next five years, ten years, and twenty years. And I hope that we're going to be able to sort of tackle some of the biggest problems, which are still sort of seemingly unreachable right now with and by AI.
and it's been speeding him up. It's taking away a lot of the drudgery. But the cynical take is that, and Scott Arison posted something similar as well.
The cynical take is that maybe laziness is is stepping in. And in some pernicious way, using AI models is actually stopping us from thinking outside the box. So it's it's encouraging us to kind of search in the neighborhood of things that are known.
And that is very useful. It's very useful to have an artifact that knows all of the experiments, all of the things that were ever done by people twenty years ago.
their their brilliance, their talent in completely new areas? So first off, it's great that these experts are already using the technology in their day to day work. Right?
And I think it's also important that really, really top level scientists try to push what's capable with these systems or squeeze out where there might be sort of black spots or stuff where you these systems can't do. Second off, I think it comes down sort of to discipline and how we raise sort of the next generation. Right?
So discipline on the personal level, like how much do you just sort of tap accept everything that's being proposed by these systems and responsibility in terms of educating the next generation in the sense that we need to sort of teach our kids that ultimately what comes out of these systems might not always be be true, that facts can be sort of subjective if you will, and that there needs to be more research about what's being given to you.
try to make the best out of. Yeah. The autopilot thing is very interesting because there is a tendency using cursor just to know, at at some point, the models are getting so quickly that you can't even read Yep.
The tokens coming at you and then you just press accept and you press accept. It's the same thing in cars that as soon as you have too strong of an autopilot, you just completely switch off. And then you see a divergence because there's something about thinking that it must be grounded on your path.
There's this path dependence. And when you start becoming parasitized by this other train of thought, then you stop thinking about your path, and then you're not in the driver's seat anymore.
are almost like drugs, right, in the sense that you become addicted, you you use up all your sort of budget, and then you need to load up again. And once you you fully reached sort of the the budget limit, you feel like, okay. What am I gonna do now?
And I think once that happens to you, you should really sort of rethink the way how you work. Right? And to me, right now, there are certain parts where, like, sort of auto accepting is acceptable, and there are certain parts where it's definitely not, and you really need to go deep into it.
And I think right now we're sort of in this weird non equilibrium state where things are moving constantly. Right? So the systems or the models are changing.
The features are changing. The sort of parts where the systems are good is chain are changing all the time, and we humans need to constantly adapt to to that. Right?
And I think it's a big cognitive challenge, and I think we just all need to be aware that there are certain problems and certain challenges that we have to adapt to. I think the best way to do so is just interact with this technology as much as you can and maybe find new research ideas for out of that experience. And how is AI scientist v two different to v one?
In v one, we we used sort of a template based approach. So we had, like, a base experiment. And then for that base experiment, we asked sort of an LLM to generate ideas sort of with semantic scholar calls and sort of literature search.
And then it implemented sort of these ideas based on the template. Right? It did basically code diffs.
And then it linearly executed like an experiment plan and wrote a paper in the end. And so what could happen was that there was an idea and that idea didn't work out. Right?
But then in the end, the paper like, the experiments were still executed linearly and you wrote a paper. And this was already impressive in the sense that it looked very much like like science. But if you think about human sort of science and, like, the scientific method, it's much more like tree search, like I said before.
Right? You sort of adapt what you're gonna execute next, and you sort of refine based on evidence that you accumulated. Right?
So this is sort of the the notion of falsificationism from from Karl Popper. Right? In the sense that we collect evidence for hypotheses and reject we reject others, and we do so in in a loop basically until we we want to publish or we find something.
And we tried to take this notion and directly build it into the agentic scaffolding for the AI scientist v two. So now it's basically like an paralyzable agentic tree search where there's no longer a template experiment needed, but this is drafted up by the LLM itself. And thereby, the AIScientist v two can be applied to many more sort of settings, if you will.
and then write a paper in the end. Go further with the American Express business gold card. Earn three times membership rewards points on flights and prepaid hotels when you book through amextravel.
com. Whether your destination is a business conference or a client meeting, your purchases will help you earn more points for future trips. Experience more on your travels with Amex Business Gold.
Terms apply. Learn more at americanexpress.com/businessgold.
Amex Business Gold Card, built for business by American Express.
Chronic migraine is 15 or more headache days a month, each lasting four hours or more.
A, prevents headaches in adults with chronic migraine before they start. It's not for those with 14 or fewer headache days a month. It prevents, on average, eight to nine headache days a month versus six to seven for placebo.
Prescription Botox is injected by your doctor. Effects of Botox may spread hours to weeks after injection causing serious symptoms. Alert your doctor right away as difficulty swallowing, speaking, breathing, eye problems, or muscle weakness can be signs of a life threatening condition.
Patients with these conditions before injection are at highest risk. Side effects may include allergic reactions, neck and injection site pain, fatigue, and headache. Allergic reactions can include rash, welts, asthma symptoms, and dizziness.
Don't receive Botox if there's a skin infection. Tell your doctor your medical history, muscle or nerve conditions, including ALS Lou Gehrig's disease, myasthenia gravis or Lambert Eaton syndrome, and medications, including botulinum toxins, as these may increase the risk of serious side effects. Why wait?
Ask your doctor, visit botoxchronicmigraine.
or call 44 to learn more.
I'm trying to say this in the most polite way possible. But a critic might say I don't want to use the word slot. But a critic might say, we are producing papers which appear like papers.
So they have figures and they have results and they have things written in a certain way but they're not grounded deep down the epistemic phylogeny which means that they have near the top of the tree we're seeing some novelty and composition happening, but it but it doesn't reflect a deep understanding.
What would you say to that charge? It's for sure that not every paper that comes out of the AI scientist v two is a nature worthy publication. Right?
That that's for sure the case. So definitely there is some amount of, let's say, slop or content that is not like a scientific big discovery being written up by the AIScientist v two. But ultimately, like, we we showed that it was possible to obtain a workshop level paper.
I do think this is sort of the first time basically where we can see that at least now we're able to fully autonomously spend compute, spend API calls to obtain some amount of scientific insights. And for me, at least right now, it's a good way to sort of prototype ideas or to investigate a certain field, get like initial starting point, initial results, and then to to work on top of it.
and essentially produce many more sort of true positives as you will. Yeah. And it might be one of these things, you know, when we moved from, GPT three to GPT four, there was just a massive increase in fidelity.
Because the thing is with with SLOP, to me, it simply means lack of deep grounded understanding. And there's no reason in principle why these things couldn't have a deep grounded understanding. They just don't have it yet.
Yeah. So it's something that could improve over time. But it's likely to improve quite slowly.
And then at some point, we might just think, oh my god. We've got an AI scientist. Yeah.
what we were discussing about before. So first off, there is a verifier in the loop. Right?
Or in the sense that, experiments are actually executed on a computer. Right? So the numerical results can be be fed back or are fed back into the system to come up with the next thing to explore.
But, like, we haven't made, like, a let's say discovery, a residual connections or something that have diffused into everything in machine learning. And I think what we really need is to make these systems be much better at sort of integrating knowledge over multiple experiments and sort of become better at sort of formulating the next hypothesis based on previous insights.
in in an efficient but scaled up way. I'm just thinking that the the first breakthrough discovery, would it resemble the AI scientist paper or would it resemble Cinke Revolve? So for example, we we could do, like, a massively scaled up Cinke Revolve, and we could say, want to discover a new architectural design.
Yeah. And would that happen? And then we would get the AI scientist paper to kind of write it up and do ablations and stuff.
May maybe that would be the the pattern of it. To a certain degree, I've been thinking a lot about how you can potentially even combine these two paradigms. Right?
The AI scientist and and and Schenka or AlphaEvolve style optimization algorithms. And I do think there is some amount of work to be done on sort of this auto verification sort of aspect to it, on the sort of problem formulation aspect to it. The paper writing part is actually the least important about the AI scientist.
Right? It's a form factor that we humans are sort of used to, and it helps anchor our mental model of, like, a scientific discovery. But, ultimately, I'm not sure if the paper is going to be the the knowledge transmission medium in, let's say, twenty years.
Right? Something else I've been thinking of a lot is whether or not we can make papers much easier, agentically accessible. Right?
In the sense that right now it's it's it's a LaTeX document, but you could imagine sort of equipping every paper with sort of several model context protocols so that every figure is reproducible, data is accessible, and essentially make it much easier for the LLM agents to essentially either replicate work or to work off of them afterwards. Right? Doing sort of absolute improvements, ablations yourself through that interface to a paper.
But to be entirely honest, I'm not sure if it's gonna happen because there have been many great ideas for improving sort of, let's say, the the format of scientific artifacts out there, and people still seem to to like the paper format which has existed for, let's say, hundreds of years. Right?
AI agents for scientific discovery. Yeah. Paper is a great human interface.
It's a similar thing with, automated driving. Right? That we could revolutionize the road network to have sensors, and we could dramatically improve the monitoring and observability and optimization.
But I'm fascinated by the idea. You're saying not just reproducibility of the experiments, but also the way that the figures are designed and the code and so on. Because then we could create this huge playground where agents can repurpose, recombine, restudy work that has been published by other scientists.
And it also made me think, like having an automated scientist, does that make peer review more or less important?
makes it more important, at least for now. Right? In the sense that we now have a mechanism or could have a mechanism that generates many, many papers.
Right? And it first increases, like, the workload on on on human reviewers, and we need some effective way for filtering and then essentially only taking the cream of the crop for human verification afterwards. Right?
So I think for now, like, the ultimate verification is still, the human and the diffusion of the result through the community, and we need better tools for doing this automatic filtering and verification. Like, we have the AI reviewer that comes with sort of the AI scientist, but you actually probably need some form of experiment execution for actually verifying everything. Yeah.
But there is, for example, work by OpenAI on on paper bench and trying to go into that direction using sort of LLM soft verification and these types of things.
I'm I'm hopeful that we're going to figure this out in the next years. Yeah. And I think one of the Rubicon moments is when the the new transformers architecture or something massive is discovered by AI, and we're all using it.
new things in science. And it's important to have, work that's openly available. Right?
I think, like, with the AI scientists in Schinka, we're really trying to to make sure that we can sort of apply the collective intelligence of all of us to to shape how this might look in the future. Amazing. Well, Rob, this has been so fantastic to have you on the show.
Sokana is hiring amazing engineers, by the way. So if if this sounds like and it it is an amazing opportunity. Get in touch with Rob and and the guys.
And I trust you're working on some exciting new things that are coming up. Yes. And I I hope to be able to talk to you in the future again about some of this.
Absolutely. Rob, thank you so much for coming. Thank you so much, Tim.
Ever spend all day fishing and catch nothing? That's what happens to hackers when Cisco duos on watch. Every login, every device, every user, protected.
Cisco Duo. Phishing season is over. Learn more at duo.
com.
Shared via Hopper