Vlad Tenev and Tudor Achim from Harmonic discuss their AI system, Aristotle, which achieved International Mathematical Olympiad gold-medal performance using formally verified Lean proofs. They delve into Aristotle's architecture, including Monte Carlo Tree Search and lemma guessing, and explore how verifiable reasoning could revolutionize mathematical practice and harden mission-critical software. The conversation also touches on the metaphysical nature of math, the future of theoretical abundance, and the importance of formal verification for trustworthy superintelligence.
Hello, and welcome back to the Cognitive Revolution. The presenting sponsor of today's episode is Granola. Regular listeners have heard me describe the blind spot finder recipe that I'm using on Granola to look back at my recent calls and help me identify angles and issues I might be neglecting.
And I love that concept, but it's also worth highlighting how granola can help raise your team's level of execution by supporting follow through on a day to day basis. This morning, for example, I had two very practical calls in which I committed to a number of things. In the past, to be honest, there's a good chance I'd have forgotten at least a couple of the things I said I'd do.
But with granola, I can easily run a to do finder recipe recipe and get a comprehensive list of everything I owe my teammates. This is the sort of bread and butter use case that has driven granola's growth and inspired investment from execution obsessed CEOs, including past guests, Guillermo Rausch of Vercel and Amjad Masad of Replit. See the link in our show notes to try my blind spot finder recipe and explore all of the ways that granola can make your raw meeting notes awesome.
Now, today, my guests are Vlad Tenev and Tudor Akim, cofounders of Harmonic, an AI research lab dedicated to building mathematical superintelligence, and also the creators of Aristotle, an AI system that achieved gold medal level performance at the twenty twenty five International Mathematical Olympiad. While OpenAI and Google DeepMind achieved similar performance by scaling reasoning in chain of thought, Harmonic stands out for their commitment to formally verifiable methods. Because it generates candidate proofs in lean, a programming language that serves as a proof checking assistant by using a trusted kernel to confirm that every single step of reasoning follows from a few explicit premises and accepted logical rules.
Aristotle's work can be automatically validated, and its performance is in principle limited only by the scale of compute available for reinforcement learning. In an effort to better ground my own intuitions for mathematical superintelligence, we begin with a metaphysical discussion about the nature of math, what it is that mathematicians do, the assumptions that underpin a lean verification, and how lean is already revolutionizing the math world by eliminating the need for traditional peer review. From there, we turn to the Aristotle architecture that delivered IMO gold performance.
It consists of a large transformer model that uses a Monte Carlo tree search strategy, reminiscent of systems like AlphaGo, to discover valid paths from point a to point b in mathematical reasoning space. Plus a lemma guessing module that helps manage context and keep things on track by generating candidate waypoints between a given starting point and a potentially distant end goal. Plus a specialized geometry module modeled on DeepMind's alpha geometry.
We also discuss the Aristotle API's informal mode, which attempts to auto formalize whatever the user asks it to prove and discuss what its responses to my admittedly silly requests that it prove propositions like all his love and Epstein didn't kill himself imply about the boundary between statements that could in principle be mathematically proved and those which are sufficiently factual or philosophical in nature so as to fall outside the scope of the system. In the final section, we discussed the role of entropy and the importance of taste to Harmonic's future plans, how the community is using Aristotle sometimes on a standalone basis and sometimes in conjunction with other frontier models to solve previously unsolved problems, how we might use systems like Aristotle and Lean to harden mission critical infrastructure and improve complex systems across society, how Harmonic's emphasis on verifiable outputs could create a superintelligence we can trust even in the absence of mechanistic understanding, and what mathematical superintelligence might look like in twenty thirty. On this last point, I have to say, with so many grandiose AI promises flying around these days, from a country of geniuses in a data center to a century of progress in five years to curing all diseases in our natural lifetimes.
It is rare that I am genuinely taken aback by a company's vision for the future. And yet, as you'll hear, Tudor did manage to leave me at least momentarily speechless when he described a future of theoretical abundance in which all physical phenomena we observe have multiple, competing, coherent explanations, which can only be separated by increasingly exotic experiments. If you're like me, you'll find this episode a useful opportunity to improve your intuition for the nature of math, an instructive preview of what's to come as RL continues to scale across the industry, and an inspiring challenge to keep thinking bigger and bolder about the nature and impact of superintelligence.
With that, I hope you enjoyed my conversation with Vlad Tenev and Tudor Akim, co founders of Harmonic. Vlad Tenev and Tudor Akim, co founders of Harmonic, makers of Aristotle, and winners with an asterisk of the IMO Gold in 2025. Welcome to the Cognitive Revolution.
Thanks for having us. Greetings and salutations.
Thank you. So this is gonna be, I think, a fascinating conversation. It's probably gonna be more metaphysical than most of our episodes, but also there's a lot of practicality because what you guys are are doing certainly has aspirations to go beyond the pursuit of mathematical superintelligence.
Maybe just for starters, how do you guys understand what math is? That was something I was really wrestling with in preparing for this. And then, know, that's obviously very metaphysical.
To make that a little bit more practical, what would you say are the core cognitive skills that people that are good at math really develop and excel at? And how does that, how do those skills do when we look at the performance of, like, the Frontier large language models that all of our listeners are familiar with today?
Yeah. Well, look, first, thanks for having us. It's really great to be here.
You know, when you ask what is math, what is it useful for, what are the core kind of skills, it gets, like, one of the core theses of our company, which is that mathematics is reasoning. So a lot of people think of mathematics as this really esoteric thing. You know, you're you're thinking maybe group theory, stuff you've seen in movies like Goodwill Hunting.
But mathematics at its core is the process by which humans understand the world by breaking their understanding down into small sequences of logical steps that other people can understand and verify for themselves. So when you're solving a physics problem or doing your taxes or thinking about what happened at the beginning of the universe, ultimately, you have to have an explanation that is self consistent, that follows from other facts, and that your colleagues or other humans can check. And so when we talk about what it takes to be good at math, the question is what it does take to be good at reasoning.
And so that's, again, that ability to break this down into steps. And it turns out math is really useful for understanding the universe and building lots of engineering things, but ultimately, it's just about reasoning.
I watched your, podcast that you did with Sequoia maybe sixteen months ago or so now. And I recall Vlad's story of like, basically, I thought that if I got good at math, then I'd probably be good at other things, and it sort of worked for me. So that's like one way to, in a very practical sense, unpack the idea that math is reasoning.
It like certainly seems to help people generalize to at least related domains and be really effective, for example, in entrepreneurship. But I'm not entirely clear still on like, are you making a more almost like platonic claim there? It seems like there's the there's the very simple notion that like, okay, should teach my kid a lot of math because then they'll be smart generally.
And again, that works for humans. But is there something that you see as like a more fundamental law of the universe sort of correspondence between what we are doing in math and what we are doing in these other domains. Because it it doesn't seem like we have the same sort of, like, verifiability in almost anything else.
Like, we do have it a little bit in computer science, but even in physics, right, we've got, like, still, like, very fundamental questions about, like, is the paradigm even right or, you know, what would it mean for it to be proven right? Like, I don't think that stuff is at all agreed upon. So maybe you guys throw up your hands at this mystery too, or maybe you feel like you have kind of an intuition for what the the answer is.
Yeah. I can I can give you my perspective? I got into math through physics.
So when when I first came to Stanford as an undergrad, I had read Brian Green's The Elegant Universe, which was sort of like the first popular string theory book. And when I was a kid, one of one of the earliest memories, one of the first like English full books that I read was Brief History of Time by Hawking. So I've always been interested in kind of the big questions, right?
What happened before the Big Bang? Like, how did the laws of physics come about? Is there just like one law, one particle, one force that eventually as the universe cooled and expanded, splintered into all the different forces we have today, like gravity, electromagnetism, strong and weak force.
Because, you know, before, back in the day, that was not obvious. You know, we thought electricity was separate from magnetism and it was just like a really big I probably one of the greatest achievements of science figure out that these two are actually, like, two sides of the same coin, really. And then and then the big question is, like, what's going on with gravity?
Is it is it the same? Right? And, you know, in in the mid in in the middle of this, we found out that, you know, the the weak force and the electromagnetic force were also splintered off of one electroweak force.
So, it kind of feels like there was just one thing at the beginning and we have to understand what that thing is. And what I found when I became a physics major at Stanford and I started asking all these questions, eventually, they'd send me over to the math department. And they're like, well, you know, if in order to understand string theory, you have to understand all of these other things.
Right? And if you wanna understand general relativity, you've got to get into differential geometry. And so that's how I became a pure math major and ended up doing a PhD.
The impetus was actually trying to understand the real world through physics. And if you think about what's the usefulness of physics, I mean, all of the big inventions that humanity has that really push us forward are kind of like physics inventions really. I mean, when when you think about flight, rocketry, computers, transistors, GPS, obviously, one of the main examples of why relativity is useful, they're physics things.
So the the real reason to do math is math is interesting and beautiful, and there's an art aspect of it, but it helps you it helps you understand physics. Physics helps you understand engineering, and then you can, like, create things that have huge value.
You know, you you were asking, how does math work in other fields where things are not as precise? I think math shows up just maybe a little more subtly than people think. So there was this physicist Eugene Wigner that wrote a famous essay called the unreasonable effectiveness of mathematics, which was commenting on a really interesting phenomenon.
So Vlad mentioned differential geometry and special relativity. It turns out that when Einstein was creating that theory, he relied on these thought experiments from the nineteenth century around how to think about certain manifolds and properties of them. And that was actually the key tool that we used to explain what special relativity is and then develop on it for general relativity.
And that's a case where I think it's a perfectly representative case because those thought experiments in the nineteenth century were almost preposterous. It made no sense to think about them because how could you possibly apply these concepts to the real three-dimensional world? And then it turns out that it's very useful for understanding the four dimensional world when you include time and curvature.
And there's there's myriad examples like this. If if you consider number theory, for a long time that was really seen as an incredibly esoteric branch of math, no practical implications, but people really pushed on that theory for a long time. And then it turns out that that's the key tool you need to create a secure digital economy.
So now essentially all of human civilization has a digital economy, which is based on this branch of math. So I think it's almost the wrong question to ask. Well, I don't know.
There's a lot of math out there. How is it useful? The point is you just do the math, and then eventually, some of it, not all of it, will be more useful than you possibly could have imagined.
So the investment in math is it's it's not just to build a a really smart system. It's to create a lot of new math that we can then figure out ways to apply later.
One interesting thing that that the conversation reminded me of when when you first asked, okay, what is math? What does it look like? I think one of the reasons we got excited about applying AI to this domain is there's lots of different things that mathematicians do.
Like, some of them are very creative, almost like artists. Right? And, you know, may maybe they're not super prolific, but they come up with something new once every five to ten years, and and that can be just like an amazing accomplishment in the field.
Yeah. Like a Gregori Perelman, for example. Others are just machines and they just can read more papers and comprehend more papers per unit time than other people.
And what they're doing is basically like synthesizing all the knowledge, figuring out all the tricks, applying those tricks quickly to new domains, and they're kind of like reusing these things. And I think we have, we're very excited about the prospect of AI accelerating the former, we think that'll happen, but the latter is something that AI is already really, really good at today and was good at to some degree when we got the idea for Harmonic. You know, you look at GPT four, which had just come out when we started, and it excelled at just, like, pulling information, you know, doing these types of needle in a haystack type things of can you just like really quickly go through all the literature and and pull things that might be relevant?
And I would say, even if like, you can be an amazing mathematician if you're in that category, and I think a lot of the work could be accelerated if you just like knew all the math that was being done and could pick out the relevant things to an unsolved problem that you have at hand. So I think the the problem itself lends itself really well to what AI is already good at.
Hey. We'll continue our interview in a moment after a word from our sponsors. One of the best pieces of advice I can give to anyone who wants to stay on top of AI capabilities is to develop your own personal private benchmarks.
Challenging but familiar tasks that allow you to quickly evaluate new models. For me, drafting the intro essays for this podcast has long been such a test. I give models a PDF containing 50 intro essays that I previously wrote, plus a transcript of the current episode and a simple prompt.
And wouldn't you know it? Claude has held the number one spot on my personal leaderboard for 99% of the days over the last couple years, saving me countless hours. But as you've probably heard, Claude is the AI for minds that don't stop at good enough.
It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. And with Claude Code, I'm now taking writing support to a whole new level.
Claude has coded up its own tools to export, store, and index the last five years of my digital history from the podcast and from sources including Gmail, Slack, and iMessage. And the result is that I can now ask Claude to draft just about anything for me. For the recent live show, I gave it 20 names of possible guests and asked it to conduct research and write outlines of questions.
Based on those, I asked it to draft a dozen personalized email invitations. And to promote the show, I asked it to draft a thread in my style featuring prominent tweets from the six guests that booked a slot. I do rewrite Claude's drafts, not because they're bad, but because it's important to me to be able to fully stand behind everything I publish.
But still, this process, which took just a couple of prompts once I had the initial setup complete, easily saved me a full day's worth of tedious information gathering work and allowed me to focus on understanding our guests' recent contributions and preparing for a meaningful conversation. Truly amazing stuff. Are you ready to tackle bigger problems?
Get started with Claude today at claude.ai/tcr. That's claude.
ai/tcr. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.
ai/tcr. AI agents may be revolutionizing software development, but most product teams are still nowhere near clearing their backlogs. Until that changes, if it ever does, designers and marketers need a way to move at the pace of the market without waiting for engineers.
That's where Framer comes in. Framer is an enterprise grade website builder that works like your team's favorite design tool, giving business teams full ownership of your.com.
With Framer's AI Wireframer and AI Workshop features, anyone can create page scaffolding and custom components without code in seconds. And with real time collaboration, a robust CMS with everything you need for SEO, built in analytics and AB testing, 99.99% uptime guarantees, and the ability to publish changes with a single click, it's no wonder that speed, design, and data obsessed companies like Perplexity, Miro, and Mixpanel run their websites on Framer.
Learn how you can get more from your.com from a Framer specialist or get started building for free today at framer.com/cognitive and get 30% off a Framer Pro annual plan.
That's framer.com/cognitive for 30% off. Framer.
com/cognitive. Rules and restrictions may apply. Okay.
That's that's quite helpful. I think coming into this, I had focused my own mind on sort of two modes of math, I guess. One being the kind of Einstein like, obviously, that's a, you know, a high level example of a kind of eureka moment of having some insight that, hey, this highly abstract and, you know, seemingly perhaps, like, very esoteric formalism can actually unlock, like, major understanding.
That's kind of amazing, very amazing. And then there's also this sort of grind it out, like, I've got this thing that I wanna prove, and I'm gonna kind of perhaps, like, kind of stumble my way even through the space of, like, possible logical moves until I finally chart a path there. And then you're adding another a third layer, which is, like, problem selection in the first place, which I guess is pretty related to the Einstein thing, but but certainly distinct in some ways.
Let's take a minute before we get into the Aristotle system and how it works and how you've trained it and all that stuff to just talk about, like, lean. Lean is basically a programming language that does this kind of very bit by bit logical maneuvering, right, where you have kind of certain assumptions coming in, you're gonna take these various steps, and the goal is to get to a certain outcome. Tell us give because I'm just learning about this, you know, in the context of preparing for this and a and a couple other podcasts, and I think most people don't know anything about it.
So maybe give us a little bit of a more intuitive understanding of what Lean is. And I'd I'd be keen to understand it on a little bit of a practical level too. Like, how many symbols are there?
How many axioms are there that we're starting with? How many rules are there that we can apply? Like, how big is the space that we are, like, manipulating our way through?
Lean, in my view, is the best programming language ever created. So in Lean, you can write any program you would write in Python or c or c plus, but you can also express essentially any logical concept. So if we're okay getting into a bit of the details, it is a dependently typed programming language, which means that at compile time, you can express very complicated properties of the program that you can check before ever running it.
So on the one hand, you have on one end of the spectrum, you have something like JavaScript where you can check basically nothing, and then on the other hand, you have Lean. But the really cool thing is you asked about axioms. So when Aristotle produces any output, it's produced as annotated lean code.
So there's programming language lean. We write theorems. We write programs.
We prove things. And there's a lot of comments explaining to the person reading it what it's doing. But when we talk about proving things, you end up relying on three axioms in addition to just the basic concept of the calculus of constructions, which is what the programming language is based on.
So two of them are extremely technical. One is propositional extensionality. One of them is something about quotient soundness, but the third one is the axiom of choice.
And just as an example to show what an means, so the axiom of choice, it's not saying anything that would be controversial. It's saying if you have a non empty set, it's possible to choose an element from it. And so from these three extremely basic axioms, it turns out you can build all of mathematics, all of computer science, all of mathematical modeling and physics, economics, stats, biology.
It's all based on this core set of axioms. And so the goal of a system that outputs lean is to find interesting statements and programs then prove things that just depend on these axioms, and that that's really where the difficulty lies. As you alluded to, sometimes you have to make big logical leaps.
Sometimes you have to grind through a lot of math, but both of those are essential. So you you can't really skip one of those steps. But the lean itself is it's just it's just incredible.
You can express so many ideas in it. You can prove so many things, and you can use it as a programming language too. So it's it's really up there for me in programming languages.
I started playing with Lean, you know, when Tudor and I started making a plan for this business and we had a pretty early decision about whether we wanted to go formal and informal. And one thing that struck me about it is, as a former mathematician, I barely used the computer when I was doing math, right? I mean, I was in my PhD in the late two thousands and the only time you'd really be using a computer when doing math is when you wanted to like type up your homework or your research paper or something, but all the thinking about it would happen on a chalkboard or a whiteboard.
All the collaboration about it would happen in person at these conferences or on a chalkboard in one's office. And for a while, it was just like maybe mathematics would always be this pure thing that would just be kind of untouched by technology. But what Lean has done is it transformed mathematics from kind of like chalkboard and couch to now it's in Versus Code, you know, you can do it in cursor, you're putting your math on GitHub where now you can run these large collaboration projects.
So, even when you subtract out AI, I think the Lean by itself without AI changed how people do mathematics because now you're seeing extremely prolific famous mathematicians running these large projects where they're collaborating with, you know, dozens of people around the world trying to do things like formalize research or formalize the proof of Fermat's last theorem. And more and more of the folks are adopting lean as like an accelerant. So I I think it's changing how mathematics is being done and and actually accelerates collaboration, accelerates progress, and sort of like removes this notion of peer review.
If you're a mathematician and, you know, you wanna prove something, a big part is getting someone to read it and actually spend the time to tell you if it's correct. And so, you know, you you have the proof of Fermat's last theorem, which took many many years to be proved and and what happened was sort of this collection of people got together and when they all agreed that the proof was complete, it it was it was sort of like ordained that that the thing was proven. And I think another thing formal does is it makes it so that that's unnecessary.
Like, if the proof checks and there's no the caveat that there's no bug in the lean kernel or how you've set up the statement, you you obviate the need for manual human verification. And the implications of that are pretty interesting too, right? You have all these potential citizen mathematicians who now with AI can solve unsolved problems, and they don't need to get anyone at a PhD program, a lean institution interested in their problem in order to tell that it's correct.
They just have to have the lean certificate and the proof is correct. So, yeah, I think that's a powerful thing. If you think about journals, journals and math exist for this.
It's like the prestige of the review board tells you whether you should read something or or trust it. So I do the the notion of trust is really changed fundamentally with with tools like Lean.
Yeah. And I I think the the open source software community has really solved this problem a long time ago. So if you go on GitHub, one can simply open a pull request on some repository.
And if it passes the tests and the author of the repository agrees with your style, they gets merged. So now you've contributed. So that element of trust is not so present.
You know, you can just run the tests. And also when you talk about impact and prestige, you can look at the number of stars you have. So if a repository is very popular, it gets forked a lot.
It gets a lot of stars. So you've disintermediated, you know, essentially any gatekeeper here. It's totally open source.
There's no more trust required, and there's a measure of impact. And so I think math is gonna start going the same way. Previously, mathematicians relied on, you know, their social networks to figure out, okay, who tends to do the right thing, who tends to not make mistakes.
But with Lean, you can have a big math project. Anybody can come and contribute a proof. And if Lean accepts it, then it's right.
And if a lot of other mathematicians start to depend on that result, we're gonna notice a lot of forks, a lot of dependency graphs, a lot of stars on it. And so then you start to measure the prestige that way. So it would it would be very interesting if Lean is the one tool that allows you to go from kind of the cathedral style of development where, you know, very close networks, etcetera, to more bizarre style development where it's kinda wild west, but lean is like the computational certificate that everything is correct.
I wish I understood a little bit better or had a had a more intuitive sense for, like, what exactly is going on with lean still? This is gonna be hard, I think. But, you know, in doing my kind of research, one thing that stands out is the kernel is really small.
So, like, in terms of what you need to trust, it's a pretty small amount of core code that, you know, has been thoroughly vetted many times by many people. So there's kind of that level of understanding. I think it would I would still love to have a little bit better sense because when you've mentioned, like, the three axioms, for example, it's a little weird for people outside the field to be like, oh, there's two that are, like, kind of bizarre and technical.
And then there's this one that's like, you have a non empty set, you can choose an element from it. I'm like, that seems like common sense, but why was that ever controversial? Is there a way to describe the the sort of space of legal moves in math or in lean in sort of I don't usually like analogies, to be honest.
I often try to set this up as an analogy free zone. But because I'm I think I and a lot of others are gonna struggle with the very literal understanding, Maybe this is a time for an exception to my no analogies rule. Is there a is there sort of like a I don't know, like a chess analogy or something where you could say, like, here's the pieces and here are the legal moves that you can make to kind of give people a little bit of a better sense of, like, what it actually means to, like, move through these spaces?
I think the chess example is perfect. So a theorem in lean is something like given this starting configuration of a chessboard, it is possible to get to this configuration. And a proof of this theorem would be listing the sequence of moves.
And what the kernel is doing in lean is saying for every single move that you claim is valid, it's checking, hey. Does this rule exist in my rule book? So the theorem says you can get from a to b.
The sequence of moves is, okay. Here's the sequence. And the kernel is just saying, This step is right.
This step is right. This step is right. And now I've confirmed that I've ended up in the target state.
So Lean is doing that, but, of course, the individual steps are different. They're mathematical steps, and they depend on one or more of these three axioms. It's the the three axioms, although they're technical, they're very short.
So if you write them down as mathematical statements, they're under I think each of them is under a tweet in length. Like, the axiom of choice definition in lean is maybe 10 characters, and the other ones are maybe a 100. So they're not very complicated.
They're just a little bit annoying to, write in math. And then people say, okay. Well, if we assume these axioms are true and they're also common sense, just like a bit more complicated, and we've checked every single step against those axioms, then we say the whole proof is correct.
And could you give, like, a few examples maybe of, like, the pieces and the moves?
I'll give mathematical a but but simpler example of a primitive. So let's consider first order logic. So the deduction rules you have are if a, then b.
So let's say you have a proof that says, if I have a and I know if a, then b, and if b, then c, the theorem says c is true. And the proof of that says a is true, and I have if a, then b, which means b is true. And then I have the step b is true, and I know that if b, then c.
And then I can conclude that c is true. So this is a first order logic, so it's not quite the same as what we're talking about in lean. You can do more advanced types of logical statements there.
But ultimately, that's what's happening. I think it's gonna be hard to essentially, the next step beyond that is just getting to lean and the calculus of constructions and these axioms.
one thing when I learned it, there's actually people are also exploring use of lean to teach math. And I think now it's sort of like practical at the high school level, but you could see a world where it extends to like, middle school and maybe even younger if someone's precocious enough. But I think mathematics education will go from sort of like the chalkboard to the computer lab.
So, there's this thing called the natural number game where you learn lean by deducing properties of like multiplication and addition, basically. So, for example, the commutative law, which is basically that, you know, a plus b equals b plus a, right, or the distributive law, right, where a times quantity b plus c equals a times b plus b times c. So you can sort of like discover and prove these fairly basic facts just using the core axioms and and the lean language.
So that's a good way, you know, if if anyone just wants to like, Alright, what is this lean thing? Why is it useful? But I'm not a research mathematician, dip your feet into it.
I think I would recommend that and that's been extended to harder things too. I think there's like the real analysis game now, which is if you wanna learn real analysis, it's very proof based and it's basically the foundation of calculus. You can start with like basic facts about what's a sequence, what's a real number, how many of these numbers are there, how big are the sets, and then you can kind of keep, proving more and more complex things.
That's a great tip. I'm definitely gonna bookmark the real numbers game and, see if I can can get my Yes. Soon to be seven year old into it.
Hey. We'll continue our interview in a moment after a word from our sponsors. Want to accelerate software development by 500%?
Meet Blitzy, the only autonomous code generation platform with infinite code context. Purpose built for large, complex, enterprise scale code bases. While other AI coding tools provide snippets of code and struggle with context, Blitzy ingests millions of lines of code and orchestrates thousands of agents that reason for hours to map every line level dependency.
With a complete contextual understanding of your codebase, Blitzy is ready to be deployed at the beginning of every sprint, creating a bespoke agent plan and then autonomously generating enterprise grade premium quality code grounded in a deep understanding of your existing code base, services, and standards. Blitzy's orchestration layer of Cooperative Agents thinks for hours to days, autonomously planning, building, improving, and validating code. It executes spec and test driven development, done at the speed of compute.
The platform completes more than 80% of the work autonomously, typically weeks to months of work, while providing a clear action plan for the remaining human development. Used for both large scale feature additions and modernization work, Blitzy is the secret weapon for Fortune 500 companies globally, unlocking 5x engineering velocity and delivering months of engineering work in a matter of days. You can hear directly about Blitzy from other Fortune 500 CTOs on the Modern CTO or CIO Classified podcasts, or meet directly with the Blitzy team by visiting blitzy.
com. That's blitzy.com.
Schedule a meeting with their AI solutions consultants to discuss enabling an AI native SDLC in your organization today. Everyone listening to this show knows that AI can answer questions, but there's a massive gap between here's how you could do it and here, I did it. Tasklet closes that gap.
Tasklet is a general purpose AI agent that connects to your tools and actually does the work. Describe what you want in plain English. Triage support emails and file tickets in linear.
Research 50 companies and draft personalized outreach. Build a live interactive dashboard pulling from Salesforce and Stripe on the fly. Whatever it is, Tasklet does it.
It connects to over 3,000 apps, any API or MCP server, and can even spin up its own computer in the cloud for anything that doesn't have an API. Set up triggers and it runs autonomously, watching your inbox, monitoring feeds, firing on a schedule, all twenty four seven even while you sleep. Wanna see it in action?
We set something up just for Cognitive Revolution listeners. Click the link in the show notes, and Tasklet will build you a personalized RSS monitor for this show. It will first ask about your interests and then notify you when relevant episodes drop.
However you prefer. Email, text, you choose. It takes just two minutes, and then it runs in the background.
Of course, that's just a small taste of what an always on AI agent can do. But I think that once you try it, you'll start imagining a lot more. Listen to my full interview with Tasklet founder and CEO Andrew Lee.
Try Tasklet for free at tasklet.ai, and use code cog rev for 50% off your first month. The activation link is in the show notes, so give it a try at tasklet.
ai.
And we haven't really talked about Mathlib, but the Lean Kernel is quite small. But there's an open source project called Mathlib, which you can kind of think of as the largest digital repository of mathematical knowledge. So all of like the a lot of the famous theorems and results can be found in MathLib and those give you almost like additional complex moves or algorithms to prove your thing.
So you can apply a theorem and it's almost like applying a function from a library and that can help you get to the goal. So MathLib you can think of as an amalgam of the knowledge that's written in English or Russian or German and all the textbooks out there and it's consolidated and lean and available open source. You can just browse it in GitHub and you could kind of browse it by category.
There's like algebra, real analysis, geometry, statistics. So so part of doing this is actually building this great repository that people are then building and, you know, proving novel results on top of.
Yeah. I think people can understand what it is better. Just think of it like every math textbook in the world merged into one in a self consistent way.
So eventually all of mathematical knowledge will be in this one repository. And if you hit build on your computer, you're gonna be able to check it all from the foundations. And if you have any question about any math concept, just search for it.
You click on go to definition. You can jump around. It's really going to be the new foundation for math in the future.
It's pretty exciting.
is certainly going to change fundamentally, like how it's done, how fast it moves. And I think to a large degree, it already has. And AI is just going to accelerate it and, you know, the great thing about our timing is Harmonic really started when both of these things matured to a level of capability where you could start doing interesting stuff.
Lean basically went from being essentially beta software, like not appropriate for real mission critical use case, which was version three, Lean three to Lean four and that was about the same month we launched the company. Also, GPT-four, which you're starting to actually see glimmers of it being really really good at synthesizing information and the starting points of reasoning that came out at around the same time. And I think both of these mature to the level where you can start putting them together and and doing really cool things.
And I think we're just the the first to see that and and that's how we came up with this concept of mathematical super intelligence, which which really is means the combination of of formal verification and formal tools with artificial intelligence.
Funny story. As I was using Aristotle a little bit to try to wrap my head around all of this, you know, I I don't have the sophistication to pose any really interesting problems. So, one challenge that I gave it was to prove that two plus two equals four, and then I had to laugh when it came back just citing something from MathLib that was like, this is already proved in MathLib for the theorem is literally like the two plus two equals four theorem.
So I was like, it's done. I was like, yeah, that's not exactly what I was looking for, but I guess I kinda got what I deserved there for asking it such a such a basic question.
were you using the web interface or the terminal UI?
I started by having Cloud Code install the terminal and then was using that a little bit. And then somehow it tipped me off to the fact that there was a web interface, so after that I moved over to the web interface.
Yeah, that came to last week. They're probably a little bit more appropriate for those types of questions. I think we wanted to roll it out for on on terminal because I think it makes it a little bit more clear what the tool is great at.
I mean, lots of things can answer two plus two equals four. But Even I can answer that. Even the calculator.
Yeah. And I think I think for a while, we were talking about, like, how do we describe this this like, what Aristotle is. I mean, it's it's kind of like an amazing calculator where you can imagine you could just talk to your calculator.
So it has it has both the reliability, like you know if your calculator gives you an answer, it's correct, but it's not very expressive. At the same time, you know that, you know, something like ChatGPT or Claude, very expressive, but sometimes you have to double check its work because it doesn't always, you know, it doesn't have the verification. But really the intent is to put those together and it turns out that the things the first things that people really want to be sure about and to verify are, like, more complicated things.
So so I think the you probably found this out, but the complicated things I think is where you really start to have moments when you're using it.
Yeah. Let's get into Aristotle and I appreciate the time spent in remedial education. I think it's beneficial not just for me, but hopefully everybody will now be able to kind of grok what we're about to get into much better with that that foundation that we've laid.
So Aristotle has three core parts. I'll just kind of sketch them, and then you can, you know, give me the double click on them. First, there is this Monte Carlo tree search type thing.
I kind of think of that as sort of an AlphaGo like structure where we are systematically exploring the space of moves. I guess that's where I got the chess analogy, right, is that I kind of was making this equivalence between Aristotle, at least that part of Aristotle and AlphaGo. And so it's kind of, well, maybe I can make this move, and then there's this learned scoring function that's like, okay, does that move seem promising?
Does this path of you know, does this branch of all possible moves that I could make, does it seem promising? Do I seem like I'm getting closer to my goal? And with that, you can kind of grind things out, run deep tree search.
Right? The second part, in some ways to me, jumped out as even more interesting and kind of I really wanna dig into the metaphysics of it a bit because this is the lemma based informal reasoning system, which I take to be sort of saying, okay, if I have some really big mountain to climb and it's maybe so big that I can't just grind my way, maybe becomes impractical to grind my way through all these small localized steps, it's sort of guessing, like, what's the base camps that I would want to get to along the way that are, like, the really good waypoints such that if I can get there, then I know I've made it somewhere. But that's really interesting because it sort of strikes me a little bit more like a like it behaves it seems like a little bit more like a language model where it's kind of guessing and not so formal.
Mean, it says in the in the technical report that it is an informal reasoning system. And then there's a third part, we maybe don't have time to go as deep on, which is specifically dedicated to geometry. And in in the technical report, you described that as being like alpha geometry, which I think DeepMind developed.
So correct any misconceptions that I have there and give you the double click on what, like, what more I should understand about how this thing works.
Sure. I I think I think you you covered the components pretty accurately. So one thing I have to say is that, you know, we revamp our systems pretty often here.
So I think Aristotle now looks quite different than Aristotle for the IMO. You know, I think a lot of things here consolidated and improved. I think that you made this point about the Monte Carlo tree search being more of a grinder.
I wouldn't quite characterize it that way. So the Monte Carlo tree search is actually doing a lot of inference on its own about high level steps. So the lepas that we're talking about, they're much closer to solve a challenging math problem than they are to prove that two squared equals four.
So there's a lot of reasoning that goes into them. In some sense, it's grinding once you get low enough in the search tree because you're just, you know, closing out cases or easy subproblems. But it's it's really solving harder problems on its own.
And so when we combined it with the informal reasoning system, you could almost think of it as a form of context management, actually. So, ultimately, you need to end up with a lean proof, and that's gonna involve big steps and small steps. And it's helpful when you're focusing on the smaller steps and don't have to remember the entire context of the bigger steps.
And so it turns out the informal reasoning system itself actually makes enormous quantities of mistakes. So one should not think of it as, oh, it's a really smart human that's laying out the steps to base camp. It's more like a system that can propose lots of things that are wrong and don't have to be formalizable or even correct, and you kinda try to assemble things from that.
So you can think of both of them as kinda doing the same thing, just at slightly different scales and complementing each other. And they're actually all LLMs. So as we described in the tech report, the tree search itself is driven by language models.
And part of the language model is proposing steps, part of it is scoring steps, but they work in concert to solve the lemmas and then eventually the full the full problems. And as you mentioned, alpha geometry, it's a slightly different system. We're, you know, we're exploring kind of high level steps and then trying to use a algorithm to grind through the rest of it.
I think if we're talking about systems grinding through a lot of math, I I would say alpha geometry and the deductive reasoning system is really a grinder. So it's really trying to every possible conclusion of a geometry diagram.
going on there. Yeah. And that's because geometry, if you think about it, it's more constrained.
You basically, you have points. You can basically start with points. If you have three points, there's only so many angles involved.
Obviously, if you go to, like, 10 or 15 points, things blow up pretty quickly, but it also then becomes hard for humans to solve. I think that's why geometry was among the first class of competition problems to fall to AI and automation. I think there's also a couple of other components that might seem simple, but are nontrivial that Aristotle, the system does and are independently improving.
One is auto formalization. So, taking input that you provide in natural language and faithfully translating it into lean in the best possible way. And I I think, you know, relative to our competitors at least, I'm not aware of anything that's as good or bad as as we are.
And also theory building, like sometimes in the in the in the way of solving something, you have to create new theories and new structures that might not exist in in Mathlib and Aristotle has the capability of actually building that on the fly and incorporating that into the proving process.
Another funny anecdote. So that you're referring to what I discovered is informal mode, right, where I can provide I think real users would not do this, but, like, you can provide anything. Right?
Any natural language input. Just something that the system will then try to prove. I asked it to prove all is love, and, it came back and said, this is a philosophical statement and outside the scope of the lean for colonel's ability to prove.
I also asked it to prove Epstein did not kill himself, and it, came back and said, this is a statement about current events. And, it's sort of outside the Lean4, ability to prove. But, yeah, I think this I mean, this does kinda get back to this sort of metaphysical question that I find so perplexing around, like, that translation from the messy real world of human affairs and intuitions to the formal definitions of, okay, this is actually the thing that we would want to prove.
I did find that very, very interesting that you had such a thing at all. And I I guess well, do you have a sense for I also do wanna get into a little bit more details of just, technically, you know, how you created the models and all that stuff. But, you know, on my spectrum from, like, two plus two equals four to all is love, is there how do you think about the intuition for, like, what the boundary is of what is inside the what because because I, again, thought when in listening to your previous interview with Sequoia folks, it seemed like you had the sense that eventually as the system and systems like this get capable enough that more and more things that are of interest to everyday people will start to become the sorts of things that they can do.
So, like, how do you think about that boundary, and and how does that boundary expand over time?
I think the the ultimate boundary of a system like Aristotle is in reasoning through any problem where people can also agree on what does it mean to be a valid sequence of reasoning steps. So right now you have math. That's one obvious one.
When we talk about mathematics being the same as reasoning, that chess example you gave is a perfect one. So you can express the logic of a chess game and then check it, right, and then reason about it. I think one one area that's really gonna touch a lot of people's lives is it turns out you can use the same reasoning approaches to think about software.
So when people write software, they write these things called unit tests and integration tests, and it's kinda having the computer just run the program and check the output against what they expect. But that's what they do after they've written the code. It turns out that when engineers are writing code, they're thinking logically.
Okay. If I have this range in my input, I can think, okay. As I go to this for loop at these if statements, it implies certain things about the output.
And that itself is logical and mathematical reasoning. And so we're starting to see API users reason about programs in the same way that they can reason about math. People are writing cryptography implementations and then checking, hey.
Is there any possibility that two inputs might give me the same output, which would be violating a certain principle of the crypto algorithm? Or they might be implementing a controller for an autopilot and saying, is there any sequence of inputs for which I'll have an unstable dead zone or something? So I think the same kind of reason it was good for math will then go to software and help take us to a bug free software future.
Now Vlad and I disagree a little bit. It's not clear to me if we'll be writing history essays or something. You know, maybe there is a way to value them objectively.
But I I I think the boundary is really in anything that's quantitative and logical in nature.
Yeah. I think in the first version of Aristotle, it would actually formalize and build a theory for your all is love example, and it would give you a correct proof that it's probably true. And I think it surprised us, people were asking all sorts of questions.
We had, you know, we have people asking biology questions, and, you know, we've had people update and ask medical questions to it. Of course, economics and financial math, and Tudor mentioned computer science. So I think it's actually surprised us how broad of a set of things it can successfully create a theory around and formalize.
And I think the constraints we put were just, know, when when you're building a product, you wanna make sure that you deliver value and at this point, I don't think we we provide the most value if you wanna write a history essay. Right? So so, we're trying to like nudge people to the point where they can discover what Aristotle is really really good at as easy as as quickly and simply as possible.
And I think over time, you should expect that the surface area increases and, you know, we start formalizing things and I don't think it's inconceivable that at some point it pulls like current events and news from the internet, out the axioms, right, puts out the axioms and can sort of like fact check and and make conclusions based on real world events. Not our focus right now, but I don't think that I don't think it's a crazy thought. I mean, I I ask a question sometimes.
I'm I'm interested in astronomy. Right? And I wanted to know when's the next full solar eclipse that I can see from within 50 miles of Palo Alto, California.
And the models usually struggle with this type of stuff because nobody's asked that identical question out on the Internet, so they can't pull it. You actually have to do some math. So you you can imagine there's a spectrum and there are questions like this that a model that can reason actually from first principles is going to be way better at.
Okay.
how your experience, lessons learned, etcetera, kind of relate to some of the the live questions more broadly in the AI space. I think you can take on faith that folks listening to this show will be familiar with things like reinforcement learning from verifiable rewards and stuff like that, and certainly understand kind of how the ability to generate synthetic data, you know, feeds into a system like that. And that's, I'm sure, part of what you're doing.
But what more can you tell us in terms of, like, would it make sense to start training something like this from some off the shelf pretrained model, or does that messiness that those, you know, LLMs start with corrupt or pollute your your the purity of the mathematical reasoning too much? Can you tell us anything about size of models, which could be parameters, could be tokens, whatever? I'm interested in things like, also, is there any role for taste in this process?
Obviously, like, mathematics mathematicians are very interested in correct proofs, but they're also interested in these eureka moments and the sort of sense of elegance of the proof. Right? There's a there's a sense of the beauty that, you know, matters as much, I think, to many people as the correctness or maybe not as much, but, you know, it's certainly heavily weighted.
And then I also noticed there's test time training that's part of this, and I think that's, you know, a huge trend that I'm kind of watching in general. So Yep. You know, you can swing or or, take any of those pitches, but, what do you think are are kind of the the most interesting next level of depth that people can use to inform their own AI worldview with?
Well, first, have to say that if your audience knows about reinforced money from verified rewards, you've got a great audience. That's not Authentic data. Yeah.
So I think that is a safe assumption.
Nobody was talking about that stuff. Right? It was like science fiction almost.
consciousness. I I wanna I wanna address the taste question because that that actually, you know, strikes at a key thing that, you know, companies can decide on. So we get gold at the gold performance at the IMO.
We have a very powerful system and it was obvious we had to give it to people. And there's two ways you can do it. So, one way is you can say, well, we're gonna keep this in house.
We're gonna recruit some great mathematicians to come in house and work in secret on problems. And that as they make progress, we say, well, Aristotle's now done x and y and z. That's one way of expressing taste in the research map.
The other way, which we ultimately decided to do, and we think it's been great for the community is we said, well, we're not gonna be the ones to decide what's important in math. We're gonna make aerosol accessible to everyone. And so we open up the API, the web interface, there's a lot of great features coming.
And then in this scenario, taste is expressed by the community by the revealed preference of what they submit to the API. So we don't choose what kind of math they do. We're not saying, hey, Navier Stokes is more important than p percent p.
It's the mathematicians that have the credits on the API to say, well, we care about x or some other thing. And that's why we've seen so much interest in computer science and crypto and certain branches of number theory. And for a while, there are people doing a lot of interesting conjectures and graph theory on the platform.
And I think that that's actually the the right way for companies to engage with the community. You know, you open the system and you let the people decide, you know, where they wanna allocate those compute resources. So, I think that's an important decision.
We've come on one side of it, but I think that's the right long term approach.
are we headed for a future where the AI labs themselves are gonna generate all the discoveries. Will cure for cancer or diabetes look like a giant AI lab with a two gigawatt data center just churning on this problem and then, you know, it comes out and they capture all the value or does it look more like millions of people empowered with these tools working independently and collaborating and, you know, in that world, they'll they'll get like the the credit and they'll get the value will largely accrue to them. And I think we believe that the second world is more interesting likely and the first one is rather dystopian and less likely.
And I think we noticed that because when we rolled out Aristotle, you know, we had one view of what people would use it for, but then we started getting all of these, you know, Erdos problem results and things like that. Yeah. And it's like, we're not gonna run on all the Erdos problems.
We're not gonna do like computational learning theory formalizations in house. So I think the amount of cool things being done with it just like explodes if you put it if you make it generally available. So, I think it's it's not only right from a business strategy standpoint, but also like I think the the world that we built, assuming this path, I think, is a is a better world that I would like to I would like to live in.
So that speaks to taste in terms of problem selection, but I was also just thinking in terms of, like, as you're training the model, you've got the correctness signal. But, I mean, maybe one sort of heuristic for elegance would be, like, just brevity, which is maybe one, you know, kind of way of trying to send an elegance like signal through a deterministic mechanism. But I would be very interested to know if there is, like, a panel of mathematicians that you guys have reviewing solutions for elegance to try to make sure that this thing is, you know, not just a pure grinder long term, but, it really has a more eureka flavor to it.
Well, if brevity if brevity is definition of elegance, then our two plus two equals four proof probably takes the cake. Right?
Yeah. I I can't get any shorter than that. I would feel bad for any mathematician whose job it was to compare AI proofs.
That's certainly not the job I've won.
across all domains. Right? Many billions spent on expert validation of AI outputs.
Yeah. We we have done essentially zero of that in in the two years we've been around. I think the the metric we optimize for is the net present value of future proofs.
So or computational cost of future proofs. And so that guards very naturally against certain phenomena. So when you're solving easy problems early on in reinforcement learning, you absolutely can solve them with grinding.
So you can see how much do brute force. But you know that if you do that, it's gonna cause issues later because you haven't learned how to do more complicated things. In contrast, if you're given two proofs that are not grinding, but one is drastically longer and more inefficient than the other, you prefer the more efficient one.
So there's a tension there because you can get more efficient by grinding. Right? But that messes you up in the future.
So it's it's a balance that our AI researchers strike based on their intuitions about what'll be helpful long term, but we have never had panels of mathematicians do AB testing on proofs or anything like that. Really, you wanna you wanna give your system as few priors as possible and just run reinforcement learning at scale. There's a famous essay called the bitter lesson, which I'm again, I'm sure your reader your viewers are familiar with.
We we really believe in that at Harmonic. To get to your question about, you know, how we started, yeah, you know, sometimes we'll start from pretrained models. Ultimately, you get you wanna do whatever optimizes that net present value of future cost of proof.
So pretrained models are great for that. I think at some point, you might ask the question, well, is that gonna bias you too much towards how humans do math? And so you wanna mix in reasoning systems that are not trained from human knowledge.
Right? And they have, like, more entropy and more complementary knowledge. So that that kind of thing we always play with, but it it hasn't really been the living factor so far.
I I think that models are a great starting point.
Cool. I guess one so Goodfire just announced today that they raised a bunch of money at a unicorn valuation. I was a very small scale early supporter of theirs.
And it has me thinking and this also kinda connects to Vlad's comment where, like, you said that the system can sort of invent new theory. So obviously, like, one big thing that people have have said AIs can't do or AIs can never do, which is always a dangerous position to take, is they can't come up with new abstractions. Right?
Sure. They can learn from what we have done and what we've encoded into language, but will they ever come up with their own abstractions? I think that's not a very increasingly, that's a hard position to defend.
But what is so interesting with Goodfire is they're now starting to look at model internals and unlock new kinds of understanding based on looking at what has the model learned. Right? So the the famous or the one they just put out is, like, new markers of Alzheimer's that people didn't know about, but the model was able to figure out, and they were able to figure out what the model had learned by looking internally.
I'm kinda wondering, you know, have you guys done any interpretability work on your models? Do you think that there is sort of a different kind of latent space that you are tapping into? And do you see sort of hybrids as part of the future?
Because one thing I could imagine happening is start starting to stitch together a mathematical superintelligence with a more, you know, kind of fuzzy, associative, understand the world superintelligence, perhaps, like, later in in post or later in the, you know, the training process to try to get the best of both worlds.
Aristotle powering a spacecraft, right, much like HAL 9,000, but benevolent one, you know, one that doesn't go crazy. So, yeah, I think I think eventually you'll see it expanding into more real world things.
I think the I don't know if you're as excited about that.
A safe 9,000 sounds like A safe Hal 9,000, I think, would be very valuable.
I think that interpretability is often used as a proxy for trustworthiness. So a lot of the reason that people explore interpretability technology is that they can make sure that the system does the right thing or aligns with the user's intent. So when it comes to trustworthiness, we made the explicit decision at the very beginning of the company to focus on lean.
By outputting our our reasoning in a formally verified way, that that is the most interpretable possible output. So the computer can check it. If the human wants to understand how the proof works, they just keep hitting go to definition.
It's almost like navigating through a code base. There's no more interpretable way to output math than in Lean, really. That that's the that's the maximal version.
So now the question is, okay. Well, how interpretable is the model? I think, you know, it's the in the context of the bitter lesson, you know, we just focus on letting the system do whatever it can to optimize for computationally cheap proofs of more and more complex things with the caveat that it has to output in a way that's verifiable.
I think down the road, we know we're we're very curious, you know, how does it do math? How is it so smart? And we'll look into that.
But for us, we we solve the trustworthiness question upfront by focusing on formally verified output.
Yeah. Okay. That's quite interesting.
I do sort of feel like I have this one kind of mental mathematicians that are famous for visualizing things. My kind of visualization of what is happening in a large model is sort of like shrink wrapping reality. Like, you're you're, you know, wrapped in plastic all of, you know, all of Internet data or all of kind of whatever domain it is that you're trying to learn at scale, and you're just sucking all the air out of it and gradually shrinking down to whatever, you know, hopefully is kind of the true structure.
It strikes me that in math in particular, that structure might be amazingly simple or, you know, there might be really interesting things to learn by running that process and then kind of cracking it open and seeing what is inside. I would expect it to be maybe a lot more interpretable internally than, you know, something that has had to learn all of Internet data and can recite Wikipedia and all that sort of stuff.
these models are doing is interesting because they're they're smashing together all of the techniques that all mathematicians have done before. And so while while I haven't seen the spark of superintelligence yet where it's some breakthrough eureka idea that it's incomprehensible, I'd say that if you push it in, you know, learning how the models do things, you can kinda ask it to solve more and more complex problems and just see, like, oh, how did it pull together these three subfields of math in a way that no human has done before? I think that'll be a lot more interpretable and comprehensible than trying to, like, ding through the weights.
I might be wrong, but that's probably where I'd start to to interpret how it does things.
Yeah. So does that mean maybe we can kind of look at different levels of difficulty of problem. We've got the air dish problems.
There's definitely a phenomenon happening right now where people are using either Aristotle by itself or I've also seen a lot of not that many, but, you know, increasingly more examples of g p t 5.2 pro to sort of generate a proof in token space, then bring it over to Aristotle for formalization. Then there's, of course, the IMO.
If I understand correctly, everybody who and I think it was just three. Right? You guys, OpenAI, and and DeepMind got the gold level performance.
I think everybody missed the same one question, which is really interesting to me. I'd be interested in your thoughts on why so consistent. Then, of course, you've got these extreme problems where you would need this Move 37 like moment to solve them.
So maybe kind of sketch out, like, where are we on this curve of problem difficulty? And are there do you think that we'll we're just gonna ride a smooth exponential, you know, meter task length style all the way up to millennium prize problems? Or do you think that there are gonna be breakpoints of some sort where you might need a new architecture, a new insight, a new, you know, learning method to get from one one range of problem difficulties to, you know, something that's, like, qualitatively different?
I mean, I I think so on the IMO, the the three labs that announced gold medal performance, you know, Us, DeepMind, and OpenAI, all missed question six. And I think that it wasn't super surprising to us because question six is probably, I don't know, five x harder even for humans. Right?
It's just a more complex question with lots of steps, and it requires this type of, like, spatial reasoning that right now is is more difficult to encode in formal. We were running our system on it quite a bit, and we felt like we saw signs of life. So it's definitely not inconceivable that before too long, question six is gonna fall and be gobbled up just like the other questions.
I mean, even, you know, one one year before questions, I don't know, three and five would have probably been well beyond reach from for most of the models. So I I think it does appear to be more or less a smooth exponential.
Yeah. I I I agree with that. I I wanna highlight that there there's two aspects of this.
So I think we're continuing to see a smooth exponential in terms of AI capabilities in math. What I think is a little more interesting actually and was less predictable before was that there I think there's now definitively been a phase transition to formal. I think years ago, if you had asked someone, hey, could you automatically formalize a number theory paper in Lean or Rock or Isabelle, these are languages, you would have been laughed out of any room of mathematicians you'd be in.
And today, we are seeing people upload the full text of the math paper and running Aristotle a few times. We're thinking of adding a Ralph button to just keep going, keep going, keep going. And then you get a formal version of it.
I think that phase transition has essentially come and gone now because of Aristotle. So in the next couple years, as AI keeps improving, the fact that we can now formalize the AI's arguments obviate the need for the humans to just be the verifiers, right, just sitting there and checking if some output is correct to ones being the tastemakers. So we're the ones studying what problems to work on if we're happy with the techniques used.
So that, I think, is the interesting transition that's happened. So smooth exponential capabilities, but I think we've gone zero to one on verification.
I think that's such a great point because I think there was some debate about this at the beginning and Yeah. Yeah. In in a way, if you look at DeepMind, they started with formal with Alpha Proof, which was the silver medal winning model back in 2024.
It was a great result at that time, and and that was a formal model. And then they went back to informal for Gemini this year, and I'm sure they ran Alpha Proof, maybe it was just that Alpha Proof didn't do as well. OpenAI, obviously, informal.
But if you if you think about okay. Let's say we go to a world five years from now. Right?
And the autonomous math being done by AIs increases. And instead of, you know, five to 10 page proofs, you're starting to produce 5,000 page proofs, which which you should you should assume, right, as as these models can autonomously reason more and get more efficient, they'll produce longer and longer output per unit time is going to be a proxy for complexity. Who's going to review that?
Nobody's reading a 5,000 page math proof.
to validate it check doesn't actually grow linearly with the complexity of the proof. Yeah. That was, I mean, that was really the founding thought experiment of Harmonic.
So we asked ourselves in 2023, okay, so these models can do high school math poorly, but they could do elementary school math poorly, right, a year ago. And so what happens in ten years if we ask it to prove the Riemann hypothesis? Any model will make an attempt at it and give you a 100,000 pages of output, which you might as well throw in the trash for two reasons.
First of all, there's probably a mistake somewhere. Second of all, you can't process like there's just nothing to do with it. No, you just can't wrap your head around what is going on in that proof.
And so there were two hypotheses, both of which have been proven out, which is that, first of all, outputting math formally makes it digestible for humans, and there's a high level of certainty and trust. And secondly, it'll lead to more efficient ways to do reinforcement learning for math, which is what we saw proved out. Right?
If you compare the resourcing we've had compared to the big labs, you know, we're punching well above our weight at the IMO. So I I think the the in in our view, the debate on formal versus informal is settled. I mean, clearly, it's gonna be formal.
One can debate, okay, what's the most efficient way to train a model? There's some aspects informal that are helpful, but I don't think we're ever going back to a world where we're like, oh, it's just gonna be informal from here on out. I think the interesting question though is to extend this to software.
Right? Because the same things actually hold for software that hold for math. Let's say AIs are getting to the point where they can autonomously work and create a software project over a period of a week or multiple weeks.
You know, who was it? The cursor team ran this and generated like a chromium compatible browser, right? It was something like one and a half million lines of code.
It was incredible. So, who's going to read that code and find all the security vulnerabilities and the bugs? And is it is that code in the future that's generated by AIs going to be in Python and Java anymore?
Like, why would it be in Python and Java? Those are just languages optimized for human readability. And, you know, if if the the answer we think to humans reading and trusting something or even an AI that the model is collaborating with checking something are the same.
You want to make the cost of verification as low as possible.
is formal as well. And more and more software will be written in formally verifiable languages. Yeah.
And I I think, know, Lean is our favorite language. It'd be amazing if everyone can write in Lean. I think that as AI writes more and more code, it will be easier for people to accept that.
But we'll see. Yeah.
stuff where bugs are much more serious and much more costly. And and there's a bunch of domains that already are doing formal verification for software, but they're doing it in a in a very artisanal way. You know, they're hiring Lean or Rock or Isabelle experts and kinda painstakingly formalizing stuff.
So I think you'll start to see it accelerating the work of those people first, but then it'll just diffuse, and you'll you'll see, like, formal vibe coding before too long.
Yeah. I love the term vibe proving, by the way. Yeah.
I think that vision is an incredibly compelling one, and, you know, it's it's also one that I'm still kind of wrapping my head around. For listeners who haven't already heard it, I did one with, an episode with Kathleen Fisher who was at RAN. I think now he's just moved to ARIA in The UK to lead their whole operation, and Byron Cook, who's, like, a legend of the formal methods field at at AWS.
And, yeah, they're kind of, right there with you, you know, envisioning this world of basically totally verified bug free software starting with mission critical stuff, but potentially extending to everything over time. I guess one so I think that is super compelling. The the one kind of nagging I don't know if it's a worry that I have or what exactly, but I'll just frame it as a question is like, we are training an AI to be superhuman at formal reasoning within the formal reasoning system that we have, how do we get new abstractions from that?
Or how do we get a sort of Einstein kind of moment where, you know, like, it seems that at some point we all sort of thought the world was just naturally three d and that was, like, obviously intuitive. Then it's kind of come to light, obviously, now that, well, that was an adaptive understanding of the world that served us well as monkeys, you know, and allowed us to survive, but it was at the end of the day, we now know that it's like a lossy approximation of true physics. And so I'm kind of like, do we have any room for doubt or or worry that the math that we have now, as sophisticated as it has become, might also at some point prove to be not quite the right paradigm?
And is there any way if you're training in this, like, purely formal way, is there any way sort of to punch your way out of the box as an Einstein did? Right? He he seems to have Yeah.
I I The fourth wall.
So he he broke the fourth wall conceptually, but the key thing to remember is that he was able to describe his theory rigorously and formally in the framework of differential geometry. So the point I was making earlier about math being reasoning is the point I'll appeal to now, which is to say that no matter what complicated theory somebody might come up with to explain how the universe works in the future, if it's gonna be based on a series of logical deductions that can be explained to someone else and checked independently, that is itself a logic that can be encoded with Lean or other languages like Lean and then verified. So, again, the axioms that Lean is based on are so minimal and just expressing just the most basic possible common sense about how reasoning should happen.
Like, one thing might follow from another or if two things look the same, they are the same. That's the level of axiom we're talking about. So I I really don't think there's any conflict here.
I think that one should just think about formal reasoning as an especially detailed version of informal reasoning that a computer can check automatically. There's no there's no limitation to it. Sometimes it might be a little more verbose than you'd want.
Right? So you want to write tactics and things to cut down on that, but there's really no fundamental tension between the two.
And I think there you also, you know, might be thinking about Gerdl's incompleteness, like the fact that in any sort of axiomatic system, are statements that are true and unprovable. And there's also statements that are undecidable, right, and independent. So there's sort of like a bunch of edge cases here, but I think it doesn't prevent us from making a lot of progress and proving actually the lion's share of useful things.
I mean, there could be things that are unprovable but true that are very, very useful to know as well. But, yeah, no no way to know unless you explore the frontiers.
Do you think there's always gonna be a role for entropy of some sort in these systems?
the Not the I can argue. Yeah. I I think I think hallucinations are a key part of a reasoning system.
Hallucinations are what allow a model to explore something that has never been encoded by a human before. So, you know, when we run Aristotle, whether it was at the IMO or now, it makes a lot of mistakes. It tries a lot of paths that don't work, but that exploration is the very thing that lets you get to the right answer after enough attempts.
So, entropy is crucial. I I I think this whole notion of seeking fundamentally hallucination free LLMs doesn't really make much sense. Now, of course, you wanna pair them with a system like Aristotle that can verify things intent.
But now I think entropy hallucinations are a key part of the training process for models like this.
You've gotta be able to pose false statements in order to
prove that they're false. Learn like humans, you know. You try a lot.
True for humans. True. Yeah.
Some of the most creative humans are the ones that hallucinate the most.
what's kind of the latest progress on the path to superintelligence? You said you and I think this is true of all good frontier AI companies, whether, you know, at the application layer or the model layer or anything any hybrid of those. You know, you're you're updating your systems frequently.
It sounds like there's kind of a convergence of some sort going on between the tree search part and the informal, lemma guesser that you described in the technical report. Can you tell us about kind of what the trends are right now?
I think a lot of the well, just to review the progress. Right? So we started in 2023, and then in 2025, goal performance at the IMO, we topped up this Virena benchmark at the end of the year with our public API.
Users started solving AirDish problems, right, which were unsolved for what, thirty, forty years. So I think there's a very clear trend, right, in in capabilities. I think the phase transition I mentioned has also happened.
So I think what's next for Harmonic and for the field at large is, you know, a couple of things. Well, we can expect Mathlib to grow. So Mathlib is the think of it like the Wikipedia for math that's computationally certified.
So as Aristotle makes it possible to auto formalize a lot of math, you can expect that users will start contributing a lot of pull requests to MathLoad, and that makes it possible to solve more and more problems on top of that base. I think when we look at how mathematicians are using our API, certainly people are starting to work on more important unsolved conjectures that a lot of people would care about. So you can kinda think about conjectures as like, okay, there's a conjecture that's technically been open, but nobody really cares about it.
So it's not like people are trying all the time. But now you might have some conjectures that, yeah, like, a mathematician might try it once or twice a year, just take a shot at it, maybe a 100 mathematicians would. And then eventually, about the millennium prize problems where, you know, any mathematician would be happy to spend years on it if they might be able to solve it.
So I think what you can expect from Aristotle and and other systems is, you know, more and more problems get picked off. It becomes easier to use. It extends to software, as I mentioned.
So we have users using it to check safety critical software, whether lean or other languages. And overall, if I had to pick on just one trend, it's really just that formal reasoning goes more and more mainstream. So as more stuff is produced with AI, I think you'll see complementarily more formal reasoning to kind of verify all of it.
And I think on the product side, you know, we've gotten a lot of feedback coming in from the folks using it. Obviously, you know, whenever you've got customers that are using a technology like this, they're very passionate, so there's lots of ways in which they're still like complaining about things and improving the ergonomics a bit, making it so that people don't have to hop between so many different tools and we could just solve their problem as simply as possible and at the lowest possible cost, you should see that continue to improve. You know, there's there's been updates to the to the system pretty much on a daily basis.
Maybe you've seen some of them just as you've been kind of experimenting yourself, but that that's gonna continue and you should expect that it's get it gets exponentially more useful over time.
So maybe a good place to close is kind of the vision for what that looks like as you succeed. I mean, obviously, one thing is is solving money on price problems, but I'd love to get a little bit more of kind of an intuitive understanding than that. I mean, one dichotomy that kinda comes to mind is this, like, very formal reasoning based paradigm versus what I think of as intuitive physics.
And it does seem like models are, like, very good at developing intuitive physics in kind of any number of spaces. Right? Like, folding a protein with a model is not something that's done in a formal way.
It's just kind of something where whatever kind of mess of heuristics, you know, they learn, they kinda learn, and they can they can do a a protein fold orders of magnitude faster than we would be able to do it if we were gonna do it through a sort of physics based simulation kind of approach. When we think of no limit to math and what does a mathematical superintelligence look like, I also think that Elijah once famously wrote, at least famous to me, that, you know, a real superintelligence in his mind could look at one still image and deduce all of physics from just the information contained in that one still image. That kind of also connects, I guess, to, like, test time training.
What is your vision, I guess, you know, to to you can bounce off any of those concepts, but what is your vision of how this thing evolves? Is it an ever bigger tower of formal statements? Is there some role of new kinds of intuition, new abstractions that that emerge out of that that aren't so strictly defined but potentially useful?
You know, what is what is this thing doing in 2030 once all the millennium prize problems are solved?
I think that by 2030, we will have theoretical explanations for everything, basically. I mean, if you look at the history of science, there's leaps of intellect and leaps of data. So the microscope comes along.
All of sudden, you can build a lot more theories of biology. Now the electron microscope comes along. You can build more theories of, like, chemistry.
I think right now, there there's really been a shortage of people that are able to reason logically at the highest level. So when you think about unifying general relativity and quantum mechanics, it's just it's a very hard thing to do. And so I think what you'll see is really, like like, anything that can be posed mathematically, which is what underlies all of science, I think we're just gonna get theories for everything that are self consistent and make sense.
I think we're then gonna go back into a regime where we're data limited. So we're gonna have maybe five theories that unify QM and GR, and we're gonna have to run very high energy experiments to figure out which one is right, and we'll have to wait a while to build those colliders. But at the very least, we're not gonna be bottlenecked anymore on, like, wondering, like, can we explain something?
We'll have a system that can explain anything perfectly correctly. So it really will be a a renaissance of science, I think. And you you just remove the the intellectual bottleneck in in everything.
So do I understand that correctly? Basically, you're envisioning, like, multiple grand unified theories that all explain all the data that we have, and then it becomes a problem for the collider experiments to figure out which one of these is in fact right. Yeah.
Because, you know, AI AI is a it's not omniscient. You know?
it's our model or others, like, they'll be able to reason about anything they can kind of ground, right, in in their own logical deduction rules. But, ultimately, I think there are aspects of the universe. You just have to run the experiment and find out how it really works.
Yeah. Wow. But I I just to be clear, think there's a lot of utility before you get there.
I I if I have to analyze asymptotically where we get to that, that's my point.
Well, I mean, that's we've heard about centuries of scientific progress collapsed into five years. That sounds like more like a a few thousand years perhaps of scientific progress.
will happen, and then you just have to get more data, but you'll have a superintelligence system that can help
Wow. Okay. That's about as grand of a vision as I've heard anywhere.
Do you guys worry about the safety of these systems? Like, we haven't talked about that really at all in this context, but I've done many explorations of different safety concerns. You know, Elijah, when he described the model, whatever some whatever AI he was kind of envisioning when he described it understanding all of physics from a single image, he also thought that was gonna be super dangerous because it would be so powerful.
How do you guys think about that aspect of this whole I mean, we're talking about a lot of stuff in the next five years.
I mean, I think right now, we're not so worried about it because the outputs of our system are constrained. I think that you're likely to see like, the the first dangers will probably look a lot like cybersecurity incidents. Right?
Because, you know, you have the models that are making API calls and running autonomously and, you know, interacting with other systems. So that both creates a p creates API level cybersecurity holes and the mechanisms to exploit those. So I think you're likely to see a lot of those.
I think for our model, since it's basically just the interface to the outside world is tightly constrained and it it's not just gonna fire off a a request to your Gmail account or the iMessage APIs. We're a little bit further away from that. But, you know, you can imagine we're gonna have to start taking that much more seriously when we do get to a point where we're connecting the model to to the outside world, and it's, you know, speaking in in the interfaces are not just sort of, like, lean files being outputted.
Yeah. I do think constrained action space is certainly one of my favorite paradigms for keeping things under control. But I mean, a full like malt book, malt log thing has been fascinating to watch.
you know, I I think we're entering a strange new world for sure. And I think the benefit is we're probably not at the danger frontier. So we'll have the opportunity to learn from others mistakes and hopefully they don't screw up too badly in order for us to learn.
Yeah. Okay. Fascinating stuff.
This has been fascinating stuff, guys. I really think the approach is really interesting. The, you know, the vision for how far we can expect or, you know, even somewhat entertain the possibility of being in 2030 is arresting and both inspiring and for me a little bit scary.
Anything else you wanna leave people with before we break?
I think for me, and you kind of see this in the values that we put on our website of what we care about, you know, we believe in a future where humans are going to be at the center of all this progress. So, I think that we'll definitely accelerate it, but the humans should be in charge and and calling the shots. And I think that's also why we care so much about putting this into people's hands and making them use it and not just kind of be a lab that runs things secretly and, you know, makes big proclamations because I think humans need to be at the center of everything and and still calling the shots.
You know, that's what we believe in in in the world that we're helping the future that we're helping bring to life.
Yeah. And I I think just to add to that, you know, for me, when I started using Aristotle, it was very different to have an experience where the output's always correct. And so I think if people haven't experienced that before, should just try it out.
It's free to sign up for.
Cool. Well, there's there's I'm sure there'll be plenty of ways to monetize mathematical superintelligence when the time comes. We might do ads.
Yeah.
Yeah. I can't wait for that. Alright.
To put those anthropic ads to life.
Fascinating stuff, guys. I really look forward to watching your progress. Thanks for the both remedial education and the grand vision today.
It's really extraordinary. What a time to be alive. Vlad Tenev and Tutor Akim, cofounders of Harmonic.
You both for being part of the cognitive revolution.
Thanks for having me. Pleasure to be with you.
If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network.
The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.
ing. And thank you to everyone who listens for being part of the cognitive revolution.
Shared via Hopper