Grant Sanderson – AI and the future of math

Dwarkesh Podcast
30 June 2026 1h 33m
0:00 --:--
Episode Description
Always so much fun to chat with Grant.AI has been making much faster progress in math than in other fields. As a result, mathematics is showing us, very concretely, what AI progress in other fields will look like. Even within mathematics, there’s a jagged landscape. What does it look like?What is the nature of the most important conceptual breakthroughs in the history of mathematics, and how different are they from what AIs are currently able to do?Does AI (on net) increase or decrease human und

Summary

This episode explores AI's rapid progress in mathematics, particularly its ability to solve complex problems like those in the International Math Olympiad and generate counterexamples to long-standing conjectures. The discussion delves into the 'spiky frontier' of AI capabilities, contrasting current AI strengths in pattern recognition and connection-finding with the human capacity for conceptual breakthroughs, definition generation, and 'mountain building' in mathematics. It also examines the future roles of human mathematicians in an AI-accelerated world, focusing on curation, explanation, and the unique advantages of AI like parallelization and systematic exploration of different approaches.

Chapters

AI Progress in Math & AGIThe discussion begins by highlighting AI's rapid progress in mathematics, questioning if achieving gold in the International Math Olympiad signifies AGI, and noting the 'spiky frontier' of AI capabilities within math itself.
Creativity and Millennium Prize ProblemsThe conversation explores different ways AI might solve problems like the Riemann Hypothesis, distinguishing between finding connections, building new conceptual 'mountains,' and brute-force computation, and how these relate to human creativity.
Next Benchmarks: Conjectures & DefinitionsThe hosts discuss the next impressive benchmarks for AI in math: generating interesting conjectures and creating new definitions or conceptualizations that unify fields, drawing parallels to historical mathematical breakthroughs like Galois's group theory.
The Role of Explanation and UnderstandingThis chapter addresses concerns that AI-generated proofs might lack human understanding, emphasizing the distinction between proof and explanation, and the potential for AI to also excel at distilling complex ideas into understandable forms.
AI Advantages: Parallelization & ContextThe discussion shifts to AI's inherent advantages beyond raw intelligence, such as parallelization, the ability to merge knowledge, and systematically exploring problems by manipulating context or biases, contrasting this with human limitations.
Grindability and Formalization (Lean)The hosts explore 'grindability' as a key driver of AI progress in math and coding, and debate the role of formal verification systems like Lean, considering its potential for autonomous mathematical exploration without human oversight.
AI and Writing: Theory of MindThe conversation touches on why AI struggles with writing compared to math, attributing it to a lack of modularity, the need for unpredictable insight, and the models' difficulty in building accurate mental models or 'theory of mind' for readers.
Future of Human MathematiciansThe episode concludes with advice for aspiring mathematicians in an AI-dominated future, suggesting roles will shift towards curation, teaching, and distilling AI-generated insights, emphasizing the social and relational aspects of human endeavor.

Topics

AI in MathematicsAGI BenchmarksInternational Math OlympiadMathematical CreativityRiemann HypothesisGroup TheoryConjecture GenerationDefinition GenerationMathematical ExplanationAI Training ParadigmsFormal VerificationAI ParallelizationTheory of MindFuture of Math Careers

People

Dwarkesh (host) Grant Sanderson (guest) Dario (mentioned) Hugh Montgomery (mentioned) Freeman Dyson (mentioned) Fermat (mentioned) Abel (mentioned) Lagrange (mentioned) Galois (mentioned) Einstein (mentioned) Liouville (mentioned) Jordan (mentioned) Gaillman (mentioned) Timothy Chow (mentioned) Minskowski (mentioned) Claude Shannon (mentioned) Feynman (mentioned) Newton (mentioned) Andy Jassy (mentioned) Karpathy (mentioned) Eric Jang (mentioned) Alex Kontorovich (mentioned) Annie Matushak (mentioned) Stephen Strogatz (mentioned) Terry Tao (mentioned)
Key Concepts (30)
Spiky Frontier of AI — AI's progress is uneven across different fields and even within specific domains like mathematics, with some areas being much easier for AI than others.
International Math Olympiad (IMO) — A benchmark for AI in mathematics, where AI has shown significant progress, particularly in geometry, but struggles with combinatorics problems.
Millennium Prize Problems — Seven highly challenging mathematical problems, whose solutions are considered significant benchmarks for advanced AI capabilities, potentially indicating AGI.
Riemann Hypothesis — A Millennium Prize Problem concerning the distribution of prime numbers, used as an example to discuss different ways AI might achieve significant mathematical breakthroughs.
Random Matrix Theory — A field of physics studying eigenvalues of random Hermitian matrices, which was unexpectedly connected to the statistical properties of Riemann zeta function zeros, illustrating cross-field connections.
Fermat's Last Theorem — A famous theorem that took centuries to prove, requiring complex mathematical machinery like elliptic curves and modular forms, used to illustrate the 'mountain building' aspect of mathematical progress.
Group Theory — A mathematical field developed by Galois, which studies symmetry and abstraction, initially not recognized for its utility but later found widespread applications in physics and cryptography.
Galois Theory — A theory that allows one to determine if a specific polynomial has roots that can be written down using radicals, based on the symmetries underlying the formulas.
Unit Distance Problem Conjecture — A mathematical conjecture that was disproven by an AI, serving as an example of AI's ability to find counterexamples and connect existing ideas.
Conjecture Generation — A proposed next benchmark for AI, where models would generate interesting and fruitful mathematical conjectures, moving beyond just proving existing ones.
Definition Generation — The 'premium tier' of mathematical intelligence, where AI would create new kinds of objects or conceptualizations that unify fields and spawn new areas of study.
RLVR Environments — Reinforcement Learning with Human Feedback environments, discussed in the context of training AIs for tasks that lack clear, quantifiable benchmarks.
Compression as Intelligence — The idea that a smaller, more predictive expression or theory feels more intelligent, suggesting a potential verifiable reward for AI in finding succinct, elegant mathematical concepts.
Kolmogorov Complexity — A measure of the computational resources needed to describe an object, suggested as a way to quantify elegance or conciseness in mathematical proofs and theories.
Continuum Hypothesis — A problem in set theory asking if there's an infinity between natural numbers and real numbers, whose answer depends on axioms and requires complex methods like 'forcing' to describe.
Forcing — A method used to describe the continuum hypothesis, known for being very difficult to understand, highlighting the gap between proof and explanation.
Unsolved Expository Problem — A concept where a problem has been proven, but the underlying reasons or clear explanation of why it's true remain elusive, emphasizing the importance of human understanding.
Theorem Economy — A concept describing how mathematical credit is often given to theorem proving, which is seen as parasitic on the more fundamental work of coming up with definitions and concepts.
Space-time Diagrams — Visualizations used to illustrate concepts in special relativity, such as length contraction and time dilation, raising the question of whether conceptualization is inseparable from the idea itself.
Minkowski Space-time — A mathematical framework that unifies space and time into a single four-dimensional continuum, providing a geometric interpretation of special relativity.
Langlands Program — A vast web of conjectures connecting different areas of mathematics, particularly number theory and representation theory, representing a research ethos focused on finding deep connections between disparate fields.
Autoregressive Chain of Thought — The sequential, token-by-token generation process of LLMs, which can make it difficult for them to draw unlikely connections or escape local contexts.
Entropy Collapse — A concern that AIs, trained similarly, might converge on the same ways of thinking, leading to a lack of diversity in ideas, which could be mitigated by systematically introducing different biases or contexts.
Grindability — The ability to repeatedly run parallel rollouts or experiments in a deterministic and verifiable environment, identified as a key driver of AI progress in math and coding, but lacking in real-world applications.
Sample Efficiency — The challenge in deep learning where a large number of parallel rollouts are needed to learn a skill, making 'grindability' crucial for efficient training.
Lean — A formal proof assistant, discussed for its role in providing verifiable process-based supervision in mathematics, allowing for automated checking of proof correctness.
Mathlib — A GitHub repository containing a large library of formalized mathematics written in Lean, envisioned as a potential environment for endlessly running AI programs to extend mathematical knowledge.
Process-based Supervision — A training method where the correctness of each step in a process is verified, rather than just the final outcome, enabling AIs to learn complex tasks like writing clean code or mathematical proofs.
Natural Language Verification — The use of meta-verifiers to check the correctness of natural language mathematical proofs, suggesting that formalization tools like Lean might not be strictly necessary for process-based supervision.
Theory of Mind — The ability to understand and attribute mental states to oneself and others, identified as a potential struggle for LLMs, impacting their ability to write effectively or understand human questions deeply.
References (8)
Polylog channel
Gemini 3.5 Live Translate product
AI Studio product
The Fall of the Theorem Economy by David Besses article
Cursor tool
Mathlib project
DeepSeek Math model
Chaos and Nonlinear Dynamics by Stephen Strogatz book
Transcript (65 segments)
Speaker 1

Today, I'm chatting with Drance Anderson, who runs the Blue one Brown and is now working on a new project documenting the progress AI is making in math. And I wanted to talk to you about this because AI has been making the fastest progress in mathematics as of any other field. So whatever is happening here and whatever we were seeing AI progress happen or not happen would tell us about what will happen to the rest of the world as AI gets better and better.

So I wanted to start with this question I asked you when I first interviewed you three years ago. And I asked you, once we have AIs that can get gold in the International Math Olympiad, wouldn't that just be AGI? Wouldn't this just be able to do anything any human can do given how hard these problems are?

And you had an answer which in retrospect turned out to be very wise and correct, is like, it'll be another benchmark like all these other benchmarks that they are passing. Obviously, AI has gotten better in general way since then, but there won't be some moment when this happens. First, I'd I think I'd be curious to get your heuristics on why that turned out to be true.

And second, I'm curious how long you think this narrowness can continue to continue to be true. So by the point that AI has solved the middle enterprise problem, do you think it's still possible that at that point, there's lots of tasks that humans are doing that AI still can't automate in the economy? It's an interesting question because it's hard to answer without knowing what the solution looks like ahead of time.

Speaker 2

that's something where I think the spirit of your question three years ago was in looking at how some of the solutions to these problems really seem to require creativity. Yeah. And the designers of these problems, they'll try to have them come up with things that you can't train for as easily.

I think the dirty secret with the IMO is that you really can train for a lot of them, and so with the with the whole AI and math project undergoing, I think, as you point out, one of the reasons it's interesting at all is that there's a spiky frontier to AI. Math is just right there in one of the spikes. But there's kind of a fractal nature to that spikiness because when you zoom into the specific progress within math, you have some things that are a lot easier than others.

So if we just think about IMO, which is old news at this point, where it's kinda like two years ago, they're really doing quite well, they would have gotten a gold in 2024 if for not the following reason. They're very good. They're just cold solved geometry, basically.

And the IMO has these four categories of problems, this geometry, number theory, algebra, and combinatorics. Geometry just solves in nineteen seconds in 2024 because it's kind of a brute force solver. The dirty secret is for students, there's also sort of a brute force way that you can go at it.

Combinatorics is the one that's the wild card of much more playful puzzly seeming problems, and there were two combinatorics problems on that year's test. There's not always. There's four categories, six different problems, so it's kind of a a toss-up which one is gonna have two questions.

Had it been more geometry questions, they would have gotten a gold that year. But it struggles on those combinatorics ones. And, you know, someone who's trying to keep that torch of the last holdout of, like, math for humanity might say, well, you know, those are the ones that require the more creativity.

Even then, though, I think the spirit of your question, like, if they're solving, you know, a Millennium Prize problem, does that also service a lot of white collar work? It suggests that whatever the rate limiter is between where we are now and that is the same as the rate limiter for making things better at white collar work. And we can maybe, like, paint a couple different ways that like, we focus on, I don't know, Riemann hypothesis.

Like, what would it look like to solve that? One possibility would be these things are extremely good at a specific domain of knowledge and just knowing it very deeply, and then knowing another domain, and knowing another domain. And you've pointed this out.

It's bizarre to have something with this superhuman breath that knows all the fields so well that's not just finding those lightning bolts that connect them. I think we're starting to see sparks of that, of actually finding connection between the things that it's an expert at. I'm sure we'll talk about it.

If the nature of the solution to the Riemann hypothesis was something like that, that feels pretty distinct to me than what's necessary to get good at white collar work. And there's a reason to believe, actually, that that that might be the nature of the solution. I don't know if you know the story of, like, Hugh Montgomery and Freeman Dyson at the IAS.

This is this is a side tangent, it's just kind of a fun story on how I don't know if it was over lunch or something like that. Basically, you have this number theorist who is pointing out, just trying to understand the statistical correlation between pairs of zeros of the Riemann's zeta function. So the Riemann hypothesis is all about, like, do all these zeros sit on a straight line?

And he's finding this, like, this quantitative question you could ask about. And he writes down a formula that looks like one over sine squared or something like that. Freeman Dyson, a physicist, is like, I know that expression.

That expression comes up in studying the eigenvalues for random Hermitian matrices, which was something that comes up in studying the energy levels of a nucleus. The idea that the statistics of those two seemingly different things were the same sort of prompted a potential exploration on, hey, are there aspects of random matrix theory that might be relevant to Riemann's zeta function? I think it's a little bit of an open question, like, there food to be had there?

But that kind of bridging together from two different fields, like if it turned out that the solution to the Riemann hypothesis was exploring an idea like that even further, that has this character of kind of how you expect LLMs to be good at math. It's like they're an expert at the quantum physics, they're an expert at the analytic number theory, They should be able to see that similarity in a way that doesn't require Montgomery and Dyson to be having lunch and happening to talk about that. That's totally different from white collar work, right, in terms of the extent to which you maybe have a hard time using an AI as an editor.

It's not because they know everything and you just need them to find that lightning bolt in between. Different possibility would be what's the right analogy? Maybe, like, if we think of Fermat's last theorem between the moment of Fermat phrasing the question and then what the solution itself looks like, where ultimately the solution involves such heavy machinery in math.

Right? So the beauty that problem is you can phrase it so simply. You ask about, you know, x to the n plus y to the n equals z to the n.

Do you have integer solutions for this when when n is bigger than three? And it's it's something you might expect there to be an elementary number theory approach to it, but just as far as we can tell, there's just not. Whereas the actual solution, you know, maybe there is something simpler, but this might be the the the what it has to be.

There's such a complicated set of ideas that build on centuries of work centered around elliptic curves, and then this other mountain of ideas centered around these things called modular forms. Both of those mountains have to be built before you can ask the right question that connects it.

Speaker 1

like it would be surprising if that didn't permeate into other aspects of the economy besides, like, just the mountain building for math itself. Yeah. Or at the very least, even if it couldn't, like, literally do every single thing white collar humans can do Yeah.

It would just have transformative effects in the way that getting gold in the IMO did not have transformative effects on the world. First of all, I do wanna point out that I'm totally moving the goalpost here because when I interviewed Dario about two, years ago, I asked this question about why haven't they been able to use their vast knowledge to connect ideas together and come up with a new discovery that way. That seems like the kind of thing, even if a moderately intelligent person knew this much information, they'd be able to come up with a medical diagnosis from the fact that this drug causes migraines and this other thing, whatever, does this.

Maybe that it's the same drug that can cure both things. And, yeah, I don't know. From an outsider's perspective, mathematics seems clearly like a field where finding this counterexample to the unit distance problem conjecture was like an example of this kind of thing.

As a total goal plus moving. But then we can ask, okay, what is the next benchmark now that AIs can do this thing that we should have thought they should be able to do? What is the next thing that would be quite impressive?

And there's a couple of candidate ideas here. So one could be coming up with interesting problems in the first place, and the other is coming up with new kinds of objects or conceptualizations that create or unify fields. On the first one, now, we just train these models to like, we have these Millennium Price problems because yeah.

I know the mathematicians have noted like, Riemann came up with this idea of this, like, Riemann's data function and because he thought that it had it would have some connection with, the density of prime numbers or if the zeros on this function would have some connection to prime numbers. And so like figuring out that there's why do we think this is an interesting thing to study in the first place? Why why are we building this object and trying to answer questions about it and answer this particular question about it?

Seems like the kind of thing that would be the next benchmark.

Speaker 2

I mean, you you highlight two pretty good examples there. The for anyone curious about the the unit distance conjecture, there's this really nice video by a math channel called Polylog where where they talk about it. And one of the people in that because all of these discussions, it causes people to reflect on the process of doing math.

They're like, oh, this thing can do this impressive stuff. What does that mean for us? Highlights this quote, how good mathematicians prove theorems, great mathematicians come up with conjectures, and the greatest mathematicians come up with definitions.

That's more or less exactly your framing here on those two we need the conjecture generator and then the definition generator. That's the premium tier mathematician. I don't understand how exactly you'd make that a benchmark in the sense that usually, when I think of the word benchmark, I'm thinking something that you have like it's a goalpost.

The the ball is through the goal or it's not. Like, you can clearly say, like, yes. This is done.

Partly to be able to do things like our LDR, but also partly just to be able to, like, know that you haven't moved the goalpost in answering. You know, OpenAI can have their headline on disproving the unit distance conjecture because it's a clear distinct. It's like, it did it.

Right? Whereas imagine trying to have a headline on three fifty five before came up with a really good conjecture. Right?

Like, we promise. Everyone thinks it's a good conjecture. It just doesn't it doesn't land the same way.

Right. But maybe that doesn't negate the fact that that's the right thing to be thinking about. I would be surprised if it ever took the form of looking like a benchmark, and we have a score saying that it's past this benchmark because we can quantify how good a conjecture it is.

Probably the nature of what it would take is that you would feel a tone shift in conversations with mathematicians about the way that it's useful to work with. Series that you referenced that is not at all produced yet and probably won't be for a couple months takes the form of us interviewing a lot of mathematicians. What's interesting is we started doing this over a year ago, and it's fun to see a little bit of a tone shift in the way that they talk about AI between mid twenty twenty five and where we are now in 2026.

In the real world, that's a very short amount of time. In the AI world, that's eons. Right?

And we're able to see over those eons, like this tone shift. I think the way that you'd measure conjecture generating ability is gonna be more subjective on that tone shift where it'll be mathematicians saying they're not just using it to solve their problems, but as they step back and decide what their research field should even be, that a conversation with such and such model was genuinely helpful for that. I think it's likely that you'd see it in the form of a headline saying that this was yet another benchmark knocked down.

Right. And so it's very interesting.

Speaker 1

are also the kinds of things, at least in the current paradigm, you can't easily train for, Because there's really no fundamental difference between a benchmark and a training environment. Yes. I think it's very easy to come up with some dichotomy of like, here's a deep reason why AI can't do a certain thing and then it turns out, well, you're just thinking about it the wrong way and actually I can do it pretty soon thereafter.

Speaker 2

You're gonna kip them with a kippel anyway?

Speaker 1

And I think that this this will probably it'll probably turn out that there's ways in which that we can train AIs to do these kinds of things in the relatively near term. But it seems like it would have to be different from current ROVR training. So thing I'm curious about and the thing it seems to me that drives a lot of the big progress in mathematics and in science generally is like coming up with a new way to think about a problem or the new way to understand the world that then unifies different fields, spawns entire new fields, solves problems we weren't even thinking where we were trying to solve in the first place.

Like, the reason Einstein was thinking about GR is not because he wanted to explain why light bends or why black holes exist. These are phenomena he didn't even need know it needed to be explained in the first place. But in mathematics, often seem okay.

A total outsider, I don't even know the details of what I'm talking about here. From the outside, it seems like there's often ways to say prove a specific problem that can motivate a new conceptualization. One which results in a whole new field, a whole new way of thinking, is immensely productive, and one which doesn't.

I think I'd be curious to hear you talk about whether Galwa coming up with group theory and distinguishing his like solution to the the quintic having no formula for the roots and Abel coming up with a different proof a few years earlier that didn't come up with group theory. But then if you wanted to do a verification loop on like, is group theory a interesting concept That was like, was something useful done here? Why is this proof better?

Potentially, that verification loop is a 100 long. Yeah. And it involves the cryptography coming around and physics making progress and the ideas in group theory being relevant and understanding like symmetries in physics and all those kinds of things.

It's like a hundred year verification loop, but why is this a productive concept in the first place? Yeah.

Speaker 2

Boy, yeah, you struck a nerve because I I had this, like, project about Galway I was gonna do in 2022 that I put on the shelf, but I spent, like, a year of my life, like, thinking a lot about what he did. So there's a risk of me accidentally talking too long on the specifics. Hold me back on.

It's a perfect example for your case because describing why it was a valuable insight does not come from immediate utility. So certainly if you're thinking about RLVR environments, it's like, okay. This is gonna be really hard to do.

But it's interesting to note how even with, like, human verifiers at the time, like, it took a really long time to recognize it as being useful. Like, I think Einstein with GR, people sort of felt. You can, like, feel this feels like a good theory right away.

Like, the what makes the Gatwa theory such an interesting example is you have literally this one hundred year segment of an idea that flows through many different people's heads before it settles into something that the math community agrees is good. So to back up a little bit, do you want the background Alright. On the problem at Well, so we all learn about the quadratic formula in school.

Speaker 1

I thought you were gonna say, we all learn about group theory in school. We all learn about group theory, the quadratic formula.

Speaker 2

So this was known in some sense, Greeks could solve quadratics, but they didn't really write things in algebra, and so it's really more like the Arabs that wrote down that formula. There's this delightful story around some dueling Italian mathematicians, not real duels, just intellectual challenges, who secretively found a formula for the cubic. And then very shortly thereafter found a formula for degree four polynomials.

So a natural open question for mathematicians is, you find a formula that solves degree five equations? Now the nature the degree four, it's monsters. It's like it's a it would be wild to write it down.

You usually don't really write it down in full. You break it up as like a procedural thing. So you might believe these things have this exponentially increasing complexity.

Many hundreds of years, nobody is really answering that question. Usually, we say Abel was the first to prove it. He was this young, precocious Norwegian mathematician, and he showed it's simply impossible.

It's not that you can find a quintic formula. He thought he found one, but he showed it's impossible. I think the real credit though, like, you have to back up a little bit and talk about Lagrange, where Lagrange found the right kind of question to ask about this.

I can go into the details if you want, but I'll give it a very high level. He he he was studying the question, and he recognized being able to solve these polynomials is actually very related to understanding, like, the way that certain algebraic expressions are, like, symmetric, like, more or less so. Like, if I write down a plus b plus c plus d, just, adding four variables, if I permute those, it doesn't change the value of the expression.

Whereas if I write a plus b multiplied by c plus d, some of the permutations don't change it, but some of them do. And he had this really, really nice insight about how if you can find expressions like this that have four free variables, but all the permutations take on three distinct values, that had this unexpected relationship with being able to reduce degree four into degree three. So he started approaching the, like, can we find a quintic polynomial by saying, I wonder if I can extend that.

To extend that method, you would have to have an expression that has five free variables such that as you permute them over all the five factorial permutations, it takes on only four values or fewer. You could put that in a puzzle book. You could put that in a brain teaser that a 12 year old could engage with, It's not too hard to find yourself feeling like that's an impossible task.

So Lagrange is sitting here saying, here's a strategy that I'm trying to solve this problem. Can I find a quintic polynomial? This strategy doesn't it seems like it might be impossible, at least from this strategy.

But that was the first time in history that people had the instinct that some kind of question about symmetry was the right way to be studying these polynomials. In his mind, was just a way. It had yet to be discovered that actually there's a tighter connection.

Also, maybe rather than searching for the formula, we should be asking the opposite question. Can you prove that it's impossible? He sort of planted that seed.

Around fifty years later, Abel definitely read Lagrange and was influenced by it. Galois, we know that he loved Lagrange when he was falling in love with math. And so it's very hard to imagine that these two young geniuses, the fact that they both come up with pretty similar insights around that problem, It's not, like, born from Lagrange.

But to your question on, like, are you are you able to verify that this was a good idea? There there wasn't any, like, result that Lagrange came to. There's never, like, he solved the problem, and therefore, we know that that was, like, the right question to ask.

He asked it. There's some intrinsically interesting thing. It also wasn't very important for math at the time.

Most people were more interested in the applications to physics. This is almost in that side, almost recreational hobbyist type thing. Abel, he started working on quintic stuff, but then he was advised to spend more of his efforts studying elliptic functions, and so more of his work was on that before he died young.

He died at 26 from tuberculosis. Then Gawah, he pushed both of those ideas in the right direction, where he really understood the nature of abstraction. He had this really nice piece that he wrote while he was in prison, actually.

We could talk all about his life story, it's pretty wild. He's like this teenager. He's in prison.

He had tried to submit his math papers, they had been rejected. And so, again, it's like verifiable reward. The, like, verifier function that is the academy at that time is rejecting what what he wrote.

Yeah. Because frankly, it was not very coherent. Like, it it wasn't a complete proof.

He wasn't giving, like, a clear thought of, like, what the theory actually was. He was just, a young fledgling mathematician getting his bearings. So it's, the verified reward there is, like, no good.

But he has some instinct that there's something there. So he's writing this diatribe on, like, the nature of, like, math being something which is it undergoes these shifts over time, and he talks about the advent of just algebra itself and going from just thinking in terms of numbers to having a certain fluency just with pure algebraic expressions where you're not tied to interpreting those expressions. He has this instinct that there is another layer of abstraction that seems like what we should be doing, where rather than thinking about the formulas themselves, thinking about what symmetries underlie those formulas.

It was still a pretty ill defined theory, so if you're trying to say, Okay, is the verified reward that he has solved a problem that other people haven't? Abel proved that quintics are unsolvable. You say, What was Galois doing?

In principle, the thing that Galois theory will let you do is take a specific polynomial, and it gives you the rules to say, does that specific polynomial have roots that you could write down? For example, like, x to the fifth minus one, you know that a solution is one, or x to the fifth minus two, you can write down fifth root of two. So it's not that every quintic polynomial you can't write down the solution, but could you find a specific one where you prove you can't write the solution using radicals?

He also didn't even solve that exactly. He has a much more abstract he didn't show for a specific example that he couldn't. Even describing what problem did he solve is very tricky.

Then he he dies. It's this very, like, romantic story of he has this duel. We can get more into it.

There's a lot of myth around, like, supposedly, he writes up all his diet ideas the night before the duel. Really, tried to get them published. Four, like, by the doesn't it's seem to be good for your It's very bad.

Yeah. Yeah. Yeah.

Yeah. If you're a young genius, don't work on the quintic. And so he he asks his brother and his close friend, like, get these notes to Gauss.

Get these notes to, like, the important mathematicians of the day because I think there's something here. Even then, it didn't really take like, so his brother and his friend, like, tried to get them out. It wasn't another twenty years until Louisville, like, sees these notes, sees that maybe there's something in them, and tries to, like, clean it up and understand, like, what was Galois getting at.

Even then, it was another twenty years or so until Jordan actually puts together something like a modern treatment of group theory that they attributed to Gatwa. You could easily imagine history turning differently where, like, these ideas were kind of coming about from other points in math, and, like, Galois could have been forgotten in history if he was the less, like, florid character. But between the time of Lagrange, like, having this inkling of maybe symmetries of roots is the right way to go to where it all looks like modern group theory.

You've got this long span. A lot of the time, it's not even passing the verified reward of human reviewers. It gets on someone's desk, they say, I don't really know if there's anything here.

Gets on someone's desk, they don't. You have to have this one person sort of recognizes it. And then even then, it's not really solving practical problems at that point.

Like, you point out cryptography and and physics and things like that. You have to get into the twentieth century before you have, like, Gaillman thinking, maybe understanding the nature of, like, how certain groups break down has this relationship with what particles are made out of. He anticipates quarks based on a purely group theoretic question, and that's one of the more interesting applications of group theory is that to even predict the existence of quarks is a group theoretic question.

That's so long after Lagrange before you have anything like that. You have to ask, what is the way of measuring progress that's not based on solving a problem? Right?

And that's that's somehow capturing what is the instinct that's inside Galois's mind when he says, I think there's something here. What's the instinct that's inside Lagrange's mind when he says, like, I think this is the right way to think about it? What's the instinct inside Louis Vuitt's mind when he says, these, like, scattered notes from this, like, long dead youngster, like, might have something to them?

It's so hard to put a finger on that, but, I mean, a different, like, series of videos I'm making right now is is about, like, that the whole compression is intelligence idea. And even though this isn't really the angle I'm taking, you know, there is something to the idea that the smaller expression that's more predictive, like, feels more intelligent. Yeah.

And so I wonder the extent to which you can give some kind of verifiable reward around not just like, did you solve it or what is it solving, but around the smallness of the concepts required to to do it. I mean, going back to Riemann hypothesis solutions, what would that look like if an AI solves it? I think a third way that it could happen is it just straight up works harder.

Right? The same way that you could maybe have an elementary proof of Fermat's Last Theorem that's just spelled out over thousands of pages that would be incoherent, but the cleaner way to view it is with elliptic curves and all that. Maybe there's some thousand page proof of Riemann hypothesis that's like, no one's really getting anything out of it.

What you actually want is like, what are the succinct compressed versions of those ideas that would then lend themselves to human understanding? Like, I don't know, Kolmogorov complexity, maybe you throw that into your attempt to quantify what you mean by elegance.

Speaker 1

have you solved a problem? It's very hard to come up with a core like, the the heuristic for science. But it's clear, like, human humans have been doing this somehow.

And, like, obviously, AIs will do it at some point.

Speaker 2

the end goal is understanding, like human understanding. And so even if you do have some, like, thousand page proof of some math thing or some, like, grand new physical theory, the goal is understanding. Yeah.

Right? May maybe if the goal is predictiveness, you can just have, automated engineers go off and, like, build rocket ships or something where, like, we have no idea how these work, but we can get between stars. But, like, there's gonna be a lot of people who wanna understand.

You're still gonna want whatever the, like, concision function is that, like, distills down. Here's this complicated way of thinking into, like, the right one, like the equivalent of the universal law of gravitation for Newton. Yeah.

Like, you would still want to train AIs to be able to do that and, like, find the the compressed representation.

Speaker 1

I grew up in India till I was eight. And so in addition to English, I also speak Gujarati. And since Google just released Gemini 3.

5 Live Translate, I thought it'd be fun to put it to the test in this mid roll. 3.5 Live Translate automatically detects more than 70 different languages and translates them in almost real time into the target language.

Speaker 2

just like it's doing right now.

Speaker 1

I visited China back in 2024, and I remember thinking at the time that this trip would have been so much more productive if I could have been able to live translate the conversations I'm having with the researchers and random people I meet on the street. Now we have that technology. So if you're building an app that needs live translation, you should 100% check out Gemini 3.

5 Live Translate. It's available now via the Gemini Live API and in AI Studio. Go to a i.

studio/live to get started. So people have this worry about mathematics in particular that, you know, the AIs will prove the Herman hypothesis and our understanding of mathematics won't be any the better for it. I have a couple of questions about this.

The first one is whether this is like a thing you should expect. Like, isn't the reason humans come up with general natural objects and sub goals and whatever when we're working on a big problem is that this is just useful when you're trying to work on the complicated important problem. And so we we can just think about like theoretically, is it would this even be a simpler way to solve the right one hypothesis as opposed to just coming up with the natural abstractions that are relevant to thinking about the problem?

And then two, empirically, is this what we observe when AIs do make progress on problems today? When the when the AI came up with that counterexample to the unit distance problem conjecture. You can just read its chain of thought and it seems it's understandable to me because I don't know anything about mathematics.

But it seems to other mathematicians, it was, like, understandable. And it made it made use of, like, known concepts in mathematics and, like, proved relations between them and all natural language. And as a result, accelerated our understanding of the connection between this object and this conjecture.

Speaker 2

So is this even a like, empirically, is this a thing we should be worried about? I think it depends on the nature of yeah. Like, again, if we sort of break down, like, the three possible ways of, like, solving the Riemann hypothesis, That one and the other big one from this year was a certain Erdos problem numbered one one nine six, but it's about these things called primitive sets.

Basically, it had that character of bringing an idea from a seemingly different field. As soon as you just present the basic idea to a mathematician, you say, if we this, try the Markov chain process where we show that this thing is one from the bottom up probabilistically rather than the top down, and use the von Bengold function? If you say that to someone in the know, they'd kind of know how to run with it.

You have this very small idea that has the form of expertise in one field, expertise in another, draw a little lightning bolt between them. Like, those are those are gonna be very human parsible. Right?

Because all you have to do is just, like, show the start and end point of what those connections are. If the character of it is mountain building, you have to put in a lot more time to understand that new mountain that was built because it's a new thread that's not just like lightning bolt between them. And then if the nature of the progress was just like raw hustle, it's just like this super long thing, no new theories, but it's just like long, long, long chain of reasoning answer, then you would have that where it's like, okay, there's this whole digestion process.

So I don't think there's one clear answer. I think it depends on what the solution there would look like. And on the mountain building side, I would actually be really interesting to see.

Is it by default a very human understandable, like the way that we see new theories from great mathematicians, or is it like an alien different kind of mountain being built where we even have to reprocess the kinds of abstractions that we engage with? Well, the closest example here would be the attempted solution of the ABC conjecture that was we maybe shouldn't get into that one, but it just is not probably not a correct solution, but basically, it's this whole new way of thinking that this otherwise reputable mathematician in Japan had come up with. And it just took mathematicians a long, long time to even parse what he was saying, but it had the feeling of just like an alien bit of mathematics that's theory building.

It's not just like long long chain of reasoning. It's like he called it like inter universal geometry or something. And so the fear that you would have is that like, yeah, it like does that.

The the biggest fear would be that it does that, and then much like the ABC conjecture, like, people work for years to go up the mountain, and they're like, dang it. This just isn't right. Right?

And, like, if there if it turns out to be wrong, but it, really looked right. But even if it was right, there's there's just a lot of effort to, like, hike up a new mountain. Yeah.

Speaker 1

If we end up in that situation, David Besses had a really great blog post called The Fall of the Theorem Economy, where he's talking about this, you know, historically, you were saying, mathematics is coming in by these definitions and problems and it's about proving theorems about them. And that really the theorem proving stuff is what gets all the credit, but it's like really a parasite on the coming up with the definition stuff. Yeah.

And historically, it's not even a problem in terms of credit apportionment because if you come up with a definition, you're probably gonna be the guy who who comes up with a theorem. But now we're in a situation where if the valuable work is the the coming up with the insight and then AI just automates the latter part. It so okay.

Imagine a scenario where we have AI comes up with like the Abul like direct arguments about a bunch of important conjectures in the world. And then we just have these proofs. And now it's up to humans or to future AIs to then consolidate I mean, I'm sure if you had access again, having no object level understanding of this argument whatsoever.

I'm sure if you had access to it, it would make it easier for you to then think about like, well, what is going on here? Is there is there some deeper way in which we can understand why this proof works that would make it easier to come up with the ideas behind group theory? Yeah.

Speaker 2

I think it would it would be hugely helpful. Right? Like, because, I mean, so much of, like, trying to discover new math is, like like, mostly being wrong.

Right? You're, like, trying to solve a problem. It, like, what it it does it doesn't feel like constantly taking the correct step up the mountain.

Right. Like, mostly, it feels like a random drunken walk where you're, like, doing a thing and then, oh, you're wrong and, like, constantly discovering. So if at the very least you know that trying to digest what you know is ultimately leading to, like, a correct solution, like, that feels like progress simply because it's it's providing, like, a a sense of knowing that it leads to a solution.

And there's plenty of plenty of, like, instances in the recent history of math where it feels like the reach has sort of exceeded the grasp where there's things that are proven, like, long before they're understood. And, I mean, one of my favorite, like, openings to a paper, it's not even like a research paper, it's more like an expository one, is from this mathematician named Timothy Chow who was trying to understand a concept called forcing. And so there's this problem called the continuum hypothesis that more or less asks, like, you have a size of infinity for the natural numbers, you have a size of infinity for the real numbers, is there something in between?

And the answer is both yes and no. It depends on your axioms. It's sort of outside the scope of our usual axiom systems, which is an interesting answer.

But the method to describe it is just really, really hard to understand. It's the thing called forcing. And in the beginning of this paper, writes, I I wanna like, everyone knows the idea of an unsolved research problem.

Like, I wanna propose the idea of an unsolved expository problem. We're like, sure, we've proven it, but we don't really know why it's true. And suddenly he proposes, like, a partial solution to that expository problem.

Can imagine why I loved that framing because this is my whole life. I don't do research math. It's just wholly about what's the most clear way to understand this even if it's proven.

There's a difference between proof and explanation. And so on that side, I think that you are basically getting to the the importance of that distinction. Yeah.

Speaker 1

for or the the incentive would have to change in not just mathematics but in other areas of science from proving things about the worlds to consolidating proofs into problems or higher level insights. But we're having a discussion earlier at lunch about like a recent talk you were giving about, you know, design and how it helps us understand things. And then in the limit, is there really a difference between the conceptualization for an idea and the idea itself?

So, you know, if if you think about special relativity and, like, space time diagrams and Minskowski space time, is it like yeah. This is like a way in which we illustrate this idea of, like, why there's length contraction and time dilation. But is that, like is it like, that is the reality.

Speaker 2

So the exposition does seem to be, like, the explanation in some sense here. Yeah. I mean, there's a couple interesting things there.

One is it seems like there's a really strong correlation between the people who come up with genuinely novel insights and also who are actually quite clear in their communication of it. Like, you you might imagine, given that the experience of a university student is often that the expert they're teaching them is not necessarily the best explainer of that topic because they are so spoiled by their expertise. But what seems, at least in some cases to be the case, is how the people who are really coming up with something quite novel, so you've got like Einstein or like Claude Shannon or something there, you read their papers, they're really lucid papers.

It doesn't feel like, Oh, this is just for the experts, you have to chop through it with a machete to get They're very good expositors. Feynman has this characteristic too, very good expositor. Maybe the same part of the brain that comes up with the correct new way of thinking about it at a research level also has this knack for good explanation.

I think this is pertinent to the AI one, where I kind of used to think that AIs will become these automated theorem provers, but the role of the mathematicians is gonna shift towards my job, explain these things. I kind of suspect that actually they'll also be quite good at doing that and probably just better than most humans are at doing the explanation half and distilling half. And that's actually not what's left for the mathematicians is, like, digesting and and explaining what was going on.

Probably the nature of how these things are going. I could have envisioned it.

Speaker 1

new idea that solves some new problem is just also good at explaining it. Yeah. That's my new like, that's a that's a way my, I think, beliefs have changed.

What's the last thing you think you'll be doing? Where are they, like, my both you and then also what with the mathematical community the human mathematical community will be doing?

Speaker 2

I will probably be doing something like what I am until I die.

Speaker 1

I have the doers right. Maybe that'll be the same. Exactly.

That's what I did. That'll be for the same reason.

Speaker 2

Yeah. Yeah. You know, it's you, like, build a man a fire, and he's warm for one night, but set a man on fire, and he's warm for the rest of his life.

So that's where I am with AI. No. I because some of the some of the, like, function of an explainer or a teacher is to, like, add clarity to a thing that someone's curious about.

That's one thing. But some of it is, like, a little bit more relational and a little bit more, like, providing, like, motivation, providing a sense of curation. Like, one interesting take that I've heard about, like, what mathematicians will end up being is actually more analogous to art museum curators than anything else where the AI solved the thing, so the art exists.

Right? They even know how to explain it really well, you know, out there. But, like, you still you still want someone to help you navigate in this, like, nearly infinite space of, like, what ideas are worth engaging with, like someone kind of doing that.

And that one, even if AIs were in some sense better at that, I think we would always still prefer a human that we had a relationship with because the way that we get motivated to be interesting interested in things is a social phenomenon. If you have some specific technology you're trying to build, you know, that might be different. You need to know there.

But I think, like, the people listening to this podcast, they sort of trust your curation on, like, what's an interesting topic in the first place. It's not that they're landing on here because whatever your next topic is, that's, what they, in a prior sense, wanted to understand. They're trusting you as a curator.

Yeah. So my role, and arguably that of other mathematicians, might actually just shift subtly into that curation direction of what ideas are worth displaying. And that's a lot of my job right now, even now.

It's basically I think people think a lot of the time for a video goes into the visuals. Sure. A little bit.

It it is. It's not immediate, but actually a lot of it is just deciding what's worth saying in the first place or what's worth putting And because that is that's just I want to engage with that, and I think I have a trust with certain people, and they are curious what I would choose to put forward, even if the AIs are better than that, in the same way that human musicians are always gonna have a role because of that social function of the story behind them, even if the objective quality of the m p three file coming out is better from some model. That's kinda what I see happening to my job.

Yeah.

Speaker 1

I wanna go back to this question of earlier, we were sort of just as AI has crossed this threshold, this important benchmark of being able to connect existing ideas to come up with a new discovery or prove

Speaker 2

or disprove something. Just as it crossed this threshold, we're like, okay. But what's the next thing?

I wanna just There's a lot more to do on that one, by the way. Like, just because a couple lightning bolts have been I still I I think there's like this flourishing future over the next couple years of like really connecting. Yeah.

Speaker 1

I don't know if this is accurate to say, but potentially, a lot of the maybe the biggest breakthroughs, look like this at some level. It's just general relativity. Oh, I've I've you could like, you're just you're just connecting together like Romanian geometry and special relativity.

Right? And so as AIs keep getting better and better at this connection thing, maybe a lot of big breakthroughs are not really of a different qualitative nature. I don't know if you have a take on that.

Speaker 2

I mean, a lot of the conversation focus has been on problem solving and that nature of math, you know, like taking off Erdich problems or something. I would say it's not even a majority of mathematicians who would maybe characterize their work as, like, really targeting the next problem to take down. Are you familiar with, like, the Langlands program?

No. Ah, okay. So this is, like, it's not even a field of math so much it is like a research ethos where Fermat's Last Theorem is one inkling of this, and you had these two different seemingly disparate things, and a connection between them led to a solution.

Lingguins was a mathematician. He has this famous letter now essentially spelling out how it seems likely that there's a lot more connections like that, and who even got a little bit more specific about the nature of the connections such that you might imagine this large map, and you've got this valley over here and this mountain over here and this set of planes over there. And there's a lot of mathematicians who would characterize their work as being part of trying to understand the threads on this map.

And the progress there, it's not even like, here's this one specific problem that we know will be solved by that connection. It's more that there's been enough time and time again cases where big problems were knocked down by finding connections that it's almost preemptively finding the connections. It's actually very interesting.

Anytime you run into a mathematician, ask them whether the character of their work is more akin to Langland's program or if it's more akin to targeting one particular problem, and you get a certain bifurcated split there. The possibility of AIs being supercharged connectors feels like it might be an amplifying tool in that pursuit. It's hard to measure though, because this cuts to what we were saying earlier.

How do you assign a score to say, Yes, you've done it? If it's knocking down a problem, you have a clear way of saying, Yes, you've done it. You can write the headline.

You can have your PR move as the AI company to say we did it. Whereas if it feels like that was the right connection drawn, you can write theorems around it, and this is the nature of what the papers in that field look like. But I think it will require a lot more human in the loop to basically say what was it the kind of connection that we're going for.

But that's my guess on what most of the useful progress from these models will look like in the next five years, is just really filling in that landscape of connections that you can draw if you're an expert in multiple fields. Like you've pointed out, it's kinda surprising we haven't already had this. Right.

And what I'd be curious like, I would be curious to know at a technical level what causes the unlock there. Because on the one hand, you can kind of paint an explanation in your head for why you could be an expert in all of these things and not be drawing those connections, which is when the thing is reasoning, like, method of reasoning is this autoregressive chain of thought phenomenon. Autoregression is actually, a really, really weird way to produce stuff, I think, if if you think about it.

Like like, you're an intelligent person. Imagine I've walked you in a box. Right?

And then the the only way that you have of interacting with the world is that you receive a slip of paper, and then someone says, can you predict what will come next? Right? And then you predict what will come next, and then your memory is wiped.

Right? And then you get another slip of paper, and you go, imagine that was done a whole bunch, and then what comes out on the other end, they're like, look at this essay that you wrote. You might look at that and be like, this is awful.

That's not the essay that I would have written. Right? Because, like, the process of, like, repeatedly, like, predicting something is just pretty different from how you would think as a writer to, like, compose it and think it through and everything.

And in particular, what would probably happen is you're sort of a slave to your context where you might be answering some question about some particular field, and so you draw on all the context around that and you're going there. The connection that actually is where all the substance is gonna come from is by its nature a very unlikely one. You can do all the RL that you want to try to get better in some way, but what's the thing that's specifically upwaiting and incentivizing making these unlikely connections when the vast majority of them aren't the predictable next token that would come in there?

And so it's like, it might be the case that you just have this intelligence that sort of lock it in there inside that box, but it's just a weird way of interacting with it. So the thing I'm curious about is, like, do you ever get any fruit by just, like, questioning the premise of how tokens are generated, like, every now and then in some way. Right?

And I don't think it would be as simple as you, like, manipulate the temperature or something like that. But, like, are there any things that you can do to take, like, the existing level of intelligence, but, like, find the right ways of sparking those connections that, like, unlocks these sorts of things that we're seeing? Or do you need just a a little bit more intelligence such that at the level of prediction, it's kind of predicting that it should be making that lightning bolt to another field?

Speaker 1

reason instead of architecture or even loss function to reason about data. Like, I don't know. We have diffusion models that do that do text, and they're, like, out of a whole the kinds of things that produce are not of a wholly different character.

They're just not being explored as much. I think the more relevant thing is what is the data on which whatever architecture, whatever loss function you have is incentivizing you to produce. And it does seem like they're getting better at like okay.

Forget about math. I mean, we did have this a couple of examples of this kind of thing. But if you just look at why they're getting better at being autonomous agents, it just I don't know.

They they have like, they're in an environment where auto regressively producing the step that says, let's step back and do a search over the whole code base. Right. And then let's step back and like assess my mistake.

It's like the the thing that works. I assume what happened in the case of progress in science or maybe in math is you have frontier math like problems which require like, mathematicians specifically designed them because they require connecting together two different fields. And there's all I'm guessing there's all kinds of clever, like, partially synthetic ways in which to make harder and harder problems like that that require these kinds of connections.

For example, by, like, eliminating assumptions and still requiring the AI to continue to get to the answer.

Speaker 2

And then, like, it doesn't really end up mattering what the loss function is. It just like it's really about can you come up with an environment which incentivizes the civility. Yeah.

It feels like you should be able to. Yeah. I can't I certainly can't speak to the correct ways of doing that, to, like, unlock all this, but it would just be pretty surprising.

Like, don't you think it would be kind of surprising if over the next three years there's not just, like, a lot more of those lightning bolts? So this, I think, is an important thing to think about, which is we often think about how smart a single system is. Yeah.

Yeah.

Speaker 1

AIs having advantages that are more the result of other facts about them. So in this context, the key fact about them is that we can just paralyze and arbitrarily scale them so that whatever level of capability they have is not just like one idiosyncratic genius in the history of mathematics who makes a few connections and then dies in a duel. Right.

It's just universally applying the waterline across all problems that are accessible at the level of capability. I feel like this is among the many advantages that digital minds inherently have that we don't think enough about. The fact that you can the the other ones being the fact that you can like they can merge all the knowledge together.

At least that there will be techniques that allow this to happen that you can like that you can spawn off copies with identical levels of knowledge. But, yeah, I feel like this parallelization is, like, quite an important property. I I'd be curious about your predictions of even if they're not as smart as mathematicians, the fact that they are just, you know, billions of because for PR reasons, the AI companies are just dumping billions and billions of dollars at this, have a quantity has a quality all of its own.

That seems in the right direction.

Speaker 2

I mean, if we take that conversation between Montgomery and Dyson at the IAS that suggests some connection between Riemann hypothesis or Riemann zeta function zeros and random matrices, that feels like the kind of thing that you could try to automate, and that you have agents representing expertise in all these, and basically having okay. We all know that an institute is smarter than an individual, and that the reason for having people all in the same geographic location is because you want those serendipitous conversations to happen. What does it look like to engineer those between agents?

It's interesting because you point out you can pool all your knowledge. I actually wonder if one of the advantages is that you can do the opposite of that, where you have sometimes when an AI is failing, it's because it sort of gets into a bad chain of thought, and it's really hard to get it out of it. So you're like, I'll just start again.

Same deal with humans. Sometimes you start thinking about it in a certain way, and actually what's required is to just back up. Maybe sometimes the form of that, know, there's stories about people trying to prove something for a long time, and then at some point they say, hang on a second.

What if I tried to prove that it's impossible? That prove And the that, like, unwinding your own context and going at it with a fresh mind, you could imagine systematizing that or, like, having multiple different agents deliberately given different pieces of context and try to, like, compare and contrast there. Like, we we don't have the same level of manipulation on our own context.

One in this AI and math series, the first episode will be about when they solve the IMO. And I wanna focus on one specific IMO problem that they failed on, which is one that a lot of very smart students failed on. Terry Tau also failed on it.

And the nature of it is basically that it people were very mad at the problem because they called it a troll problem. I almost don't wanna spoil it because I I wanna construct the episode around, like, leading someone in with without knowing that it turns out to have a simple solution, because you can really empathize with what it's like to be a student solving this. Basically, there's a really elegant way of going down what you really feel like is going to be the solution based on the context of being the International Math Olympiad problem positioned as it is.

The character of the solution is really enticing, but it's hard to prove that it's the best. The reason is that it's not. There's this almost brain dead solution that is the best.

The relevance of that to the whole AI story is for a human, what's required to answer that question is to escape your context. Escape the context that you're in the IMO. Escape the context of the way you've been trained to solve these contest math problems.

If you just approached it like a brain teaser that I throw someone off the street, they'd probably answer it well. You want the same sometimes for human research other contexts, where sometimes just being able to say refresh your thinking, come at it completely differently. So of all the advantages that digital minds have, that might actually be one of them.

A little bit more of a systematic, what does it look like to refresh your thinking, try answering two separate questions, like spin off two agents, one who's trying to prove it, one who's trying to disprove it, one who tries it like this way, one who tries and they deliberately have different contexts.

Speaker 1

erasing the context previously, like, trying a bunch of different things as opposed to merging the results of, like, a bunch of different It is incredibly interesting because a common concern people have about AIs is this entropy collapse, where they all think the same way because they're trained in similar ways. This is why they're bad at writing. They kind of just like go down the same path and have similar patterns of speaking and so forth.

But maybe actually the key advantage AIs have is that you can systematically It sounded like one of the reasons the unit distance problem conjecture took so long to be disproven was because people assumed the conjecture was actually true. So mostly they were trying to figure out ways in which to prove it. And so maybe one of the key advantages the AIs will have is actually to increase the entropy by systematically trying out both the negation and trying to prove the positive of any given statement or being able to systematically give different agents different biases.

That's a good point. It seems like an important thing in the history of human science is that Einstein is just really motivated by this bias, that things should look the same in different reference frames. Then he had multiple other biases like these.

But that is just a very formative in his thinking. And you can just systematically survey a bunch of heuristics and see which ones are being productive at a given problem.

Speaker 2

Yeah. And so you would suggest basically systematically increasing entropy at the prompt level even though you have this inevitable collapse at the autoregression level. Yeah.

That mean and and, I mean, Einstein would be an interesting example because it's like he's got this bias towards things should be real. He also has a bias towards God should not play dice. Right?

And it's almost like you you wanna make sure that you don't accidentally have all of your LLMs or Einstein because you might halt on quantum mechanics progress. Right. Which actually goes to show you that you there's not a correct heuristic Exactly.

For science. Exactly. You actually just need multiple independent research programs with their own heuristics.

Yeah. And that feels like old school software, right? As long as you're able to describe that in some way, you have old school software that amplifies that entropy in some way, and if you're able to put a clear ontology to the distinct ways of thinking that you want to prompt, you explore that full ontology, and then each individual one runs off doing what it is.

But think there's a certain design question there on how exactly do you describe the different approaches. The easy one is, are you trying to prove it or disprove it? The harder one would be to say, what are all the tactics that you could take to prove this?

And make sure that you're, like, sufficiently applying sufficient breadth to exploring that.

Speaker 1

I don't think people appreciate the kinds of things that these models can just go handle for you when you equip them with a good harness like Cursor. For example, I started publishing my episodes on Bilibili for a hopefully burgeoning Chinese audience, but everything I upload there needs the sponsored segments cut out. Normally, that would have meant that I would have to ask my editors to go back through all the old episodes, cut out the ads, and re export everything.

But in about just as much time as it would have taken me to send them that Slack message, I can just tell Cursor to do it instead and spare them. And for research for the podcast, I have a whole repo that I've set up where I've just put every single book and paper that's been relevant to prepping for any of the recent episodes. And I've been able to hodgepodge everything because the Cursor harness is just extremely good at helping the model figure out exactly what information to pull, whether that's from my repo or from the web, in order to answer the questions I have while I'm doing research.

So whatever you happen to be working on right now, just try pointing cursor at it. Go to cursor.com/floorcache to get started.

Obviously, AI for Mac is making a lot faster progress than everything else, and people point to verifiability of the domain as the key reason this is happening. I think that's one of the two important reasons, but I don't think I I think people really neglect the other one. And I'm I'm outside the labs.

I don't know what's actually going on, but this this is a, you know, totally naive theory. Okay. A tangential question to why AI is making so much progress in math.

Why has it been so slow to computer use? Which is what you would you know, a computer is actually very verifiable. It's like, you know, is my Etsy package coming or like, is my event booked, you know, whatever.

These are extremely verifiable things to survey. What computer use lacks is grindability. So because websites have like bot detectors and also it takes a tremendous amount of compute to run parallel rollouts, it's very hard to just run like, a thousand parallel rollouts at the same checkout flow on Amazon because you'll get, like, shut down by Andy Jassy.

Right?

Speaker 2

Presses the, like, red x on DoorDash button. Exactly.

Speaker 1

And so you can try to build clothes every single website. This is very labor intensive and slows you down. And the reason, by the way, you need to do so many parallel rollouts in order to learn a skill currently with deep learning is that we haven't solved sample efficiency.

Speaker 2

Sucking supervision to a straw, Of what he

Speaker 1

course, people are working on many different techniques, but fundamentally, there's this big problem and there's this big constraint in the way we train AIs that we just with code also, you can containerize a given level of progress in a repository and then just pair spin out thousands of parallel containers or hundreds of parallel containers and say like, try to implement this feature. And it's totally deterministic. And because it's deterministic, you can solve the credit assignment problem because you know that whatever caused this rollout to succeed and this one to fail, the diff is the thing that like worked.

This way solve the credit assignment problem. If you have situations that are starting off at different starting points, this credit assignment problem becomes much harder to solve. But most of the things in the real world are just very hard to containerize in the same way.

Like coding and math are exceptions to this rule. But if you're just trying to figure out, how do I build a new business that succeeds? How do I like go trade in the markets for a day and like make money?

You can't like, the fact that had to interact with the real world and like things change day after day means that you can't keep replaying and grinding and farming the simulator. But the the math of course is the exception and I I feel like this is actually an important driver of progress in this domain and also in coding. It's not just verifiability, it has to be grindable.

The third reason that people point out that AI is making fast progress is they focus a lot on lean and formalization. Again, I have literally no idea what's going on in the labs. I feel like lean just doesn't matter that much for the current level of progress in AI.

Or why is AI able to solve the unit distance problem? Or sorry, disprove the conjecture by the unit distance problem. They released a chain of thought or released a rewrite of the chain of thought.

Didn't have any Lean in it. I think it just like the process based supervision that Lean provides where you know each step is correct seems like less relevant than just having this grindable outcome that is verifiable.

Speaker 2

grindability mattering more. I guess I will say on the yeah. Okay.

So naively, you might think Lean provides something unique for math because you're able to see if it can prove it. You have old school software that can tell you yes or no. You use that as your VR.

I mean, what so what would corroborate your point is the idea that, like, the initial attempts again, I'll just circle back to IMO. It's like initially, DeepMind basically does that. It's like everything in Lean, and then the next year, it's all in natural language.

So to your point, not needed. I do I think there is a a yet to be explored benefit of that formalization domain, which is at the moment, you still need you know, ultimately, like, human is is reviewing that counterexample to the unit distance conjecture to say, looks good. And that that provides a certain bound on how, like, endlessly explorable things are.

Like, if you consider, like, AlphaGo, AlphaZero style stuff where they're just, like, off in their own universe, just like playing a bunch of Go and exploring themselves, just completely going potentially off the rails of what any human needs to look at, but they still have this automated verifiable reward. It's not just that, hey, you can do RL on that. It's also you basically never have to check-in, and you can just, like, pour compute at them, like, exploring the universe of Go.

What stands to be interesting like, maybe this won't pan out, but I think the the jury should still be out on, like, whether this will yield anything. With Lean, you could imagine having a basically endlessly running program that's constantly trying to extend Mathlib. So Mathlib, it's this GitHub repository that's basically all of math written in code.

It's very far from all of math, but they want it to be all of math. Written in code that you can ask, like, is this proof correct? It's very labor intensive to write these proofs.

There's like a whole sub community around it. But you could imagine, what if you just had an AI where you say, simply try to extend Mathlib. Maybe it's a fork of it so it doesn't have trash in it because people have certain taste for what they want to be in there.

So you have your fork of the pure AI Mathlib, and it just goes, and it just doesn't stop. It doesn't need anybody to check-in on it. Right?

It could just keep going. It might come up with its own conjectures. It might come up with its own theories and, like, different definitions.

Maybe many of them are useless, but it just has this infinite tree that it can, like, grow out. That's a very unique thing that math has that nothing else has, where you could press go and then just pour compute at it and look away for ten years and then come back and say, what do you have? And there's gonna be something.

And then there's a question, is it useful or not? How do you suss that out? That's just an interesting thing to be able to do.

Yeah. Yeah. It would be very surprising if that didn't yield, like, some sort of interesting mathematical insight from it.

Right? So I think, like, that's the real case for okay. There there's there's like two different ways that Lean is important in this story.

That's the first one of them, basically, is how it's like, you could let go, not even check-in, and progress will be made. You can do that with Go. I don't think you can do that with natural language math.

Speaker 1

That's very interesting. Did you see Karpathy's auto research He wrote this basically one Python file that does basic LLM training and then just had a repo where agents would try to make modifications to the file if it sped up the speed run, the modification stays. Eric Jang, who came on to explain how AlphaGo works, did a similar thing when he was trying to build in a very strong gobot.

And he had interesting observations about the kinds of like, it's it's really gonna just go running an experiment and going down that path, but it's bad at stopping at dead ends and just doing extremely parallel things. Anyways, this will probably be this this will change in the future. It's very interesting to think about what what it looks like in the limit.

I mean, is fundamentally like what the human institution of mathematical research is. Right? It's just like this is a library extended it in interesting and useful ways.

And this way you don't have any outcome based supervision. No. There's no outcome that you're trying to incentivize but you have a process.

Speaker 2

You know the steps are correct, you just don't know if it's going in an interesting direction. But yeah, if you were doing that, you don't wanna completely go off the rails and, like, do a random walk through the space of logic. You'd probably want some, like, supervisor model that's trying to provide heuristics on whether it's useful or not.

But, yeah, something of that character. I mean, you know people are working on it, and, like, that's one of those five years from now. I'd be curious to be able to get the future version of us talking about whether maybe that goes nowhere, but Terry Tau was talking about one research project that's basically try to exhaustively search the space of possible algebras.

You could imagine different axioms that you apply to algebraic systems, and so when we come up with group theory, there's a certain axiom system that has this flavor of they kind of look like arbitrary rules unless you know the motivation. But it's basically like, what if you tried all of them? Do any of these yield useful things?

And, like, the vast majority of them is just trash in some way. Like, it all collapses to, like, no interesting results. But, like, every now and then, there would be this little island of, like, a completely different type of Acxiom system that at the very least seems rich in terms of the number of theorems that can come out of it, and that's bread and butter for what you would imagine automated provers being good for.

It's exploring that space and seeing which one of them turns out to be something. Maybe one of those islands actually turns out to be something you can retroactively put motivation on to say this is the kind of structure that's trying to get at in the same way that you could imagine looking at the axioms for a group, not knowing that it's about symmetry, but retroactively realizing, wow, this is very relevant to studying symmetry. So you could imagine results of that flavor, but instead of just exploring possible algebra systems, it's like all possible, like, logical consequences of any kind of axiom.

Speaker 1

without lean, So DeepSeek had their DeepSeek math model that and they released a paper on how they trained it. And it was quite interesting. So they have the problem with having natural language proofs is you don't know if it's correct or not.

And so they have a verifier. And then the verifier is trained by a meta verifier that makes sure that any all all the problems that they're training this model to solve and like the art of problem solving, that the verifier is getting good feedback on that. And it like it works.

And so it's just interesting Yeah. Natural language verification with some sort of meta verification kind of work. It at least seems to work so far in the published literature.

And also it seems to work in the published products that we're using. Like, if you look at coding agents Mhmm. They're getting better and better at, like, writing clean code and refactoring code and stuff like that.

And I'm sure that that there's process based, like, LLMS judge kinds of things which are saying trying to provide taste and say, hey, is this, a clean way to write this function? Are we like are there are there duplicates of the same kind of modular forms and so forth? I feel like that should also work for mathematics.

Right?

Speaker 2

for math than anything else, even if you're only working in natural language that you could trust a verifier. I mean, you and I were talking earlier about why they're bad at writing, and, you know, I I was asking, like, why you can't just have like, seem to be good judges. If I give them two essays that, like, students write, they they'd be able to say which one's more, like, accurate and insightful.

So why can't you just have a verifier saying, is this a good piece of writing or not? And maybe the ultimate failure there is even if they're good at discriminating between a B essay and an A essay, they're not actually good at discriminating between an A essay and a thing you actually want to read that would be followable on substack and insightful and all of that. They actually end up preferring uninsightful pieces of writing.

And so on the math front, I guess the question would be like, that step to simply know, like, this a correct proof or not? That lends itself to, like, an automated verifier even in natural language. You could probably still make a ton of the progress.

It still doesn't like, I still like the sort of tree of logic out of lean front just in that you can really go off the rails. There's just no constraint on the previous way that things had been phrased before in the same way that everyone talks about move 37 and alpha go and such. What is the thing that lends itself to just going outside the the prior heuristics?

And it seems productive to have a disconnection from the rest of the world in that exploration as, like, a complementary research pursuit to the natural language math front. I mean, the other the other relevance of Lean there would be like, okay. Let's say you have your pure natural language RL environments, and you have a pure natural language set of proofs.

And people have just said like, proceed, AI mathematicians, and they go and they generate, like, 10 papers a day that produce a bunch of stuff. If the error rate if there's, like, any error rate to that at all so Alex Kontorovich has talked about this. It becomes insufferable, like, as a mathematician because you would basically be like, every single time I see one of these, I don't know if it's worth my time.

Even if 99 out of 100 of them are right, I don't know if it's worth my time to even go through it because it's really labor intensive to find what that error would be. It's really frustrating if it turns out you spent all your time on a paper that was trashed. Having anything that's able to give you that green check mark that says, even if this is gonna be complicated to understand, even if it's gonna be a pain, you at the very least know it is correct.

Every other field would kill for that. Right? And math has that if if the models are also able to take their natural language proofs and formalize them.

And so that seems huge. Right? The ability to have that.

Every field would love to have something like that. And so I think you are right that Lean is maybe overrated on the side of the importance of it being used as a VR environment for any kind of, like, just progress in math generally. But I I I definitely wouldn't write it out of the story.

Yeah. Yeah.

Speaker 1

a metaphor for, like, what's gonna happen to our civilization pretty soon. For sure. Yeah.

Right? It's just like for millennia, humanity is building this, like, corpus of knowledge and understanding and everything that we have now distilled into these models. And at some point, the models will just like extend that arbitrarily.

By the way, on the writing front, I actually have I I I have a theory of why writing is making worse progress than these other domains. So I think one of one of them is what you said that they're bad at judging not only a versus b, but they get like this totally derailed by d star Okay. Which is this like a shitty essay that just hits all the all the bells and whistles that, like, a is supposed to hit.

And then so the reward hack thing just, like, totally goes off the rails. But I think the other important thing is that writing is not modular in the same way that code and math are. Like, you know, you can write a function many different ways and they kind of do the same thing and, course, you want it to be very clean and stuff.

But, like, at the end of the day, it works. It works. Same with, like, lemmas and mathematics.

And then, you know, you can, have some end product that is different from the way it is produced. So the code is the thing that produces some end product and you are you want a functional end product. Whereas in writing, the end product is directly the thing the AI is producing and each paragraph, sentence, word matters because that is a thing that is like like that is the substance.

It's not like some separate thing that is produced out of the writing. And so it any it's a it can't just be it can't like be slop. It had the in the way that, like, code can be slopped and still produce some outcome that you want.

Speaker 2

we've gotten much better at agents writing not just functional code, but clean code. Why is it not the case that the same progress that allows you to go from merely functional to, like, clean and and, like, a mergeable PR doesn't also result in, like, clearer writing? Yeah.

That's a good point.

Speaker 1

has it not? Like, I agree there's many ways in which there are terrible writers, but for a lot of writing I consume, I find it's better to just copy paste it into an LLM and just say, like, explain this to me. The explanation will be better than the thing that is produced by the human.

So it's funny that we say, like, these are such terrible writers. And also my revealed preference is just like, can I just have another one explain it? Even when I'm talking to a human expert like live on a call Mhmm.

If it's a piece of knowledge they have that only they have that's not encoded in the distribution, I want them to explain it to me. But then if in order to understand that I need to understand a more basic concept, I would prefer if it was socially acceptable for me to just be able to say, let's pause there. I'm just gonna ask NLM how that works, then we can come back to your your your special piece of knowledge.

Speaker 2

it sounds I mean, that's distillation, right, and explanation. And so if if you're if I'm thinking, like, quality of view as an essay writer, if it's that I give you a book to read and I want a book report, right, then I might believe that, okay, the LLM maybe gives me a better book report. But I think what we what people are really getting at when they say it's better.

Right? Like, what is writing? It's not just distillation of preexisting ideas.

It's not just, like, how do you explain clearly, because they are good explainers. It's like, what is the insight? And this is where it gets like just autoregression is a very weird way to generate stuff because when you're writing, you sort of know in order for it to be good, have to have an element of the unpredictable.

And it's it's not just like increasing temperature in your mind or something. Right? It's like knowing exactly the correct point when you want to make an unpredictable move, and that that's gonna be what's more insightful.

And so even if it's, like, better at explaining a preexisting thing, it's like, what generated that book that you wanted distilled in the first place? Yeah. Right?

It wasn't it wasn't an LLM that, like, generated it, and you just needed it. It's like some author who who, through a lot of exploration of ideas in the world and then deciding what aspects of it were interesting and which ways of presenting it were, like, the the coherent, well motivated narrative. It's like they put that all together in some way.

And, you know, if they're a good author, it's probably one that actually you would on the side of reading their book instead of the distillation. But still, what makes it worthwhile to, like, explore at all in the first place, and you're uploading it at all, I think it's all of that side of it that's the when when people will cite them being bad at writing.

Speaker 1

very directly contradictory to, like, the way that things are being produced. Yeah. That's a good point.

I think they're also really bad at building really good mental models of people, which I think is a very important skill in writing. So Annie Matushak and another collaborator, whose name I'm forgetting right now, did interesting report where they tried to teach LLMs to write good space repetition prompts. Mhmm.

And I really like this because even though it seems like a really totally random skill, it's it just like people are talking about recursive self improvement in ER. Yeah. And we can't get these things to write good flash cards.

And what's going on there? Right? Right.

And they tried many different kinds of techniques, and they're like, you know, sophisticated people. Like, they tried to RL open source models. They tried all kinds of including chain of thought and the big prompt they sent to the best closed source model, etcetera.

And the key constraint it seemed to me was that writing a good card is about projecting somebody's mind in three months. And what is the way in which they will associate the quest like, what what kind of answer we'll be thinking about the moment? And is that is the is the elicitation that inspires the detail you actually want to take away from the passage you're trying to make cards about?

I think writing also is similar to this where if you're writing something, you're like, the reason it's such a innovating process that takes so long is each word you should be think or each pair of sentence you should be thinking, what is happening in my reader's mind right now? Yeah. Even if I flip the phrasing around, where so the end phrase goes to the beginning and, like, this is the first image that comes your mind before you read the rest of the sentence.

That kind of maybe autoregression is is bad at that kind of there's maybe a more diffusion like property of considering the whole rather than going sentence by sentence. But also, think that that requires a lot of mentalizing, which these models weirdly struggle at. Well, I mean, interesting question.

Like, is it weird that they struggle at that?

Speaker 2

I might butcher this. This you know how when you, like, cite studies that you once read, and it's like, may maybe the study wasn't real or something. There's one very memorable one on okay.

So let's say you wanna quiz people's EQ. Like, you show a a flashcard of someone's, like, facial expression, and someone's trying to describe, like, what's that emotion. So I believe there's really good tests online that'll have, a face and then four possible emotions, and it's surprisingly hard to describe exactly the correct emotion, but you also get the sense there really is a correct answer.

If you try this with people in your life, you'll notice that the ones who actually are pretty plugged in socially do really well on it, and the ones who are a little bit more, like, left brain, like, dumb. Okay. So that is the kind of test you can do.

I vaguely remember an experiment to this effect where they took people who had freshly gotten, like, Botox in some way, and they did, like, a pretest and a posttest. And posttest, they were just much worse at reading people's expressions. That feels kind of weird.

Wait. They got Botox. So the person taking the test.

Do the test, and then you go and you get Botox, and your face is all frozen. And now you are worse at understanding the emotions of what you see. And the thought is that part of Interesting.

Part of understanding this emotion that you're looking at is doing it yourself. That's crazy. At a facial level, you're moving your facial muscles, and it's like you see that, you mimic that, you're like, oh, yeah.

That's anxiety. Right? At some very subconscious level.

So in that sense, if it is the case that models have bad theory of mind, sure, they know everything because they've read what everyone wrote. But at a level of actually able to put themselves in your shoes in the same way my face muscles are mimicking your face muscles. That's what helps me understand how you feel.

Not surprising at all. They don't have face muscles. Their brain works completely different.

It's just like it's like an alien trying to empathize. How how could it have theory of mind? It would be this very emergent thing to have theory of mind.

Whereas we can just plug it into our own minds. It's like we've got the ready made hardware to just place it in. And so That's very interesting.

It's not that from that lens, it's not that surprising. Okay. Grant, we are both partners with James Street.

I'm sure over the years, you've interacted with a lot of James Streeters. What have you found that's unique about them or their culture? I mean, was I did this interview with them this year that partly was interesting because they don't usually have anything outward facing.

I mean, in the industry, they're known as having, like, a pretty wild retention rate. Like, people just stay there and it's, like, getting an inside view of that. I remember one of the comments someone was saying, even though the people have role titles, like you know, researcher or trader or engineer, they often don't know what their colleague's actual role is because everyone's doing a little bit of everything else.

Like even if you're officially a trader, you're doing a lot of research. Even if you're officially a researcher, you're doing a lot of coding.

Speaker 1

anyone who wants to be growing, they just have the chance to do a lot of different kinds of things. Alright, Grant. I'll do the plug for you this time.

If you wanna watch this full sit down interview that Grant did with some of the folks there, go to 3b1b.co/janestreet. Alright, Grant.

Let's talk more about AI and math. What advice do you have about using LLMs to learn? I I so as I was describing for a lot of well known concepts, find them very helpful.

And but often, it just a couple of further messages down and I'm trying to understand something. And I just they're they're so confused themselves or confusing me and they don't explain it the right way. And then I'm just I know that talking to the right human could clear up my confusion in three minutes.

I don't know. And then I feel like more and more we're gonna want to use these things as somebody's talked a lot about education Yes. And, you know, representation stuff.

We're gonna wanna use these things to learn things.

Speaker 2

yeah, have you have you noticed the ways to use them more productively to understand concepts? I'm curious to hear your take on I mean, I'll give mine. I even pre LLM, I feel like a relevant insight in learning was recognizing that, like, who matters more than what.

Mhmm. So, like, advice to any college student when they're choosing what courses to take, care a little bit less about your preexisting interests because they're kind of arbitrary right now, and care a little bit more about whether the person teaching it is a good educator and someone you resonate with. I think in choosing what to read, like what books to read, like who the author is maybe matters more than if it's a a prior interest.

So if there's a book you've liked before, read what else that author has written rather than reading another thing on that subject. On and I I'm getting to, like, LLMs on this. So, like, there's a there's a difference in feel for trying to learn something if you look at a Wikipedia page of it versus if you look at let's say, like, it's a philosophy topic and you go to the Stanford Encyclopedia of Philosophy, or if it's a math topic, you go to the, like, Princeton Compendium of Math, The difference there is the articles are deliberately written by one individual who tries to actually craft a motivation around it and everything, whereas Wikipedia, it's this local minimum that's reached where basically every sentence has to be correct.

And I think a good exposition you care a little bit less about correctness on the way, but you can deliberately craft things that are a little bit wrong that you correct along the way that gets edited out in a crowdsourced environment. That LLM So explanations feel to me at the moment a lot like Wikipedia. Yeah.

Which is to say, amazing. Right? Like, imagine world before Wikipedia, like, how how long it would take to, like, find and, like, suss in and everything.

But nevertheless, what's the most useful part of a Wikipedia page? It's often just the references at the bottom. Right?

You look at the key references and you go to them and you read them, and it's like, actually, sometimes that gives a much better overview of it. So often I like to just ask an LLM, who should I read? Right?

Like and and maybe I can even give some specifics on ways I wanna learn. I actually got gaslit by this once where I remember trying to learn about, like, semiconductors or something. I was like, this feels very visual.

This is all, like, text. I'm like, is there any really good, like, well visualized math video or not math. Sorry.

A well visualized video kind of like explaining the concepts that you're getting at in Claude? It's like, yeah. Here's a couple in the top one.

It was like, here's one from three Bloem Brown. I'm like, I can guarantee that there's not. I go ahead.

And it was an actual video, an actual link, but it just had to misattribute it to someone else's I mean, it was good. And I like, I had a much better experience clicking over and watching that video to learn about the thing rather than, like, trying to proceed forward with questions there. So in that sense, basically using it like a very souped up version of Google on, like, zero in on the right human written resource.

What about you? Like, what you you you engage with these a lot. What's the best way to I think you put your finger on it.

Speaker 1

some artifact that a human's produced, whether it's an article, a book, a video, that organizes the relevant concepts in the correct way and builds up the motivation of why building up the next idea would be relevant to solving the next problem you did encounter and the next idea and the next idea. And then using the LLMs to just do a little bit pruning around this this this branch that the book has identified.

Speaker 2

textbook on The Chaos One? Yeah. Chaos and Nonlinear Dynamics.

Yeah. I love that book.

Speaker 1

it was it was like bliss. It was like your videos in like a book form. He's so good.

It was super fun. And the way I was learning it is like I'd have on one third of the screen, his like lecturer from university. On one third of the screen, I'd have that part of the textbook, and on one third of the screen, I have an LLM.

And I was actually thinking, if I was back in college and watching this lecture live, it would just totally go over my head. Like, these kids must be really smart. Because I'm, like, pausing and like reading the textbook and talking about LLMs and then restarting again.

But with him curating what is the right order to understand concepts, what is the right problem to motivate understanding a concept. Oh, also another thing LLMs are really bad at is a thing a really good human can do is when you ask a question, they say like, actually, you're just like not really thinking about this topic the correct way. Yeah.

Like, the question you wanna be asking Yeah. The correct way to organize these concepts is x. Yeah.

And LMs just can't really do that. Yeah. It it's it's a little too placate.

Speaker 2

you know, that's very, like, oh, what an insightful question, you know, that kind of thing. You wanna you wanna strip that down? That's a good point.

And I think that cuts to theory of mind a little bit. Yeah. Recognizing that to ask a certain kind of question reveals that the mental structures are not at least they're not the same as what the explainer has.

And sometimes people do this to a fault. I think a really good teacher let's say you have a middle school math classroom or something. If a student asks a question that suggests they're thinking about it in a different way, it's actually really hard to take seriously in the moment, hang on, could you get to a right answer with that before you say, oh, instead of that, let's do this.

The really good teachers are able to jujitsu the creative way that the student was thinking about it and and and and bring it in. I mean, LLMs aren't doing that, right, when they are not reframing your question. Instead, they kind of, like, run off.

Right. But the very least, it it feels like there's three levels here. And so, like, LLMs at one, good explainer is at another.

But then, like, the the A plus explainer is the one who can, like, jujitsu your way of thinking, and say, like, oh, that's that's where that's useful. And so maybe there is a certain, you know, cycle all the way around where, again, five years from now, the LLMs will still be doing that, but in a better way. Mhmm.

Speaker 1

students who I'm sure email you this question all the time? Look, I want I was curious about doing mathematics. I'm really passionate about the subject.

But seeing all the progress the guys are making, it doesn't I don't know if it makes sense for me to pursue this as a career. And this is not relevant not only to people in mathematics, I'm sure to people who are noticing that their field is more and more getting productivity gains or whatever from AI. So coding is very adjacent to this.

Yeah. What advice do have for people?

Speaker 2

I wouldn't trust any advice that I give. It would maybe be how I'd, like, couch it. But even pre AI, it feels very important for any job that you're gonna go into to really understand like, if we're talking about a job, right, we're not talking about, like, you're a gentleman scientist and you want engage with the math world or something.

You should understand where the money's coming from and what value you're actually adding and the connection between those two. And I think often a surprisingly small amount of thought is put towards that. Especially students, they're in this environment where they probably want to go into math because they've always been good at it, and they've just been rewarded in life for proceeding through the next hoop correctly and next step.

And when they think they want to be a mathematician, it's because it's a version of getting to continue to engage with that. It's like, oh, I'll go, like, where do people get to do this? Rather than thinking, like, what value am I adding to other people, and to what extent is that, like, the reason that that, like, salary is flowing in my direction?

It's actually quite different in different cases. Like, in some cases, it's a very prestigious mathematician, and their presence at a university lends a certain brand value, and that's why the university wants them. In some cases, it's like the NSF grant is given because you've got this public good belief that we have that basic science has, and we've got this institution around that, and there's going be this whole bureaucracy around trying to act as a proxy for what we think that public good is, and a whole song and dance around how to correctly make them predict that your progress will be in the spirit of that funding.

Sometimes it's just straight up teaching. People like to send their kids to an institute that has experts teaching them, and that's what you're doing, and you are providing the brand value by being an expert, and then the direct value by being a teacher. So regardless of whether AIs are proving theorems or not, or whether we're talking in 2016 or 2026, that is a thing that not enough students thinking I wanna be a mathematician think about, but I think it's worth thinking about.

Like, for me, I think that, you know, it's I just, like, wasn't necessarily thinking about it and kind of stumbled into this career path where basically math exploration can be monetized as entertainment. Right? And I, like, stumbled into that.

I'm, like, very grateful that I did, but it was an accident. It wasn't this deliberate thing, and I think I could have avoided relying on serendipity and maybe done that a little bit more by design had I been thinking critically about it. So to your question, if it's the case that you have almost automated theorem proving, And then let's say it's the case that they're also really good explainers, so it's like even to get the human understanding.

I think a lot of the social role that mathematicians serve actually doesn't change that much. You still have a sense of, as a public, we sort of feel like there's value to basic science, and we're trusting in the judgment of mathematicians to determine where their time has best been. The prestige comes from within that community.

It's like other members saying that this was a really good result more than it is like the grant writer who really understands algebraic number theory to understand that it's a good result. And so there's gonna be some inner culture of what constitutes valuable contributions. Maybe it shifts away from theorem proving, and maybe it shifts towards good definition writing.

Maybe it's that museum curator idea, but you're gonna have that same community. As long as society as a whole is still valuing the premise of basic science, and if we're in the abundance world of what AI brings, probably there's more funding in that direction in some sense. On the side of prestige to institutions for who their lecturers are, I actually think teaching is one of the most stable post AGI jobs that there is because it's so relational.

This is where parents want to spend their money if they have an abundance of wealth, is on good teaching and good educating, and it goes so far beyond explanations. Even if LLMs are good explainers, the thing that a teacher is doing is such a social coaching mentor type thing that that's probably one of the most stable careers that's going to exist over the next fifty years. So insofar as what a lot of mathematicians' role overlaps with that, you as the prospective student going into it, you could lean into that.

Actually, I think a lot more students should think about and pay credence to the idea of being just a math educator and the value that that can serve towards the next generation. So I'll I'll couch again on I don't think I'm the one to say, here, prospective young mathematician, here's how you should think about the future, because I'm like a YouTuber. Right?

I'm someone who is is not in the institution that they are thinking of going into, and so I'm speaking as an outsider looking in. But it feels like generally good universal advice. Know where the money is coming from.

Know where you plug into that. And, like, if you're just asking those questions, you're actually already, like, steps ahead of all of the other like fledgling prospective mathematicians.

Speaker 1

Yeah. And and in fact, I think in the crazy world, in the world where within five, ten years, the AIs are coming up with not only solutions to the the millennium prize problems, but coming up with like, just totally novel problems to be solving in the first place and the novel mathematical fields and objects and stuff. It is in that world where, first of all, there's a ton of abundance and two, the the things that AI minds will have, like, gone furthest in, where they will have seen furthest beyond our horizons will be mathematics.

And there will be so much demand of what have the AI seen? Can you explain it to us? Yeah.

I feel like in the in that world, if there's any jobs whatsoever, surely distilling what the AIs have learned will be one of them.

Speaker 2

all of this sort of presumes that it's useless. Right? Like, we're not talking about the actual practical applications of what math is being is being done.

So insofar as there's any economic utility to it, you would imagine that the people who understand it and are able to, like, make the decision of where it should point, like, they actually have a lot more economic value by, like, being able to make that judgment as curator and point this, like, behemoth of, new math, like, pointed in a useful direction. Yeah. Like, suddenly, that's a much more levered move to make than it had been previously.

Can I just ask you about that? Yeah.

Speaker 1

the one question for AI for math is not only can it do it, but is it any good? Yeah. Or is it any good for anything?

You were describing all the ways in mission group theory, we're trying to solve this we're trying to figure out random facts about the roots of different kinds of functions. And now there's all these different applications that are practical across many different fields.

Speaker 2

fields? Or I think there's some fields that probably will I mean, it's it's super spiky. Right?

I think progress in algebraic number theory, it feels unlikely that that then unlocks something. But I I don't know. I remember talking to this mathematician who does more like dynamics and and and, like, PDE solving type stuff, and he was referencing basically, like, group had some ideas that at least if I summarize this right.

It's like the way that Boeing would make planes is they would make it, and then they would do a bunch of tests, they had to disassemble it and reassemble it based on those tests. They essentially had some insights on how to do more things in simulation such that you don't have to deconstruct and rebuild it. It saved Boeing just billions of dollars or something, and then they just started funding that group, which is so that's it's much more obviously application adjacent because PDEs just sort of are that.

So progress in that domain, you would you would imagine, like, actually do unlock some things. And I don't know if it's these, like, step changes, but maybe it's more on the side of, like, engine design becomes just a little bit more fluid or, you know, like, coming up with the right wing shape instead of running a whole bunch of complicated, like, CFD, or maybe you're able to speed up your CFD simulations because of certain pure math insights that makes those more efficient. I bet you'd just see a lot of great incremental improvement there.

It seems less likely that the massive breakthroughs in math immediately turn into this massive economic breakthrough. You solve the Navier Stokes problems, and then that unlocks an ability to simulate more things. But you probably will see at those fringes just some meaningful leakage outside of the pure math insights into other things.

Also, there's a ton of people working on things like AI engineers, like physical engineers, like material science and things like that. You have to imagine that they would be in a good position to look at the AI math insights and decide if they're relevant in some way or not. It's another one of these things where I'm not gonna sit here and put a flag in the sand predicting that there will be.

It'll be a little bit disappointing and a little bit surprising if there weren't over the next five years, economically valuable improvements that were made that were directly, like, referable to the, like, AI progress in math.

Speaker 1

you know, it it it wasn't doing any of the math that actually directly touches physical world. Yeah. I mean, to your to your point about, well, a lot of history in mathematics is about, like, building up these, like, piles of concepts and connections and whatever.

Yeah. And sometimes the the piles connect with each other or they're you discover an application somewhere else. At the very least, you just build up this huge pile.

Speaker 2

hopefully are useful in other parts of the world. I mean, yeah. It it like I said, one of the interesting things about what's happening is it causes people to step back and ask, like, what is math?

And maybe one of the awkward conclusions of it will be a re revealing, like, oh, man. Over the last like, it's just become wholly useless. Yeah.

Yeah. Yeah. Like, the kind of questions being asked have become, like, so divorced from things that are physically applicable that, like, that's one of the things mathematicians have to come to terms with where everyone will look and be like, hang a second.

Like, are you guys supposed to like, if there's so much that's like 10x progress there. Like, why aren't we seeing it over here? And then that trench is like, ugh.

Every time we wrote those grant proposals and said, like, trust us. Like, the elliptic curve progress is gonna help with, like, cryptography. Like, it, like, shines a light on the fact that, like, maybe it doesn't.

So that's that's one possibility.

Speaker 1

Grant, this is super fun. Thanks so much for doing it. Absolutely.

My pleasure.

Shared via Hopper