Terence Tao – Kepler, Newton, and the true nature of mathematical discovery

Dwarkesh Podcast
20 March 2026 1h 23m
0:00 --:--
Episode Description
We begin the episode with the absolutely ingenious and surprising way in which Kepler discovered the laws of planetary motion.People sometimes say that AI will make especially fast progress at scientific discovery because of tight verification loops.But the story of how we discovered the shape of our solar system shows how the verification loop for correct ideas can be decades (or even millennia) long.During this time, what we know today as the better theory can actually make worse predictions.A

Summary

This episode explores the history of scientific discovery, particularly Kepler's laws of planetary motion, as a lens to understand the future of AI in scientific and mathematical research. Terence Tao discusses how AI excels at idea generation and broad exploration, shifting the bottleneck to verification and the identification of truly fruitful concepts. The conversation also delves into the evolving nature of scientific progress, the distinction between artificial cleverness and intelligence, and the potential for a highly complementary human-AI future in mathematics.

Chapters

Kepler's Discovery of Planetary LawsTerence Tao recounts Kepler's journey from his beautiful but incorrect Platonic solids theory to the empirical discovery of his three laws of planetary motion, leveraging Tycho Brahe's precise data.
AI and Scientific Idea GenerationThe discussion draws an analogy between Kepler's trial-and-error approach and the potential for LLMs to rapidly generate hypotheses, highlighting that verification, not idea generation, is becoming the new bottleneck in science.
Evolution of Scientific ParadigmsThe conversation traces the shift in scientific methodology from theory and experiment to numerical simulation and, more recently, data-driven analysis, with Kepler being an early example of a data scientist.
Challenges of AI-Driven ProgressWith AI driving down the cost of idea generation, the challenge now lies in effectively verifying, evaluating, and identifying genuinely valuable ideas amidst a flood of AI-generated content, which human reviewers are struggling to manage.
Nature of Scientific ProgressThe episode explores how scientific ideas often gain acceptance through the 'test of time' and future applications, and how correct theories can initially appear less accurate or plausible than established but incorrect ones, requiring conceptual leaps.
AI in Mathematics: Current StateTerence Tao discusses the current state of AI in solving mathematical problems, noting that while AI has solved many 'low-hanging fruit' problems, it struggles with partial progress and depth, excelling instead at breadth and applying existing techniques.
Artificial Cleverness vs. IntelligenceTao differentiates between artificial cleverness, which involves brute-force trial and error, and true artificial intelligence, characterized by cumulative, adaptive, and interactive problem-solving that builds on prior understanding.
Formalizing Mathematical StrategiesThe discussion highlights the need for a formal or semi-formal language for mathematical strategies and conjectures, similar to how Lean formalizes proofs, to enable AI to better assess plausibility and contribute to heuristic-driven fields like number theory.
Serendipity and Future of MathTao reflects on the importance of serendipity and 'inefficiency' in fostering unexpected discoveries and collaborations, suggesting that while AI will revolutionize many aspects of mathematics, a hybrid human-AI approach will likely dominate for the foreseeable future.
Career Advice in AI EraTerence Tao advises aspiring mathematicians to cultivate an adaptable mindset, embrace change, pursue curiosity, and be open to non-traditional learning and new ways of doing science in an increasingly unpredictable era.

Topics

Planetary motion lawsAI scientific discoveryData-driven scienceMathematical proof automationHuman-AI collaborationScientific communicationMathematical strategiesSerendipity in scienceFuture of mathematics

People

Terence Tao (guest) Dwarkesh (host) Kepler (mentioned) Copernicus (mentioned) Aristarchus (mentioned) Tycho Brahe (mentioned) Newton (mentioned) Johannes Bode (mentioned) Herschel (mentioned) Leibniz (mentioned) Einstein (mentioned) Aristotle (mentioned) Darwin (mentioned) Edward Dahlnick (mentioned) Thomas Huxley (mentioned) Lucretius (mentioned) Sean (mentioned) Jane Street (mentioned) Gauss (mentioned) Descartes (mentioned)
Key Concepts (16)
Heliocentric model — The astronomical model in which the Earth and planets revolve around the Sun at the center of the Solar System, famously proposed by Copernicus.
Platonic solids theory — Kepler's initial, beautiful but incorrect theory that the ratios of planetary orbits were related to the five perfect Platonic solids inscribed within spheres.
Kepler's laws of motion — Three empirical laws describing planetary motion: elliptical orbits, equal areas in equal times, and the cube of the orbital period being proportional to the square of the semi-major axis.
Inverse square law — Newton's law of universal gravitation, stating that gravitational force is inversely proportional to the square of the distance between two objects, which explained Kepler's laws.
AI idea generation — The concept that AI can rapidly generate a vast number of hypotheses or potential solutions for scientific problems, driving down the cost of this aspect of research.
Verification bottleneck — The new challenge in science where the ability to verify and evaluate the massive number of ideas generated by AI becomes the limiting factor, rather than the generation of ideas themselves.
Scientific paradigm shifts — Major changes in the fundamental concepts and experimental practices of a scientific discipline, often involving the deletion of old assumptions rather than just adding new theories.
Cognitive Copernican revolution — The idea that humanity is currently undergoing a shift in understanding intelligence, moving away from human intelligence as the sole or central form, recognizing diverse types of intelligence with different strengths.
Deductive overhang — The potential for significant undiscovered knowledge that could be derived from existing data or insights if the right approach or 'insight about how to study a problem' were found.
Artificial cleverness — A type of AI capability characterized by trial and error, brute force, and repetition, which can solve problems but lacks cumulative, adaptive, or interactive improvement of understanding.
Artificial intelligence (Tao's view) — A form of intelligence that involves collaborative, adaptive, and cumulative problem-solving, where understanding evolves through discussion and systematic mapping of what works and doesn't.
Formal mathematical strategies — A hypothetical formal or semi-formal language for expressing mathematical strategies, conjectures, and plausibility assessments, which could enable AI to contribute more effectively to the heuristic side of mathematics.
Prime number theorem — Gauss's conjecture, later proven, that the number of primes up to x is approximately x divided by the natural logarithm of x, representing a statistical pattern in prime numbers.
Random model of primes — A conceptual framework in number theory that treats prime numbers as if they were generated randomly with a certain density, allowing for accurate statistical predictions and conjectures, despite their deterministic nature.
Twin prime conjecture — An unproven conjecture in number theory stating that there are infinitely many pairs of prime numbers that differ by 2 (e.g., 11 and 13).
Riemann Hypothesis — A famous unsolved problem in mathematics concerning the distribution of prime numbers, which if proven false, would seriously challenge the 'random model of primes'.
References (8)
The Clock of Universe by Edward Dahlnick book
The Origin of Species by Darwin book
Principia Mathematica by Newton book
Cosmic Distance Ladder by Terence Tao series
Arisnet project
Lean tool
Mathematica tool
WolframAlpha tool
Transcript

Speaker 1: Okay. Today, I'm chatting with Tao, who needs no interaction. Terence, I wanna begin by Speaker 2: having you retell the story of how Kepler discovered the laws of planetary motion because I think this will be a great jumping off point to talk about AI for math. Okay. Yeah. So I've always had amateur interest in astronomy, and so I've loved stories of how the early astronomers worked out the nature of the universe. Keppola was building on the work of Copernicus, was himself building on the work of Paterostarchus. So Copernicus very famously proposed the heliocentric model that instead of planets and the going around the Earth, that the sun was at the center of the solar system and the other plans were going around the sun. And Copernicus proposed that the orbits of the planets were perfect circles. And his theory kind of fits the observations that the Greeks and the Arabs and Indians had worked out over centuries. I think a couple of got interested. He learned about these theories in his studies, and he made this observation that the ratios of the size of the orbits that communicate predicted seem to have some geometric meaning. I think he started proposing that, you know, if you take, say, orbit say, the Earth and you enclose it in, I think, maybe a cube, the outer sphere of that encloses the cube almost matched perfectly the orbit of Mars and so forth. And there were six planets, none of the time, five gaps between them, and there were five perfect platonic solids, the cube, the tetrahedron, isoprotection, octahedron, and dodecation. And so he had this this theory which he thought was absolutely beautiful, that could inscribe these solids between the spheres of the planets, and seemed to fit. And it seemed to him like, you know, God's design of the planets was was matching this mathematical perfection of the platonic solids. So he needed data to, confirm this theory. And at time, there was only one really high quality dataset, almost in existence, okay, which was the so, Takubrahi, this Danish astronomer, very wealthy eccentric astronomer, had managed to convince the Danish government to fund this extremely expensive observatory. This in fact, an entire island where he had taken decades of observations of all the planets, Mars, Jupiter, every night of at least every night for which the weather was clear. With a eye, actually, this is he was the last of the of the naked eye astronomers. And so he had all this data which Kepler could use to confirm his theory. And so Kepler started working with with Tycho, but Tycho was very jealous of the data. He only gave a little bit bits of it at time. And I think Kepler eventually just stole the data. He he copied it and and had to have a fight with with Brahi's descendants. But he did work out he did get the data, and then he worked out to kind of his disappointment that his beautiful theory didn't quite work. Like, the data was sort of off from his platonic solid theory by, you about 10% or something. And he had all kinds of fudges moving the circles around and things that it it didn't quite work. But he worked on this problem for for for years and years. And, eventually, he figured out how to use the data to to work out the actual orbits of, of the planets. And that was incredibly clever, genius amount of data analysis, actually. Yeah, and then he eventually worked out that the the also, actually, ellipses, not circles, which was shocking for him. And then he worked out so he worked out the two laws of planetary two laws of planetary motion of ellipses, also equal areas, super out equal times. And then ten years later, after collecting a lot of data, the furthest planets like Saturn and Jupiter were the hardest for him to work out. Then he finally worked out this third law also, that the orbits, the time it takes for a planet to complete this orbit was proportional to some power of the distance to the Sun. And these are the three famous capital laws of motion and he had no explanation them. For it was just all driven by by experiment. And it took Newton a century late later to give a theory that explained all three laws at once. Speaker 1: The take I wanna try on you Mhmm. Is that Kepler was a high temperature alum. Where Newton comes up with this explanation of why the three laws of planetary motion must be true. And of the way that Kepler discovers the laws of planetary motion or figures out the relative orbits of the different planets is, as you say, a work of genius. Then, know, he's through his career, he's just trying random relationships. And in fact, in the book in which he writes down the third law of planetary motion, it's sort of an aside on the harmonics of the world, is this book about all these different planets have these different harmonies and the reason there's so much famine and mystery on Earth is because the Earth is mi fa mi, that's the note of Earth. So all this random astrology, but in there is the cube square law, tells you what relationship period has to a planet's distance from the sun, which is, as you're detailing, if you add that to Newton's f equals m a and then the equation for centripetal acceleration, you get the inverse square law. Mhmm. And so Newton works that out. But the reason I I think this is an interesting story is I feel like LLMs could do the kind of thing of, like, twenty years. Let's try random relationships, some of which make no sense. As long as there's a verifiable data bank like Brahi's dataset Mhmm. Where, okay, I'm gonna try out random things about, like, musical notes. I'm gonna try random things about platonic objects. I'm gonna all these different geometries. I have this bias that is there's some important thing about the geometry of these orbits. And then one thing works. And as long as you can verify it, it can then draw these empirical irregularities can then drive actual deep scientific progress. Speaker 2: Traditionally, when we talk about the history of science, idea generation has always been kind of the prestige part of science. So, I mean, a scientific problem comes with there's many steps. You know? You have to identify a problem, and then you have identify a good problem to work on, a fruitful problem. And then you need collect data. You need to figure out a strategy to analyze the data, to make a hypothesis. And at this point, you need to propose a good hypothesis, then you need to validate. Yeah. So this and then you need to write things up explain. And There's a there's a a dozen different components. But, yeah, the the ones we celebrate are these of eureka genius moments of of ID generation. And yeah. So so Kepler certainly had to to, as say, cycle through many ideas and and several which didn't work, and and and I bet many that he didn't even publish at all because, yeah, they they just didn't fit. And that's an important part of the process, trying all kinds random things and seeing if they worked. But as you say, it have to be matched by an equal amount of verification. Otherwise, it's slopped. I mean, we we celebrate Kepler, but we should also celebrate Brahi for for his his his asidious data collection with which was 10 times more precise than than any previous observation. And it would that extra decimal point of accuracy was actually essential for for Kepler to get, his, his his results. And, you know, and he was using, you know, equidistant geometry and and and, like like, the most advanced mathematics he could, use at the time to to match his his models with the data. So, like, all aspects had to be in play. You know, the the data and the theory and the hypothesis generation. I'm I'm not sure nowadays that hypothesis generation is the bottleneck anymore. Sciences has changed in in in the centuries since. So classically, sort of the the two big paradigms for for science for theory and experiment. And then in the twentieth century, numerical simulation came along. And so you can also do do computer simulations of of of of to test theories. But then finally, in the late twentieth century, we had big data. Now we we had the the error of data analysis. And so a lot of new progress is actually driven now by analyzing massive datasets first, collecting large datasets, and then drawing the patterns from them to to do slots, which is a little bit different from how science used to work where you you make a few observations or you just have one out of blue idea, and then you collect data to test your idea. That's the classic scientific method. Now it's almost reverse. You collect big data first, and then you try to to get hypotheses from it. I mean, Kepler was maybe one of the first early data scientists, but but even even he didn't start with Tycho's dataset and and analyze it. He he had had some preconceived theories first. But it's it seems that this is less and less the way we make progress in in in yeah. Just because, Speaker 1: yeah, the data is is just so much more massive. It's just much more useful. Oh, interesting. I I actually feel like the mold of twenty eighth century science that you're describing is actually very well described with how in Kepler, where he did have these ideas. 1595 and '96 is where he comes up with first polygons and then platonic objects theory, but they were wrong. And then a few years later, gets Brahe's data. And it's only after twenty years of just trying random things that he gets this empirical regularity. And so it actually feels a closer to Brahe's data as analogous to some massive data vanco simulations, and then we he now he now that you've got the data, you can keep trying random things. But if it wasn't, Kepler would be out there just writing books about harmonics and platonic objects, Speaker 2: there would be nothing to actually verify against. Yeah. Yep. Yeah. So the the data was extremely important, but the distinction I was trying to make was that sort of traditionally, you make a hypothesis, and then you test it against data. Yeah. But now with machine learning and data analysis and statistics and somebody, you can you can start with data and who say statistics work out, laws that, were not present before. And so so Kepler's third law is a little bit like this, except that, for the third law, instead of having a thousand data points that Brahi had, Kepler had, like, six data points. Like, every planet, you knew the length of the orbit, and the distance of the sun, and there was, like, five or six data points. And he did, what we would now call regression. You know? He he could fit a curve to these six data points, and he got a square coupler, which was amazing. But, actually, he was quite lucky, I mean, that these six data points gave him the right conclusion. You know, it's, that's not enough data to be really reliable. There was a later astronomer, Johannes Bode, who took the same the same data, actually, the the distances to to the planets. And inspired by Kepler, I think, he had a prediction that the the distances of the planets formed basically a shift to geometric progression. That he also fit a curve. Except that there was one was one point missing. Alright? So there was a big gap between Mars and Jupiter. His law predicted that there was a missing planet. So it was a kind of a a crank theory, except when Uranus was discovered by Herschel, the the distance Uranus fit exactly this this pattern. And then Ceres was discovered, this asteroid between I think in in the asteroid belt, and it also fit the pattern. And people got really excited that that board had discovered this this amazing new law of nature. But then Neptune was discovered, and it was it was completely, like, way off. And, you know, and and basically, it was just a numerical fluke. You know? There was there six six data points. Yeah. So maybe one reason why Kepler didn't highlight his third law as much as the first two laws is that maybe instinctively, even though we didn't have modern statistics, he kind of knew that with six data points, he had to be somewhat tentative with with the conclusions. Speaker 1: Maybe to ask the question about the analogy more explicitly. Does this analogy make sense to if we have, you know, in the future, we'll have smarter and smarter AIs, and we'll have millions of them. And then they can go out and hunt for all these empirical regularities. It sounds like you don't think the bottleneck in science is finding more things that are for each given field, their equivalent of the third law of planetary motion so that then later on somebody can say, oh, we need a way to explain this. Let's work out the math. Here here's the inverse square law of gravity. Right. So I think AI has basically driven the cost of idea generation down to almost zero. Yeah. In a very similar way to how the Internet drove the cost of communication down to almost zero. Yeah. Speaker 2: Is an amazing thing, but it, you know, it it doesn't make it doesn't create abundance by itself. Yeah. So now the bottleneck is is different. So we're now in a situation where suddenly people can generate thousands of theories for a a given scientific problem. And now we have to to verify them, evaluate them. And this is something which we we have to to change our structures of science to actually sort this out. So, you know, in fact, traditionally, we we build walls. You know? So in in the past, you know, before we had AI Slop, you know, we we had sort of amateur scientists, you know, create you have their own theories of the universe, many of which were basically of very little value. Yeah. And so we brought these, like, you know, peer reviewed publication systems and things to kinda filter out and try to isolate the high signal ideas to to test. But but now that we can generate these these these possible explanations at massive scale, and some of are good and a lot are terrible. I mean, human reviewers, we just it's just they're already being overwhelmed, actually. I mean, many, many journals are reporting AI during submissions. Just I just I just flooding their their submissions. So it's great that we can generate all kinds of things now with AI, but it it means that we have to the rest of the rest of the aspects of science have to catch up. Yeah. So verification, validation, and and assessing what ideas actually move the subject forward and and what which ones are dead ends or or red herrings. And that's that's not something where we've we know how to do at scale. You know, for each individual paper, we can discuss it with, you know, have a debate among scientists and get to consensus in a few years. But when we're generating, you know, a thousand of these every day, it's yeah. This doesn't work. Yeah. So I think there is this incredibly interesting question of you have billions of AI scientists. Mhmm. Speaker 1: Not only how do you gauge which ones are real progress, but how do you I mean, this is actually a question that human sciences had to face, and we've solved somehow. And I'm I actually am not sure how we solve this, but in any given field, let's say in their nineteen forties and there's if you're at Bell Labs or if you're just generally trying to there's these new technologies coming out. Pulse code modulation, basically, do you transfer signals? How do digitize signals? How do you transfer them over analog wires? And then but there's, like, all these papers about the engineering constraints there and the details, and then there's one which is, like, comes up with the idea of the bit Mhmm. Which has implications across many different fields. And you need some system which can then look at that and say, okay. We need to apply this to probability. We need to apply this to computer science, etcetera. And for in the future, the AIs are coming up with, you know, the next version of this kind of unifying concept, and how would you identify it among millions of papers which might actually constitute progress, which have much less general Speaker 2: Right. Unifying ideas. So a lot of it is the test of time. So so many great ideas didn't actually get a great reception at the time that they were first proposed. It was only after some other scientists realized that that they could take it further and apply them to their own. You deep learning itself was actually a niche area of AI for a long time. The the idea of of getting answers entirely through training on data and and not through first principles, you know, reasoning was was was very controversial, and then they would just took a long time before it actually started bearing fruit. You know, you mentioned the bit. You know, I mean, there are there are other proposals for computer architectures than the zero one that is universal today. I think there there there are trits, you zero one, know, three valued logic. And, you know, in an alternate universe, maybe a different paradigm would have would have showed up. People have argued that, you know, that the transformer, for example, is is the foundation of all modern large language models. And it was the first deep learning architecture that really was was sophisticated enough to capture language, but it didn't have to be that way. There there could been some other architecture that was the first to do it. And once that was adopted, it would become the standard. So I think one reason why it's hard to assess whether a given idea is gonna be fruitful is that it it depends on the future. It it depends on and it it it depends on on also on the culture and society. Like like, which ones get adopted, which ones don't. You know? The base 10 neural system in in mathematics, extremely useful, much better than the Roman neural system, for instance. But, again, there's nothing special about 10. It's it's it's a system that we it's useful for us because everyone else uses it. And we've standardized it, and we've built all our computers and our and our number of representation systems around it. So we're stuck with it now, Some people occasionally push for other systems than decimal, but there's too much inertia. You So you can't look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context both in the the past and the future. And so it it may never be something that you can just reinforce and learn the same way that that you can for much sort of more localized problems. Yeah. Speaker 1: It seems often in the history of science when what when a new theory comes up that in retrospect we realize is correct, it seems to make implications that just either make no sense because they're wrong Mhmm. And we realize later on why they're wrong or they're correct but seem wildly plausible at time. So in as you've talked about, Aristarchus had heliocentrism in the third century BC, and then the ancient Athenians were like, this can't be because it would if the Earth is going around the sun, we should see the relative position of the stars change as we're going around the sun. And the only way that wouldn't be the case is if they're so far away that that you don't notice any parallax, which is actually the correct implication. But there's times when the actually, the implication isn't correct, and we just need to graduate to a better level of understanding. So Leibniz would, you know, chide Newton and disagree with nuancy or gravity on the basis that it implied action at a distance. Mhmm. And then there's we don't know the mechanism and Newton himself was sort of stunned that inertial mass and gravitational mass were the same quantity. So all these things were they were which were resolved by Einstein. Yes. Yes. But it was still progress. And so the question for system of peer review for AI would be, even if you can falsify a theory, how would you notice that it still constitutes progress relative to the thing before? Yeah. So Speaker 2: often, actually, the the ultimately correct theory initially is is worse in many ways. Yeah. So Copernicus' theory of of the planets, it was less accurate than Tomuli's theory. You So so geocentrism had been developed for for, you know, a millennium by that point, and they had they had made many, many tweaks and and had very increasingly complicated ad fixes to to make it more and more accurate. And Copernicus' was a lot simpler, but but much as accurate. There was only Kepler that made it more accurate than Tomlin's theory. I mean, science is always a work in progress. You know? So, yeah, so when you only get part of of the solution, it it looks worse than than a a theory which is incorrect, but somehow you've has been completed to the point where it it it kind of answers all the questions. As you say, know, Newton's theory had, yeah, had big mysteries, you know, the equivalence of mass and action at distance, which were only resolved with a very conceptually different approach centuries afterwards. Often, progress has been made, I can not by adding more theories, but by deleting some assumptions that you you have in in in in your mind. So, you know, one reason why geocentrism held on for so long is is we we had this idea that that objects naturally want to stay addressed. This is the Aristotlean notion of physics. And so the idea that the earth was moving, you know, how come we want all sort of all falling over? Know, once you have neutrals of motion, you know, object motion remains in motion and so forth, then then it makes sense. But you had to so conceptually, it it's it's a very big conceptual leap to to realize that that that the earth is is in motion. It doesn't feel like it's in motion. And, like, the biggest advances, you know, Darwin's, theory of evolution, you know, is the the idea that that species are are not static. But, you know, it's it's it's not obvious because you you you don't see evolution in in your lifetime. Well, now we actually can, but but but, you know, it's it's it it it seems it seems permanent and static. You know, right now, we're going through an an cognitive version of the Copernican revolution where we used think that human intelligence is the center of the universe. And now we're actually seeing that there's there's very different types of intelligence that that that are out there with very different strengths and weaknesses. And so our assess assessment of which tasks require intelligence, which ones don't, has to be reordered quite a bit. And so, you know, trying to fit AI into sort of our theories of scientific progress and and what is hard and what is easy. We're struggling quite a lot. You we have to ask questions that we've never really had ask before or maybe the philosophers had, but now we all have to deal with it. This actually Speaker 1: brings up a topic I've been very curious about. So you mentioned Darwin's The Year of Evolution. There's this book, The Clock of Universe by Edward Dahlnick, covers a lot of this era of history we're talking about. Mhmm. And he has this interesting observation in there that the origin of species is published in 1859. Mhmm. The Principia Mathematica is published in 1687. Mhmm. So the origin of species comes out basically two centuries after the Principia. And conceptually, seems like Darwin's theory is simpler. There a contemporaneous biologist to Darwin who reads the origin of species Thomas actually and he says, how stupid not to have thought of that. And nobody ever says about Friendshipia that shining themselves are not having beaten union to gravity. And so there's a question of well, why did it take longer? It seems like a big part of the reason is that the evidence for natural selection cumulative and retrospective, whereas Newton can just like, here's here's many equations. Let me see the moon's orbital period and its distance. And if it lines up, then we've made progress. And so Lucretius actually had the idea, this idea that species adapted their environment in the first century BC, but nobody ever, like, really talks about it until Darwin because there's Lucretius can't run some experiment that people are, like, forced to pay attention. And so I wonder if we'll, in retrospect, end up seeing much more progress in domains which are have this kind of tight data loop where you can verify them quite easily even though they're conceptually much more difficult. Speaker 2: I think one one aspect of science is is not just creating new theory and validating it, but communicating it to others. So so Darwin was actually an amazing science communicator. He wrote in English in natural in no. Natural language. I'm speaking like it. So in In Nauleen. Okay. My my yeah. Okay. I have I have to sort of get out of my my technical mindset. Yeah. Okay. He's spoken plain English. You know, didn't use equations. And he synthesized a lot of, you know, dispatch. Yeah. So, you know, little pieces of evolution had been worked out in the past, but he had this very compelling vision. And and, again, still missing things. Like, he he didn't know the the mechanism for for for for her editor. He didn't have DNA. Okay. And yeah. But his writing style was persuasive, and that that helped a lot. Newton wrote in Latin. He he he had invented, you know, entire new new areas of mathematics just to explain what he was doing. He was also from an era which was where scientists were much more secretive and competitive. So, you know, academia is still competitive, it was even worse back in Newton's day. So he he held back some of his best insights because he didn't want his rivals to get any advantage. He was also, like, a somewhat unpleasant person from what I what I what I gather, actually. So, it was actually only a couple decades after Newton where other scientists explained his work in much simpler terms that they became widespread. So, yeah, the the art of exposition and making a case and creating a narrative is is also a very important part of science. And if you have the data and the it helps, but but people need to be convinced. Otherwise, they will not push it further or they wanna take initial investment to learn your theory and really and really explore it. And that's another thing which is really hard to reinforce and learn on. Yeah. How can you score how persuasive you are? Okay. Well, okay. There's the entire marketing departments who trying to do this. So maybe it's good that AI not are yet optimized to be persuasive. So, yeah, there's there's a social aspect to science. You'd you'd so even though we've tried us also having an objective side to it where there's data and there's experiment and validation, we still have to tell stories and convince our fellow scientists. And that's a soft squishy thing. Like, it it's, you know, it's it's a combination of data and and painting a narrative. And and it's a bit of gap, you know. I mean, as you know, so so even Darwin, as said, there there are pieces of his theory he could not explain, but he could still make a case that, you know, in the future, people would would find transitional forms, that they would find the mechanism of inheritance, and they did. Yeah. I don't know how you can quantify that such a precise way that that you can start to reinforce on learning. Right. Maybe that was to be forever the human side of science. One Speaker 1: takeaway I had from reading and watching yourself on the Cosmic By way, I highly, highly, highly recommend people watch your series with Through the One Round on the cosmic Distance distance ladder. But one takeaway was that the deductive overhang in many fields could be so much bigger than people realize where if if you just had the right insight about how to study a problem, you might be surprised at how much more you could learn about the world. And I wonder if you think that's sort of a product of astronomy at the particular times in history that you're studying, or is this that based on the data that is incident on the earth right now, we could actually divine a lot more than we happen to know. Right. So astronomy was one of the first sciences to Speaker 2: really embrace data analysis and and squeezing every last possible drop of information out of information they had because because data was the bottleneck. I mean, it still is the I mean, it's it's really hard to to collect astronomical data. So astronomers are the best, you know, almost world class in extracting, know, almost like Sherlock, know, it's just like extracting all kinds of conclusions from little traces of data. I hear a lot of hedge funds, they're preferred hires in this primary PhD. I understand they also are very interested for other reasons in extracting signals from Speaker 1: various random bits of data. Speaking of clever ideas, of my listeners, Sean, solved the puzzle that Jane made for my audience and posted a great walkthrough on X. For context, Jane Street trained Arisnet and then shuffled all 96 layers and then challenged people to put them back in the right order using only the model's outputs and training data. You can't brute force this. There's more possible orderings than atoms in the universe. So, broke the problem into two different parts. The layers into 48 different blocks, and second, put those blocks in the right order. For pairing, realized that in a well trained resonant, the product of two weight matrices in a residual block should have a distinctive negative diagonal pattern. And this arises as a way for the model to keep the residual stream from growing out of control. From this insight, was able recover the right pairings. For ordering, Sean noticed that the model seemed improve if he sorted the blocks by the size of their residual contributions. Starting with that rough approximation, combined a clever ranking heuristic with local swaps to recover the exact right order. His full walkthrough is linked in the description. Don't worry if you didn't get to this puzzle in time though. There's still one up about backdoor LLMs that even Jane Street doesn't know how to solve. You can find it at janestreet.com/thwarcash. Speaker 2: Back to Terrence. We we we do under explore sort of how to extract extra information from from various signals. Like, I I just to to to pick one random study. I I remember reading once that that people have discovered we're trying to to measure how often scientists actually read these citations that the papers that excite. So how do you measure this? Okay. You could try to survey different scientists, but had some clever trick. So many citations have little typos like you know, a number is wrong or or punctuation symbol is wrong. And they they measured how often a a type word got copied from one reference to to the next, and and they could infer whether an author was actually just copying it, cutting it, pasting a reference without actually checking it. And so from that, they they were able to infer some measure of of sort of how much attention people were paying. So there are also clever tricks to extract. You know? So these questions you you posed earlier of, you know, how can we assess whether scientific development is fruitful or or interesting or or represents real progress? You know, maybe there are really useful metrics and footprints of this of this of this this phenomenon in in a data data you know, we can we can examine citations and and, like, how often something is mentioned in a conference or something. And and maybe that there there's there's there's a lot of social sociology of science research to be done that could actually detect these things. Maybe we usually get some astronomers on the case actually. Speaker 1: So I think this brings us nicely to the progress that from the outside it seems like AI for math is making. Mhmm. And I think you had a post recently where you pointed out that over the last few months, programs have solved 50 out the 1,100 odd burdos problems, but then I think I don't know if it's still correct, but as of a month ago, you said that there had been a pause because the low hanging fruit had been picked. Speaker 2: First of all, I'm curious if actually that is still the case that we have picked the low hanging fruit, and now we're now we're at this plateau currently? It it does seem so. I mean, there's still activity at the earliest yeah. So so 50 odd problems have been solved with AI systems, which is great, but there's, like, 600 to go. Right. And people are still chipping away at one or two of these right now. We are seeing a lot fewer sort of pure AI solutions now where the AI just one shots the problem. So so there was a month where that happened and that has stopped. Not for lack of trying. I know three separate attempts to get Frontier Model AI to just attack every single one the problems simultaneously. They pick up some minor observations or maybe they found some problems I already saw from the literature, there hasn't been any further AI purely powered solution yet. People are using AI a lot currently. So someone might use AI to generate a possible proof strategy, and then another person will use a separate AI tool to critique it or rewrite it or generate some numerical data for it or do a literature survey. And and some problems have been solved by ongoing conversation between lots of humans and lots of AI tools. But it it it does seem like it it was this this one off thing. So maybe one analogy to for these problems is, like, imagine, like, there's there's there's all these that you're in some sort of mountain range with all kinds of of cliffs and walls. Maybe there's little wall which is maybe like three feet high and one that's six feet high and then there's 15 feet high and then there's some Mao high cliffs. And you're trying to climb as many of these cliffs as possible, but it's in the dark. We don't know which ones are tall, which ones are short. So, you know, we try light some candles and make some maps and slowly we kind of figure out some them are climbable. Some of them identify some partial track in the wall that you can reach first. And then these AI tools, they're of like these jumping machines that can kind of jump two meters in the air, you know, higher than any human. And sometimes they jump in the wrong direction, sometimes they they crash, but sometimes they they they can reach the tops of of of the lowest, you know, walls that we couldn't reach before. And so we've just basically set them loose in this mountain range hopping around and then there's this exciting period where they could actually find all the low ones and they could reach them. But then there's been no I mean, maybe if the next time there's a big advance in the models, then they will try it again and maybe a few more will be breached. But it it it's a different style of doing mathematics than sort of the you know? So normally, we would hill climb, you know, we would we would make little markers and try identify partial things. And, you know, these tools, they either succeed or they fail. And they they've been really bad at creating sort of partial progress or identifying intermediate stages that you you should focus on first. Again, going back to this previous discussion, we don't have a way of evaluating partial progress. Same way we can evaluate a one shot successful Speaker 1: failure of solving a problem. So there's two different ways to think through what you've just said. And one of them is more bearish on AI progress and one of them is more bullish and bearish on being. Oh, they're only getting to a certain height of wall, which is not as high as humans are reaching. And the second is that, well, they have this powerful property that once they achieve a certain waterline, they can fill every single problem that is available at that waterline, which we simply can't do with humans where we can't make a million copies of you and give each of them a million dollars of inference compute and have you do a hundred years of subjective time research on a 100 different problems at the same time or a million different problems at same time. But once AI's reached Terence Tower level, they could do that. And then once they reach intermediate levels, they could do they could do the intermediate version of that. So the same reason that we should be bearish now is the reason we should be especially bullish, not even when they achieve superhuman intelligence, but just when they achieve human level intelligence because their human level intelligence is qualitatively Speaker 2: wider and more powerful than our human level intelligence. I agree. Yeah. So they at breadth, and humans excel at depth, and human experts at least. Yeah. So, I think they're very complementary, but our current, way of doing math and science is focus on depth because that that's where the human, expertise because humans can't do breath. But, yeah. So we we have to redesign, the way we do science to take full advantage of of this breath capability that we now have. So as I said, we do we should have a lot more effort in creating very broad class of problems to work on rather than than one or two really deep important problems. I mean, we should still have the deep important problems, and humans should still be working on them. But but now now we we have this other way of of of of doing of doing science. You know? I mean, we can explore entire new fields of science by by first getting the these broad, moderately competent AIs to sort of map it out and clear out all the the easy make all the easy observations, okay, and then identify certain islands of difficulty, which, you know, then human experts can come and work and on. So I I I see very much a future of very complementary science. Eventually, you would hope to get both breadth and depth, you know, and and somehow get the both best best best of both worlds. But I think we we need practice with the breadth side because it's too new. We don't even have the paradigms really to make full advantage of it, but we will. Then science to be unrecognizable after that, I think. Speaker 1: To this point about complementarity, programmers have noticed that they're way more productive as a result of these AI tools. And I don't know if you as a mathematician feel the same way, but it does seem like one big difference between vibe coding and vibe researching is that with software, the whole point of the thing is to have some effect on the world through your work. And if it leads to you better understanding a problem or you coming up with some clean abstraction to embody in your code, that is instrumental to the end goal. Whereas maybe with research, the reason we care about solving the Millennium Privacy Problem is presumably that in the process of solving them, our we discover new mathematical objects or better new techniques and those who understand our civilization's understanding of mathematics. And so the proof is sort of instrumental to the intermediate work. I don't know if you agree with that Speaker 2: dichotomy or if that in any way will explain the relative uplift we'll see in software versus research. Right. Yeah. So so certainly in in math, the process is often more important than the problem itself. The problem is kind of a proxy for for measuring your progress. I think even in software, there's there's different types of software tasks. I mean, the you know, like, if you're just trying to create a web page that does the same thing that 1,000 other web pages do, there's there's sort of no skill to be learned. Well, there's still some skill maybe that the individual programmer could pick up. But, you know, for for kind of a boilerplate type code, definitely, you know, it's something that you should definitely offload offload to AI. But, you know, sometimes once you make the code, you know, you still maintain it and and there's issues with upgrading it and making it compatible with other things. And and that, I think, I've I've heard that that program is our reporting, you know, that even if if if an AI can create the first prototype of of a a tool, making it mesh with everything else and and making it interact with the real world in way they want. I mean, it's that's an ongoing process. And if you didn't have the skills of that you pick up from from from writing the code, that that may that may may impact your ability to maintain it down the road. So, yep, certainly, mathematicians, we you know, we've we've used problems to build intuition and to train people to have a good idea as what's true, what to expect, what is provable, what is so just getting the answers right away may actually inhibit that process. I mean, so as I made distinction between theory and experiment before. So in most sciences, there's an equal division between there's a theoretical side and experimental side. But in math has been almost unique. It's almost entirely theoretical. We we pay the premium on sort of trying to to have coherent, clean theories of of things are true and and false. And we haven't done much experiments as to the like, you know, maybe we have two different ways to solve a problem, which one is is more effective. We have we have some intuition, but we haven't done large scale studies where we take a thousand problems, and we and we we just test them. But we can do that now. So I think AI type tools, we really will actually revolutionize the ex the experimental side of math where where you don't care so much about individual problems and and the process of solving them, but, yeah, you you wanna gather just large scale data about about what things work, what things don't. Know, same way that if if you want to do if you're a software company and and you wanna to roll out a thousand pieces software, you know, you don't really wanna handcraft each one and learn lessons from each. You just wanna find what are the workflows that let you scale. So we we don't yet we we we the idea of doing mathematics at scale is at its infancy, but that's where AI is really gonna revolutionize the subject. Interesting. Speaker 1: I feel like a big crux in these conversations about how much how good AI will be for science is I think you said this. It's like, oh, they they're using existing techniques and modifying them. And it would be interesting to understand how much progress one can make simply from using existing techniques. Like, how much of if I looked at the top mat journals, how many of them are how many of the papers are coming up with whatever coming up with the technique means doing that versus using existing techniques and and new problems and what the overhang is where if you just applied every known technique to every open problem, would that just constitute a humongous uplift in our civilization's knowledge or would that not be that impressive and useful? Speaker 2: It's this is a great question. We don't have the data to fully answer it yet. Certainly, a lot of work that human mathematicians do, you know, when you when you take a new problem, one of first things we do is we just find we we look at all the standard things that have worked on similar problems in the past, and we try them one by one. And sometimes that works, and that's still worth publishing sometimes because the the question was important. Sometimes they almost work, you have to add one more wrinkle to it, and that's also interesting. But then, you know, the papers that go in the top journals are usually ones where you, you know, the existing methods can kind of solve, you know, 80% of the problem, but then that because this is is 20%, which is resistant. And and a new technique has to be invented to to fill the gaps. It's very, very rare now that a problem gets solved with sort no reliance on past literature where all the ideas come out of nowhere. That was more common in the past, math is so mature now that it's just so much of a handicap to to not use the literature first. So, yeah, AI tools are really good at are getting really good at the first part of that, just trying all the standard clicks on a problem, often now actually making fewer mistakes implementing them than than humans. It's it's they still make mistakes, but but I've tested these tools, you know, on on on on, like, little tasks that I can do. And and sometimes they pick up errors that I make. Sometimes I pick up errors that they make. It's it's about a tie right now for but, yeah, I I haven't yet seen them take the next step. You know? So so when there are holes in in in the argument where none of the things are working to to then what do you do? And then they can kind suggest random things, and it it but it it often I find that trying to chase them down and make them work and finding they don't work, it wastes more time than saves. It Yeah. So now so I think some fraction of problems that we currently think are hard will will fall from this this method. I mean, especially the ones that haven't received enough attention. So, like, with the Irish problems, you know, like, almost all of the 50 problems that were solved by AIs were ones for which basically there was no literature. I mean, Irish put postpone once or twice. Think maybe some people tried it casually and they they couldn't do it, but they never wrote anything. But it turned out that there was a solution and it was just, you know, maybe combining with this one obscure technique that that not many people know about with some other result in the literature. And that's the kind of love the median level of what AI can accomplish. And that that's really great. It clears out 50 of these problems. So I think you'll see some isolated successes. But this is but what we found so people have have done large scale sweeps of these early problems. And, like, if you only focus on the success stories, the ones that they get they get broadcast on social media that looks amazing, you know, like, all these problems that haven't been sold before for decades, now that now they're falling. But whenever we do a systematic study, any given problem, an AI tool has a success rate of maybe one or 2%. It's that just that they can buy a scale, and and if you just pick the winners, it looks great. So I think it'll be a similar thing happening with, you know, there there are hundreds of of of really prestigious difficult math problems out there. A couple may make, you know, some AI may get lucky and I keep solve them. And there was there some some backdoor to solve the problem that that that everyone else missed, and that will get a lot of publicity. But then people will try these fancy tools on their own favorite problem, and they will again experience the one to two percent success rate. Right. So there'll be a lot of noise amongst the signal of sort of when they're working, when they're not. We have to do yeah. It's increasing well, it'll increasingly important to collect these really standardized datasets. You know, there are efforts now to create a standard set of challenge problems for for AI to solve and not just rely on the AI companies to only publish their wins and and disclose the negative results. Speaker 1: So that will maybe give more clarity as to where where we're actually at. Well, I think it's worth emphasizing how much progress in AI constitutes already to have models that are capable of applying some technique that nobody Speaker 2: Yeah. Had written down as applicable to this particular problem. The progress is simultaneously amazing and disappointing. It it is it is a very strange feeling to to to see these tools in action and and that you know? But also be acclimatized really quickly. You know, I remember when when when Google's web search came out twenty years ago Right. And it just blew all the others all the searches out of water. Like, you're just getting relevant hits on the front page, perfectly, almost, you know, exactly what you wanted. And it was amazing. And then after a few years, you just took over granted that that you could just Google anything. And yeah, so a lot of, yeah, I mean, 2026 level AI would be stunning in 2021. And a lot of it, you know, face recognition, natural speech, you had to doing, Speaker 1: know, college level math problems, we just take for granted now. Right. Yeah. Okay. So speaking of 2026, yeah, you made a prediction in 2023. Mhmm. I think by 2026, what was it that it would would be like like a colleague in mathematics or? Yeah, I trustworthy co author if used correctly. Is looking pretty good in retrospect. Yeah, I am pretty pleased. Let us see if can continue this streak. You personally are 2x more productive as a result of AI. What year would you say that? Speaker 2: Yeah. So productivity, I think, is not quite a one dimensional quantity. Like, I'm definitely noticing that the style in which I do mathematics is changing quite a bit and the type of things I do. So for example, my papers now have a lot more code, a lot more pictures. I because it's so easy to to generate these things now. So some plot which have taken me hours to do now, I can I can do in minutes? But in the past, I just wouldn't have put the plot in. It be in the first place. I I would just talk about it in words. So it's hard to merge to measure what two x means. So, yeah, on the one hand, you know, I think the type of papers that I would write today, if I had to do them without AI assistance, they would definitely take five times longer. Interesting. But I would not write my papers that way. Five x? Yeah. But but it's it's because but the the the these are sort of auxiliary I mean, it you know, the you know, so so things that yeah. Things like like like doing a much deeper literature search, supplying a lot more numerics. Yeah. I mean, they they the paper. So, yeah, the the core of what I do, like, not actually solving the most difficult part of of a math problem, that hasn't changed too much. I still use pen paper for that. But, you know, there's lots of there's lots of silly things. I I I use an an AI agent now to to reformat. Like, sometimes, oh, my parentheses are not quite the right size. You know, I used manually change them in my hand, I I can get an AI agent to sort of do all that quite nicely now in the background. So, yeah, they they they really sped up lots of secondary tasks. They haven't yet sort of sped up the the core thing that I do, but it it's allowed me to sort of add more things to to to my papers. Yeah. But by the same token, like, if I were to write a paper I wrote in 2020 again and not add all these extra features, but just have something of the same level of functionality, yeah, then that he doesn't hasn't saved that that much, to be honest. Yeah. So it's made made the papers sort of richer and broader, but not necessarily deeper. Speaker 1: You made this distinction between artificial cleverness and artificial intelligence. And I would like to better understand those concepts. What is an example of intelligence that is not just cleverness? Speaker 2: It's intelligence so is famously hard to define. It's one of these things that you you kind of know it when you see it. But when I when I when I talk to someone and we're trying to collaboratively solve the math problem together, There's this conversation where, know, neither of us knows how to solve the problem initially, but one of us has some idea and and it looks promising. And and so then then we have some sort of prototype strategy, and then we test it, and then it doesn't work, but then we we we modify it. And there's some adaptivity and and continual improvement of of of the idea over time. And, eventually, you know, we sort of we've systematically mapped out what doesn't work, what does work, and and and we can kinda see a path forward, but it's evolving with our discussion. And this isn't not quite what the AIs the AIs can kind of mimic this a little bit. So to go back to this analogy of of these jumping robots, you know, so, you know, they can jump in and jump in and and jump in fail. But but what they can't do is they kind of jump a little bit and they reach some handhold with and then but then they sort of stay there and then they pull other people up and then the titers jump from there. Right. There there isn't this cumulative process which is sort of built up interactively. It it it seems to be a lot more trial and error and just repetition brute force, you know, which can you it scales and it it can work amazingly well in in certain contexts. Yeah, this idea is sort building up cumulatively Speaker 1: partial progress is kind of what's still not quite there yet. Interesting. You were saying, if Gemini three or Claude Speaker 2: 4.5, whatever, solves a problem Yeah. It it is not the case that its own understanding of math has progressed. Or even if it works on a problem without solving it, it's not that it's own understanding of Yeah. It Math has progressed. Yeah. You you run a new session and it's forgotten what what it just did. Right. It had it, you know, it has no new skills to to attach to to to build on on on related problems. Maybe what you just did is part of one zero point zero zero one percent of the training data for the next generation. So maybe eventually, somebody gets absorbed, but yeah. So Terence talks about the importance of decomposing particularly gnarly problems into a series of easier chunks. Even if this doesn't result in the full solution, Speaker 1: approaching problems in this way helps you build up the intuitions and practice the techniques that you'll need to keep making progress. But models today tend to struggle with these kinds of problem solving techniques. That's where Labelbox comes in. Helps you train models not just to get the right answer, to think the right way. They've operationalized these reasoning behaviors into rubrics, giving you the ability to evaluate every important dimension of a model's output. These rubrics go beyond simple correctness. The model reach for the right tools? Did it check its own work and explore alternative paths? How clear was its response? These skills are useful across domains: math, physics, finance, psychology, and more. And they're becoming increasingly important as models take on harder, open ended problems, some of which have multiple solutions and some of which we don't even know the solutions too. Labelbox can get your rubrics tailored to your domain, helping you systematically measure and shape how your models think. Learn more at labelbox.com/dwarkash. One big question I have is, how plausible is it that if we just keep training AIs that get better and better at, you know, solving problems in lean, that they will continue to solve more and more impressive problems, and then we will, in retrospect, be surprised at how little insight be got from some lean solution to proving the rebound hypothesis or something? Or do you it is a necessary condition also in the rebound hypothesis even by an AI that is, like, totally doing it in lean that the constructions which are made, the definitions which are created even in the the lean program have to advance our understanding of mathematics? Or do think it could just be assembly code Google to Gook? Speaker 2: Yeah. We don't know. I mean, some problems have been basically solved by pure brute force. A full color theorem is is a famous example. We have still not found a conceptually elegant proof of this theorem. It it basically and and maybe we never will. I mean, some problems may only be solvable by just splitting into some enormous number of cases and and doing a brute force, unincible computer analysis on on each case. I mean, part of the reason that we we prize problems like the human hypothesis is we're pretty sure that something amazing has to a new type of mathematics has to be created or a new connection between two previously unconnected areas of mathematics has to be discovered to to make this work. We we don't even know what the shape of the solution is, but it doesn't feel like a problem that will be solved just by exhaustively checking cases or something. I mean, it could be false actually. So we we could actually okay. There is an unlikely scenario that that hypothesis is false, and there's this you can just compute, oh, he's a zero off off the line, and a massive computer calculation verifies it. That would be very disappointing. I don't know. I I I I do feel that, you know, fully autonomous one shot approaches are not the right approach for these problems. I mean, I think you you will get a lot more mileage of interplay between between humans collaborating with these tools. And I can see one of these problems being solved by by some smart humans assisted by some extremely powerful AI tools, but the exact dynamic may be very different from what we envision right now. I mean, it could be a collaborate collaboration of type that we just doesn't exist yet. Yeah. I mean, we there may be a way to to generate, you know, a million variance of the human data function and do some data analysis, AI assisted data analysis, and we discovered some pattern between connecting them, which we didn't know about before, and this lets you transform the problem into a different area of mathematics. Mean, there could be all kinds Speaker 1: scenarios. So suppose the AI figures it out and latent in the lean is some brand new construction, which, you know, if you realize the significance would we would be able to apply it in all these different situations. How how do you recognize it? Right? Like, if if you just again, a very naive question, but you if you if you come up with the equivalent of, like, Descartes comes with this idea, oh, you can have this coordinate system where can unify algebra and geometry. But in Lean Code, would just look like r to r, it would look that significant or something. Or similarly, Speaker 2: I'm sure there's other constructions which have this kind of property. Well, the beauty of formalizing a proof in something like Lean is that just you can take any piece of it and study it atomically. Know, so when I read a paper written by humans with which shows some some difficult problem, you know, there's some some big sequence of lemmas and theorems and things. So ideally, the author will talk talk their way through, you know, what's important, what's not. But but sometimes they don't reveal what what steps were the important ones and which ones are just kind of boilerplate standard steps. But you can study each lemma in isolation, and some of them I can say, oh, this looks fairly standard. This this this resemble something I'm I'm familiar with. I'm pretty sure there's nothing interesting going on here. But this lemma, oh, that's something I haven't seen before. And I could see why if you could if you had this result, that would really help prove the main result. Like, you could you know, you can assess whether some things are are really sort of key to your to your to your argument or not. And Lean really facilitates that. You know, you can you can you can you know, the individual steps are identified really precisely. I think in the future, know, there'll be entire professions of mathematicians who might take a giant, lean generated proof and maybe do some ablation on it or something, and try to remove steps of parts of it and try to find it find more elegant ways, you know, maybe some other AIs to sort of do some reinforcement learning. How can you make the proof more elegant and maybe other AIs will grade whether this is this proof looks better or not. One thing that will change quite a bit in in the near future is is that until recently, papers was the most time consuming and expensive part of of the job. And so you did you did it very rarely. You know, you you only wrote up your results once everything was all the other parts of your argument were were checked out and and things because you just rewriting it again, refactoring was just a total pain. But that's one thing that's become a lot easier now with modern AI tools. So, you know, you don't have to have just one version of of your paper. You, you you can you have one, you know, people can generate hundreds more. So, yeah, one giant messy lean proof may not be very meaningful or understandable on own, but but other people can can can refactor it and do all kinds of of things with them. We have seen if with the Irish problem website, you know, that people will will an AA will generate a proof, and then he has 3,000 lines of code that that verify the proof. But then we people got other AIs to summarize the proof, and and and people write their own proofs. There's actually post processing. Once you actually have one proof, we actually have a of tools now to to deconstruct it and and interpret it. It's a very nascent area of science or mathematics, I'm not as worried about, so some people are concerned, what if the real analysis is proof of a completely incomprehensible proof? I think once we have the artifact of a proof, we can do a lot of Speaker 1: analysis on it. You posted recently that it would be helpful to have a formal or semi formal language for mathematical strategies as opposed just mathematical proofs, which what Lean specializes in. I would love to learn more about what that would involve or look like. Speaker 2: We don't really know. I mean, we've been very lucky in mathematics that that we have worked out the laws of logic and mathematics, but this is actually a fairly recent accomplishment. I mean, it was started by Euclid, you know, millennia ago, but but only in, like, the early twentieth century did we finally list out here the axioms of of mathematics or also the standard axioms of or or ZFC and the axioms of first order logic. And this is what a proof is, and and this we've managed to automate and and have and a formal language for. But there could be some way to assess plausibility of certain you know, so you you have a conjecture that something is true. You you test a few examples and it works out. Like, how does this increase your your confidence that the conjecture is true? We have a few sort of mathematical ways to to model this, Bayesian probability, for example. But they're not but you you often have to, they often, you have to set certain base assumptions and, and people, it's, it's, it's, there's a of subjectivity still in, in these tasks. So it is, it's it's not clear if I I mean, it's this is more of a wish than than a than than a plan to develop these languages, but just seeing how successful having a formal framework in place like Lean has made deductive proofs so much easier to automate and and train AI on. If there was some similar framework yeah. So the the bottleneck for using AI to to to create strategies and and and make conjectures is we have to rely on human experts to, and the test of time to validate whether something's plausible or not. If there was some semi formal framework where this could be done semi automatically in a way that isn't sort of easily hackable to, you know, it is of of course, yeah, the it's really important with these formal proof assistants that that that there are just no there's no backdoors or exploits that that you can do to somehow get your your certified proof without actually proving it because reinforcement learning is just so so good at finding these these the these backdoors. But, yeah, if if a strong framework that sort of mimics how scientists talk to each other in a semi formal way, you know, using data and argument, but also, you know, constructing narratives and and there's some sub there's some subjective act aspect of science that we don't know how to capture in a way that that we can insert AI into them in in a useful way. Interesting. So, yeah, this is a is a future problem. I mean, there are research efforts to, you know, to try to create automated conjectures and and maybe there are ways to benchmark these and and get some some way to simulate this, but this is, it's all very, very new science. Can you help me get some intuition Speaker 1: for, I have two sub questions. One, it would very helpful to have tangible sense of It would be helpful to have a specific example what something like this would look like, the way scientists communicate that we can't formalize yet. And two, it seems almost definitionally paradoxical to say, building up some narrative or building up some natural language explanation, and then also having something which you could have formalized. I'm sure there's some intuition behind where that overlap is, and I'd love to understand that better. Speaker 2: Alright. So so an example of of a conjecture so Gauss interested in the prime numbers, and he computed he he created one of the first mathematical data sets. He just computed the first 100,000 prime numbers or so, hoping to find patterns. And he did find a pattern, but maybe not not the pattern he was expecting. He he found a statistical pattern in the primes that that if you count how many primes there are up to 100, 1,000, 1,000,000, and forth, they get sparser and sparser, but the the drop off in in in the density was inversely proportional to the natural logarithm of of of of of range of numbers. So he conjectured what we're not now for the prime number theorem. The number of primes up to x is like x divided by the natural log of x. And he had no way to prove this. It was it was data driven. So this this was a conjecture. It was revolutionary for its time because, it was maybe the first really important conjecture of of math that was statistical in nature. You know? So normally, you you talk about patterns like maybe the spacing between the primes has a certain regularity or something, but, yeah, but this was really something which it didn't tell you exactly how many primes there were in any given range. It just gave an approximate approximation that got better and better as you went further and further out. But it it helped or so it started the field of what we call analytic number theory. But it was the first in many conjectures like this, many of which got proved, which sort of started consolidating the idea that the prime numbers actually didn't really have a pattern, that they behaved like random random sets of numbers with a certain density. I mean, they had some patterns, like they they're almost all odd. Okay. So there's there's some and and they're not actually random. They're what's called pseudorandom. I mean, there's no random number generation involved in creating the prime numbers. But over time, it became more and more productive to think of the primes as as if they were just generated by some some some god rolling dice all the time and just creating this this random set. And this allowed us to make all these other predictions. So there's a still open conjecture in in in number theory called the the twin prime conjecture that there should be infinitely many pairs of primes that are twins. This is two apart, eleven and thirteen. We can't prove that, and there's actually good reasons why we can't prove it. But, but because of this statistical random model of the primes, we are absolutely convinced it's true. We we know that if if the primes were sort of generated by flipping coins or something that we would by random charge, just like infinite monkeys at a typewriter, we would see, twin primes appear over and over again. And we have over time developed this very accurate conceptual model of what the primes should behave like based on statistics and probability, but it's all mostly heuristic and non rigorous, but extremely accurate. So the few times when we actually can prove things about the primes, it has matched up with the predictions of this what we call the random model of of the primes. So we have this conjectural concept framework for understanding the primes that we, everyone believes in. And it's the same reason why we believe the real hypothesis is true, why we believe that cryptography based on the primes is mathematically secure, things like It's that. All part of this this belief. In fact, one reason why we care about the hypothesis is that if the Reuben hypothesis failed, we knew it was false. It means that it would it would be a serious blow to this model that that this it would mean there's a secret pattern for the primes that we were not aware of. And I think we would very rapidly abandon any cryptography based on the primes because if there one pattern that we didn't know about, there's probably more. And these patterns can lead to exploits in in crypto, and, yeah, it's it's gonna be it'd be a big, big shock. So we really want to make sure that that doesn't happen. So, yeah, it's it's so we've been convinced things like the women hypothesis and things over time, some of it is experimental evidence, some is the few times we've been able to make theoretical results, they've always aligned. It is possible that the consensus is wrong, and we've all just missed something very basic. There have been paradigm shifts in the past in scientific history. But we don't really have a way of measuring this. I think probably because we don't have enough data on on on how mathematicians develops. We we have one timeline of history, and, you know, we we have, like, you know, 100 stories of turning points in history. If if if we had access to a million alien civilizations and each of the the different development of history and and of science in different orders, then maybe we actually have a have a have a decent shot at at at an understanding of how do we measure what is progress and and and what is a good strategy. And we could maybe start formalizing it and and actually having a a framework. Maybe if what we need do is actually start creating lots of mini universes or simulations of AI solving very basic problems, you know, in arithmetic or whatever, but coming over their own strategies for doing these things and having these little laboratories to test. I mean, there are who people investigate like trying to, what's the smallest neural network that can do 10 digit application and things like that. I think we could actually learn a lot just from evolving small AIs Speaker 1: on simple problems. We could learn a lot. I was super excited when reached out about sponsoring the podcast because I've been banking with them for years. I think I opened my first account with them in 2023. Something I've come to appreciate over the last few years is that Mercury is constantly updating things and adding new features. Take their newest feature, Insights. Summarizes your money in and out, showing you your biggest transactions and calling out anything that deserves extra attention. Like maybe your revenue from a particular partner has gone down, or you've got a big uncategorized purchase that needs to be investigated. It's a super low friction way for me to keep tabs on my business and make quick decisions. For example, I tried to invest any cash that I don't need on hand to keep running the business. With insights, with just a couple of clicks, I was able to see exactly how much money I spent in each month of 2025. And that lets me know exactly how much cash I'll need for the next year or so of operations. And then I can go invest the rest. Mercury just keeps adding new features like this. Go to mercury.com to check it out. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Columna, members FDIC. You have to learn about new fields, only very rapidly, deeply enough to contribute to the frontier. So in some sense you're also one of the world's greatest autodidacts. What is your process of learning about a new subfield in math? What does that look like? Speaker 2: I certainly identify with kind of yeah, as talked the, depth and breadth before. And it's it's not purely human AI distinction. I mean, humans also split and so it's I think it was a group split limited hedgehogs and foxes. And it's the hedgehog knows one thing very, well, and a fox knows a little bit about everything. So I definitely I didn't you know, I I I think of myself as a fox. You know, I mean, I work with hedgehogs a lot, and sometimes I can be a hedgehog if need be. But yeah. So I've I've always had a little bit of obsessive streak. If if there's something which I read about, which I feel like I should understand, I I I have the capability to understand this, but I don't understand why it works. There's there's some magic in it that, you know, so someone was able to use it, a type of mathematics I'm not I'm not familiar with and get over that, which I would like to prove, and I can't do it by myself, but they could do it by by their method. Then I wanted find out what was their trick. It bugs me that they someone else can can do something, which I think I I can do, but but I I can't. So I've always had that kind of obsessive completionist type type streak. I've weed myself off computer games because I I I I I start a game. Wanna play it to completionist to all the levels. And so that's one one way in which I I learn new fields. I collaborate with a lot of people who has taught me other types of mathematics. I just make friends with other mathematician who's working on another area of mathematics, and I find their problems interesting, but needed to but they have to teach me some of the basic tricks and and what's known and what's not known, and learn a lot from that. I found that the writing about my what I've learned, I have a blog where I sometimes record things that I've learned. Because in the past, when I was younger, I would learn something and do this call trick and say, okay. I'm gonna remember this. And then six months later, I I forgotten I I I remember remembering it, but I don't but I can't reconstruct my arguments. And it the first few times, was so frustrating to have understood something and then lost it. That it sort resolved. I should always write down anything cool that I've learned. And that's this is part of why how this blog came about. Speaker 1: How long does take you to write a blog post? Speaker 2: It's something I often do when I don't want do other work. You know, like like, there's some referee report or something. There's there's there's something that that it feels slightly unpleasant for me to do at the time. And so writing a blog, it feels creative and fun. Like, it's something that I I do for myself. So maybe depending on on the topic, it could be a quick, you know, half an hour or several hours. But I it doesn't because it's something that I do sort of voluntarily, it doesn't feel like it it doesn't feel time flies when when write these things. As opposed to sort of doing something which I have to do for administrative reasons, but it's just a it's it's drudgery. Okay. Those are tasks that AI is really helping with now, exactly. Speaker 1: Is it if, like, civilization could could from first principles decide how to use Terry Tao's time. You know, it's like a limited resource. How how how what the biggest diff between within the if the veil of ignorance got to decide how to use Terry Tao's time versus what it does now. Speaker 2: So this podcast wouldn't be happening. Yeah. So I could the as much as complain about certain tasks that I don't want to do, but I have to do it. So as as you get more senior in in academia, you get more more responsibilities, and get some more committees and and and whatever. But I have also found that a lot of events that I kind of reluctantly went to because I was obliged to for one reason or another. Because it's outside my comfort zone, I often find interactions with people who wouldn't normally talk to, like like you, for instance. And I've I would learn interesting things and have interesting experiences, and I I would have opportunities to to to to then network with other people that I would never have have done before. So I do believe a lot in serendipity. I mean, I I do optimize my time in in in when I so there's some portions of of my of a day where I do schedule very carefully. But I I have been willing to sort of leave some some portions just, okay. I'm gonna do something which is which is not my usual thing, and then maybe it'll be a waste of my time, but maybe I'll I will learn something. And more often than not, it it's I've I feel like I've gotten a positive experience, which is not something I would have planned for. So I believe a lot on serendipity. And maybe there's a danger actually that, know, in the modern societies, not it's just AI, we've become really good at optimizing everything. And and maybe we are optimizing we're not optimizing a level of optimization that you know, with with with COVID, for example, we we switched, like, we we switched a lot to remote meetings, and so everything was scheduled now. And so we kept busy, at least in in academia. You know, we we met almost the same number of people that we met in person, but everything had to be planned. You had to schedule things in advance. And what we lost out on was sort of the casual, like, know, knocking on a hallway, just meeting someone for you know, while getting a coffee. And there's yeah, serendipitous interactions that you may think are not optimal, but actually are really important. You know, when I was a grad student, I would go down to the library to look I had to look for a journal article. Yeah. Had physically go down to the library, check out the journal, and read your article. And and sometimes the next article, you know, you can just browse through, and and the next article is also interesting. Sometimes it wasn't, but but you could accidentally find interesting things, which is something which has basically been lost now because you can just type in you know, if you if you want to access an article now, you just type it into to a search engine or even an AI, and you can get instantly what you want. But you don't get sort of the accidental things that that you you you might have have gotten if you've it more inefficiently. So, yeah, there've been times when I'm in I I spent a year once at the Institute for Advanced Study, which is a great place to you know, there's no distractions. You you you're there to just do research. And, like, the first few weeks you're there, like, it's great. You're getting all these papers written up that you've been wanting do for a long time. You've been thinking about problems for blocks or hours of a time. But I find if I'd stay there for more than several months, like, I'd I run out of of inspiration somehow. Like, I get bored. I actually serve internet a lot more. You actually do need a certain level of distraction in your life. It somehow adds enough randomness high temperature if need. So, yeah, I don't know the optimal way to schedule my life. It just seems to work. Speaker 1: I'm very curious when you expect AIs that can actually do Speaker 2: frontier math better than at least good as well as the best human mathematician. I mean, in some ways they're already doing frontier math that is super intelligent that humans can't do, but it's a different frontier from what we're used to. Right. I mean, could argue that calculators were doing frontier math Right. That humans could not accomplish, but it wasn't number crunching. Right. But replacing Terry Tao completely. Mean, what do you want me for? Speaker 1: You'll just go on all the podcasts after. Speaker 2: I'm not sure if it might not be the right question to ask. I think within a decade, lot of things that mathematicians currently do, where we spend a lot of the bulk of our time doing it and a lot of stuff we put in our papers today can be done by AI. But we will find that that actually wasn't the most important part of what we do. Know, a hundred years ago, a lot of mathematicians were just solving differential equations. Physicists Like, needed some exact solution to some system, just, they hired a mathematician to logologically go through the calculus and work out the solution to this fluid equation, whatever. A lot of nineteenth what century mathematician would do, you could make a call to Mathematica or WolframAlpha or a computer algebra package now more recently in AI and it would just solve the problem in a few minutes. But we we moved on. We did we worked on different types of problems after that. You know, once computers came along, you know, computers used be human. Right? People used to laborously create log tables and work out primes as Gauss did, And that has all been outsourced to computers, but but we moved on. In genetics, you know, to to sequence at the the genome of a single organism, that was an entire PhD of a geneticist. You know, so carefully, you know, separating all the chromosomes and one whatever. And now you can just spend a thousand dollars and send it to a sequencer and get it done, but genetics is not dead as a subject. Move to a different scale. You maybe you study whole ecosystems rather than individuals. Speaker 1: I take your point, but on the question of, well, when is most mathematical progress, almost all mathematical progress happening by AI? So if you find out, oh, year a millennium price problem has been solved, would put, know, a 95% odds that an AI did it autonomously. Surely, there will be such a year. Speaker 2: I guess. I mean, I I I do believe that that hybrid human plus AIs will will dominate mathematics for a lot longer. It it's it will depend it will require some additional breakthroughs be beyond what we already have. So it's it's gonna be sarcastic. You know, I think, you know, AIs currently are very good certain things, but but really tear award others. And and while you can sort of add more and more frameworks on top to kind of reduce the error rates and and and make them work with each other a bit more and so forth, I I it feels like we are we don't have all the ingredients to, like, really have a truly satisfactory sort of replacement for all intellectual tasks. It is complimentary currently. It's not a replacement. But maybe the I mean, because the current level of AIs will accelerate science in so many ways, hopefully, you know, new discoveries, new breakthroughs will happen quickly. Mean, it's possible that also by somehow destroying serendipity, we actually inhibit certain types of progress. Anything is possible really at this point. I think the world is very, very unpredictable at point in time. What is your advice to Speaker 1: somebody who would consider a career in math or is early in a career in math? Especially in light of AI progress. How should they be thinking about the career differently if at all as a result of AI progress? Yeah. So we live in a time of change. Speaker 2: It is as I said, it it we live in a particularly unpredictable era. And I think in like, things that we've taken for granted for centuries may not hold anymore. So, yeah, the way we do everything and not just mathematics will change. And, you know, so I I think which is you know, I mean, in many ways, I would prefer the much more boring, quiet era where things are much the same as they were ten years ago or twenty years ago. But so I think one just has to embrace this that this it we're there's gonna be a lot of change and that, you know, the things that you study, some of them may may become obsolete or revolutionized, but but some things will be retained. And so you somehow always have to keep an eye on yeah. Like, yep. There'll be a lot of opportunities for things that you you wouldn't be able to do before. So, you know, I mean, in in math, you know, you previously had to basically go through years and years of education with math PhD before you could contribute to the frontier of of research. But now it's quite possible at the high school level or whatever that that you could get involved in math project and actually make a real contribution because of all these AI tools and and and lean and everything else. So there'll be a lot of nontraditional opportunities to to learn. So you need a very adaptable mindset. Yeah. There'll pursuing more things just for curiosity for playing playing around. And and, I mean, you still need to get your credentials for I mean, thank you for a while. It'll still be important to to sort of still go through traditional education and and learn math and science so forth the old fashioned way for a while. Yeah, but you should also be open to very, very different ways of doing science, some of which don't exist yet. Yeah. So it's it's it's a scary time, also very exciting. Yeah. Awesome. That's a great note to close on. Thanks. Thanks so much. Yeah. Thank Pleasure.

Shared via Hopper