He won a Nobel here for AlphaFold. Then he left. - John Jumper

Machine Learning Street Talk (MLST)
22 June 2026 53 min
0:00 --:--
Episode Description
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at notion.com/mlstProtein folding stalled biology for fifty years. A sequence of amino acids dictates a three-dimensional shape, but reading that shape meant a year and roughly $100,000 of crystallography per structure. Then AlphaFold 2 won CASP14 so decisively the organizers called the problem essentially solved.In this documentary cut, John Jumper, who shared the 2024 Nobel Prize in Chemistry and has since

Summary

This episode features John Jumper, co-creator of AlphaFold, discussing the breakthrough in protein structure prediction that won the 2024 Nobel Prize in Chemistry. Jumper explains the technical innovations behind AlphaFold, its impact on biology and drug discovery, and his recent move to Anthropic. The episode also highlights how AlphaFold is empowering researchers globally, including efforts to train scientists in Africa.

Chapters

Introduction and AlphaFold ImpactOverview of the protein folding problem, AlphaFold's breakthrough at CASP14, and its transformative impact on biology.
AlphaFold's Scientific ContextJohn Jumper explains the biological importance of protein folding and the challenges of experimental structure determination.
AlphaFold Architecture and InnovationsDetailed discussion of AlphaFold's machine learning architecture, including Evoformer, SE(3) invariance, and iterative refinement.
AlphaFold Versions and Diffusion ModelsComparison of AlphaFold versions, clarifying misconceptions about AlphaFold 3 as a diffusion model and its geometric refinement approach.
Machine Learning and UnderstandingExploration of prediction, control, and understanding in AI, and the role of data versus code in AlphaFold's success.
AlphaFold's Broader Impact and FutureDiscussion on AlphaFold's role in drug discovery, its limitations, and the future of hybrid AI models in science.
Training African Scientists with AlphaFoldInterview with Emmanuel Ny about using AlphaFold to accelerate research in Africa and training scientists to utilize the technology.
Closing Remarks and SponsorshipFinal thoughts on AlphaFold's significance, John Jumper's move to Anthropic, and sponsor messages.

Topics

Protein foldingAlphaFold architectureStructural biologyMachine learningDrug discoverySE(3) invarianceDiffusion modelsAI for scienceHybrid AI modelsScientific predictionBiological dataTraining scientistsGlobal research access

People

John Jumper (guest) Demis Hassabis (mentioned) David Baker (mentioned) Emmanuel Ny (guest) Tim Scarfe (host)
Key Concepts (10)
Protein folding problem — The challenge of predicting a protein's 3D structure from its amino acid sequence, a problem that stalled biology for decades until AlphaFold's breakthrough.
CASP competition — A biennial scientific experiment evaluating protein structure prediction methods, where AlphaFold 2 decisively outperformed competitors in 2020.
Evoformer architecture — A core component of AlphaFold 2 combining axial attention mechanisms to integrate evolutionary and geometric information for protein folding prediction.
SE(3) invariance and equivariance — Geometric deep learning techniques used in AlphaFold to respect 3D spatial symmetries, improving the model's ability to predict protein structures accurately.
Frame Aligned Point Error (FAPE) — A novel loss function developed for AlphaFold that measures error in local reference frames of protein residues to improve geometric accuracy.
Iterative recycling — AlphaFold's mechanism of repeatedly refining structure predictions through multiple passes to improve accuracy and convergence.
AlphaFold 3 diffusion model — An extension of AlphaFold incorporating diffusion techniques to handle protein complexes and ligand binding, focusing on geometric refinement rather than image-like diffusion.
Prediction vs. understanding in AI — Distinction between AI models predicting outcomes and humans deriving mechanistic understanding, with AlphaFold excelling at prediction but not full biological understanding.
Data vs. code in machine learning — The debate on how much of a model's success comes from its architecture (code) versus the training data, with AlphaFold showing the critical role of both.
Hybrid AI models — Customized AI systems combining domain knowledge and engineering, exemplified by AlphaFold, as opposed to generic foundation models.
References (8)
AlphaFold by John Jumper et al. project
CASP (Critical Assessment of Structure Prediction) by CASP organizers project
AlphaFold Nature Paper by John Jumper et al. paper
Computational Protein Design by David Baker field
Cryo-Electron Microscopy by Various technique
Isomorphic Labs by Alphabet company
Google DeepMind by Google company
Anthropic by Anthropic company
Transcript (37 segments)
Speaker 1

When you need to build up your team to handle the growing chaos at work, use Indeed Sponsored Jobs. It gives your job post the boost it needs to be seen and helps reach people with the right skills, certifications, and more. Spend less time searching and more time actually interviewing candidates who check all your boxes.

Listeners of this show will get a $75 sponsored job credit at indeed.com/podcast. That's indeed.

com/podcast. Terms and conditions apply. Need a hiring hero?

This is a job for Indeed sponsored jobs.

Speaker 2

to athletic ish stuff like a half mile stroll. Get those steps in. Head to Sierra or sierra.

com for the brands you want at prices that let you do it all. From athletic to athletic, Sierra's got it. I don't really love the bitter lesson as people try and apply it.

In fact, AlphaFold two is the opposite of that.

Speaker 5

type problems in biology.

Speaker 4

We predict nature level science with the press of a button in a very narrow category of nature level science of the structure of a specific protein.

Speaker 6

the system that predicted 200,000,000 protein structures. In 2024, he won the Nobel Prize for chemistry. And now Jumper is leaving DeepMind.

But what did AlphaFold solve? What remains unsolved? And could AlphaFold be the template for AI for science?

Speaker 4

We are not trying to tell you everything. We are not a model of the entire cell. You try it.

You measure. Nine times out of 10, you find out you're wrong. Right?

If you're wrong nine times out of 10, you're a very successful machine learner. You're incredibly productive. So for half a century, structural biology had a massive bottleneck.

Speaker 6

DNA was easy to read, but protein structures were not. A protein structure begins as a chain of amino acids and then often with help from the cell they settle into a three-dimensional shape. And that shape determines what it binds, what chemistry it catalyzes, where it sits in the cell and whether it even works at all.

But from a machine learning perspective, if you only have the sequence, can you predict the fold?

Speaker 7

Can you predict the structure? We've discovered more about the world than any other civilization before us. But we have been stuck on this one problem.

How do proteins fold up?

Speaker 6

So every two years, there's a big scientific experiment called CASP. Essentially, teams from around the world gather to see if they can predict protein structure from sequences based on recently done but not yet publicly available experiments. So for many decades, the progress was incremental until 2020 when John Jumper's team, AlphaFold, they produced a result which was significantly better than the competition.

For many single chain targets, the predictions from AlphaFold were so close to the targets that the organizers of the event said that the problem had been essentially solved. So a protein structure that might have taken a year of specialist work can now be predicted and operationalized in minutes.

Speaker 5

Everyone, the whole team. It's been an incredible effort.

Speaker 4

Congratulations on this work. It is really outstanding.

Speaker 5

AlphaFold represents a huge leap forward that I hope will really accelerate drug discovery and help us to better understand disease.

Speaker 7

It's mind blowing. You know, these results were, for me, having worked on this problem so long after many, many stops and starts and will this ever get there, suddenly, this is a solution. We'd solve the problem.

Speaker 6

And fair play to DeepMind. So they could have kept this close to their chest, but they decided to release it. They released a database with over 200,000,000 predicted protein structures.

Now it's important to emphasize the word predicted. These are not experiments, but these are the basis for a lot of interesting new search work that the world of biology can now perform. And today, AlphaFold is now used by more than 3,000,000 people in over a 190 countries.

So in 2024, the Nobel Prize Committee, made a formal verdict. Half of the chemistry prize went to David Baker for computational protein design, and the other half went to Demis Hassabis and John Jumper for protein structure prediction. So AI had become a new tool for chemists, a way of seeing molecular structure at a level of resolution that structural biologists couldn't even have dreamed of before.

Speaker 4

So it's it's absolutely wonderful to be here. It's it's truly an extraordinary honor to tell you about this work, to tell you about the work of our team in protein structure prediction. So by about 10:30, I said, Oh, well, I guess not this year.

And I told my wife, and she goes, No, no, wait. And as she's telling me to wait, my phone lights up with a phone call from Sweden. And thankfully, it was not the world's meanest prank call.

Speaker 6

And all of this makes John's next move very interesting. So just a few days ago, he announced his departure at Google, and he's going to Anthropic. It's important to note that John Jumper was not building generic prediction architectures like Claude or like Gemini.

It was extremely, structured, designed, and engineered for the purpose of doing a specific thing. Why might that be interesting to Anthropic? We can only guess.

So this isn't just about Google DeepMind and, you know, the CASP competition and and whatnot. There are now structural biologists around the world that can innovate and build new products and save lives potentially because they have access to this protein database. I spoke with Emmanuel Ny.

He is a structural biologist based in Africa.

Speaker 8

went back and did just one protein purification, collected the data and we are for fold.

Speaker 6

two, three months. So, yeah, we're talking about potentially years of work compressed into several months. That's pretty good.

Agents are getting smarter every day but even the best agents get stuck without good context. And this is where Notion comes in. With the recent launch of Custom Agents, Notion becomes the collaborative AI platform where agents and humans work side by side.

And now their new developer platform is turning that into infrastructure that developers can work on. The way I think about it is it is the kind of curated materialized memory plane for all of my work. And the great thing about Notion is it's so easy to work with programmatically.

It has a CLI. It has an MCP server. It has agents built into it.

Right? So that means using my phone, I can ask an agent to go and do some research or to put some information in there. I can sync things to my calendar.

Without Notion, I honestly don't think I'd be able to do anything that I do on MLST. So it's really, really good. So I highly recommend you give Notion's platform a go.

You can sign up at notion.com/mlst, and you'll be supporting the show if you do. And now let's get back to John Jumper.

Now we filmed this before the Anthropic announcement, so it was really cool to speak with John. He's such an inspirational guy. I was lucky enough to have dinner with him the night before, so we had a bit of a warm up conversation and, you know, we drilled in to the various topics we wanted to discuss.

One thing that struck me is John is unusually careful about what AlphaFold does and does not solve. So he doesn't sell it as a model of life or, you know, as a model of curing disease. He sells it as something a little bit more narrow and possibly more radical, actually.

A machine that predicts one class of structural biology measurement well enough to change what scientists can do next.

Speaker 4

Jon Jumper. I mean, AlphaFold itself is this kind of, I guess, I now say landmark in AI and science, but it's really about how do we use AI to solve problems that humans can't, that are really hard, that we go and we do years long experiments. And in the case of AlphaFold, it's this problem of protein structure prediction.

This, I guess it's machine learning street talk, not biology street talk. So, you know, DNA is the instruction manual for life, but what does it actually tell you what to build? And it tells you one of the many things it tells you how to build are proteins.

And these are little nanomachines, couple thousand atoms in the cell that actually do the work of the cell. And so three letters of your DNA tell you how to add one extra piece to this protein. This protein is kind of a long string, and there's a tiny machine itself made of proteins and RNA that's built one kind of string at a time.

And you you make this rope of 20 types of chemical groups. So it's kind of 20 types of letters, and people, of course, use the alphabet for these things. And each of them are quite different.

Right? My my PhD supervisor could tell you lovingly about what's special of each one. But you build this kind of rope of the protein, and then it assembles itself.

It twists. It curls. It folds up into a really kind of compact and interesting shape.

It has these helices, sheets, all these things, and that's actually what works. And the the analogy I always kind of like to say is it's like you have an IKEA bookshelf, and you open the box, and it builds itself. And so this really, really central problem for maybe seventy plus years in in biology is, okay.

How do I I can read DNA. In fact, can read DNA really well now. You know, you can probably many people in your in your listeners have had their DNA sequenced.

But understanding the structure of even one protein is extraordinarily difficult. Right? That's a worthy PhD project.

I would say maybe a typical kind of time frame is a year. If you wanna put money on it, maybe a $100,000 to get one answer. And this is really important to biology because we wanna understand how these proteins work.

When they misfold. Sometimes it's disease. Even when they work, they are the parts of the cell.

They do all the parts the things of the cell. You know? They're beautiful proteins.

The reason that, you know, cells can move, right, or is this giant protein machine whirling around driving the force to move cells? All of the functions of the cell basically are in these proteins. Humans have about 20,000 different types in different locations in your genome.

And so what scientists have done is they've gone to these enormous, enormous machines, synchrotrons normally, the size of small towns, in order to produce extraordinarily bright X rays. And even that, only after they've done really, really hard experiments trying to figure how to what's called crystallize a protein, and this takes years and years. And once they do that, and then they solve another mathematical problem that maybe we'll talk about, maybe we won't, they get one picture of a protein.

And they get kind of this progress and often this whole wealth of understanding of, oh, okay. I can understand how this DNA change that was found in the population might affect Parkinson's because look, it's right here on this protein, and that makes so much more sense. And so people have studied this problem for a long time.

There have been almost innumerable Nobel Prizes given for individual proteins, right, the ribosome, many others. People have an incredible societal investment, collected around 200,000 of these structures, about a 140,000 at the time, we did AlphaFold. Each one still extraordinarily difficult.

Right? Each one still I remember seeing people talk about their PhD and give their one of their talks near the end of their PhD progress toward whatever protein. Right?

So I did my I'm gonna be doctor, and I probably am not gonna crystallize this protein. I guess I'm I'm telling you all about proteins and nothing about what we did, but what we did was develop a new deep learning system from the publicly available experimental data, so all very public data, that was vastly more accurate at predicting protein structures. So predicts it to something like within the radius of an atom, right, in in typical accuracy and an accuracy that starts to rival at least some experimental methods, but more importantly than that, is, you know, extraordinarily fast.

So it takes five, ten minutes to get the structure of a protein instead of a year. I should, at some point, figure out what that ratio is in terms of time. But then also, of course, it's incredibly scalable.

So we've predicted the structure of 200,000,000 proteins, basically every protein from an organism whose genome has been sequenced.

Speaker 6

Right? We've made this widely available, and scientists are using it like crazy. It's absolutely amazing.

You have released a database of all of these proteins and the map lit up. So now scientists from all around the world, they can access these protein structures for many downstream tasks. But to bring this to life, you know, we have proteins doing things in the body, and we can use these structures, and we could do things like drug discovery and and and whatnot.

But what what's the gap? So so what can people do now that they have these structures? I think the right way to think about this is it's a starting point for biological research.

Speaker 4

what people do, what are some beautiful studies that people have done, we see it all the way. One that just came out was scientists trying to understand how cholesterol is moved about in the body, right? What actually is the thing that takes cholesterol and moves it from one place to another?

How might mutations in that affect high cholesterol, heart disease, etcetera, there's this beautiful weird protein that kind of wraps around it in a shape that we really didn't know until a few months ago when this paper came out. And what they were able to do actually is it's one of the ways in which scientists, I think, really commonly use AlphaFold is they use both experimental techniques and AlphaFold. So they used an experimental technique, cryo electron microscopy, to take an incredibly blobby picture.

You know, they used to call cryo EM blobology. It's gotten much better, but it's incredibly kind of rough picture. And they don't really know the atomic details.

And then they also run AlphaFold, and they see, well, actually, AlphaFold has this shape that almost exactly fits within this kind of blob. And so you get both confirmation and more detail, and suddenly you have this beautiful atomic model where you can start to then go and say, now how where are the changes in this protein? What might it affect?

How might it affect how it takes cholesterol from one place to another? Then you have to go figure out, now how do I make drugs for this? Do I do I bind to this protein?

I think when you see it in drug development, there's actually kind of two or three ways in which it's used. I mean, the first I will say is that the hardest part of drug development is that we do not know how biology works very well. Right?

That it's not you know, the thing preventing us from curing, say, I guess, autism, right, is not that we know exactly one protein. And if we just had its structure, then autism would be cured. It's a huge disease that involves the whole body.

And so we kind of are trying to unwrap and unravel biology well enough to figure out which proteins. How do these proteins interact? How does that ultimately contribute out to these phenotypes?

And so people do biology across all these length scales. And the contribution of AlphaFold is to say, this protein, for example, that you didn't even know was important. Like, there was oh, one study from a few years ago was on a pro was you know, there's all sorts of recycling mechanisms in the body that take proteins it doesn't need anymore and gets rid of them or doesn't want anymore.

And there were hundreds of genes, in fact, that people found were turned off at a certain phase in cell development, and they didn't exactly know what protein was involved. They did some genetics, and they found this protein that had essentially never been studied before, a human protein called mydalin. If you knocked it down, then these proteins didn't get recycled.

And that's kind of more or less what they knew, and they knew it didn't work in the standard way. And they ran AlphaFold, and they looked at it, and they saw some pieces that were suggestive. And then they ran AlphaFold together with, you know, I think it was almost all 500 proteins that were not that were kind of responsive to knocking this protein down.

Right? So change. So this is kind of how biologists develop evidence.

And they found in about 40% of these, when they ran AlphaFold this very, very specific pattern where one part of that protein was trapped between two parts of myndolin kind of grabbing it like clamps. And they could find and then they would go and they would do experiments. Right?

Because and then they would say, well, what happens if I take this bit of protein and I remove the place where AlphaFold says it's clamped by middolent? And suddenly that protein doesn't drop down in the cell. Right?

So and they found this on maybe the 10 examples. Nine of them worked exactly this way. One of them only partially was reduced in how much it's knocked down, but then they looked at the AlphaFold predictions and found that AlphaFold actually put it two places.

And so if they take out that second place as well, then the degradation is completely abolished. So now they have this mechanistic understanding of this new protein they had never thought about before, and now they know exactly how it recognizes what's developed in this really important stage of cell division. Now the question becomes, okay, now how do you take that knowledge and do drug development?

That's where so AlphaFold two was what came out now five years ago. What we've done more recently about a year ago is AlphaFold three, which says, well, let's let's not just do proteins. Let's do the protein cinematic universe.

And so, you know, I said proteins bind cholesterol. Right? So this is a nonprotein kind of fatty molecule.

More well, not more importantly, but very importantly, they also bind drugs. Drugs are small molecules, maybe, you know, twenty, fifty atoms that stick to proteins and change how they behave. And you couldn't even ask this question to AlphaFold two.

You couldn't say, how does this drug stick? It's like, well, you better give me a protein. Only if your drug is a protein, which some are.

But AlphaFold three, we expanded it to kind of do the whole universe of things that appear in the PDB, the whole universe of cells. And now we can say, well, this is exactly where that drug sticks. And then people around the world are using these ideas, building others, For example, isomorphic labs inside alphabet, developed inspired from the AlphaFold breakthrough are trying to say, okay.

Let's really use this to start to do drug design. Let's start to take these technologies that finally work, that are finally predictive about this, and now let's see if I can design a small molecule with it or I can design a drug that binds, that changes how this machine works, and then in a way that hopefully makes someone healthy. And I think the the best kind of analogy for how you should think about drug development really now or maybe the way to think about how hard it is is this old joke.

You know? Do you know this joke that there's this giant factory, and one of the most important machines in this factory has stopped working? And they call in a technician who comes.

He looks at it. He goes to some screw or some some nut and turns it a quarter turn. The factory roars back to life.

And they said, that's wonderful. Thank you so much. Can we have a bill?

And he says, $10,000. Yeah. And they say, what Knowing what to turn.

Yeah. Yeah. So there's, you know, knowing what to turn or turning this 50¢, knowing what to turn, all the rest.

And I think this is the right analogy that we are both learning how to turn this, right, how to build drugs, and also learning in this big complex factory of the cell what do we need to do to cure disease. Yes.

Speaker 6

But it's so incredibly complex, isn't it? The human body is alive and there is a symphony of complex adaptive compensatory, mechanisms. I guess the the idea here is that we're proposing a mechanistic understanding of how this works, which means we can design interventions that are very effective.

But in machine learning, we've kind of learned the opposite lesson, which is that all of our intuitions about how things work don't don't really work, and we need lots of data, and we need to test lots of things. Could it be a similar thing here that, you know, it's like whack a mole. You you kind of you you you do one thing and then something else compensates.

Speaker 4

is almost the humility of AlphaFold in that, you know, people say, you know, we are trying to predict what this experiment will give you. We are not trying to tell you everything. We are not a model of the entire cell.

We are a predictor of this experiment that you did all the time and took you a year. And so in a certain sense, I think and so we have validity in that. I can characterize very well how well we're we we will reproduce that experiment.

And then people figure out how to take this machine and use it in other ways that we didn't expect to find out new you know, discover new mechanisms, to try thousands of alpha fold predictions, to find two proteins that stick together, and find this unknown component of this complex system. So people are finding ways to push this further. But in a certain sense, we are narrow, or we predict the result of a scientific paper.

We predict the result of a scientific paper that often appears in Nature and Science and Cell and these big journals. Right? We predict nature level science with the press of a button in a very narrow category of nature level science of the structure of a specific protein.

But there's this enormous wide universe of biology that ultimately we're gonna have to figure out and understand what data will we pin ourselves to, what experiments will we predict, and predict really, really well, such that you know, I mean, maybe the other story of machine learning is that predicting things okay is alright. Predicting things extraordinarily well starts to produce amazing machines. We see this, of course, in language models, in image generation, but also in protein.

Speaker 6

kind of a narrow predictor, but we are doing something truly useful. Can we talk through the predictive architectures of of the different versions of AlphaFold? So, you know, the first version was was a CNN.

The last version is a diffusion model. The second version, we spoke about this last night, it had a structure component. And, you know, obviously, geometric deep learning is spoken about a lot.

And I think people misattributed the the benefit of having these, you know, and kind of symmetries. It did these s e three symmetries. And just just talk me through that process because you were kind of saying at the beginning, you were really trying to imbue your human understanding of this and then kind of experience told you differently.

I think there's two or three things.

Speaker 4

but kind of thematically to AlphaFold three as a diffusion model. Boxes. We love to have the highest level bit be the answer for why these things work.

Yes. Right? Oh, they switched from CNN to I think the answer is really okay.

AlphaFold one really was as a network. It predicted a subpart of the problem. It started from kind of biological data, evolutionary correlations.

It ended in kind of ish data, distance between atoms. In between was a CNN. Right?

It was actually an off the shelf CNN from a computer vision that someone else had done. Okay. That was a CNN.

But after AlphaFold one and then kind of all the protein specific bits were kind of wrapped around the machine learning. And so the I would say AlphaFold two was, let's build the science instead of building the science of image recognition and then applying it to proteins because, you know, human visual the human visual system is exactly what we needed to fold proteins is not something true. Right?

Humans were bad are bad at at predicting protein structures. How are we going to actually build all the pieces? Now there was an s e three piece.

In fact, AlphaFold two was built iteratively. There were many stages. Actually, the s e three piece was the first part of AlphaFold built.

But AlphaFold two at the end was really this giant trunk of an architecture we called Evoformer, which is axial attention plus a bunch of other stuff. And that is 90 plus percent of the compute and the accuracy. And then but it produces this kind of end by end.

So you start off with two pieces of data. You start off with the protein sequence, and then you find the sequence of every protein evolutionarily related. And protein structure changes slowly.

The the structures of my proteins are in most cases similar to the structure of proteins in yeast, sometimes even out in E. Coli. So you grab many related structures.

You often have hundreds or thousands. You provide this information, and we have this this specialty architecture called Evoformer, which had two forms of axial attention that were kind of having a conversation between what we believed about geometry and what we believed about evolution. We had these two representations.

And we end up we take the geometric bit, the n by n, which we have actually as an intermediate loss said, what are the distances between these atoms and made categorical predictions? And then we hand it to what we call the structure module, and it's best thought of as a geometrization engine. Right?

If you you have n squared predictions about n positions, you're somebody's gonna have to harmonize this thing. And this used a s e three I guess it was invariant in the sense we collapse it on every layer. S e three invariant attention.

This was actually one. This was one of mine that was kind of even starting at DeepMind. I'm like, oh, we should probably put points in, or I was thinking about pro even then protein residues.

Right? So you have this backbone, which has three atoms. You can align a frame to it.

It's extraordinarily rigid. And I know the business in the places where all these where all these residues differ is kind of just off that. So if you align them to reference frames then and you operate in points in those reference frames, then it's natural, and you can just take attention, and you can let it project points in local frame.

You can transform it. Then you can use the distance of those points as a way in order to bias your attention, and this is invariant point attention is what we called it. At the end, it's kind of fun.

More important than that probably was this defining of frames, actually, almost certainly. Defining of frames let us write down a really interesting loss function. So we call it frame align, point error, or FAPE.

And this was kind of saying, in the reference frame of the I th residue, where is everyone else? And it's kind of locally registered, and then you have n squared errors, and then you average them together. That, I think, was really, really important.

I think that was one of the breakthroughs early on was this loss. But, of course, the really fun part is an SE three invariance. And but remember, we didn't start with any geometric data.

We started only with nongeometric data. So our geometry emerged in the middle. We started with what we would call black hole initialization, where we just stick all the residues on top of each other in the world's least physical structure, it was important also that we disrespected known symmetries of a protein.

For example, the known symmetries of a protein are these residues are separated the atoms this atom and that atom in a residue are separated by 1.3 angstroms plus or minus point o one five. And in fact, even in AlphaFold one, when we would do the optimization, we would actually use turning kind of like a jointed robot arm to optimize these, say, typically 300 residues.

And so you would do all this twisting. You would actually have a very ugly geometry. The geometry of a 300 joint or actually, no.

Sorry. Wouldn't be 300. It would be a 900 joint robot arm is really bad, and so that means that your optimizer has to take many steps.

So one of the important things is let's just break it up. Let's just treat them as a residue gas, we called it, so that this can proceed in, say, four steps, eight steps instead of the number of steps of this twisty geometry. And then we used equivariance, and it helped.

But one of the things that was really surprising, think, is maybe because of the early talk or maybe people were working on equivariance. Geometric deep learning has been very popular, and people said, ah, they mentioned my keyword. That must be the reason it worked.

And I remember being a little bit confused, and I thought, okay. But we'll we'll very carefully ablate this. We did quite a few ablations for AlphaFold two, and we'll publish the paper, and everyone will realize.

And we published a paper, and I remember the fifth row was called no IPA. So AlphaFold two was about 30 points on the GDT scale better than AlphaFold one. Right?

So so that's that's the kind of 30 points is is your thing. And we did these ablations. They were all small, almost all small.

Removing the invariance, the equivariance cost about two points. Right? And you could measure it, maybe two and a half.

Right? So it it contributed, but it contributed two and a half out of 30. And I thought that would put it to bed, and it put it didn't even put it to bed at all.

People still talked about AlphaFold two as the great victory of equivariance. They they never talk about FAPE, and they talk about, you know, an equivariant transformer that they think came from others. And, actually, we did this equivariant transformer in, like, 2018.

I actually remember it was October 2018. So it was like, we did one. We tried a little bit to improve it.

It didn't make AlphaFold better to try and improve that part, so we went on to the next thing. Right? We're kind of ruthlessly empirical about it, but it's a very cool thing.

And so I think people really hooked on to it. And I think what it really happens is I think we we were talking about it kind of at dinner last night at this awful dinner, but the equivariance is one like, global s e three symmetry is not a very powerful symmetry. It's not nearly as kind of big and powerful as a symmetry like, oh, all the residues are permutation invariant.

Right? So we do still have permutation invariance as probably the big symmetry of alpha fold. Right?

We have a transformer that is position is relative position coded only. We clip the relative position coatings. But I think this particular symmetry, it's not like physics where you write down the symmetry group and then you derive the laws of physics from your big symmetry group and you get the standard model.

This is this is asymmetry of a messy real world problem that probably doesn't pin it down so much. So I think it's good, but we shouldn't obsess about one good thing, or you can't you don't wanna valorize things. My favorite review of of AlphaFold two, we got the reviews back when we submit the paper.

And one of them said, this is six or seven papers worth of ideas. Right? And I think I think that was that was right.

There are many, many ideas that added up to be a transformative system and many you know, to use a baseball analogy, it's not one or two home runs. It's, you know, 18 doubles. Right?

That it it's really you know, these midsize wins stacked together and together make a transformative system. Now we would sometimes find in our ablations, we ran a double ablation. I think it was no recycling and no IPA.

We turned off two things, and performance cratered. Right? And I think it was kind of there are many problems we needed to solve.

We solved most of them two ways because it was better than solving one. And if you knock out both things, then your building maybe collapses. Or this was maybe a 12 or 15, which was our biggest ablation, which was still only half the gap to AlphaFold one.

Right? We I remember doing the ablations and saying, guys, we've never crossed alpha fold one performance. But a lot of those ablations actually went into alpha fold three.

So we said, okay. Well, equivariance isn't super important. We had another ablation that if we take out, you know, giving the raw genetic information and give the pairwise correlations, that's one or two worse.

So maybe this fact that we're processing these all the time is not so good. We made this kind of interpretability kind of projection of each layer into a structure and made movies and could see that most of AlphaFold's capacity was spent optimizing the structure geometrically. It's much more a geometry engine than an evolution engine outside the first few layers.

So we said, okay. Why don't we just cut back this Evoformer to just operate a few layers, and then we'll do a much simpler version called a Pairformer, and that improved performance. And we basically use these ablations to say this is what the machine learning is telling us.

Whatever we may, you know, feel like being a machine learner is all about, you know, you think about the data, you look at it, you come up with hypotheses, maybe equivariance is important. You try it, you measure nine times out of 10. You find out you're wrong.

Right? If you're wrong nine times out of 10, you're a very successful machine learner. You're incredibly productive.

And and you build this local intuition. You build this notion of what the problem needs and how it works, and you build, in my view, a kind of science local to your area, kind of a local manifold of ideas in protein structure prediction near these architectures. We would develop intuitions like thou shalt not put a one d representation rather than a two d representation anywhere near the Evoformer or your performance will go down.

Maybe a year, six months in alpha fold two, at some point, we had a mix of axial attention and convolutions in our pairwise processing. It was somewhat different architecture. And I remember someone did an experiment where they just deleted the convolutional layers, not like added attention layers, deleted the convolutions, added no parameters, just strictly fewer parameters, and the model got more accurate.

The validation loss improved. And that doesn't normally happen in machine learning. They don't say remove parameters and your generalization will improve.

But convolutions were probably actively harmful to learning what we wanted to learn, and I have kind of a hypothesis of what that is. But all of these kind of lessons and explorations were about how does deep learning generalize in proteins. And you have to build that knowledge and that expertise.

And what's really, really special in AlphaFold in AlphaFold two, and then I'll I guess I should talk about AlphaFold three in a minute, is that we have kind of the mix of biological hypothesis, physical geometric hypothesis, and experience and the kind of, you know, tactile feel of years of banging our head against this particular problem and dataset. Actually, AlphaFold one and AlphaFold two were the exact same data. We decided to just have more eval data and not bump our training data at all.

And and the effect of that was really, really large. There was a a wonderful study by the Al Qureshi lab, which was retraining AlphaFold twos and train them on 1% of the PDB. So instead, you know, so instead of a 100 a 150,000 structures, it was something like 15,000.

And they found AlphaFold two on 1% of the PDB was more accurate than AlphaFold one.

Speaker 6

and training ideas that we put into AlphaFold two were worth a clean 100 x in data. We were speaking last night about all of the tacit knowledge that you guys acquired during, but maybe we'll we'll park that just for a second. But the the enterprise of machine learning is about building models of understanding for things that we don't understand.

So, you know, we're we're building these these alien artifacts. And you gave a wonderful example of, you know, imagine I'm generating some text, And we we might think naively that just like the way it's rendered around the edges is actually the thing doing the heavy lifting, but actually the the the load bearing thing might be something completely different. So many many of our intuitions don't don't really work.

But for me, and I know you're allergic to the word understanding, but for me, understanding is is possession of a generative model, which can do the thing. So if you could create a physics simulator of some phenomenon, assuming that it's it's not lossy, I would say that you understood that thing. And we think of machine learning as modeling extant examples, which means rather than how it's constructed, how it's built, it's modeling the thing at the end.

But you were describing, you you know, like, the Game of Life is a great example of this. It's kind of like it's modeling it's it's it's learning the path, not the destination. But you were describing something fascinating last night, which is that we have, for example, this recycling mechanism in AlphaFold, and you can place things through the structure many, many times.

And what seems to happen is at the beginning, it solves the most complex problem, and then it's kind of refining. It's kind of refining. And this is a little bit like the game of life.

It's not it's not like it's simulating the creation. It's almost like it's at any point, learning how to refine and optimize the structure. Okay.

So we I think we should distinguish three things. Predict, control, understand Yes. First.

Speaker 4

you say, I'm gonna do a thing. What am I gonna what will be this value of my machine? What will appear on my computer screen in the future?

That is predict. Control is I want to measure this thing in the future, and I want it to come out 17. Right?

That's control. Understand is a lot like predict, except there's a human in the loop. Understand means that I have such a small collection of facts that you will predict, and you will do it with facts that I can communicate to another human in kind of this compact fix fits on an index card.

That's almost understand. And so I think these machines let us predict. They let us control.

We have to derive our own understanding at this moment. Right? We can experiment now on the artifact.

We can look at the 200,000,000 predicted structures, not just the 200,000 experimental structures in order to help us understand, but it doesn't do the act of understanding for us. It does the act of predict and maybe control. Now though, then there's maybe one other thing.

There is the algorithm, and it's really important, I think, concept to machine learning. There's the algorithm you program and the algorithm you get or, you know, machine learning as code meets data produces weights. And so one of the all one of the kind of lasting debates in machine learning, how much work is done by the code, how much work is done by the data that ends up in the weights.

And so what I think you what we see in AlphaFold in a certain sense is a very beautifully intuitive algorithm, an algorithm we can, in some sense, understand, right, that it does successive geometric refinement. I communicated that to you in a few words. You probably almost saw it in your head even though I don't think you've seen the these videos.

I mean, they're in the supplement of our nature paper. But but that is an algorithm that humans already came up with. Maybe we should almost do, you know, gradient descent and some empirical model that makes each thing more correct.

Maybe that's how AlphaFold should work, but that wasn't what we programmed. But we also still thought about it in a certain sense, or we were thinking about things like recycling in terms of, wow. Isn't it weird that AlphaFold in in layers has to give an answer for no matter how hard this problem is?

And maybe we should give it some more layers. And my GPU is out of memory, so maybe I should just, you know, run it back through so I don't have to have more memory. But even without that, I think AlphaFold, even without recycling, was learning this kind of iteration.

And then we put in a kind of code idea, architectural idea to help this process that it was gonna learn from the data. Going back to the earlier thing about exactly how far residues are apart, we didn't tell AlphaFold that. We knew that the data would scream at it, that I and I plus one were 1.

3 angstroms apart. So I think when we think about our human understanding, I think one of the you know, I don't really love the bitter lesson as people try and apply it. In fact, AlphaFold two is the opposite of that.

We did a whole bunch of specialty stuff because our data is not finite. And in fact, now that we've gone to language models, we found our data is still finite. The Internet is finite.

So, yeah, I think, you know, don't do architectural research is the wrong thing to draw from it. But have some humility about which things go into your code and which things will be derived from your data. Look at what's missing.

Understand the algorithm that deep learning that the deep learning is trying to learn, how can you accelerate it, how can you add hypotheses, and where you especially, you add kind of communication. The most important thing we would do within the architecture is modify which units communicated and how. I think all of these have been kind of how we drive understanding to ultimately make an iterative process.

And it should shock no one that if you're trying to make an intricate geometric object that you are going to iterate. Or similarly, if you think about generating text. Right?

And one kind of naive assumption that people will make is that these are next word generators, so they have no idea what's gonna happen in two words ahead or three words ahead. But, of course, you can't think of the you can't write down the next word without I don't I don't start a sentence not knowing how it's gonna end most of the time. Right?

At some points, I change, but I think ahead a little bit in order to accomplish my task. And so the understanding that we see built into these models are kind of the structures that we want sometimes emerge and sometimes don't. And I think we valorize the high level ideas that impose, for example, an AlphaFold three coming back to AlphaFold three.

Right? You said it is a diffusion model. But I would argue it's a different diffusion model than an image model.

Maybe well, there's some different for one thing, there's a huge trunk that is not in any way a diffusion model that's only run once. That trunk is probably where the structure is actually determined, and the diffusion is just like the structure module was a geometrization engine that took a set of really quite good constraints that had very clear notion of the structure within those constraints and solved up the details. I think AlphaFold three diffusion is similar, and it's especially similar because, in fact, in images, okay, you start generating an image and you see especially these early trained diffusion models generate kind of colored blobs, and they start to decide what those colored blobs mean.

And they pretty clearly kind of decide what those colored blobs will mean later because you could stop them in the middle of the process and run them again and get a somewhat different interpretation of those colored blobs. In AlphaFold three, you actually have an interesting thing that if you look at AlphaFold two, we can kind of, through this process of projecting out intermediate layers, see what it solves first. And it basically solves local details, local pieces.

It starts to put local pieces together. It's agglomerative in how it solves a structure as is kind of natural. The easiest thing to predict is your local structure.

The hardest thing to predict is your largest scale structure. That's how AlphaFold two works. If you look at AlphaFold three and you take coordinates, which you've added a very large amount of noise to, well, the very first thing you have to solve is how say, you have two proteins, how do they associate it?

Where are their two blobs relative to each other? What are their Gaussians? So the the problem that alpha fold two is solving last is the problem that alpha fold three's diffusion has to realize first.

And how does it do it? The answer is not that it comes up with an orientation and builds the protein around it because, of course, it's going for one correct answer or at least a very narrow distribution. The answer is really the big network before it plus the first pass through the diffusion network is solving the overall structure, and then the diffusion is realizing in any details it couldn't solve before.

It's basically sampling among. So it is diffusion technically, but it's much closer to AlphaFold two. I think there's no reason that it was kind of very specific technical reasons around kind of laziness and geometry that made diffusion a really good choice for AlphaFold three.

It made it easier to handle ligands. It handled some bond distances and local things. But it's not like diffusion in the same way as, oh, it's drawing the blobs and deciding what they mean at the end.

So I think all of these are people like to think of these like to say, this works because it's a transformer. And this works because it's a transformer doesn't explain why chat models have gotten vastly better in the last three, four years. It doesn't explain all the research.

It doesn't explain what researchers do every day. All of these details are far more important than this high level bit of, is it a transformer, is it a diffusion model, that we wanna talk about? And then also even these diffusion mechanisms don't work in the way of kind of progressive refinement that makes sense for images.

Right? Maybe you'll make colored blobs, and you'll decide what those colored blobs mean.

Speaker 6

not entirely the story, but it's definitely not the story for proteins because that's the hardest problem is the large scale structure. I mean, in a sense, this is leaning towards this idea of constructive complexity that we were talking about before. And I'd love to get your your general take on on what this means for artificial general intelligence.

Because with language models, for example, we train them basically with behavior cloning. So, you know, we have this this rich adaptive generative process, and we we generate all of this language, and we train language models on them. And for me, intelligence is the adaptive acquisition of coarse grained representations.

Culture and language is changing all of the time. So we invent the word unalive to get around the, the filters on social media platforms, and that's an that's an example of ling linguistic agency. Language models, we we noticed that when we do this iterative adaptive refining with active, you know, active fine tuning and adaptation, they become a bit intelligent.

They they learn new representations and and they adapt. And in a way, what they're doing is even though they're ungrounded from the path, they can they can take a code solution like AlphaEvolve and they can refine it and they can refine it. And it seems to work really, really well.

But are we in this regime, do you think, that we're not necessarily building artifacts that have the same type of generality? I mean, what what do you think about intelligence in general?

Speaker 4

and far less important than people believe five years ago in the explicit way. So just like we were talking about the things that AlphaFold does and the things that AlphaFold is forced to do by its code or or, you know, obviously, everything that's forced to do by its code, it does, but many things it does, it does without being forced. Because it had to learn it to make a good predictive model of the data.

It had to find good intermediate representations. So in a certain sense, I think the most seductive idea in in machine learning is always there's this thing I know will have to be in the end there in the end, so I'm gonna have to have a u I'm gonna have to have a place in my code that is named that and then forces the mechanism to high level concept builder thingamajigger. Right?

And that was a very popular kind of I'll make the concepts units. I shall force disentangled representations via this law sometimes on the intermediate layer. This is where it will store those.

And that's reasonable to go test, but what we've seen a lot of is that a lot of the things that you would imagine needed or needed for intelligence are developed by desperately trying to predict the next token really, really well. And they're not they're not developed because you predict next tokens at all. They develop because you do a really, really good job at it.

And so these kind of generalized spaces, representations, understanding of concepts is forced very slowly with data. Right? Pretty much all the kind of you know, there's a lot of log linears or my you know, everyone's least favorite functional is right.

The things go up as the things go up linearly with the exponent of effort that we see all the time in our scaling laws. But we do get these concepts and representations, and what we don't really, I think, have an answer for is how do we get them cheaper? Now sometimes we can get them via programming.

Right? We get memory like things. Now we have language models writing notes for itself and then retrieving those notes, or we find out it's better to keep reminding agents what they're doing so they don't forget over long trajectories.

So we we build weights. We build artifacts. We find deficiencies.

We can often paper over those deficiencies in some kind of software harnesses, but then we don't yet know we don't that doesn't immediately drive back into exactly your machine learning, or you don't, like, put a harness with external memory and then distill it back into the network and have amazing memory things that no longer need this harness. That we haven't figured out how to do.

Speaker 6

John, we we have we have to wrap it. But doctor John Jumper, it's been an an honor to have you on the show. Thank you so much for joining us today.

Been tremendous fun. Thank you. So as I said earlier, Emmanuel Nie, he's based in Africa, and he's actually training scientists, not just he's not just giving them access to AlphaFold, but he's training them how to use it, how to interpret the results, and how to help scientists build experiments using the database.

Speaker 8

Yes. So my research focus on on drug discovery for malaria and enteric bacteria. And then I'm also involved in capacity building for Africa based researchers, using tools like AlphaFold.

Initially, African scientists didn't have access to expensive structural biology tools. With AlphaFood, these researchers can now do complex experiments that were not possible before and tackle diseases such as malaria, HIV, and other, antibiotic resistant infection. In my own research on on drug discovery, so I use AlphaFold in terms of solving structures of cryo EM data, and I also utilize that to to to map out the mechanisms of the proteins.

Speaker 6

So for him, AlphaFold was so impactful. Like, if you think about the before and after, we're living in a different world now.

Speaker 8

to face a protein was, like I said, was really, really difficult. And so I tried several years, close to four, five years, and it wasn't successful. And with AlphaFold, imagine this is more than ten years ago, with AlphaFold, I went back and did just one protein purification, collected the data, and we alpha fold.

Speaker 6

two, three months. And now it's his goal to train as many scientists as he can how to use this technology for the betterment of humankind.

Speaker 8

This year, with funding from Google DeepMind and Swedish Research Council, we have scaled up to 100, and there's no drop in the quality of the training. In fact, it was there was an improvement. So based on this based on this, we want to train 100 scientists every year for the next ten years.

So we're targeting close to 1,000 African scientists in the next decade to be able to utilize this tool effectively, and then we want to form an emerging community of structural biology practitioners working on prevalent diseases in Africa.

Speaker 6

So that was the AlphaFold show. Thank you very much to John and Emmanuel. Yeah.

The the conversation with John was very interesting. He's he's so inspiring because I think he is testament to the fact that even though we talk about all of these general purpose foundation models, to really advance the frontier and to build cutting edge applications in science, we need to do a lot of engineering. We need, you know, domain knowledge.

We need serious expertise. And a lot of our models will actually look quite hybrid. They'll look quite customized.

And I think AlphaFold is a kind of proof of existence for the types of hybrid models that we can deploy to further the field of science. John, I wish you the very best of luck in your new position, Adanthropic, and thanks for watching the show.

Speaker 3

This episode is brought to you by Google Chrome. You think you know a browser, but Gemini in Chrome? That's new.

It can help you with practically anything on the web, like restoring a motorcycle from a 50 page restoration block or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it. Ready to make anything online make sense?

There's no place like Chrome. Check responses, set up required compatibility, and availability varies 18 plus.

Shared via Hopper