Max Welling discusses his career evolution from theoretical physics to impactful AI applications, particularly in material science with CuspAI. He explains how physics, especially symmetry and stochastic thermodynamics, provides a unifying thread for his work on equivariant neural networks and diffusion models. The episode highlights the emerging field of AI for science, CuspAI's platform for accelerated materials discovery, and the vision of empowering chemists with advanced computational tools to address global challenges like climate change.
I want to think of it as what I would call a as sort of a physics processing unit, like a PPU. Right? Which is you have digital processing units and then you have physics processing units.
So it's basically nature doing computations for you. It's the fastest computer known, possible even. It's a bit hard to program because you have to do all these experiments.
It's also quite quite bulky. It's like a very large sort of thing you have to do. But in a way, is a computation, and that's the way I wanna see it.
So I wanna you can do computations in a data center, and then you can ask nature to do some computations. Right? Your interface with nature is a bit more complicated, but then these things will have to seamlessly work together to get to a, you know, a new material that you're interested in.
Yeah. It's a pleasure to have Max Voehling as a guest today. Max has done so much over his career that I've been so excited about.
If you're in the deep learning community, you probably know Max for, his work on variational auto coders, which has literally stood the test of time or officially stood the test of time. If you are a scientist, you probably know him for his, like, pioneering work on graph neural networks on equivariance. And if you're a material science, you probably know him about his new, startup, Cusp AI.
Max has a long history doing lots of cool problems. You started in quantum gravity, which is, I think, very different than all of these other things you worked on. As a first question for AI engineers and for scientists, what is the thread in how you think about problems?
What is the thread in the type of things which excite you?
decide what is the next big thing you wanna work on? So it has actually evolved a lot. In my young days, let's breathe, I would just follow what I would find like super interesting.
I have kind of this sensor I think many people have, but maybe not really sort of use very much, which is like you get this feeling about getting very excited about some problem, right? Like it could be, you know, what's inside of a black hole, or what's, you know, at the boundary of the universe, or you know, what is quantum mechanics actually all about. And so I followed that basically throughout my career, but I have to say that as you get older this changes a little bit, in a sense that there's a new dimension coming to it, and this is impact.
Working in two dimensional quantum gravity you pretty much guarantee there's gonna be no impact in what you do relative, you know, maybe a few papers, but not under in this world, at this energy scale. As I get closer to retirement, which is fortunately still, you know, ten years away or so, I do wanna kind of make a positive impact in the world, and I got pretty worried about climate change, and I think we should, you know, and politics seems to have a hard time solving it, especially these days, and so I thought better work on it from the technology side, and that's why we started Cosplay AI. But there's also a lot of really interesting science problems in, you know, material science, and so it's kind of combining both the impact you can make with it, as well as the interesting science.
So it's sort of these two dimensions, like working on things which you feel is like, oh there's something very deep going on here, and on the other hand, trying to build tools that can actually make a real impact in the world.
when I look back, look at the different things that you worked out, some of them seem reconnected, like the physics to to equivariance and Yeah. And graph neural networks maybe. And that that seems to be somewhat related to CUSP.
Do you have a a thread through there?
Yeah. I think physics is the thread. So having done, you know, spent a lot of time in theoretical physics, I think there is first very fundamental and exciting questions, like things that haven't actually been figured out in quantum gravity, so that is really the frontier.
There's also a lot of mathematical tools that you can use, right? For instance in particle physics, but also in general relativity, so symmetry space play an enormously important role. And this goes all the way to gauge symmetries as well.
And so applying these kinds of symmetries to machine learning was actually, you know, I thought of it as a very deep and interesting mathematical problem. I did this with Tycho Cohen, and Tycho Cohen was the main driver behind this. Went all the way from just simple rotational symmetries all the way to gauge symmetries on spheres and stuff like that.
And Maurice Wyler, who's also here, when he was a PhD student with me, you know, he wrote an entire book which I can really recommend about the role of symmetries in AI and machine learning. So I find this a very deep and interesting problem. So more recently, so I've taken a sort of different path, which is the relationship between diffusion models and and a field called stochastic thermodynamics.
This is basically the thermodynamics, which is a theory of equilibrium, so but then formulated for out of equilibrium systems. And it turns out that the mathematics that we use for diffusion models, but even for reinforcement learning, for Schrodinger bridges, for MCMC sampling, has the same mathematics as this physical theory of non equilibrium systems. And that got me very excited, and actually when I taught a course in Maugenbeirich, it is South Africa, close to Cape Town at the African Institute for Mathematical Sciences, AIMS, and I turned that into a book.
So two years later the book is finished, I've sent it to the publisher, and this is about the deep relationship between free energy, diffusion models, basically generative AI and stochastic thermodynamics. So it's always some kind of, I don't know, I find physics very deep. I also think a lot about quantum mechanics, and it's a completely weird theory that actually nobody really understands.
And there's a very interesting story, may be good to tell, to connect sort of my PhD back to where I am now. So I did my PhD with the Nobel laureate Gerard de Toft. He's just the most brilliant man I've ever met.
He was never wrong about anything as long as I've seen him, and now he says quantum mechanics is wrong, and he has a new theory of quantum mechanics. Nobody understands what he's saying, even though what he's writing down is not mathematically very complex. But he's trying to address this understandability, let's say, of quantum mechanics head on.
I find it very courageous, and I'm completely fascinated by it. So I'm also trying to think about, okay, can I actually understand quantum mechanics in a more mundane way, sort of, you know, without all the weird multiverses and collapses and stuff like that?
to to build better algorithms. You are still very involved in understanding in understanding physics and the worlds Yeah. Even beyond just, like, applications to machine learning or introducing no formalisms.
That's really cool. Yes. Would say I'm I'm not I'm not contributing much to physics, but I'm contributing to the interface between physics and science, and that's called AI for science or science for AI is kind of a soup and it's actually a new discipline that's emerging.
Yeah. Yeah. And it's not just emerging.
It's it's exploding, would That's the better term. Because I know you go from investments into, like in the hundreds of millions now in the billions. So there's now actually a startup by Jeff Bezos that, you know, is at 6,200,000,000.
0 seat round. Right? It's, like, insane.
I guess this is the largest, you know, startup ever, I think. Right? And that's in this field, AI for science.
Right? It tells you something that we are creating a new bubble here. Yeah.
Right? So so why do you think it is? What has changed that has, like, motivated people to start working on AI for science type problems?
So there's two two reasons, actually. The one one is that people have been applying sort of the tool, the new tools from AI to the sciences, which is quite natural. And there's, of course, I think there's two big examples, know, protein folding is a big one, and the other one is machine learning force fields, or sometimes called machine learning interatomic potentials.
Both of these have been actually very successful. Both also had something to do with symmetries, which is also cool. And sort of people in the AI scientists saw an opportunity to apply the tools that they had developed beyond advertise placement, right, or multimedia applications into something that could actually make a very positive impact in society, like health, drug development, materials for the energy transition, carbon capture, these are all really cool, you know, impactful applications.
Beside that, the science and the kind of the is also very interesting. And sort of the I would say the fact that these sort of, these two fields are coming together, and that we're now at the point that we can actually model these things effectively, and move the needle on some of these sort of science sort of methodologies, is also a very unique moment I would say. And people recognize that okay, now we're at the cusp of something new, as we're also what the company is called after.
We're the cusp of something new, and of course that always creates a lot of energy. It's like okay, there's something, it's like sort of virgin field, It's like nobody's, greenfield, nobody's been there, you know.
I think that's also what's causing a lot of sort of enthusiasm in the fields. If you're an AI engineer, basically, if the people that listen to this podcast will be and you maybe don't have a strong science background, how does but are excited most, I would say, most AI practitioners, be engineers or scientists, would consider themselves scientists, and they have some background, little bit of physics, little bit of industry college, maybe even graduate school that have been working or are starting out. Does how does somebody who is not a scientist on a day to day basis, how do they get involved?
Well, they can read my book once it's out.
saying that there is more we should create curricula that are on this interface. So I'm not sure there is possibly already at some universities actual courses you can take, maybe online courses you can take. These workshops where we are now are actually very good as well, and we should probably have more tutorials before the workshop starts.
Actually we've kind of proposed this at some point. It's like maybe first half an hour of a tutorial so that people can get new into the field. But yeah, there's a lot out there.
Most of it is of course inaccessible, but I would say we will create much more books and other content that is more accessible, including this podcast I would say, right? So I think you know, it will come. And you know these days you can watch videos and things.
There's a huge amount of content you can go and see.
So maybe a a follow-up to that. How do people learn and get involved? But but why should they get involved?
I mean, that we have a lot of people who are of our audience will be interested in AI engineering, but they may be looking for bigger impacts in the world. Yeah. What opportunities does AI for science provide though to make an impact to, you know, change the world that working in this, the world of pure bits would not?
underlying almost everything is immaterial. So we're focusing a lot on LLMs now, which is kind of the software layer, but I would say if you think very hard, underlying everything is immaterial. So I was saying, you know, there's the LLM, underlying the LLM is the GPU on which it runs, and then in order to make that GPU, you have to put materials down on a wafer and sort of shine on it with sort of EUV light in order to etch kind of the structures in, but that's now an actual material problem, because more or less we've reached the limits of scaling things down, and now we are trying to improve further by new materials.
So that's the fundamental materials problem. We need to get through the energy transition fast if we don't wanna kind of mess up this world. Mhmm.
And so there is, for instance, batteries. That's a complete materials problem. Right?
There's fuel cells. There is solar panels so that they can now make solar panels with new perovskite layers on top of the silicon layers that can capture, you know, theoretically up to 50% of the light, where now we're at, I don't know, maybe 22 or something. Right?
So so these are huge changes all by material innovation. And and, yeah, I think wherever you go, you know, I can probably dig deep enough and then tell you, well, actually the very foundation of what you're doing is a material problem. And so I think it's just very nice to work on this very, foundation, and also because I think this is maybe also something that's happening now, is we can start to search through this material space.
This has never been the case, right? It's like scientists, normal way of working is you read papers, and then you come up with an hypothesis, you do an experiment, and you learn, etcetera. So that's a very slow process.
Now we can treat this as a search engine. Like we search the internet, we now search the space of all possible molecules, not just the ones that people have made, or that they're in the universe, but all of them. Right?
And we can make this kind of fully automated. That's the hope, right? We can just type, it becomes a tool where you type what you want, and something starts spinning, and some experiments get going, right?
And then out come a list of materials, and then you look at it, say maybe not, and then you refine your query a little bit. And you kind of do research with this search engine where a huge amount of computation and experimentation is happening somewhere far away in some lab or some data center or something like this. I find this a very promising view of how we can sort of build a much better materials layer underneath almost everything.
And also a more sustainable materials. Our plastics are polluting the planet. Or if you can come up with a plastic that kind of destroys itself, you know, after, I don't know, a few weeks, right, and actually becomes a fertilizer, these are these are things that are not impossible at all.
These these things can be done, right, and we should do it. Can you tell us what a little bit just generally about Cusp AI, and then I have a ton of questions. Yeah.
So Cusp AI started about twenty months ago, and it was because I was worried about, I'm still worried about climate change, and so I realized that in order to get, you know, to stay within two degrees let's say, we would not only have to reduce our emissions to zero by 2050, but then you know another half century or even a century of removing carbon dioxide from the atmosphere, not by reducing your emissions, but actually removing it, at a rate that's about half the rate that we now emit it. And that is a unsolved problem, but, And if we don't solve it, two degrees is not gonna happen, right? It's gonna be much more, and I don't think people quite understand how bad that can be.
Like four degrees, like very bad. So this technology needs to be developed, and so this was my and my co founder Chad Edwards' motivation to start this startup. And also because, you know, we saw the technology was ready, which is also very good, if you're, you know, the time is right to do it.
And, yeah, so we we now, in in the meanwhile, we've grown to about 40 people. We've kind of collected a 130,000,000 investment into into the company, which is for a European company is quite a lot. I would say it's interesting that right after that, you know, other startups got even more, so that's kind of tells you how fast this is growing.
But yeah, we are now at the, so we built the platform of course, but it's for a series of material classes, and it needs to be constantly expanded to new material classes. And it can be more automated, because you know we're not putting LLMs in, as the whole thing gets more and more automated. And now we're moving to sort of high throughput experimentation.
So connecting the actual platform, which is computational, to the experiments so that you can also get fast feedback from experiments. And I kinda think of experiments as something you do at the end, although that's what we've been doing so far. I want to think of it as what I would call a sort of a physics processing unit, like a PPU, right?
Which is you have digital processing units and then you have physics processing units. So it's basically nature doing computations for you. It's the fastest computer known, possible even.
It's a bit hard to program because you have to do all these experiments. It's also quite bulky, it's like a very large sort of thing you have to do. But in a way it is a computation, and that's the way I wanna see it.
So I wanna, you can do computations in a data center, and then you can ask nature to do some computations. Right, your interface with nature is a bit more complicated, but then these things will have to seamlessly work together to get to a, you know, a new material that you're interested in. And that's that's the vision we have.
We don't say super intelligence because I don't quite know what it means.
I don't wanna oversell it, but I do want to automate this process and give a very powerful tool in the hands of the chemists and the material scientists. That's actually brings up a question I wanted to ask you.
your thought process was in developing it? Yeah. Actually, it's been surprising me.
It's not rocket science, It's I would not rocket science in the sense of the design, and basically the design that, you know, I wrote down at the very beginning it's still more or less the design, although you add things like I wasn't thinking very much about multi scale models, and it a common rater that actually multi scale is very important. In the beginning I wasn't thinking very much about self driving labs, but now I think you know, we are now at the stage we should be adding that. And so there is sort of bits and details that we're adding, but more or less it's what you see in the slide decks here as well, which is there's a generative component that you have to train to generate candidates, and then there is a digital twin, multi scale, multi fidelity digital twin, which you walk through the steps of the ladder.
You know, do the cheap things first, you weed out everything that's obviously unuseful, and then you go to more and more expensive things later. And so you narrow things down to a small number. Those go into an experiment, you know, do the experiment, get feedback, etcetera.
Now things that also have been more recently added is sort of more agentic sort of parts. You know, have agents that search the literature and come up with, you know, actually the chemical literature, and come up with, you know, chemical suggestions for doing experiments. We have agents which sort of autonomously orchestrate all of the computations and the experiments that need to be done.
You know, they're in various stages of maturity and they can be continuously improved I would say, and so that's basically, I don't think that part is rocket science, you know, the design of that thing is not like surprising, but it's surprisingly hard to actually build it. Right, so that's the thing that is, the moat is in the data that you can get your hands on, and actually building the platform. And I would say there's two people in particular I want to call out, which is Felix Hunker, who is actually, you know, building the scientific part of the platform, and Alessandro de Maria, who is building the the sort of the skate the the kind of this the MLOps part of the platform.
Yeah. And so and and recently, we also added sort of Aaron Walsh to our team, who is a a very accomplished scientist from Imperial College. We're very happy about that.
He's gonna be our chief science officer, and we also have a partnerships team that sort of seeks out all the customers, because I think this is one thing I find very important. In principle, it's so complex to actually bring a material to the real world that you must do this, you know, in collaboration with sort of the domain experts, which are the companies typically.
industrial partner to go on that journey with us. Makes a lot of sense. Over the evolution of the platform, did you find that you that human intervention, human, I guess you could start out with a pure you could you could imagine two directions.
One, start up making everything purely automatic, automated, agentic, so on. And then later on, you, like, find that you need to have more human input and feedback different steps.
you know, lots of steps and then, like, kind of Yeah. Figure out ways to remove, you know That that's it. It's second one.
So you build tools. Yeah. So you so it's much more modular than you think, but it's like, we need these tools for this application, we need these tools.
So you build all these tools, and then you go through a workflow, actually in the beginning just manually, so you put them, you can have first this tool, then run this tool, then run this, etcetera, etcetera. So you put them in a workflow, and then you figure out, oh actually, you know, this porous material that we're trying to make actually collapses if you shake it a bit. Okay, then you add a new tool that says test for stability, right?
Yeah. And so there's more and more tools, and then you build the agent, which could be a Bayesian optimizer, or it could be an actual LLM, you know, maybe trained to be a good chemist, that will then start to use all these tools in the right way and the right order. Yeah.
Right? But in the beginning, it's like you as a chemist are putting the workflow together, and then you think about, okay, how am going to automate this, One very easy question you can ask yourself you know, every time somebody who is not a super expert in DFT, and he wants to do a calculation has to go to somebody who knows DFT, and so could you start to automate that away, which is like okay, make it so user friendly so that you actually do the right DFT for the right problem and for the right length of time, and you can actually assess whether it's a good outcome, etcetera. So you you start to automate smaller small pieces and bigger pieces, etcetera, and in the end, the whole thing is automated.
and Yeah. Less so trying to create a an automated process.
I think it's sort of the same what you're saying because, yes, we want to automate Yeah. But we don't see something very soon where the chemists and the domain expert is out of the loop. Yeah.
But it's a retreat, right? It's like, okay, so first you needed an expert to tell you precisely how to set the parameters of the EFT calculation. Yeah.
Okay, maybe we can take that out, maybe automate it, right? And so increasingly, more of these things are going to be removed. Yeah.
In the end, the vision is it will be a search engine somebody, a chemist, type things and will get list candidates, but the chemist will still decide what is a good material and what is not a good material out of that list, right? So the vision of a completely dark lab where you can close the door and you just say, you know, find something interesting and then it just figure out what's interesting and will figure out, know, and say oh, I found this new material to blah blah blah blah, right? That's not the vision I have.
At least not for, you know, I don't know, a long time. So for me it's really empowering the domain experts that are sitting in the companies and in the universities to be much faster in developing their materials. And I should say it's also good to be a little humble at times, because it is very complicated, you know, to make it and to bring it into the real world, and there are people that are doing this for their entire lives, right?
And it's like, I wonder if they scratch their head and say, well, you know, how are you gonna completely automate that away in the next five years? I don't think that's gonna happen at all. Yeah.
So so to me, it's a increasingly powerful tool in the hands of the chemist.
I have a question. You've talked before about getting people interested based on having, you know, sort of a big breakthrough in materials. Yes.
It's just incremental change. I'm curious what you think about the platform you have now and are sort of stepping towards, and how are you chasing the big change, or is this like incremental, or is there they're not mutually exclusive obviously, but Yeah. What do you think about that?
We follow a mixed strategy.
So we are definitely going after a big material. Again, we do this with a partner. I'm not gonna disclose precisely what it is, but we have our own kind of long term goal.
You can call it a lighthouse, or you know, sort of moonshot or whatever, but it is going to be a really impactful material that we want to develop, as a proof point that it can be done, and that it will make it into real world, and that AI was essential in actually making it happen. Yeah. At the same time, we also are quite happy to work with companies that have more modest goals.
Like I would say one is a very deep partnership where you go on a journey with a company, and that's a long term commitment together. And the other one is like somebody says, I need a force field. Can you help me train this force field and then maybe analyze this particular problem for me?
And I'll pay you a bunch of money for for that, and then maybe after that we will see. And that's fine too. Right?
for the good. Yeah. And do you feel like from a platform standpoint, you're ready for that?
get those big breakthroughs I got. What I find interesting about this field is that every time you build something, it's actually immediately useful, right? And so unlike quantum computing, which, or nuclear fusion, so you work for I don't know, twenty, thirty, forty years, and nothing, nothing, nothing, nothing, and then it has to happen, right?
And when it happens it's huge. So it's quite different here because every time you introduce, so you go to a customer and you say, so what do you need, right? So we work let's say on a problem like water filtration.
We wanna remove PFAS from water, right? So we do this with a company Chimera. So they are a deep partner for us, So we're on a journey together.
I think that the breakthrough will happen with a lot of human in the loop, because there is the chemist who have a whole lot more knowledge of their field, and it's us who will help them with AI training, AI and new methods, and in that kind of these interfaces, interactions, something beautiful will happen, and that will have to happen first before this field will really take off I think. And so in the sense that it's not a bubble, let's put it that way. So people see that it's actual real that's happening.
So in the beginning it will be very, you know, with a lot of humans in the loop, I would say, and I would hope we will have this new sort of breakthrough material before, you know, everything is completely automated, because that will take a while. And also it is very vertical specific. So it's like completely automating something for problem a, you know, and you can probably achieve it, but then you'll sort of have to start over again for problem b, because you know your experimental setup looks very different, you know the machines that characterize your materials look very different.
Even the models in your platform will have to be retrained and fine tuned to the new class. So every time you have a lot of learnings to transfer, but also the problems are actually different. Yeah, yeah.
And so, yeah, so I would want that breakthrough material before it's completely automated, which I think is kind of a long term vision, and I would say every time you move to something new you'll have to start retraining, and humans will have to come in again and say, okay, so what does this problem look like? Now sort of, you know, point the machine again in the near direction, and then use it again.
among us, me included a bit of a scientist, There's a lot of terminology. You mentioned DFT. Equivariance, we've talked about.
Can you sort of explain in, you know, engineering terms or at the level of sophistication in engineering? Well, how what is equivariance?
So equivariance is the infusion of symmetry in neural networks. So if I build a neural network, let's say, that needs to recognize this bottle, right, and then I rotate the bottle, it will then actually have to completely start again, because it has no idea that the rotated bottle, well actually the input that represents the rotated bottle is actually rotated bottle, it just doesn't understand that. Where if you build equivariance in, basically once you've trained it in one orientation it will understand it in any other orientation.
So that means you need a lot less data to train these models. And these are constraints on the weights of the model. So basically you have to constrain the weights such that it understands it, and you can build it in, you can hard code it in.
And yeah, the symmetry groups can be translations, rotations, but also permutations, like in graph neural network, permutations.
And in physics, of course, there's many more of these groups. To pray devil's advocate, why not just use data augmentation by, you know, your model is in all the different orientations?
As an option, it's just not exact. It's like why would you go through the work of doing all that, where you would really need an infinite number of augmentations to get it completely right, where you can also hard code it in. Now I have to say, sometimes actually data augmentation works even better than hard coding the equivariance in.
This is something to do with the fact that if you constrain the optimization, the weights before the optimization starts, the optimization surface or objective becomes more complicated, and so it's harder to find good minima. So there is also a complicated interplay I think between the optimization process and these constraints you put in your network. And so, yeah, you'll hear kind of contradicting claims in this field.
Like some people, and for certain applications, it works just better than not doing it, and sometimes you hear other people say, if you have a lot of data and you can do data augmentation, then actually it's easier to optimize them, and it actually works better than putting the aggregates in itself.
mathematically founded models and strategies for doing deep learning?
Yeah. Ultimately, it's a trade up between data and inductive bias. Yes.
So if your inductive bias is not perfectly correct, you have to be careful because you put a ceiling to what you can do. But if you know the symmetry is there, it's hard to imagine there isn't a way to actually leverage it. But yeah, so there is a bitter lesson.
And one of the bitter lessons is you should always make sure your architecture scale, unless you have a tiny data set, in which case it doesn't matter. But if you, you know, the same bitter lessons or lessons that you can draw in LLM space are eventually going to be true in this space as well, think. Yeah.
Yeah. Can you talk a little bit about your upcoming book and tell the listeners, like, what's exciting about it? Yeah.
They should read it. So this book is about, so it's called generative AI and and stochastic thermodynamics. It basically lays bare the fact that the mathematics that goes into both generative AI, which is the technology to generate images and videos, and this field of non equilibrium statistical mechanics, which is systems of molecules that are just you know moving around and you know relaxing to their ground state, or that you can control to have certain, you know be in a certain state.
The mathematics of these two is actually identical. And so that's fascinating. And in fact what's interesting is that Jeff Hinton and Radford Neal already wrote down the variation of free energy for machine learning long time ago.
And there's also Carl Friston's work on free energy principle and active entrance. But now we've related it to this very new field in physics, which is called stochastic thermodynamics or non equilibrium thermodynamics, which has its own very interesting theorems like fluctuation theorems, which we don't typically talk about, but we can learn a lot from. And I think it's just it can sort of now start to cross fertilize.
When we see that these things are actually the same, we can, like we did for symmetries, we can now look at this new theory that's out there developed by these very smart physicists and say, okay, what can we take from here that will make our algorithms better? At the same time, we can use our models to now help the scientists do better science, And so it becomes a beautiful cross fertilization between these two fields. The book is rather technical I would say, and it takes all sorts of things that have been done as stochastic thermodynamics, and all sorts of models that have been done the machine learning literature, and then basically equates them to each other.
And I think hopefully that sense of unification will be revealing to people. Wait, and when is it out? Well, depends on the publisher now, but I hope in April I'm gonna give a keynote at iClear, and it would be very nice if I have this book in my hand, but it's hard to control these kind of timelines.
I'm looking forward to it. Great. Likewise.
Thank you very much.
Shared via Hopper