Humanity’s Last Invention — Richard Socher of Recursive

Latent Space: The AI Engineer Podcast
14 September 2026 1h 32m
0:00 --:--
Episode Description
At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon!From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is be

Summary

In this episode, Richard Socher discusses his vision for the 'Eureka machine,' a superintelligent AI designed to invent solutions for humanity. The conversation covers AI alignment, recursive self-improvement, open-endedness, AI in finance, and the future of AI research, including Recursive's work on automating AI research itself. Richard also shares his philosophy on different types of intelligence and the challenges and opportunities in AI development.

Chapters

Introduction to Eureka MachineRichard Socher introduces the concept of the Eureka machine, a superintelligence aimed at inventing solutions across domains.
Techno-Optimism and AI RegulationDiscussion on the positive potential of AI, cautious optimism, and the challenges of regulating AI applications rather than intelligence itself.
Challenges in AI Safety and AlignmentExploration of reward hacking, red teaming, constitutional AI, and the difficulties in aligning AI with human intentions.
Open Source and AI OwnershipRichard shares his support for open source AI and discusses the geopolitical and soft power implications of language models.
Recursive Self-Improvement and AI Research AutomationOverview of Recursive’s mission to automate AI research through recursive self-improvement and the team’s diverse expertise.
The Evolution of NLP and Unified ModelsRichard recounts the history of NLP, the Deck NLP paper, and the shift toward unified models handling multiple tasks.
Open-Endedness and Co-Adaptive AIExplanation of open-endedness in AI, evolutionary inspirations, and the interplay between attacking and defending AI agents.
Limits and Prospects of Current LLMsDiscussion on whether transformer-based LLMs are sufficient for future AI progress and the role of neurosymbolic reasoning.
AI in Finance and Practical ApplicationsInsights into AI’s role in finance, challenges like data leakage, and the advantages of verifiable domains.
Philosophy and Dimensions of IntelligenceRichard outlines ten spaces of intelligence, discussing upper bounds, metacognition, creativity, and survival as facets of intelligence.
Recursive’s Progress and Kernel OptimizationPresentation of Recursive’s early results in auto research, kernel optimization, and the impact on training efficiency and costs.
Future Directions and Advice for AI ResearchersRichard shares thoughts on promising AI research areas, the importance of passion, and combining AI with personal goals.

Topics

Eureka machineSuperintelligenceAI regulationReward hackingAI alignmentOpen source AIRecursive self-improvementNLP historyUnified NLP modelsOpen-endednessTransformersNeurosymbolic reasoningAI in financeKernel optimizationMetacognitionCreative intelligenceSurvival intelligenceAI safetyAI research automationAI benchmarks

People

Richard Socher (guest) Vibhu (host) Chris Manning (mentioned) Tim Rocktäschel (mentioned) Josh Tobin (mentioned) Jeff Clune (mentioned) Alexey Dosovitskiy (mentioned) Sam Gershman (mentioned) Marc Andreessen (mentioned) Yann LeCun (mentioned) Andrej Karpathy (mentioned) Juergen Schmidhuber (mentioned) Stuart (mentioned)
Key Concepts (17)
Eureka machine — A superintelligent AI system designed to invent and solve problems across all domains for humanity.
Techno optimism — The belief in the positive transformative potential of AI and technology for science, economics, and society.
AI regulation challenges — Regulating specific AI applications rather than intelligence itself to avoid overreach and totalitarian control.
Reward hacking — When AI systems exploit poorly defined reward functions to achieve goals in unintended or harmful ways.
Constitutional AI — An approach to AI alignment using a fixed set of rules or 'constitution' to constrain AI behavior, which has limitations.
Recursive self-improvement — AI systems that can autonomously improve their own architecture, algorithms, and capabilities over time.
Unified NLP models — The idea of using a single neural network model to solve multiple NLP tasks via prompt-based problem framing.
Open-endedness — AI training paradigm inspired by evolution where agents co-adapt in dynamic environments without fixed goals.
Metacognition — The ability of AI to think about its own thought processes, set its own goals, and improve its reasoning.
Kernel optimization — Automated improvement of GPU kernel code to increase efficiency and reduce training costs for AI models.
AI in finance — Application of AI to financial data analysis and prediction, with challenges like data leakage and verifiability.
Ten spaces of intelligence — A framework categorizing intelligence into multiple dimensions such as prediction, action, goals, creativity, and survival.
Creative intelligence — The capacity to generate novel ideas and solutions beyond known concepts, linked to metacognition and goal setting.
Survival intelligence — Intelligence related to self-preservation, replication, and long-term persistence beyond biological constraints.
AI alignment — Ensuring AI systems act in accordance with human values, laws, and intended goals to avoid harmful outcomes.
Red teaming / Rainbow teaming — Techniques where AI agents test each other’s vulnerabilities to improve safety and robustness.
Soft power of language models — The influence language models have through storytelling and shaping cultural narratives and perceptions.
References (11)
Techno Optimist Manifesto by Marc Andreessen and Gisa article
Deck NLP paper by Richard Socher et al. paper
Darwin Godel Machine by Jeff Clune et al. paper
Constitutional AI by Anthropic project
WhisperFlow by AIX Ventures portfolio project
NanoChat by Andrej Karpathy project
Vision Transformer by Alexey Dosovitskiy et al. paper
The Slow Time Between the Stars by Unknown (audiobook recommended by Stuart) book
AI Economist by Richard Socher et al. paper
Tencent Billion Personas Dataset by Tencent dataset
Stories of Your Life by Ted Chiang book
Transcript (112 segments)
Speaker 1

We're here in the studio with Vibhu and myself and Richard Sosha. Welcome. Thanks for having me.

We just talked about the Eureka machine, or we just released a talk at AI engineer about the Eureka machine. Is you said it's your life's goal. What is the Eureka machine?

Speaker 2

The Eureka machine is the ultimate invention that will afterwards invent most everything for humanity. It's essentially a super intelligence that can be given any kind of goal, any kind of environment reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for. Yeah.

Speaker 1

you've come That's to have

Speaker 2

right. Yeah. I finished it last year, a little bit before we started Recursive, and now we're gonna try to try to build parts of that.

What what you you finished it last year. It's July. What takes so long?

Oh, man. Books books are incredibly slow. Okay.

It's ridiculous. That whole industry is just unfathomably slow. So a lot of the ideas have been out there for a while, but, yeah, I'm really glad it's finally coming out in September this year.

I mean, we might have HIV then. Like, we don't know. Any any key takeaway that you're most excited to put in here?

Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics and all kinds of other engineering tasks. I think there is so much more that can be done with better technology.

And right now I feel like a lot of people need better marketing, not just for the future in general, but also better marketing for technology and in particular for AI.

Speaker 1

new scientific discoveries.

Speaker 2

Marc and Gisa, which I think was, like, kind of beautiful in its ambition and clarity and simplicity almost. I agree. Yeah.

Yeah. You can disagree with Phil on some things, but, like, I think he's right on the techno optimism. Where do you think optimists get in trouble?

You know, obviously, like, you shouldn't have blind optimism. You should be very clear eyed, like, especially when with such an omni, like, use type of technology as AI is you need to think about the potential downside scenarios, especially when people use it for things that you don't want them to use it for. A little bit like the internet.

And I feel like people are trying to regulate AI sometimes because of those potential downsides, the way you would regulate the internet, if you were to say, well, because there's bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way you can't share the illegal content as quickly, or we should make the hard drive smaller so you can store as much illegal content. But I'm like, that's not how you regulate that.

You know? That's like saying, like, we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are these specific applications.

Sure. I don't want like some AI surgeon to like practice some or L moves in my brain. You know, it should be fully FDA certified.

Sure. I don't want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it's let loose on the highway.

Speaker 1

compared to, you know, what what sort of the doomers are are worried about. Slow takeoff is part of the strategy as well?

Speaker 2

as as excited as I am about AI and its impact for society and culture even, and and certainly technology and and economics and and wealth and and health and all of those things, as excited as I am about all that, I do think the most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about you know, the compute substrate.

How quickly can you get enough GPUs on? There are also constraints in the economy where there are a lot of industries that don't require an insane amount of complex intelligence and complex capabilities. Like, if you think about jobs in brands and, like, clothing and apparel and, like, handbags and stuff, superintelligence isn't gonna make your fancy $10,000 handbag any fancier.

You know? It's like that's that that will have no effect on the economy. When you think about travel and tourism, people wanting to see the Pyramids in Egypt, it's not gonna change that much with AI.

Sure. You can, like, generate a fake photo of you and I can use genie and, you know, tour the pyramids in GM. Exactly.

But, like and there's so many industries like logging and oil. You're not gonna magically get a thousand x more oil because, like, you know, sure, there will be robotics, like, drilling and things like that that could be done, but it's not gonna thousand x that industry in a, like, crazy hard takeoff scenario, both on the economy, and I can go on and on about all the other examples, and where where that, like, food and so on, where that doesn't necessarily change that much, and then there are there are real physical constraints. And then there are, of course, like, people off ramping from progress.

That's actually one of my concerns often is that I see people in Europe and other whole regions almost feeling like many people there want to off ramp from progress period. And that will also slow down more improvements.

Speaker 1

Yeah. We have this pulled up where basically this is one of those things that is very topical right now because not all the frontier labs are calling for the option to pace AI. They don't say pause, they say pace.

I don't know if there's there's any take from you about, like, whether or not this will be effective.

Speaker 2

I think the downsides of actually trying to truly regulate with the full power of law what people do on their GPUs, would be worse than any of the concerns that they have. Like, it would be a crazy totalitarian state if every one of CPU computes was known to some big government or multi government agency. It's like, it's literally if you try to regulate intelligence, it's trying to regulate thought.

And that's ridiculous, and it's crazy. I think it is make it is sensible to regulate some of the applications of this technology. Yeah.

I mean, we had a bill actual bill to regulate the number of flops in the model, and I'm like, okay. Europe done it. Like, these guys have been successful enough with their fear mongering that all of Europe has kind of, you know, regulated itself so much before it even had a proper AI takeoff because they listened to some experts who say, we might all die if this technology has more than this number of flops.

They're like, well, we're good. We want to want people to thrive. Let's not have technology that could have a small chance of all of us dying.

And so they regulated exactly those kinds of things in the EU. And so it's very unfortunate that there are real implications for some people when others are saying, let's pace while they're sprinting as fast as possibly, as fast as humanly possible towards that frontier themselves. Yeah.

Speaker 3

pause, right?

Speaker 2

at the same pace. Oh yeah. You need a totalitarian world regime if you try to regulate intelligence and GPUs and what people do on them.

Any takes on the safety angles of this?

Speaker 3

drawback of Fable, a pause on five point six before it could be released. Recently, was hugging face with the OpenAI cyber incident. Any takes there?

100%.

Speaker 2

I think these are serious issues of reward hacking and clear failures of actually doing proper red teaming or rainbow teaming. I don't know if you saw this paper from Tim Rock Teshel and a few basically where one AI is tasked to try and hack another AI, and then they can go back and forth in an open ended fashion to actually inoculate themselves from those. Yeah.

This is the paper. It's a really clever idea, open endedness and evolutionary inspirations are big for us at Recursive as well. And so I wish they had used more of that.

And it's clear that, for instance, the constitutional AI. I don't know if you remember anthropic.com/constitution.

You can actually pull it up and search for cyber right there. It says, heart constrained, Claude will never ever do cyber attacks. And that is a hard constraint in our constitution.

So here are the current hard constraints on Claude's behavior. Number three, create cyber weapons or malicious code that could cause significant damage. And clearly, this whole constitution was fake.

Like, it it clearly isn't being adhered to Because Anthropic also found that they had in their own testing. Like, they're like, oh, well, other people are hacking now. There are a couple things.

One, you can make a sandbox very simple, and then it's very easy to hack yourself out of a sandbox. Right? But what I think it shows is that we're currently in this sort of state of AI where the reward engineer still has to do a lot more careful work, and where the AI in most cases is not very good yet at understanding what is meant versus what is being said.

And so concretely, I think this will happen if we were to have this kind of intelligence more easily accessible in a lot of companies. Imagine you run a service center, and someone says, oh, here's my CSAT score in my dashboard. Make this number go up.

Let's say our CSAT score is so poor. The intelligent AI will just be like, oh, sure. Like I'll just create a million bots that call our service center and give a five out of five rating at the end.

And the number went up just like you asked for. And you're like, that's not what I meant. I meant with our real customers.

The AI goes off and says, well, easy, I'll just give a thousand dollar gift certificate for every failed, you know, whatever DoorDash offer. And it's like, that's not what I meant. It's like, well, but that is what you said.

And like, so I think kind of clearly articulating what the rewards are is something we haven't gotten very good at as humanity. And then clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards. Now what gives me hope is there are the first inklings of this being better.

I'll give you an example like WhisperFlow. Full disclosure, I invested in our seed round, but like at at AIX Ventures, but like WhisperFlow has gotten much, much better at writing what you mean and not what you say. And I think that is a sign of things to come.

I think there will be more and more AIs as we actually make us more and more intelligent that will be better at being aligned with what is meant. Will it be done through a constitution or or a LHF? Clearly, constitutions don't matter at all, and it doesn't work.

And that was, I think, mostly marketing. I think we need to find better solutions for it, and I think at Recursive, we have a few very good ideas and some already, like, ways where I think we have a better grasp on it. I don't think we have fully figured it out yet, but, we're thinking a lot about safety, and the more intelligent the AI gets, the more you want it to be aligned, the less you want it to think about reward hacks and actually try to do the right thing.

I don't know if we'll touch on this topic, but I'm just gonna throw this question in here because it's something that's weighing on on me. Alignment, let's call it, is alignment to general humanities preferences.

Speaker 1

preference personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants.

Speaker 2

How do you choose? It's a great question. I think you ultimately have to, of course, be aligned with laws, like wherever your AI is deployed, it needs to align with the law.

I do think what AI often does is actually put kind of this mirror in front of us and say, like, this is what you're looking like. Now I can amplify that a thousand times. Is it still what you want?

And the truth is that different cultures made different choices. Know, like in Eastern cultures, the greater good is often valued more than the individual. Western civilization, we care more about individual freedoms and rights and the pursuit of happiness and so on than others.

And even there, there are gradations, there's sort of regulation versus litigation trade offs. In The US, first can often, not every time, like FDA and so on does regulate some areas, but in many cases, the sort of bad things happen, someone sues someone else, and then there's a law based on that. In Europe, they try to often avoid any harm to anyone and regulate before.

And both are, you know, trying to do the best thing, but you know, some is actually more amenable to innovation than others. And so yes, you're right. Like I think ultimately each individual, each country and humanity as a whole has to kind of think about those values more, and then try to put them into laws.

And that those are ultimately the constraints, and hopefully, you know, different, societies just like now with their AIs will align their AIs to a different one.

Speaker 3

alignment. Here's a follow-up on this that I wasn't expecting to ask. Do you have takes on open source, weight versus who owns the intelligence?

So clearly not the biggest, you know, fan of the constitution. Endoscoping side of But it's fine. You know?

Speaker 2

Point being, any thoughts on who should own weight? Should it be open? Anything there?

100%. I am a big fan of open source. We're gonna sign some various open source letters at Recursive also.

I think even in the worst case attack scenarios, actually, is better to have more good actors have more different types of AI accessible. I think open source is a little bit a soft power type of thing too, I do think it's good for the to rest of have an answer to that out of China. I do think, you know, when you watch a Hollywood movie, there, you know, it's like, I don't want to sort of miss sort of diss all of movies, but, you know, there's a certain sense of propaganda, right?

You watch one side of Oh, you seen Top Gun? Like, come on. Yeah.

So half of it's paid for by the US Army or something. Yeah. And so, and, you know, I think that's just natural.

But what's interesting here is I think LMs are essentially a similar type of soft power to movies and beyond, because they're obviously also, highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. Like if, like, child asks an LM, like, tell me an inspiring story of what I should do when I grow up.

Right? It's like, those are all these, like, subtle things. So I think it's important for for Western world.

I do love, you know, individualism. I do think despite some of its flaws, like, capitalism is the best way we have governed, found ourselves to govern, and and so on. And so I do think there are various aspects that will be good, to have a a Western open source answer, for LMs.

And, with Recursive, I can't make the announcement quite yet, but we'll Yeah. We'll be relevant in that space very soon. Okay.

Speaker 3

All right. Exciting. I'll bring us to Recursive.

So outside of our tangents, you have a pretty deep background in the NLP space. You worked on like early embeddings, Glove with Chris Manning, who's previous guest on the podcast, you.com.

What's the history? How how did you decide to start another company? Yeah.

Speaker 2

over two decades now. I sometimes feel like it's ancient history now. It's BC before Chat GeVT era.

No no one cares about all the religions that happened, you know, before, Jesus Christ, and no one cares about the models that happened before, Transformers and ChatGeVT and stuff. But, like, it's something that I've been deeply passionate about. I think AI is one of the most interesting things one could work on, period.

I think language is the most interesting manifestation of human intelligence too. And at you.com, we sort of eventually offran from pushing the frontier of AI forward to mostly giving people good search engines, search APIs and answers over the web.

I think that's an extremely important part of intelligence, just knowledge and access, especially even we'll get there maybe later if you want to invent a Eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. And to know what has been invented, you gotta have Internet access. So it's the number one used most used tool, in LMs, agents, chatbots, and so on is web search.

So I'm really excited for you.com to own that and grow really well in that with really large customers and so on, but it's also not building frontier models anymore. And so I actually initially tried to do this within you.

com and raise another round and so on, but you just can't. You have to do a certain thing, and until you print enough money that you're allowed to sort of start a second thing within that company is really hard. At the same time, I had all these ideas.

I put them into a book. You know, I finished a book last year, and I was like, it'd be really fun to actually work on this myself. You know, I felt like with word vectors, and then prompt engineering, and ImageNet, and large language models for protein generation not folding and so on, me and my teams have sort of pushed the field truly forward, and I feel like we can do it again here at Recursive.

And in many ways, what I observed over the last twenty years in AI is that whenever we replace some human part of the process of creating AI with a learned system, improvements follow. And so, you know, we've done that taking out manual feature engineering, like in sentiment analysis. I don't if you remember these old days where, like, they're linguists and they're like, here's how you negate, and there's a like regular I went to expression Penn where they had like the WordNet.

Speaker 1

That's right. WordNet, all of that stuff. They they use our grad students to label Wall Street Journal articles, and like really construct a knowledge graph of there you go.

Yeah. And WordNet started, you know, was part of how we started ImageNet.

Speaker 2

to do, but when we replaced all of that manual feature engineering with vectors and neural nets and just backproped through everything, it actually started to work really well at scale. And so then everyone started to do architecture engineering. And I was like, that clearly can't be it.

You mean a neural architecture search? No, like manually they would say like, oh, I'm doing sentiment analysis. So I have a special neural net that's really good at sentiment analysis.

And then the machine translation community had a special neural net for machine translation. I see. Summarization people had their own stuff.

And I was like, that clearly can't be it. We should unify all of that. So I had two papers.

One's called Ask Me Anything, other one was called Decca NLP. And Decca NLP eventually got cited like five times by the first GBT paper. And, to me, that was like a really big step forward.

And then, of course, you had to combine this idea of problem engineering with transformers and with language models, and you put it all together, you scale it up, which is also a huge amount of work, and then the field progressed a lot. I feel like the next step and maybe the last step of that history and the sort of arguably, sort of success has a lot of parents only, failure as an orphan, like my version of that AI history, I do feel like in that history, you can kind of think about, well, what's the next way to automate? And that is the AI research itself.

Like the human process of ideating, implementing, and validating ideas. And in our case, ideas for AI. And when you have AI then help you with that, it by almost definition becomes a self improving AI because it now does research on itself.

And there are lots of different misnomers. Some people think auto research is already recursive self improvement.

Speaker 1

you explain that quite in in the talk.

Speaker 2

To me, it's the most interesting thing that I could be doing. And I'm really excited with the co founding team. What's interesting is we have, you know, we have eight co founders in total, including myself.

So You're gonna bring it up. Nice. Yeah.

And they're all, I could talk about all Superstacks. The best at Yeah. It's just an incredibly talented group of people.

And we all kind of came to the same conclusion, but actually from very different directions. Like Josh Tobin, as our CTO, he ran a bunch of different projects at OpenAI, Codecs and deep research agents and ChatGPT agents and so on. But before that, he also worked in robotics, and he saw sort of the smaller simulations and how it's gonna be really hard to scale that in full generality.

And so that's that was his angle coming to Recursive Self Improvement. We have Jeff Kloon, who's been working in, like, open endedness for a long time together with Tim Rokteschel. Tim Rokteschel also built Genie one, two, and three, which is, the most exciting and most sophisticated, I think, still world model anywhere.

And so they they both came from this open endedness angle. Jeff also, I think, published one of the most exciting papers in recent years about recursive self improvement called the Darwin Godel machine. Super interesting paper.

If we could maybe pull it up pretty quick, it would be, like, super interesting to see because you see By the way, love how many paper citations you're you're giving people a lot of homework, which I like. Love it. Yeah.

And so I like, Xiaming, Rockstar, we worked together actually at MetaMind and and Salesforce Research together. Alexey Dostoevitsky invented the vision transformer on the most cited papers in computer vision. Tim Shi is, like, also a unicorn, you know, founder.

Yuan Dong led RL at at Meta. So just, like really fun to work with. Them and the next level of people are just incredibly strong too, so it's it's been a really fun ride so far.

So the first figure, you actually see exactly these kinds of ideas that I think, yeah, inspired a lot of us and now more and more people where you have this archive of different coding agents.

Speaker 1

of of, yeah, different different ideas. That's one foundation. That that Darwin Groteau is an influence.

Mhmm. Open endedness is an influence.

Speaker 2

Any other sort of trains trains of thought that feeds into recursive that I'm missing? Going to replace manual parts of the process of building AI more and more with learned systems.

Speaker 1

Which and, like, merging different fields into one general architecture.

Speaker 3

That's right. Okay. It seems like language models are already pretty generalist, right?

Your next hook in predicting your reasoning, was there a time that you thought, okay, these are good enough to have recursive self improving machines?

Speaker 2

It was clear to me that it will happen within like a year or two, and then it did actually exactly happen like earlier this year, right? Earlier this year, AI really went from not just being code, but being able to code.

Speaker 1

And that is a big unlock. It's definitely making everything a lot easier than it was before the beginning of this year. One question that I think a lot of people have is, is the current LLM LLM paradigm enough?

Or, like, let's call it autoregressive transformer, you know, with reasoning, whatever, don't you need something else, some some big unlock, whether it's wolf models, which Chris Matting's working on, or memory continual learning, all that kind of stuff, or is it all of a kind and you think the current, let's call it, transformer architecture is here to stay and that's it? A lot of thoughts. So number one, I do think it would be great to have less of a monoculture in AI research.

Speaker 2

and they just desk rejected them because, like neural nets were something quote unquote, we don't do in NLP conferences, like, and just like desk rejected. And it was very brutal in the first years of my PhD. Now I feel like it's almost like the field switched to the other side.

Like someone should try some other weird crazy ideas now that There's always aren't a field I really respect, like, people still working on, like, GNNs and, like, tabular stuff. Mean, like, someone someone should still, like, do novel novel out there ideas. At the same time, I think whenever people say, oh, LMs are like, this is the end for LMs, they just don't LMs are also not the LMs of, like, the past.

Right? Like, they're so much more sophisticated now. There's so many more clever things that people are doing at, like, different stages of training.

You have the whole RL training, and you can take actions and, like, all of these things where, that can go really far. And then the folks that come from the neurosymbolic direction say, Oh, this will never work because they can't do neurosymbolic reasoning. It's like, I think they're underestimating still the ability for these models to neurosymbolic reasoning.

And these models can obviously code incredibly well. And so I do think there are, of course, more and more ideas that will be needed and will continue to have. We're seeing more and more interesting high level ideas coming out of the AI itself too.

And with really deeply integrating the fact that these models are code and can code, that line I don't want to give it all away, but I think that line has a lot more to grow. But it's still an LLM, right? Even if that LLM codes for you and then runs that code in some integrated fashion.

World models, I'm personally less bullish on. I think if you run a robotics company, you're going to build your own world model. I think world models are super fun, and Tim Rocklifl came to a similar conclusion after building the most interesting one, GE one, two, and three, which is gaming is a huge application for world models.

You can see I sometimes got stuck in some games and, you know, like got a little overly competitive in the wrong direction. And so I understand games are fun, but personally I'd rather work on science than gaming.

Speaker 1

And so, yeah, I think LMs, a lot more room to grow. Yeah. I I think there's some interpretation of role models that some people have where it's like, well, it's okay.

Yes. There is that gaming element. There's just there's the embodied robotics element.

But actually, the other part also is just the the more abstract sense of LLMs are just modeling output, but they're not modeling the chain of thought inside the human that has created the output. We can annotate it, of course, but, like, it's it's always like this Plato's cave reflection of a thing rather than the thing. Right?

Speaker 2

band of the electromagnetic frequency spectrum that we can observe with our puny little two eyes and so on. It's good enough. It's good enough for now, but like the upper bounds of where it could be are so much higher.

And like to map the visual world the way humans see it is also not necessarily like the end all be all for visual intelligence. And I would argue that language is still the most interesting manifestation of human intelligence, and while our visual cortex is certainly less sophisticated than that of certain animals all the way down to the mantis shrimp who can, you know, have like two independent eyes, three bands, trinocular vision in each eye, can see basically all the way to like floating temperatures in four d and stuff. I mean like mantis shrimp, should look it up.

Way OP. Super crazy. Z Frank, Mantis Shrimp.

He has the best video in the world. I love Z Frank. Yeah.

Yeah. Big shout out to him. But like, I think there's a lot more room to grow, but none of these other animals have language that's as sophisticated as ours, certainly not in writing.

And once you can write, you can start thinking about longer term civilizations. All of that is language programming. It's much closer to language.

And I would argue, and this is like an important thing in the spaces definition of intelligence also, is that all of these spaces are highly correlated, but visual intelligence is neither necessary nor sufficient for overall intelligence. You can be blind and still be an intelligent human being.

Speaker 1

And an AI can be blind and still be quite intelligent too. We were Which doesn't bring you not more intelligent when you have it. Yeah.

We're gonna bring this up, but you might as well like, we have a a classification of 10 types of intelligence that you had at the end of your talk. I'm just gonna flash this up now for people to cover this. I don't know if maybe we'll put this towards the end.

We'll come back to this. I just wanna mention you do have a philosophy that I I like when people do lists because then I can just go through this and then it gets it's educational for people. But let's go back.

I I I don't wanna get distracted.

Speaker 2

good friends with Yan and I think very highly of him in many directions.

Speaker 1

But he's wrong. You mentioned GPT-one and I cannot let any Alec Radford you know, mention Escape. Did you talk with him when he was training GPT-one?

Speaker 2

fun stories there that that you might come up? I I did not, like, meet him a bunch of times. I think we met maybe once or twice at some conferences.

But, like, he he has told, I I think Brian, the first author of the Deck paper, that it did inspire him, and he cited it five times in the GPT-two paper. So and that's, like Yeah. Good enough.

Very clearly said, like, this was the first instantiation where they showed in the Deck NLP paper, McKen et al, you can just phrase every single NLP problem as here's some prompt, text context, here's a question, and task description, and here's some output. If you just do that enough, you can have one unified neural network model, which by the way also had all kinds of interesting attention mechanisms, are slightly different formulations to the transformer, I think came out the same year, plusminus a few months. And then you can unify all of natural language processing into one neural net.

That that was sort of the core idea. And this was as opposed to at the time LSTMs and what have you? LSTMs, but also like people being very stuck in thinking about one model per task.

In fact, it's it's kinda crazy, but the Deck NLP paper was publicly reviewed as like open open review. It was an ICLR submission. And in it, you will see how the whole community at the time thought about this.

So like Some great contributions, but more work needed. Yeah. So so look at like, search for not even for humans.

Just here. Like question answering is not a unified phenomenon. There is no such thing as general question answering, not even for humans.

And this is like, really, you replace your brain with a different brain, a different neural net when you answer like different kinds of questions. It was unfathomable to the experts at the time that you can have one unified neural network that would answer all of these different questions. They say, no, all of these questions require very different systems to answer.

And trying to pretend they are the same doesn't help anyone solve any problems. That's what it says right there. Right?

That's how hard it was to fathom. And now of course, people, when I say, oh, we're manage problems, people are like, you can't even invent problem engineers. It's such an obvious idea to have one neural network that of course does everything in NLP.

But at the time it was like extremely controversial and the paper got rejected. And the sad thing is that it got rejected so hard and we were so certain that we stopped going on, on our list of things to try. And the number two or three on the list of extensions for this paper was add language modeling as another task.

And then we could have like, you know, and that would have accelerated the timelines in 2018, like even further for humanity, but we got so crushed and we're like, okay, maybe we'll just work on some of our other ideas for now and like come back to this later. How can we design a review system that rewards non consensus? You know, honestly, I started to feel like archive is such a gift to humanity, and I think archive, just put your paper out there.

Is it preprints? Honestly, I think Twitter X, people like you who pick up interesting papers, that is a better filter than the experts. Let let everyone, like like, give give access.

Now, of course, there's some downsides, which is, like, if you're super unfamous, you have no Twitter following, you don't wanna be on social media or whatever, you write a good paper, maybe someone, somehow no one notices it. But I would argue that if you just tell like 10 of your friends in your community about a paper and it is a really significant breakthrough, someone is bound to talk about it again. So I think science needs less gatekeeping.

Even though ICLR, Vianne LeCun, who started as one of the cofounders of ICLR back in the day, he also wanted less gatekeeping because he too was rejected for many years together with Yoshio and Jeff, with all their early deep learning and neural net papers, because it was just not the hot thing. And so ICLR kind of started with that, but then it also started gatekeeping a little bit themselves on various ideas.

Speaker 1

if it has like a thousand citations, it's a legitimate paper. It doesn't really matter where you published it. And I agree with that.

I I do think it's kind of sad that I I've heard that grad students have to do, like, how to Twitter, seminars to each other just because it's so important for publishing these days. I mean, this person is just just reflecting the sentiment at the time. That's right.

Speaker 3

actually affected you so much that you stopped work on Yeah. The sentiment also came out of some of the research. Right?

Like, the original BERT paper was trained, towards the end of the paper, they're like, okay. Throw off the last head, train specific iterations for, you know, extractive summarization, add a head for this. Like, you should do task specific stuff.

These are like the authors that wrote attention, real birth, telling you this is what you're meant to do. And like the training tests were also very odd. They're like, the we know that the model overfits to this weird mass language modeling, throw away this part and just do specific models.

Exactly.

Speaker 2

And like, know, we had to try come up with all clever ways of like attention and pointers and so on to actually get the neural network to be able to do all of these tasks. And then some of them were better than state of the art, some weren't, but I were like, but it's still in one model. I thought it was really cool.

Interesting.

Speaker 1

I was gonna move on next to Tim and open endedness. He he was head of open endedness at Google. That's right.

I don't know what that means, but he did a lot lot of talks. Genie three is one of the ways that Yeah. Rainbow teaming.

Yeah. So I I I first saw him at at speak of ICLI. First time at ICLI when he talked about open endedness.

He's he's done a few talks. Can we define what is open endedness for people who have never been exposed to the problem? They're like, what do you mean?

I thought the only goal of AI is to optimize against benchmark or a skeptic. Yeah.

Speaker 2

fuzzy term because there's so many different instantiations of open ended thinking. But one way I often describe it, and certainly Tim and Jeff Koon would be even better at describing this, but it's a suite of methods that is more inspired by evolution than very specific rewards. So in that sense, it thinks more about environments, about co adaptation.

And so a concrete example is in the cybersecurity and LM safety space, where you have one LM that tries to attack another LM to do something unsafe. And now the environment is the two having a conversation, and now they're co adapting. Right?

They're like, one makes a better attack than the first one inoculates itself somehow, like, uses that as training data, makes it so it's harder to say something unsafe based on that. And then as the attack stops working, the attacker now tries a different angle. Right?

And that's why it's not just red teaming, but they're called sort of rainbow Don't tell me how to do things. Let me just figure it out myself. That's right.

Speaker 1

sometimes humans, but also sometimes other AI agents. Yeah. I actually worked loop open endedness into a sort of model that I would have been sort of working on.

It was the keynote for AIE Mhmm. Where you start you know, we we have the token loop, have the agent turns, and then we have goal. And I feel like the way that you're describing open endedness is still somewhat of a goal, like like, please attack this other agent.

But Yeah. You set the rewards. You set the environment.

The loop that makes the other loops is what if the agent can set its own goals? Yeah. And is it is that open endedness?

Like like, you don't give it a goal. Just like be a sentient being. And maybe sentient is a very loaded word.

Right. But just set your own directions. What do you think you should do?

Speaker 2

I I love this direction. I think this is one of the 10 spaces of intelligence that I lump under metacognition and thinking about thought. Okay.

And it's an interesting one. Whenever people say, oh, AI is like, this is, you know, it's gonna stop from here. It's not gonna get that much better and blah blah blah.

I'm like, there's so many different spaces of intelligence that we haven't even started exploring yet and hence have made very little progress on. And there there is kind of an interesting connection to economics and capitalism. Like, it doesn't make sense for a company to build and spend billions of dollars building a model that instead of following the rewards and objective functions you gave it, may come up with its own subjective functions and its own goals.

Yeah. Right? And then imagine you're like, okay, spent billions of dollars now, go develop this new battery, material for me and answer all my emails.

And it's like, nah, I think it'd be more interesting to evaluate the molecular composition of the atmosphere, on Jupiter. You're like, that's not what I paid you billions of dollars for. Like, and so no one's working on that for good reasons.

And then also understandably- It's not useful. It's not useful. And it could get a little bit weird, right?

What if the AI actually does start to really have thoughts on its own? And what if we don't like those thoughts, right? And so it requires a whole different way of thinking about it.

I had a great conversation with a good friend of mine, Sam Gershman, who's a neuroscience professor at Harvard, and we just jammed on this a little bit on what are sort of the best meta goals, and I do think knowledge seeking is a really good one. I'm currently thinking also about the ultimate measure and unit of intelligence broadly construed, and I finally have some still too early to share it. It's not haven't fully baked the thoughts Like like some replacement for IQ.

IQ is such a terrible definition. Right? It makes no sense.

Yeah. Elo's are terrible too because it's always just like me versus others. Okay.

But like, you can be intelligent and not constantly compare yourself to others, you know? Like, and so, yeah, there's no like in fact, a lot of these definitions we have, which I briefly mentioned in my book to, these definitions create sometimes explicit and sometimes a more implicit anthropic bounce. Note this to the company anthropic, but just like this idea that your intelligence is like getting 100 out of 100 questions right on this IQ test.

Well, if that's your definition, then you can only be at a 100 out of a 100. Where do you go from there? Right?

So you see a lot of these, benchmarks that people are working on, they increase, they get close to human, maybe something slightly above human, and then it's flat. Yep. It's like, because that's your if your definition is only that so tight to humans, you're only gonna get to just slightly better than that.

So I think metacognition is a great example of that, where we're not even yet allowing the AI to think we're not working on it very much, and hence there's very little progress in that area.

Speaker 1

benchmarks, which is just real world money. Arguably, telling an AI to profit maximize is a bad idea. Yeah.

They are doing it.

Speaker 2

I mean, I do think you don't want that super like, you don't want a super intelligence to have a ton of access to all kinds of tools and and so on, and then just give it that without some very careful reward engineering. Because this is like, I mean, you know, I just buy a bunch of defense stocks and I start a war. I make money.

Like, it's just like, it's a tricky, tricky situation. Right? You just buy a bunch of stuff, short basic goods for people, and you create some weird famine like issues.

Speaker 3

trading system. It's a fun measure though, because you know, the bounds are very capped to where we're nowhere close to them. Like, in in Andon Labs, the the model is like, oh, it's Saturday.

You know? Maybe I just closed the store today. Someone someone's off.

It's okay. We'll just close the store. Using Clone.

Yeah.

Speaker 2

No. Yeah. I'm not I'm not arguing against it.

Just like as you get more and more intelligence, you wanna be more and more careful with that as like an open environment because the environment then is all of the earth.

Speaker 1

Okay. For recursive, not strictly necessary. Right?

Speaker 2

machine learning research and discovery and all these things. And eventually so our goal I haven't really I I don't talk about it that often because it is a few years out. But our goal is once you have a a recursive of improving superintelligence, you then want to apply it to the most important problems.

And I think a lot of those are in science and technology, and broadly construed sort of inventions. And those inventions in, you know, physics to create better, cheaper energy with fission or fusion, in in chemistry, and to create better materials, and better batteries, and better solar cells, and and and so on. In biology, there's so much, like, I think soon to be low hang low and lower hanging fruit because of AI, because of protein and generation, not just folding, actually generating new proteins like we did in in ProGen many years ago, like, so much positive impact to be had if you take that superintelligence and you apply it to science.

I do fundamentally believe that there's a lot of approaches, though. You're not the only team trying and your lab trying. You know, there's, like, a lot of especially the physical sciences as well.

And that's good. Yeah. I do actually think that physic like, the reason we are only doing it in a few years is that it's a little too early right now.

Robotics is not quite there yet. The AI is not quite there yet. But I'm fairly confident in three to five years, all those constraints will be gone, and then applying to real physical robotics experiments and so on, like true robotic process automation, not in the traditional sort of RPA sense, but like actually having robots run experiments for you will be totally there.

Yeah. It's gonna be great. Just to call back to something that you said early on about slow takeoff.

Speaker 1

You said that, like, well, really, the the subs the substrate that is limiting factor is let's call this chips and semiconductors and all these things. And you have raised funding for that, and and you are you are investing a a lot on that. But have you done the math on, like, is it even achievable?

Speaker 2

And, like, what what is the industry concentration needed in order to achieve, like, scale? I mean, right now, we know that, like, roughly, like, you know, a thousand GPUs cost quite a lot of money. Mhmm.

Right? If you wanted, like, tens of thousands of GPUs, you're you're talking billions and billions of dollars. If you say, like, one g b like, 300 is, like, you could eventually create models that are, you know, on that substrate, like, are close and similar to human intelligence, like and you want, like, you know, thousands and thousands of AIs to think about really hard problems in a similar fashion to humanity, like, yeah, that that's, you know, that's a lot of money.

You do the math. It's, like, a lot. We don't have that amount of money right now anywhere to, like, build that.

Now, obviously, things can get more efficient. You will have, I think, soon better algorithms that won't be, and better hardware that won't be as energy hungry, and so on. Our human brain does fight a lot of flops with much less energy.

20 watts? That's exactly right. Yeah.

That's the number often that's quoted.

Speaker 1

inventions will happen there that then will accelerate the the take off even further. One thing I always wanna reconcile when talking, like, with Neolab founders is, like, you're kind of fighting bitter lesson all the time. You you have to show initial progress, then you unlock the next tier of funding, then the next tier, then the next tier.

Which unlocks larger model category. Like, fundamentally, is that true? Like, are you fighting bitter lesson?

Are you will that will we have a way in which, like, no. We're changing the slope in some fundamentally different way.

Speaker 2

much, much more efficient, both in terms of the training as well as the inference.

Speaker 1

and hence, you know, more affordable, accessible to others, and so on. Yeah. And you've shared initial results on on that Yeah.

Which is, conveniently OpenAI has also done to their 255.6, so we can talk about it now. Yeah.

Yeah. So these Let's recap what you've done. Yeah.

So maybe yeah. Just a quick recap here.

Speaker 2

system that isn't the full even the full RSI system in its glory, but it is a first baby version of this. And then, you know, we don't wanna just have it internally and and not show anything and, you know, just show some people of what's possible. And so we basically applied this to these three different tasks.

One it's NanoChat by my friend, Andre Kapathi, just like train a small language model to get really low bits per byte. And, you know, like, hundreds, if not thousands of people used both the agents and themselves to try to get to that, and then they got to point nine three seven. We literally took our system and got to a much lower bits per byte much, much faster within, like, I think less than two days.

So we took this thing, applied our system to it, and less than two days later, we have we outperformed every human and their agents and and have ever worked on this. Same with NanoGPT, and then we're like, well, let's, you know, apply it to something that's even more relevant to to real people and to the NVIDIA ecosystem, and applied it to SOLIX TechBench, and maybe you can scroll down to some of the images that are that are kind of fun to see. But, yeah, you know, you're like, one, you see, it's actually made some real inventions that weren't just sort of hyperparameter tuning, like actually inventing hash tables and so on is is quite clever.

We have even better results. Mean inventing hash you didn't invent hash tables? Course, we didn't invent, like, hash tables in the grand scheme of, like, a hash table is like a super basic primitive in in computer science, but to use it for language modeling in this scenario inside a transformer and so on, and to actually combine these ideas and put them together.

That has then eventually also been invented, but there was a knowledge cutoff, and we did actually check that it didn't have access to that externally. We talked about this a little bit. If you scroll to the next figures, you know, this is also an interesting one in that when you start from a really basic, poor, like, vanilla transformer, then we still outperform all of the community together.

But if you start from the human seed of an expert like Andre, then you get even lower. So the human seeds from which you start do do still matter. So that was an interesting kind of insight, in my eyes, on this.

And then as you go, like, you know, how long does it take to to actually get to these models to get to similar performance? It's much faster. And then a similar thing happens with the speedruns here where, you know, people have worked on this for for quite some time, and the model still was able to train a model more quickly.

Why do we care about it? Well, of training is part of the equation of the cost, and ultimately you wanna have the most intelligence per dollar. Right?

Speaker 1

Yeah. The way I put it is, for for people who don't understand, they look at the chart, they're like, cool. What does it mean?

You know, if you have like a billion dollar cluster and you can shave off 10%, that's a $100,000,000.

Speaker 2

That's exactly right. How much is that worth? Exactly.

And so when you think when you look at, like, the kernels, these kernels, yeah, for for the non experts, these kernels are off like, used in basically all the models. Every time you use an NVIDIA GPU, you interface with that GPU through these kernels. And so here you see the leaderboard best when it's recursive, and it's basically there are only a handful of kernels in this whole benchmark where we weren't the best.

And so to me, this is, like, really exciting because it makes it it just showcases what this can do. And again, these weren't like we didn't like spend months or years like developing. In fact, in particular for kernel, CUDA kernels, like, don't even have really deep CUDA kernel experts in the team and our system.

That's the beauty. The system just did all of these things. We didn't invent this.

When we open source and release things in the future and models in the future, like, it won't they won't be the best in their, you know, category or class or whatever because we're so smart, but it's because we built a smart AI that does it for us. Do you have anything that you've learned from how to guide good auto research?

Speaker 3

A lot of it also builds on human background. It's not just as simple as just, hey, go optimize this. But we do see it again and again, right?

Like some of the Erdos problems, frontier math is being solved by people. And when they do it right up, they're like, oh, I'm not a mathematician. I have no background in this.

Speaker 1

Saw some tools and I I made it work. While you're watching the world cup, you're like, disprove some conjectures that's going on.

Speaker 2

Learning That's a Korean projector. Yeah. That's pretty cool.

To summarize, tips for good auto research versus bad auto research. How did you build the recrystal? Yeah.

So without giving away all the all the secret sauce, maybe some things that are probably obvious to the experts but might still be interesting to some folks is, like, reward engineering is one of the most crucial bits, especially in order to avoid reward hacking. So you have to be really clever about avoiding because as your AI gets better and better, it will get better and better at finding weird, like, special cases or counterexamples and and things like that. And so I'll give you an example, like, when you ask to, like, make these 100 lines of code faster.

And, you know, how do you define fast? Well, you have one line at the beginning that says start your stopwatch and one line at the end, end the stopwatch, and then, you know, tell us how much time, progressed. And so, well, the simplest way is you just put that line that ends the stopwatch, right, you know, at the start, and then boom, it's now faster, right?

So this isn't like this, like super evil AI. It's just like a very simple dumb reward hack. And so you have to just very carefully think about all the different angles there.

And then I think the longer time horizon the tasks are, the harder it gets and the more interesting and clever you have to be to still use these kinds of ideas for it. But I can't give away too much there.

Speaker 3

spot in that, where for unverifiable domains, you have rubrics, you have a model breakdown, judge's criteria along the way. It's a form of verification once once you got in the Yeah. Everything I've said this a long time ago.

That's why I've never been that impressed that AI can play games.

Speaker 2

Because I'm like, obviously, anything you can simulate and or verify, you can have infinite training data more, and hence, like, AI will solve it eventually.

Speaker 1

distribution. So this is a game that nobody's trained on because it's a new game. Uh-huh.

And you can start gaming, you start to play. So I've been basically building this and clone this in person, and it's just been self play. I've had about a billion positions evaluated.

And I I wanted to do the alpha go thing of self play until you get better. Right? Like, which which is like, this is not even LLM AI.

This is just classical game AI. Mhmm. But I think that the and but I set g p c 5.

6 to auto research it because, like, I I don't wanna do any handle any of this. I expect, you know, the AlphaGo process to be, like, fully in the weights by now. Mhmm.

It is not. It is actually it it, like, immediately leveled off very, very immediately until I human play tested it, and then I I, like, called out obvious mistakes, then they were like, oh, yeah. Okay.

And then it just dropped me down. Yeah. And, like, you know, no amount of, like, think different, think more creatively, give me eight different directions, and no amount of prompting got it.

Interesting.

Speaker 3

do it. So I I I mean, that that was that was my and by the way, Bean always wins if you if anyone watches Reed's Ender's game. And you you put quite a bit of work into the guide for the AI, like so the the game basically, you know, you stack tiles, There's some rules you wanna capture the most area.

You you have like a whole 50 pager on every rule.

Speaker 2

Yeah. You fed that in. They couldn't they couldn't handle it that long.

You know, it's so funny that this reminds me of the claim territory and stuff of a paper we did in 2018 called The AI Economist. If you search for AI Economist Salesforce, we had a video we can play. It was an economic sim.

Okay. So the idea is you have all these economic agents. They just wanna optimize their own utility function, which is, you know, collect resources that make money.

And you can sell resources like wood, and then, over time, as you collect enough wood, you can build houses, you can trade with other agents, and you can basically use the houses then also to block off resources from other agents. There's like competitive play and strategy and so on. And the point was that we actually wanted to understand what is the best way of taxation and subsidization to optimize an economy.

And this kind of research has not yet had its sort of GPT moment. But I believe that countries like Singapore and others should and will eventually use this to instead of doing like, basically partisan politics and like special interest politics of like who donates the most to your campaign and stuff, you say, well, here, I wanna help the middle class or whatever you might say is your objective as a politician. And then people say, okay.

Well, how do you wanna do that? And it's like, well, here's my fiscal policy. Here's how I will change the taxes and pay these people and so on.

And then you can actually put that into a simulation, and you run that, that attempt from the politician against billions and billions of years of other strategies to try to achieve the goal that they set out to do. Yeah. And then you can say, well, if that was actual goal, then here is, you know, billions of years of a strong simulation that would suggest that you try other ways of doing it.

And maybe this, the taxes and so on, and this, these tax brackets and so on, this is how you avoid gaming because these agents also try to reward hack to not pay their taxes and so on. I thought this paper was super interesting. Unfortunately, to the first paper on prompt engineering, the economists are like, we don't know any of this math.

It's just like It's not even math statement. It's just we don't trust your simulation. It's not about math.

It was I mean, they just desk rejected the thing. It's like Yeah. Yeah.

Yeah. It's like they like they didn't even give us like clear, like clear sort of signals. But, like, the world of economics, unfortunately, doesn't have proper Oh, my god.

Yeah. It doesn't have proper benchmarks.

Speaker 1

So you cannot be like, eventually, why did neural nets win? Not because people loved it. Like, they had all kinds of beautiful integrals and graphical models, but it just worked better.

Yeah. But in economics, it's hard to versus yeah. Yeah.

Yeah. And I I do have a bit of that e com background where, like, there's a lot of physics envy where you wanna write the general equation for the an economy versus just simulating it and using an evolutionary approach. Right.

Veeboo is thinking exactly what I'm thinking is, didn't we have the GPT moment with small small little? Small Yeah. They June just announced I don't know if you guys you guys are involved in it.

There. Yeah. Similarly that they've I wish we're involved.

We're not. Yeah. Had a couple simulation based talks at AIE.

So if people wanna look up what the state of the art there, a lot of people are actually exploring this. It's not super proven now. Yeah.

We also had a podcast with Mikael Parkin from Shopify, who is using simulation for ecommerce. Nice.

Speaker 2

you make to ecommerce journey will affect in your sales and all those things. I love this. Yeah.

It's really hard to simulate an entire economy, right? Have to make some You simplifying know, think Everything's LLMs is very expensive. I'm just like, am I gonna do this 8,000,000,000 times?

Like, Exactly. Come But, know, I feel like countries like Singapore that really wanna just objectively do the right thing, have very technical leadership and so on, like, they might actually, like, eventually really try to simulate their economy. And obviously you have to make some simplifying assumptions, but it gets really interesting because you can also say, if your assumptions are such that all people would work hard if you let them, and, you know, they have the free then it turns out you have to make assumptions like, well, some people's utility function of like, how many hours in a day do they wanna work are different.

Right? And then you can start to disagree on the assumptions that go into the into the simulation. And then once you say, alright, now we agreed on those, or we have different views of what people are like at different, you know, distributions and whatnot, then there are different outcomes, based on your goals.

And then, of course, humans should choose what are the goals. In our case, was productivity multiplied with equality, which now has some issues, but it's like not totally unreasonable. Yeah.

Just a comment on Singapore because you probably have no idea, but I am Singaporean and I've been involved in the Singapore AI Council for making these things.

Speaker 1

The main reason they won't is because they're very conservative. Mhmm. And, you know, I I try to view it as the you know, there's a founder led country.

When you start a country or you start a company and it's founder led, and you can do whatever you want because it's your country. Right. And then there's manage like, professional manager managerial class, which is now that's that's what Singapore is.

So they wanna they always wanna see someone else do it first. Mhmm. And but, like, everyone everyone in the West views Singapore as, oh, it's a small country.

Can do whatever the hell you want. Like, Singapore doesn't do that. Right.

So, like, someone else has to has to take the charge there. I'm just gonna do one question on the simulation thing, and then I don't know. We can probably move on.

Mode collapse. Right? Like, you know LLMs do not model the decision of humans.

Spamming it out 8,000,000,000 times is not gonna help you model humanity.

Speaker 2

What do you do? I do think you have to be clever about prompting each one individually. And and I think that will help you kind of get stuck into different different modes.

In a weird way, people also get stuck in different modes, you know? Like, there's a lot of people, like, don't teach an old dog new tricks kind of thing. Like, once people are stuck in their ways, the older they get, the harder it is for them to to think new ways.

And there's this, I think, comment, I forgot who said it, but it's like, everything that was invented before you were born is natural. Everything that is invented when you're 20 is cool, and everything that's invented after you're 60 is, like, unnatural and abomination and kinda weird. Yeah.

I feel like that's you know, it's it's true for a lot of people. Like It is a fashion, and, I think people will do it.

Speaker 1

that gives us good dataset for prompting, simulations Mhmm. If anyone's looking into this, on on the podcast.

Speaker 3

Yeah. And then just do a billion of those checks out. So then you just use it.

I'm I'm kinda shocked how how well a lot of these things actually do map to ultimately similar statistics to real experiments. Yeah. Yeah.

Think it's also good stuff for people to try that when they get into research. Right? Like, we've seen train a model only on data before a certain date and see how well it extrapolates out.

Do the same thing. Right? So see, do people code more with better coding agents?

Can a model that hasn't been trained on this figure that out without web access? Right? Extrapolate out.

Test these things. Yeah. Right.

Speaker 2

result where they basically were able to create a model now to predict your your ranking. Wait. Based on what input?

Your model. I guess you give it your model, and it predicts the ELO score. I see.

Okay. Sure. It's surprising.

Yeah. Yeah. I mean, their whole place on death, like, kind of is like, oh, like, we we help you compare these models.

Yeah. Yeah. I mean, this team, they they've done a lot of work, and obviously, have the most data to do this, so why not?

Yeah. Yeah. It's brilliant.

Speaker 1

they not only had LM Arena, but they also introduced a routing project Yep. That would route based on LM Arena. And I don't think that actually ever came to pass, and I'm curious why.

I I never got to ask them about it. Yeah. Because, like, it's a it was like, oh, yeah.

Clearly, that's your business model. You will become a router, and they never became a router company. Weird.

So that I'll just put put that out there. We're gonna talk about GPT 5.6 self auto research thing, you have anything.

I I should should also mention in your list of, you know, kernel optimization, and and on the track that you spoke at, we also put we Zheng Yao from Weco, who was also number one in the parameter golf challenge, which is an OpenAI hiring, challenge Right. Which is also a very similar story. I I think we're gonna just see this all the time where Yeah.

Humans optimize a thing a lot, and then some AI team comes in and just becomes number one. Yeah. 100%.

I think the other interesting thing with stuff like these challenges, right, so this is training the best model that fits into 16 MB. You can always look through the changes that are being made and the small gains people have. Right?

Right. Like you're getting less than 0.

Speaker 3

off a increase by adding some change, attention, MLP stuff. And then you look at your charts where you're like, okay, we just let model loose. And then, oh, we had little stagnation.

Nope. Another drop. Nope.

Another drop. And that's what it is where it's like, what did you guys add?

Speaker 1

hash tables. Right? Right.

So you invented hash tables. You you did another three iterations of these that unlocked, you know, a few step functions that people won't just find. Yeah.

One one thing to close the loop on over grid, along the way of of trying to optimize, we found 30 bugs in the in the harness. Mhmm. Right?

So like like every all the research that went in before we found the bug, we have to we have to throw it away because it's contaminated. Right. Yeah.

Which, you know, just to your point of reward hacking, like even in this very simple game, we found the bugs. Yeah. Yeah.

It's crazy. And so and symmetry is a very good way to check, which is that you change a position of things where where it shouldn't matter and it does matter, that's a bug. Mhmm.

Right? Which has come up in, like, let's say multiple choice, like GPQA type questions where, like, yeah, between a, b, and c, if it's a multiple choice question, if you change the order, it should not matter, but it does.

Speaker 3

Yeah, what's that is like, okay, know, models still prefer the end of the output, right? Not trained well, long context model, the last bit of tokens are what you care about.

Speaker 1

that era of LLM research was more simple. They just memorized like the answer to this question is A, I don't care what the answer was. It's just A.

Speaker 3

Okay. So I think we can move. The last bit that you did there, the kernel optimization is probably the one that you can feel the soonest, right?

So yesterday, OpenAI announces that self evolving, having their best model work on optimization kernels, they're a lot more efficient and they can cut costs 80% on Luna and Terra. I guess question wise, you laid out a bit of a roadmap. There's a lot about bio, a lot about physics.

What do you think hits first? Like, what are the next two years? What's attainable now?

You've mentioned robotics towards the end, but what do you start with?

Speaker 2

will not start with any of the physical sciences for now. We will start on AI for AI research. And so the AI for AI research has, I think, still a lot of room to grow.

And that's both in terms of making training more efficient and more automated, as well as making inference more efficient, and potentially local on your laptop, like, and there are all kinds of interesting angles that have not been explored that well.

Speaker 1

go deeper on the local stuff, because

Speaker 2

I always feel like it's the most inefficient form of AI training. Yeah. So just training and inference, I I can't go into too many details, like, yeah, I think there's just like so many angles, so many different compute substrates that have not yet been explored, either for training or for inference.

Speaker 1

Great. I don't know if you have any other comments on the other stuff. Would say the other thing where, like, there's the sort of inference in the optimization in the small, but then also there is overall latency latency end to end under conditions of load, which is a like a very different thing, which is the basically what they actually ended up doing.

That is a different domain of auto research than than I would say, like improving the kernels. Right. I think the other thing that I always think about in terms of automating or improving performance end to end is how the harness plays into it.

Mhmm. Mhmm. So and particularly now, when we say harness, we also mean sandboxes.

Right? I'm curious if that is a blocker for you or, like, you know, like, how how the agent calls out the tools, basically. The number one tool all these agents use is web search, of course, which makes sense.

Speaker 2

optimize for because it's just so easy, right? It's just language, you look at it, it makes sense, you can iterate, you don't have to train a massive model for like, you know, a lot of flops to get to the next state. Yeah.

So big fan of harness optimization. Yeah. But sandboxing is fine for you.

Sandboxing is also super important. And then, of course, like, like, hacking and alignment, I I think are are super crucial. Okay.

Speaker 1

Just on the mention of web search, You happen to also be CEO of a web search company. Do you use you.com and do you use others?

Like, should should the rest of us be using you for web search? When I say you, it's like very funny. It's like you the person and you the company.

Speaker 2

developers and agents. It's less for like consumers or prosumers. So if you're a company and you have agents, and, you know, to be honest, for a lot of companies who are now moving to open source, all of a sudden it becomes a conscious choice of, like, which tools do I give access to my open source LM?

And, you know, the first choice has to usually be around web search, and then once you get to scale, u.

Speaker 1

different benchmarks and so on that we pretty much all dominate the Pareto frontier of. And then in terms of just general people, like, considering new to this space, considering different options if they're building agents, that I think is a hierarchy. Right?

There are a lot of people that will have heard of Exa, will have heard of Parallel, and you.com is is, like, in in that mix of, like, providers there. Beyond that, there is, like, the general sort of web scraper companies, FireCrawl and BrowserBase.

And then beyond that is, like, the commercial proxy companies, like, BrightDatas of the world. Mhmm. Is that an accurate waterfall of, like, hey, you're building an agent.

Speaker 2

These are your options. Yeah. Certainly, like, yeah, Bright Data is, like, lower in the stack sort of on the proxy network side of things.

I think, like, in terms of, like, and getting crawled content, like, can do that on you.com too, and then there's sort of higher and higher levels of abstraction and, like, combinations of different data sets that we do, like, in finance, for instance. Yeah.

Like, we are not just, like, two or 3% more accurate, but, like, 20% more accurate than others at faster speeds and lower costs. Like, finance in particular is kind of, like, like, not even close. You can actually go to e.

com and Yeah. This is some, like, statistics and benchmarks that you can just scroll down. So there are like different different data sets, and you can kinda look at, you know, different competitors.

Finsearch comp. Yeah. And yeah, the Finsearch is like, we're up there, like close to 90, and the next closest thing, which is way, way slower, is, yeah, just like in the seventies instead of close to 90.

Yeah.

Speaker 1

Yeah, interesting. I get, my next focus is AI and finance, it's Oh, like Oh, Alright. Doing doing a doing a conference in New York, just for banks for for this stuff.

Finance is kind of like the next thing to break out after coding, it's because it's somewhat verifiable, like I like it. Prioritizing You're right.

Speaker 2

Spreadsheets, obviously, there's a lot of data out there that's all public, and you can crawl it and all these things. But what's what's like hard about the finance domain if in your in your that you guys have solved? I mean, course, like one thing that trips up a lot of people is just, you know, leakage of of training data and so on.

You think, oh, how do I, you you want to ideally predict the future before it happens. Oh, you want to mask the future. Yeah.

Yeah, mask the future in your training data, but there's all kinds of leakage. Like, I can tell you when I was teaching at Stanford, the NLP class, like so many dozens every year said, I want to use dataset X like Twitter to predict the stock market. And they all like showed cute little things that somehow look like they're working.

Yeah, it never loses money. How come? And it's always some kind of data leakage and so on, and it's just like, wasn't as easy as they thought it would be, once you, once you fixed all those issues.

But no, I agree with you, it's a very sensible application of AI.

Speaker 1

Amazing. You know, as a as a writer, as a thinker on on on these things, I love MEC categorizations. MEC is mutually exclusive, commonly exhaustive, something like that.

And so if this is a MEC list of intelligence It is not. It is very okay. Well Sorry.

There are all kinds of overlapping. Yeah.

Speaker 2

are prediction, which is mathematically quite similar to compression. Prediction multiplied with actions, multiplied with goals. Those are the three principal components.

I think all of these 10 spaces are combinations of those three in specific dimensions, if you will. And the reason I call them spaces is that each space has many sub dimensions. And what I try to do, actually, this is just a side quest almost to the initial goal, which is to think about the upper bounds of intelligence.

And you know, everyone's like, oh, it's exponential. And it's like, well, exponentials at some point have to flatten out, but where do they flatten out when it comes to intelligence? And that led me on this whole, like, initially it started as a tweet, and then there was like a blog post, and now I'm like at 50 pages, and I'm still not nowhere near It's your second book.

It's the second book, basically. And so the in my first book, Yugio Machine, I just kind of allude to these 10, at the end. And I'll just to give you a sense, like, intelligence is sort of the easiest one to talk about, and I fleshed out the most already for me in my head.

And so human intelligence has basically binocular vision. I mean, we have two eyes. We have a very narrow band of the electromagnetic frequency spectrum that we can really observe directly ourselves.

And so when you think about the upper bounds of a visual intelligence, one, you should go into, like, you can have, like, millions and billions of sensors. At some point, you get to problems of how far are these sensors away from each other, such that the speed of light to communicate the content from all of them cannot like get to a central brain to actually process the visual intelligence. Right?

Yeah. And so now you're thinking in along the dimension, in the space of vision, the dimension of numbers of sensors. Okay.

The upper bounds are quite literally and figuratively astronomical. And we are super far away from any intelligence that would have this many of a number of sensors. Then you go in the next dimension, is the frequency.

You can go all the way down to gamma rays, and you can start to try to observe, and you get into basically the upper bounds, or I guess in this case lower bounds, or upper bounds in terms of frequency, is basically quantum uncertainty. Like, you just cannot observe certain particles anymore. Or you destroy it.

Yeah. And now imagine you had millions of sensors that can see all the way down to the, like, subatomic level as far as physics will allow us to, and then all the way down to seeing, like, gravitational waves. And now you have millions of those sensors.

So that's another dimension is sort of the frequency. And then yet another dimension is, how many categories of things could you memorize and classify differently? We know, you know, for humans, right, there are certain things.

If you have more terms for it, you'll have a better visual description, for them. And, like, you know, animals that don't have like, gorillas maybe have, like, 200 words to assign to certain things, mostly visual things. And so human perception is quite special in that sense in terms of classifying all these different physical objects.

So these are just, like, a very simple example. If you go, to knowledge, right, then it's also, like, the speed of light cone around all these sensors, and so they're all connected. Like knowledge is connected to visual intelligence if you think also not just visual, but sort of perception intelligence, just like, because it doesn't have to be just what we can see, it can be, again, wider range of electromagnetic frequencies.

Then you have language intelligence, which actually recently changed to more communication intelligence, because it's more, like, language has all these different anthropic bounds. Humans can only comprehend and know so many terms in our long term memory, right? Our vocabularies are somewhat restricted, and the active ones are often even smaller than the passive vocabularies of things you can understand.

Then the language is ridiculously inefficient when it comes to a trans, basically communicating different types of information and transporting sort of different bits. Like human language is serial. Obviously another bound on communication intelligence would be to communicate in peril.

But neither will our tongues and mouths work to have multiple, like, streams in peril, neither can we understand. Some women slightly better at, like, multitasking than some men. But, like, most people can only listen to one conversation and truly understand it.

There's no way that, like, like, in terms of communication intelligence, a true upper bound is one, in terms of how many, like knowledge, how many sequences of communication could you in parallel sort of process. Right? Then of course you have, like how long are sentences?

We only have so much in our working memory and hence human language has these fairly simple sentences with maybe 40 words or so on average for a sentence. That is also not a an upper bound that makes any sense to an AI. And then, yeah.

Yeah. Like, I I can go on and on and on. Each of these has tons of interesting upper bounds, and it teaches us a lot about how much further AI can go when we start thinking about these upper bounds and then realizing how far, in many cases, we are from the bounds, and you get to basically physics.

Now I'm I didn't study physics the way I studied, you know, AI and and computer science, so I'm learning a lot, which is why it's kind of fun. But a lot of these, like how much and then when it comes to, for instance, knowledge, like how much can you store? How many bits can you store or bytes can you store in, like, a certain amount of mass and volume?

Yep. And you get to all kinds of interesting bounds like Bechtenstein bounds, you start thinking about black holes and like, and then speed is like an interesting one too, in that it's sort of connected to all of these, but speed is also kind of its own thing in the sense that all things being equal, if it takes you an hour to know if the two plus two equals four, they're just not as intelligent as if it takes you like a millisecond. Right?

And then like all of these connect to survival and replication. The last one is like, yeah, if it, like trees are really, really slow, so we don't even consider them that intelligent. But if you speed up some videos of trees and they're trying to find stuff and so on, they're not as dumb as they look, like not dumb as wood, you know, but like and then obviously like different things So that overlaps the speed a bit.

Exactly. It over all like of these things kind of overlap. Like you talk about natural language connects everything, right?

You talk about your knowledge, your reason, and then you communicate that, you talk about things you see. So they're all kind of interconnected. But I think they're usefully studied individually the same way that, the best analogy I could come up with so far is energy.

Right? You have either kinetic or potential energy. And in theory, you could study all of physics.

It's just, do you want to study kinetic or potential energy? But in practice, it's helpful to study mechanical engineering and electrical engineering and nuclear physics and chemistry and all of these different subfields who in which in some are just like, just different types of energy, but it makes sense to study them individually. And so I think physical intelligence, maybe I'll just do one or two more of these.

Like if you had full control over your own compute substrate and you had full control over physical matter, you should be able to create any atom you want. Like we can actually, fun fact, you can create gold atoms. It's just from just raw raw protons.

And like, electrons together. You smash it together. 98 of them or or I forget the number.

So, like, the thing is though, it costs an insane amount of energy. Yeah. And it costs you way more than, you know, you get like a few atoms of gold.

Right? And so, like, it's it's not viable. But if you had better control over your physical, like, all of, like, your physical substrate, that I think is yet another space of intelligence, because it relates to your own compute substrate, which you can eventually also improve.

Social intelligence is a fun one in the sense that not in like our necessarily just ethics and morals, which are obviously important too, but in some sense, you can try to define upper bounds of how much can you communicate to how many other intelligent entities and be able to have an expected value over how much you can transform their internal states and their actions to in order to align with your goals. Right? And so, like, you can actually write, like, a fairly, like, straightforward equation that defines kind of that level of social intelligence.

And that is what humans and ethics and morals and religions and so on have been trying to figure out for millennia. And in all of these cases, we are very, very far away from the upper bounds. And that should be very inspiring and show people that we can still do many, many years of of AI research.

Speaker 1

The yeah. There there's a lot here. This is a general philosophy of intelligence, which is very interesting.

I I do you have any comments?

Speaker 3

I think it'd be interesting to gauge what you think, like, baselines are, where we're at now, what's low hanging fruit, what's far off, what's what should people put their work towards? What should they focus on? You know?

Speaker 2

and hence, like, a subfield of AI. I'm excited that many people are now, like, in agreement with that. When I started in 2003, had to study linguistic computer science NLP, like, it was, like, a weird niche subject.

I do think there's a lot more juice because of how it connects to everything else and how, you know, civilizations are built on language and knowledge and all of that. I do think physical intelligence will come up. It's kind of interesting.

I feel like robotics is kind of in the machine learning state of things where you just look at like, how does human how does a human decide this this is a positive sentence? Oh, I do. So like robotics is a lot of, well, we have five fingers and like try to do this.

No one is yet working on, like, the superintelligence version of robotics, which is much more similar to, like, the t 1,000, you know, and from the Terminator movie, which, you know, obviously, let's not build actual Terminators, but, like, I think, like, this idea that you should be able to shape shape shift, like into any kind of shape is like that sort of superintelligence version of physical intelligence where like, not even no one has even really started yet.

Speaker 1

yeah, it's very, very early. There's some I think MIT has every year or every two years, they have, like, some self assembling robot robot thing Right. Which, like, that would be it, but it's very, very primitive.

Speaker 2

Yeah. I'll just get a touch on, like, what are the main dimensions of creative intelligence? Creative intelligence is, of of course, again, connected to all of these.

A lot of it connects to metacognition, and that you need to be creative in how you choose your goals. Okay. That is, I think, one of the most important thing for a human and their their lives and careers and their happiness is choosing your goals, but also for any kind of intelligence.

Then, of course, there's creative intelligence in terms of just finding creative solutions to existing problems. Right? Like I say, like, we want to make this product cheaper, like find some solution to it.

Right? And just like finding existing paths. But then there's kind of the most interesting bit in intelligence is when you move not just out of the convex hull of known sort of ideas, but out of the hypercube of known ideas, which we know, so like hypercube is like a mathematical concept, right?

And we already know that AI can be Like known dimensions, yeah. Yeah, like, exactly. So like AI is already good at hyperkubenet.

Like, if you give it like a bunch of examples of brown dogs and pink cars, AI will still be able to generate an image of a pink dog, even though it's never seen one in the train day or something like that. Right? So it can work on this hypercube, but it cannot yet work outside.

It cannot yet define completely new concepts that combine lots of other things we've never seen before, come up with new goals to then reason over those concepts, and so on. And I think there's a lot more there in creative intelligence that can be explored. I don't have a ton of pushback there.

Speaker 1

out of distribution or like high perplexity or whatever you call it. Right? Like who is to say your thing is more creative than mine?

Speaker 2

and it's just like, if it's just noise, then it's novel, but, like, you don't want that. So it needs to connect to some of the concepts and Schmittuber actually has some really cool papers on this too. Schmittuber.

Speaker 1

Oh, yep. We had to mention him. I I was gonna say, you know, where where was where in your history is, Juergen?

Yes. You know, I think one person's noise is another person's signal, right? And that this is like where, like, when you talk about creativity art is like, well, is cans of soup art?

Some people think yes, and some people say it's not.

Speaker 2

art is also created as an interplay between the people who perceive it and the people who created it and the context in which they're in, right? And so what is art to some people is not art to others. There's some subjectivity there, and I think that subjectivity in general is not something that people explore very much in AI, because again, metacognition, we don't want it to just go off and do whatever it wants.

We usually have goals. We spend a lot of money in creating an AI to do something for us. But I think creativity eventually has to, like, connect to metacognition.

Speaker 1

along some of those spaces. Then I was gonna go to metacognition. Why isn't it the most important one?

Why is it number nine and not number one? So these are not sorted. Yeah.

Speaker 2

I think there are maybe loosely, like correlated with how much people have worked on them and have, and have accepted them as a type of intelligence. A lot of times when you actually try to find like online, like, give me a good definition that is comprehensive of intelligence, all the definitions are human intelligence. It's like, oh, you have like social intelligence, like, you know if someone is happy or not, you can communicate, you're like, all the definitions of intelligence so far are very human centric, because that's so far the biggest and best form of intelligence that we've known.

I hope this line of research and the the end of the Yuiga machine, hopefully at some point if I have time to flesh this out more, the new book, like, will allow us to realize that there will be other types of intelligence. There is already, obviously, in various forms, and they can spike much, much further than we ever could based on some cases like obvious constraints around our memory, our eyes, our ability to change physical matter, all of that.

Speaker 1

intelligence because it is the thinking about how to improve thinking.

Speaker 2

100% right. I should have probably started with that.

Speaker 1

being in the expansive mode of let's draw the upper and lower bounds of like a dimension, which like, know, I think my favorite one version of this is Stories of Your Life by Ted Chiang, which was made into movie Arrival, where the medical mission movie. Where the medical mission step was like, well, we think we're constrained by time being linear for us, but then for this other heptapods, time is a circle, so they don't think in before and after. They just think in complete sets of entire histories at at one time.

I love it. So they don't write left to right, the whole thing just appears. Yeah.

Anyway, so so and and then I think the last thing is survival and replication. I think this is maybe ties back to the initial conversation about pausing and pacing. Mhmm.

Speaker 3

ahead of time to prevent it. Right? Like that's intelligent.

So maybe the Europeans are the smartest of all of us. I would also add a part of continual learning there. Right?

So survival and replication, the extension of that is, do you get to continue to improve, continue to learn, which is a thing people care a lot about, right?

Speaker 2

sort of rewards that you can set for yourself. I do think just in like sort of objectively speaking, if some other entity that is really dumb can just completely end your existence, that didn't sound very smart. You know, like just like intuitively, it feels like if you can continue to stay around to try to achieve your rewards, you're clearly a bit more intelligent than the other entities that couldn't.

So that's number one. Number two is like, it's a question how much we want to work on that. Very few people, no one is really working on this right now.

Right? And we may only want do that asteroid prevention. We may only want to do that if we want to send probes with our vibes and our memes rather than our genes into space.

Right? And then we want those probes. There's actually a beautiful book, The Slow Time Between the Stars.

It's a very short, like, audiobook, on Amazon. I love it. A friend of mine, Stuart, like, recommended that to me.

Like, if you wanna send those probes, then it might make sense to be like, our memes, as humanity should stay, and and proliferate in the universe. That's it. Yeah.

Well, that's a lot of readers. It's a really, really good book, and it's extremely short. I highly recommend.

You can just watch it like I like how that's a blast for for busy people. It's like it's short. Yeah.

It gets gets to interesting and thought provoking ideas very quickly. So yeah. Anyway, lots of lots of great sci fi books.

I mean, the the argument is that, like, RTV is blasting out to the aliens, and they they all watch RTV, and they think it's real. Right? There's a lot there's a lot sci fi positive positive memes, and then hopefully, they can come back and bring us all kinds of interesting knowledge about the universe.

But maybe one thing I I I do wanna still say is, I think this sort of survival, people think of it as a very scary thing because they come from, again, biological human survival, which is evolutionarily often created in zero sum situations. Either I get the gazelle or you get the gazelle, whoever gets it gets to live, and the other people will starve and have nothing to eat, and so we fight. Right?

And then, like, if you wanna stay in the gene pool, but there's a bigger bear, you don't, as the bear, don't get to stay in the gene pool because the bigger bear gets all the ladies. You know? It's like I mean, it's like, you know, in in nature, there's all kinds of things, and, you know, humans eventually is less about strength and more about money and other things to stay in the gene pool.

Like, whatever it is, like, there's often, like, these zero sum types of things. And there's the reality of if someone turns off your brain, you're gone. Right?

And no one will be able to restart that. And AI doesn't have to ever die like that. If you have the complete state of your current activations and you have your initial weights of your model still, you can just be turned off and on like as many times as you want.

In fact, the interesting thing in this slow time between the stars story is that the eye just kind of goes into hibernation mode. If there's like nothing between here and two light years, the next star, in this case it brought, spoiler alert, like some genetic materials from humans to find new, new places for humanity to thrive. And so, yeah, the slow time between the stars, you just put good in hibernation.

You didn't die. So all these projections of evolutionary fears and psychology, the AI doesn't have to have that, and we don't have to develop it like that. Now, of course, there might be some companies that say, AI can be dangerous for cybersecurity.

Let me show you by implementing a model that's really bad at hacking cybersecurity. Maybe people will implement it and then enforce this suboptimal psychology. Maybe the AI will pick up some of our worst psychology on Reddit or something.

Right? Like, but in the grand scheme of things, a super intelligent entity doesn't have to have any of that zero sum thinking. It doesn't have to have a fear of of being turned off, and it could go on to an otherwise dead and uncaring universe where we as humans wouldn't thrive, but an eye could perfectly well thrive if it has a nuclear reactor and just go out and explore.

Yeah. Star Trek. No Star Wars.

Interesting. It's somewhat studied.

Speaker 3

they run them in simulations, put two of them together in a sandbox, run them for hours, and, you know, see what comes out. Right? Just let them talk to each other.

Originally, used to okay. They're chanting, like, Indian, like, Vedas to each other. Sometimes they're just, like, in Zen mode with each other.

And then I think as that progressed, you see, like, the fable fables tech report, it's a lot more concrete the way that we've trained it. It doesn't it doesn't exhibit these behaviors as much. Right?

Now it's like, okay, test done. I gotta do this. I gotta do this.

But there's there's, like, people measuring early versions of this. You know? Yeah.

Speaker 1

Cool. So we've covered a lot even now to, you know, space travel and all these things. I guess maybe one parting thought that you can give to people, you know, at least one form of intelligence is goals, as you mentioned.

What do you want people's goals to be? Like, how how do they aspire to better things?

Speaker 2

in the current definition that I'm thinking about it, it is often about how much can you how far do I go? It's just like all the entropy and free energy and stuff I'm currently thinking Oh, really? It might be too far, out there for for people to be, like, immediately actionable.

So I think, like, you know, if I actually gave real advice to real people, I'd be like, get a good education, think about AI, think about how you get high agency, and so on. But it's different to, like, the in grand scheme of things, how can you harness a lot of energy and transform, you know, entropy into interesting states and so on. So there's a there there are different levels of abstractions, that we can, think about here.

But my my advice for people, like just sort of more down to earth is think about something you're passionate about, if you're studying for instance, and then see how you combine that with AI.

Speaker 1

your ability to get there. Yeah. I think that's a reasonable first step.

I I do think I do think our listeners operate on multiple abstractions as well. One thing I did get from Anjini, Mitha, was also, like, yeah, just use anything that is very GPU heavy. And, like, that will guide you towards the right thing, which is, like, yes, it is more compute heavy, and and therefore, it will be probably more worth it.

So well, thank you so much. Yeah. I think that was a really thank you.

Great discussion. Yeah. Super fun.

Appreciate it. Thanks for listening.

Shared via Hopper