Information Theory for Language Models: Jack Morris

Latent Space: The AI Engineer Podcast
2 July 2025 1h 18m
0:00 --:--
Episode Description
Our last AI PhD grad student feature was Shunyu Yao, who happened to focus on Language Agents for his thesis and immediately went to work on them for OpenAI. Our pick this year is Jack Morris, who bucks the “hot” trends by -not- working on agents, benchmarks, or VS Code forks, but is rather known for his work on the information theoretic understanding of LLMs, starting from embedding models and latent space representations (always close to our heart).Jack is an unusual combination of doing under

Summary

This episode features Jack Morris, a Cornell Tech PhD student, discussing his information-theoretic approach to understanding large language models. He shares insights from his research on extracting text from embeddings, the universal geometry of embedding spaces, and measuring language model capacity, emphasizing the critical role of data in AI paradigm shifts. The conversation also touches on the evolving landscape for AI grad students and the practical implications of his work for model privacy and efficiency.

Chapters

AI Grad Student JourneyJack Morris reflects on his journey into AI research, from early interest in BERT to the dramatic shifts in the field since ChatGPT's release.
Information Theory for LLMsJack introduces his 'new type of information theory' focusing on 'usable information under computational constraints' (v-information), contrasting it with Shannon's theory.
Embedding Inversion & PrivacyHe details his work on reverse engineering text from embeddings, demonstrating high recovery rates and its implications for vector database privacy and debiasing models.
Universal Geometry of EmbeddingsJack explains the 'Platonic Representation Hypothesis' and how his research uses CycleGAN-like methods to align different model embeddings, suggesting a shared underlying semantic space.
Language Model CapacityThe discussion shifts to measuring how much information language models can store in their weights, revealing a current capacity of 3.6-3.9 bits per parameter and the challenge of improving this.
Approximating Training DataJack describes a method to infer a model's training data by analyzing the difference between its base and fine-tuned model weights, with implications for understanding proprietary datasets.
No New Ideas, Only New DatasetsJack presents his thesis, inspired by Thomas Kuhn, that major AI paradigm shifts are fundamentally driven by the introduction of new, scaled, or specialized datasets.
Future of AI ResearchThe episode concludes with a meta-discussion on the next potential paradigm shift in AI, the challenges of predicting future innovations, and advice for aspiring researchers.

Topics

AI research careerLarge language modelsInformation theoryEmbedding modelsLatent space representationsModel privacyData compressionModel capacityTraining data extractionAI paradigm shiftsDistributed trainingGPU infrastructureMechanistic interpretabilityOn-device AI

People

Swyx (host) Jack Morris (guest) Sascha Rosh (mentioned) Shen Yue (mentioned) Chris Lattner (mentioned) Nicholas Carlini (mentioned) Andre (Karpathy) (mentioned) Noam Brown (mentioned) Emmanuel (mentioned) Oscar (mentioned) Thomas Kuhn (mentioned)
Key Concepts (14)
AI Grad Student Market — The highly competitive and rapidly changing landscape for AI PhD students, influenced by the insane market for AI talent and shifts in power dynamics from academia to companies.
Paradigm Shift (AI) — A significant, transformative innovation in AI, often followed by a period of rapid small innovations and reapplication of techniques, as described by Thomas Kuhn's 'The Structure of Scientific Revolutions'.
Emergence in LLMs — The phenomenon where models, particularly at larger scales (e.g., 8 billion parameters), suddenly exhibit new capabilities like knowing facts about the world, which were absent in smaller models (e.g., 100 million parameters).
HPC Training for Grad Students — The lack of formal training in High-Performance Computing (HPC) for grad students, requiring them to learn distributed training techniques (like multi-node FSTP) independently from online resources and communities.
V-Information — A theoretical framework for 'usable information under computational constraints,' proposing that information should be measured by how much is extractable given a certain computational power, rather than just raw bits.
Deep Learning as Compression — The idea that deep learning models act as compressors, taking a large dataset and compressing it into a smaller model that can generalize and retain acceptable information loss.
Embedding Inversion — The process of reverse engineering the original text from its numerical embedding vector, demonstrating that embeddings can reveal a significant portion of their source text.
Debiasing Embeddings — A technique to remove latent features (e.g., gender information) from embeddings, allowing for the creation of more neutral representations, which can be verified by inverting the debiased embeddings back to text.
Platonic Representation Hypothesis — The theory that as AI models scale in data and size, they converge to learn similar underlying representations of the world, implying a universal semantic structure.
Model Capacity (Memorization Rate) — A measure of how much information a language model can store, quantified by its rate of memorization when trained on random data, approximated at 3.6-3.9 bits per parameter for GPT-like architectures.
Cognitive Core — Andre Karpathy's concept of a 'dumbest possible model' that knows nothing but is smart enough for tool use and reasoning, enabling efficient on-device AI.
CycleGAN for Embedding Alignment — A repurposing of the CycleGAN computer vision technique to map between different embedding spaces (e.g., BERT and GPT embeddings) without explicit paired data, revealing a shared 'universal geometry' between models.
Approximating Training Data from Weights — A method to infer a model's training data by analyzing the difference between its base and fine-tuned model weights, useful for understanding proprietary datasets.
No New Ideas in AI, Only New Datasets — Jack Morris's thesis that major AI paradigm shifts (like deep neural networks, transformers, instruction tuning, reasoning models) are fundamentally driven by the introduction of new, scaled, or specialized datasets.
References (45)
AlphaGo project
BERT project
GPT-2 project
GPT-1 project
Google AI residency program project
GPT-3 project
InstruqtGPT project
ChatGPT product
o one project
GPU mode Discord tool
Fast AI company
DeepSpeed tool
CUDA tool
Mojo tool
LLVM project
Swift project
VLLM tool
SGLANG tool
CDE (Contextual Document Embeddings) paper
A theory of usable information under computational constraints paper
Text Embeddings Reveal Almost as Much as Text paper
Harnessing the Universal Geometry of Embeddings paper
The Platonic Representation Hypothesis paper
Claude project
GPT-4 project
MS MARCO dataset
Gemma 3n project
Lama 4 project
GPT 4.1 project
CycleGAN project
GTR project
GTE project
Mistral project
Hugging Face company
How much can language models memorize paper
Approximating language model training data from weights paper
DeepSeek project
There are no new ideas in AI, only new datasets by Jack Morris article
The Structure of Scientific Revolutions by Thomas Kuhn book
AlexNet project
ImageNet dataset
Attention is All You Need paper
GPT project
Mamba project
Muon project
Transcript (80 segments)
Speaker 1

Hello. This is Laid and Space Jest Swix today with our special guest, Jack Morris. A guest from Columbia, that's your affiliation right now?

Cornell.

Speaker 2

It's actually confusing because I go I'm in a the New York City outpost of of Cornell. So you have the city. Right?

Speaker 1

Cornell campus in New York. I just you're you're a student of Sascha Rosh who teaches at Cornell, so I I should have made that connection. Okay.

Yeah. I'm sorry. Wow.

That's that's a horrible mistake to make right off the bat. But you're one of look. You're one of the there there are not that many PhD students that make an impact with their research.

The last time someone like this happens was Shen Yue from Princeton, and he joined the OpenAI operator team quite shortly after he graduated. So, like, you're one of those, like, high profile PhD students at least that's, like, coming out of the program. And, like, I figured, like, it was a good time to just, like, talk about your work and also the fact that you're looking for, like, which lab you're you're gonna join.

That's like a whole interesting meta discussion, especially with, like, the insane market for AI talent these days. What's it like to be an AI grad student these days? Yeah.

And thanks for having me.

Speaker 2

first started or, like like, put yourself in my shoes. In 2017, 2018, I really learned a lot about machine learning. And at my I went to a state university.

It's a good school, but they didn't have, like, a deep learning research department or anything. They had people doing it, but it was just not as big at that time. But I was getting really interested in those topics, especially as applied to language.

And then in 2019, I kind of was starting to do research. And I think thinking about my career, I mean, at that point, was 2021. I was thinking about, like, where do I wanna be career wise or, like, who's doing the coolest stuff right now?

Like, looking at like, what kind of stuff is coming out of that time? Mean, I think AlphaGo. I thought AlphaGo was really good.

At that time, I was playing a lot with, like, BERT and BERT based models. So, like, you know, Google, DeepMind, they're doing great work. GPT two, GPT one from OpenAI were, like, interesting, but I think most people were into BERT at that time.

I still have a soft spot for, like, that parameter class of, like, 100,000,000 to 1,000,000,000 scale models. But this is all to say, I think at that time, I felt like the people doing a lot of the most impactful work were, like, professors and PhD students. Like, just a ton of, like, interesting ideas being explored and cool opportunities in academia.

So ended up applying to grad school. Well, at first, I did this Google AI residency program, which was mostly during the pandemic, like, 2020 and then 2021. And then I was also applying to grad school.

Started grad school in 2021. That's still what was going on at that time. Like, around when I guess, GPT three, one hundred and seventy five billion had been released, but not InstruqtGPT.

So, like, we had pretraining and sort of the science of pretraining was emerging, but that's where the models were. And I still think, like, I'm glad that I went to grad school and, like, I had a great experience, but the last five years have changed a lot. Like, the whole meta has shifted, you know?

Like, the kind of power dynamics are completely different. The ideas are coming from different places. Most stuff is open.

Now most stuff is not open. The types of questions people are asking are different. And so, yeah, I mean, for better or for worse, I did go to do the full grad school thing, and and here I am.

It's been really interesting perspective watching the science kind of emerge with the products. Like, the biggest thing that happened by far was, like, ChatGPT coming out and which was right in the middle. Like, what?

2022 before Christmas, like November? I remember that year, like, all like, my grandma was asking me about it. And that's when it hit me like, oh, this is actually becoming like a real area that people will know about and understand.

Like, I was trying to explain it to my parents. And that's when I think things really started to change in terms of the types of questions you wanted to ask can't always be answered with academic resources. So a lot of the, like, fundamental kind of, like, boundary pushing and AI science moved into companies.

Speaker 1

That was the year when, like, you know, just around Europe's as well, everyone in in N L NLP and deep learning were were were, like, very confused at, I think some people were, like, kind of expecting this already in a sense that they had they were obviously more clued into large language models. But I think that the sheer amount of consumer level interests that had that that was around at the time in 2022, that completely changed the world. Like, now now we're just like in a different sphere.

Did you have to pivot your research? Or were you already you just went from BERT to like other stuff.

Speaker 2

You you've done a lot of embeddings work. I mean, you're always heads down working on a problem. So I don't think most people in academia are the type to say, oh, look at this new product that came out.

I'm gonna abandon everything I'm doing. That can be the right move, you know? Oh, it definitely can.

Honestly, if if if I were to give advice to a younger grad student, I think the way to do it would be literally just, like, sit and wait until the next kind of paradigm shift and then just immediately start working as fast as you can to, like, reimplement it. Like, I I don't think that's, like, maybe the best way to do science, but it's probably the best way to play the sort of academic game in the in the days of AI. Like, you've seen that so many times.

Most recently, probably with the reasoning models, like, o one came out of OpenAI September 2024 last year. And then there's just been this explosion of, like, you build, like, abstraction ladders on top of that. Like, first, it was reimplementation.

Like, how do we even do this? And now it's, like, a lot about the data. What's the right data?

What are the right evals? What are the right training schemes? Like, there are so many different axes you can test and publish research in.

And, like, I think the easiest way to do that probably is just work in a field like that that, like, has it only existed for less than one year?

Speaker 1

And so no one has any, like, big advantage, I guess. That is mostly correct. I think anyone who jumped on reasoning and RL for LLabs is doing super well.

I just saw this morning that one of the recent Stanford grad students who worked on RL, they just started their company, and they're worth 500,000,000. It's it's, like, absolutely bonkers right now. Like, just, like, no products, just three dudes, you know, sitting in some basement somewhere.

I mean, undoubtedly cracked, but, like, also not worth $500.

Speaker 2

Yeah. But maybe it's not paying for the product. Right?

It's like the the The potential. Ideas behind it or the yeah. Yeah.

There was this big shift from in scale of working with a 100,000,000 parameter models. Really what happened is, like, I think the company has invested a ton more into training and infra. And, like, we all kind of had to catch up.

Like, you know, me, I go to Cornell, work with a professor there. He has to buy GPUs. Like, should he buy last year's GPUs or this year's GPUs?

How many should he get? That we were kind of, like, trying to figure that out. And there was there was, like, a big lag, I think, where basically the the seven and eight billion parameter scale like, there's a huge difference between the BERT size models, which are a 125,000,000 parameters to to 200, and then, like, the 8,000,000,000 parameters.

I mean, obviously, it's two orders of magnitude. But just, like, this idea of emergence, like, if you're talking to a model that's a 100,000,000 parameters, no matter how well it's trained, it knows nothing. Like, if you ask it, like, what's the capital of a state?

Or, like, if you ask it who's who was president of The United States in 1990 or whatever, it'll just always say George Washington because it just associates the words like president of United States with George Washington. And then when you get to the 8,000,000,000 parameter scale, suddenly it knows every single president. It knows every single capital of every single country.

And I really do think that changes the type of research you can do. And so, like, it took us a while, I think, in academia to catch up, like getting good 7,000,000,000 parameter models and then running them and getting GPUs to run them. Now I think things have stabilized a lot.

Like, we have access to compute, and we can kind of, like, fine tune and inference that scale of models, and that's, like, kind of fine.

Speaker 1

But there was, like, kind of two years where everyone in academia was working on, like, smaller models, and none of it really mattered. I think sort of branch that discussion in two ways. And we should we should sort of go to your research at some point, but I'm enjoying Because I think, like, we don't get to talk about this on the podcast too often.

One is that there's there's an often there's an often bit of advice from the industry people to to grad students, which is give up, don't work on models, just do benchmarks. Right? Like, a really good benchmarks will will get our attention, and then we'll hire you, and then you can switch to models later.

You have, for better or worse, avoided that, which is cool. And we can talk about that as well. But the other thing I think is that roundabout seven, eight b, maybe four b is when you start switching from like a single GPU setup to like a distributed setup.

And I'm wondering, like, do grad students get HPC training? How much do they teach you of, like, just how to work with, like, large clusters of stuff?

Speaker 2

Oh, to be clear, they don't teach you anything. Like, anything like, if you see a paper coming out from even, you know, Stanford, probably the best school in AI if if you had to choose. And it's not like they're learning how to do, like, multi node distributed FSTP training, like, with whatever deep speed.

You have to learn that from the Internet and from other people. And, like, there's no classes that really do that. I mean, it's that's hard to facilitate, like, as one person.

I would say most grad students are doing stuff on single GPU. Some people are doing multi GPU training.

Speaker 1

probably basically no grad students doing multi node training. I mean, there's probably a few, especially if they have, like, company affiliations, but that's really unusual, I think. Okay.

For grad students who are looking to get up to speed on that, I will recommend the GPU mode Discord, where basically the PyTorch team is hanging out in there and just waiting to help you. And then the other one would be the Fast AI team. If you have some kind of thing, Jeremy Howard will basically help you out, and they they have some distributed training.

Honestly, try to reach out to the deep speed team at Microsoft. Like, actually, they're reasonably accessible. Nobody talks to them.

Like, it's so so funny. I I, like, met them at Europe's, and, they had nobody at their like, they was presenting these speech three. I was the only one asking questions.

Like yeah. Yeah. That's good that's good advice.

Listen to this guy. Yeah. I mean, just basically, like, the people are there if you wanna ask.

This is very, very valuable experience. Once you're, like, a GPU god, like, you're basically, you know, in a, like, a different tier as a researcher because you don't rely on someone else helping you out. Like, you can just kinda be your own research engineer.

You know? Yeah.

Speaker 2

if someone has been listening to this and also following me online for a while, I think I've made a couple comments, like, saying something like you shouldn't learn about CUDA or things to that nature. And I'll I'll give some more color to that. So it's definitely a great idea to learn CUDA if you can.

I think my point was that if you're trying to enter this space, like learn about the models, learn about how they're trained, what the data looks like, what the compute looks like, axis of that is how to do more efficient training and inference. And one part of doing more efficient training and inference is studying the hardware, which is GPUs. So, like, I think that's a very small subset of all possible knowledge that you could acquire, and it's probably not the best place for a lot of people to start.

That said, if you do it, you're you've gotta be one of the most hireable people in the world. Like, if you, like, really deeply understand the architecture of the new GPUs coming out and and how to control it, you're in a very small handful of people and, like, everyone will want to hire you. But actually, the sweet spot is not even CUDA right now.

I would say, actually, it is Mojo. I don't know if you've been paying attention to modular Mojo. Oh, I listened to your podcast, man.

You had that guy on the other day.

Speaker 1

LLVM, Swift, all these things. And now he's turned his attention to the Python CUDA relationship. Right?

And he wants to basically create a viable CUDA replacement. It's basically Python married with Rust. For the last two and a half years, he was basically kinda stealth, not ready for production.

When he came on our podcast, he was basically announcing to the world, like, we're open for business. Like, you can use us now for for most models. And, like, we actually are faster than, like, the native, like, sometimes the PTX implementation.

I don't know how that works precisely, but he's a compiler languages god. I think there's a there's one of those windows now, like like you said, like, you know, bet early on something that's that's a that's a shift. It's one of those windows now where you try to implement things.

You basically, like, you know, modulates a 100 people. If you run into issues, you'll get Chris's personal help on things. Like, I'm not promising it, but, like, probably.

You know? Like, because he wants to work on improving the toolkit. And all you have to do is just, like, it's not really about becoming a CUDA god because, obviously, like, once you wrap up on on the the general concepts and principles, you can probably translate ecosystems pretty effectively.

A lot of people switch from, like, JAKs to to CUDA. But, like, the the thing is just, like, being able to experiment very quickly on a limited budget. Like, efficiency is not just because you are trying to be an efficiency guru, and that's your career, and that's kinda boring.

But it's really also just about being able to experiment very quickly and finding these these ideas.

Speaker 2

I also think VLLM and SGLANG seem, like, really good and important and here to stay. Like, they'll probably just get larger and more complex to accommodate future systems. But if I were, like, a starting out grad student, I and working in that area, I'd probably, like, wanna learn more about how they work.

Awesome. Let's go to your research.

Speaker 1

the contextual document embeddings paper. You can tell me the story about that. But I just wanna show you proof that, you know, I get one slot per day to highlight the number one AI story, and you were the slot of the day for October 5.

Oh, no way. I mean, obviously, you were producing work before that. But, like, I thought CDE was a really cool exploration of, like, oh, yeah.

You know, embedding models are kinda, like, stuck in a rut. Like, here's actually how to make them very efficient by just doing it in two stages.

Speaker 2

simple insight that was done very well. But you have a general, maybe, information theory thing that maybe we should start with, and then we can sort of create an insight our way. Yeah.

Sure. That sounds good. So we can we can circle back on that.

That's that's really cool that you wrote about it. What was that? Almost coming up on two years ago.

Yeah. This is the post I wrote. I I called it a new type of information theory.

We don't need to go into the there there's this paper about a concept called the information. Maybe I'll give, like, the most simple explanation, which is if you say you have two text files. One text file contains a paragraph of information about New York City, and then the other text file con contains the same text but encrypted with, like, whatever encryption algorithm.

So it looks like random letters. But if you decrypt it, it has the same text as the first text file. From the perspective of, like, Shannon's information theory, these two files contain the same information content.

Like, relative to everything, they have the same number of bits. But it's it's very clear to the observer that the first text file, which is plain English text, is like much easier to read and easier to process even though they have the same information. And so there's this theoretical framework proposed in this paper, which is a theory of usable information under computational constraints from 2020.

It really doesn't have that much press. There not aren't as many citations as you would think. But I think it's a really, really neat idea.

It's like we should measure information with computational power as a constraint. So, like, they have this idea, they call v information, of how much information is extractable from a given, like, file or or code. So in that case, we could say the left text file actually has more extractable information than the the right text file.

I think that's, like, really good. That captures a lot of our ideas of how these deep learning systems work. Like, why does pretraining work?

Like, if you have two sets of weights and you you wanna train on some downstream dataset, one set of weights is pretrained, one set of weights is randomly initialized. Why is the pretrained model better at all even though it's never seen your data? Maybe one way of looking at that is that it has, like it makes the information, like, more extractable somehow.

Like, there's this concept of, like, computational processing that you can almost, like, store up. I like this as a just like a lens to view problems with, like, how much information is stored where. Like, if you if you get a a set of model weights or, like, an activation vector and you open it up, like, print some tensor or NumPy array, it looks like random numbers.

Right? Like, there's nothing human intelligible about that. But, really, it's this complex combination of, like, the training data and the training algorithm, which get compressed into model weights and then the actual computation that the model is doing, which involves, like, manipulating these numbers in ways that we don't understand.

So it's like this really highly compressed nonlinear combination of all these information sources mixed with, like, computation. And I just think we don't have, like, the right words of of discussing this. I think I like the information theory analogy because back in the day, you know, we had phones and, like, telegraphs, and and people were just sort of, like, building the phone system with these crazy heuristics to, like, send information across the country or send telegraphs across the Atlantic.

People were just, like, trying stuff. And then we kinda found stuff that worked, and we we ran with it, but it wasn't really optimal. And it wasn't until someone came along and proposed this concept of, like, a bit, like a one or a zero that tells you something.

And once we have a bit, we can do all these things. We can, like, count the amount of information in a signal. We can do really good error correction.

We can measure properties of distributions of things, and and we can build, like, a really good system for for phones and then eventually which led to computers. I'm bringing this all up because I don't think we have I don't think we know what a bit is yet in terms of, like, deep learning models. I'm gonna graduate from my PhD this year, but I didn't figure it out.

So if you're listening to this, maybe you can, like, I don't know, spend more time on it or you're smarter than me or you have bet you know, a group of collaborators. You can all get together and figure out what the right lens to look at this stuff is. But even by just asking these questions, I think I was able to conduct this research agenda that I'm kinda still working on, actually.

Speaker 1

Yeah. What do you call this field?

Speaker 2

I don't know. I don't know. I called the post a new type of information theory.

I I don't think it exists yet, I guess. So maybe it'll it'll get a name once someone actually comes up with the right set of definitions. I think v information is a is a really good start.

Speaker 1

There's a couple related threads. So first of all, you don't know this, but I actually have been trying to accumulate in data about Shannon like, this is just like a Shannon information theory view of language models. I have a a lot of notes.

This is actually on my GitHub for people who are watching along. But, you know, at, like, at the limit, if a language model has a 175,000,000,000 parameters using 16 bit, you could it would take up 350 gigabytes. You can compare that to Wikipedia.

Wikipedia is about a 150 gigabytes. You know, let's say g p c three can store two Wikipedias. But, like, is is that a relevant measure of information storage?

Right? It is not because you can compress Wikipedia a lot. There's a lot of repeated patterns.

Tokenization is is, like, the first form of compression. But I think there's a there's a related talk from Elias Sotzkever about how deep learning is machine learning, it kind of is is is compression. Like, it you have a dataset, you compress it into a a model that is smaller than a dataset, but generalizes and, like, has, like, you know, some some amount of acceptable loss.

I think that one of your commenters on on the on the post made this direct comparison with com Komogorov complexity, which is what Ilya how Ilya sees it. So I think, like, people have this information theory idea or approach to language models. It is just not precise because exactly what you say, like, it's we don't know what a bit really means.

We don't know what, like, the most legible legibility is is a word that comes to mind in terms of, like like, it it matters to us that it's human readable. Like, even if it's SHA one, SHA two, six, I don't care. But, like, that is less readable.

And therefore, there's more, I guess I don't know. Entropy is not the word because it's it's directly convertible. But it's just less useful.

Speaker 2

Yeah. Yeah. Yeah.

Useful is a good word. I think maybe useful information or usable information is is the right lens. And Kolmogorov is a really interesting connection.

Like, Kolmogorov complexity. I think that's a really good concept for for computer scientists. So I'm not sure exactly about this specific talk or, like, what he was trying to say, but I I think that we have a very good understanding of language model pre training, and there's a deep connection between language models and and compression.

Actually, may maybe let's let's start with the embedding so we can come back to that. Okay. So is this are we going to the first paper?

Actually, let's go to your Wikipedia numbers if if you still have access to that. So this 50 gigabytes for text of Wikipedia, that sounds like a pretty high to me.

Speaker 1

like text files? I don't know. I I grabbed it from Andrew Ng, so I don't know.

Speaker 2

Okay. Okay. No.

No. I'm probably off. I just sort of have the sense that, like, when you store text, it's generally, like, very, very small, especially when you use it.

Maybe he's including all the languages, all the edits. I don't know. Yeah.

Yeah. That could make sense. That could make sense.

Because I I guess what you you say from you know, if you wanna do apples to apples comparisons, g p three can sort two Wikipedias. Is that right? Two two point three Wikipedias?

Two point something. Yeah. So I thought it would be a lot more.

And this is actually an experiment that you could do. You could, like, just train a model on Wikipedia and keep training it until you can perfectly extract all of Wikipedia. And that would be, like, a good way of knowing, like, how many Wikipedias can CPG store.

I like I like that idea. But I think this type of, like, back of the envelope math is is really useful for thinking about problems and, like, grounding yourself in the real world even if you can never quite answer the questions you wanna answer, at least, like, in four years. If we think about embeddings, you know, vectors that people use for search, we can do the exact same kind of math.

So if you use the OpenAI embeddings, which last time I checked, I think have 1,536 dimensions. So that if you say there's 16 bits per dimension and, like, half precision floating point, it's something like 20 kilobytes of information in in a vector. And if you wanna store 20 kilobytes of text, that's a lot of text.

Like, many, many paragraphs that you can perfectly compress into 20 kilobytes. And so I think this is kind of like the idea we had. I'll give you the practical explanation, which was I'm well, first all, of I'm a second year grad student.

I'm, like, going to these conferences, seeing all these other things people working on and thinking, you know, like, what the heck? Like, how am I gonna, like, have my own little area to do work in that no one else is working in already? And so I spent a lot of time coming up with bad ideas.

And my adviser would say, no. Like, that's not a good idea to work on. Many times this happened.

And, like, even my first year and a half of grad school was, a lot of exploration and a lot of, like, coming up with bad ideas. And then, honestly, I'd be interested to see how he remembers it. But I think I wrote a sequence of proposals about different projects, and then I came up with this idea.

I was like, oh, we should just try to do as well as we can to reverse engineer the text that's in embeddings. I then we were talking about it. He was like, oh, yeah, you should just do that.

And then that was the end of the proposals. And then that was just working on that problem for a long time, which at the time, was really motivated by that because I was like, cool. Like, my first as a grad student, my first sort of like official, like, sign off on, like, coming up with a good research idea.

And at the same time, there was this big rise of this startup business model called like a vector database. And there are all these companies popping up, raising money, raising money, getting, like, crazy funding. And then actual applications being built that do something where instead of exchanging customer data, they exchange vectors.

So we had this, like, very grounded question of, like, what data are they actually sending when they when they send the vectors? Like, first of all, you have this information theoretic argument that when you send one vector, there should be a lot of text recoverable just in terms of, like, a lot of these things represent very short documents, but they actually have many, many bits. So, like, the problem seems tractable.

And then second of all, had this justification of how the product is actually being used. Like, if if someone hacks into a vector database, what do they actually find, if that makes sense?

Speaker 1

So we were working on that for a while. I think I have the the talk that you did from that Sasha highlighted. It was this one.

Oh, yeah. Yeah. Maybe that has the graphic that that would kind of Oh, go one before, I think.

One before? This is actually yeah. This one's good.

Yeah. I like having visual aid. I like how I like giving people break rooms to follow-up if they if they're interested in digging more.

Speaker 2

hot area of research at the time. And there's been some really interesting follow ups. Like, we we ended up building a system that can do this quite well.

Like, taking and embedding, and I think our highlight number is like at a certain length, like a long sentence length, we can get 90% of the text back exactly. And a lot of people were able to do stuff with that. Like, they can for example, I I know these people that work on a problem of, like, debiasing embeddings.

And, like, in one dataset, they do something. They have a procedure for, like, removing all latent features that correlate with gender so they can produce, like, useful embeddings that from some perspective have no, like, information or usable information about gender. And they'd been doing that for a while, and then they actually just used our tool.

And they so, like, they would put in a sentence like, this woman is a doctor at Weill Cornell Hospital in New York. Or say this woman is a doctor. She works at Weill Cornell.

And then they would run their procedure, and then they run our embedding to text model. And now it would say, like, this person is a doctor. They work at Weill Cornell, which is pretty cool.

So they have, like, sort of text based evidence that their method is actually removing gender features. But let me talk for a second about the research phase here because I thought it would be I mean, I know if if you ever heard me talk about this, I probably told you about it. But just for a wider audience, I like thinking back on this because it was probably my, in some sense, like my greatest victory of grad school was like working on this embedding inversion problem for a while, for quite a while and, and proposing a lot of approaches and, like, testing stuff.

I think sometimes you do stuff and it's clear it was a bad idea. Sometimes you think you should have figured it out earlier. And then sometimes you do stuff and you kind of realize it's really complicated and and probably not worth it.

So I was testing different decoding algorithms for embeddings that are closer or text that's closer to the text that's in embeddings. And I was testing these kind of, like, inference time adaptation models for samplers. I think we tried a lot of architecture and, like, kind of training tweaks.

We should have tried RL. I think that would work. But, well, finally, we found something that ended up working.

And I guess I'm just saying this all because I thought it was, like, so rewarding. Like, we were just banging our heads against this the wall. I would have biweekly meetings with my adviser who kinda suggest things.

Sometimes we would agree we were mutually stuck. Sometimes I would get feedback one way or another and and try something new or try a couple things. And and we had this idea that it was possible from the information theory arguments and this other thing where we would we would kind of like take our best guess at what the text was and re embed it and see that it was kind of far from the true embedding.

So we had this proof that, like, a better method could, like, leverage this kind of information. And then when we finally solved it, it was it was awesome. Like, we had this number that was, like, 30 for months.

I think at one point, I got it to 35. And, actually, I think I was like, oh, I'm done. Like, I got it to 35.

And and my adviser told me, I'm not like, that's you can't really just propose a new problem and show you push a metric from 30 to 35. That's, like, confusing and probably not that meaningful to people. And I think I was you know, that was kind of like a local minimum for me where I was, like, bummed.

But then we ended up getting the number to, like, 97 or something, which neither of us knew were possible. We were all just we were just kinda staring at this graph. Like, oh my god.

Like, who knew you could get this much information from an embedding? And that was, like, so great. Like, just sort of this it was so rewarding.

And so it was invigorating, honestly, like, that research process of, like, we picked a good problem, and then we spent so long trying stuff that didn't work, which I'm probably forgetting how frustrating that was. I'm sure it was terrible. But then, like, actually solving or at least, like, coming up with a much better way of solving the problem.

I don't know if I'd say we solved it, but we definitely learned a lot from where we started was, like, was great. And it completely solidified for me the fact that I should have gone to grad school to have this, like, life experience and, like, makes me wanna do research forever.

Speaker 1

You're clearly clearly sort of in love with the the the journey, which I think is is important because this is what keeps you going through the the tough parts. Is this a good time to talk about the universal geometry side then? Yeah.

Yeah. Yeah. Let's let's do that next.

I think that's a good idea.

Speaker 2

so we have this more recent follow-up. And the so the first part I was talking about ended up in this paper called text embeddings reveal almost as much as text Yep. Which was published in 2023.

And then we recently had a paper come out on archive, which will hopefully be published at some point, and it's called harnessing the universal geometry of embeddings, which was also that was probably, like, the only other time I felt like we've made, like maybe there have been two more times, but that that was probably the the second of three times where I felt like we made, like, a real discovery about, like, the unknown. And it was, like, very rewarding just for its own intrinsic kind of elusiveness. And I'll start from explaining it in terms of the prior paper.

So we we built a system that can, you know, do embeddings to text, and it and it works very well. And we're all we're very pleased with ourselves. And then we we went to a conference.

We talked to people about it. We talked to, like, the vector databases. I think some of them changed their privacy policies, which was, like, somewhat gratifying.

And then we kept getting this perpetual question, which is like, well, you're just assuming we use the OpenAI model, or you're just assuming we use the most popular text embedding model if they fine tune their own model or if they use a model that you're not training an adversary for, then you can't solve the problem, which is, like, true. Like, none of the vector to text stuff works unless you have this assumption of, like, knowing the encoder and also being able to make a lot of queries to it. But we had this kind of underlying theory that all of the models learn very similar things.

Like, we have some preliminary evidence for that. Like, certain models that are fine tuned from the same base, you can kind of swap their representations without doing much. Or if you look at the nearest neighbors, a lot of the models will give you the exact same nearest neighbors even though they have completely different training bases.

And then there's this paper that came out last year called the Platonic Representation Hypothesis from some folks at MIT, which is really, really compelling and I think just, like, great intersection of philosophy, representation learning, deep learning research. Like, I I love this paper, and that it's it's such a beautiful idea, which is something like all models are trained on data from the world, and there's only one world. And so as the models get better by scaling data and scaling model size, they're sort of converging to learn the exact same thing.

And in this paper, they have evidence based on correlations for doing this with vision and language models. It's very neat. And so we saw this.

So so, basically, think about, you know, you're us. You see this platonic representation hypothesis paper. A lot of people have this shared idea.

Like, you know, Claude and g b t four probably do a lot of very similar internal computation because both of them are trained on trillions of tokens of human written text. Even if they have different architectures, like, maybe, you know, the actual basis or, like, the the numbers, if you look at them, look different. But in some way, they're, like, kind of computing the same thing.

And I think it's even more true with these, like, embedding models, which have, like, really only one objective that works, and they're probably all trained on, like, MS MARCO, which is a really popular dataset, and pretrained maybe on Wikipedia. But we wanted to basically combine this platonic representation hypothesis idea with the vector text thing and produce a system that can, like, align models so that we can do embedding inversion.

Speaker 1

But, you know, it's it's valuable for more than just embedding inversion. Like, you can use this to kinda Yeah. Yeah.

Glue together models. Like, that's what actually got me, like, super excited. And by the way, like, I think there's a few related threads.

I think we did an episode with Nicholas Carlini where he had an extraction attack on on one of the the GPT models, and they they got it fixed. The other thing I wanna was just reading our spell out for people just in case they're not thinking it through. Being able to invert embeddings also means that you can you can back out, like, secret prompts or context that might leak customer information.

That's potentially harmful and, like, obviously, a a tag vector issue. I think one of the things I I had a question about was whether or not position embedding does affect it, and, like, the extension of first position embedding is affected. Because obviously that like, contexts are gonna get longer and longer.

Your ability to invert will obviously decrease with longer context. What now? You know?

Maybe not that important. No. No.

No. No. You're totally right.

Speaker 2

we're operating in this space in in our work where the sequences are relatively short and the embeddings are relatively large. Like, I think we're kind of at a great advantage from that perspective. And you're definitely right.

Like, if you embed an entire book to a 500 dimensional vector Yeah. No way. There's just no way you could get the entire book back.

Like, there must be this these kind of collisions. Like, you know, in information theory, like, if you have lossy compression, two different inputs mapped to the same code, which means that you can never determine which input formed the code. And I think that's probably what will start to happen.

Like, if you have two books and you swap just one word and you embed them, I don't know. Someone can try this. You'll probably get, like, a perfect collision.

And in that case, inversion is impossible. And even, like, when you don't take it to the limit, it probably just gets very, very hard. Like, things get super compressed.

So I don't know how well this work scales. Like, it's a great question, like, exactly how much information you can sort of cram into one of these vectors, and I I don't have a sense of where the boundary is. It'd be interesting to talk to some one of the, like, linear algebra people from, the math department on, like, how literally can we take inversion?

Speaker 1

Like, you know, how like, what what measures of a matrix do they have where we can, like, kind of run that and, like, try to get some meaningful information out of that? This is, like, where information theory starts to collide with linear algebra and all the all the other stuff. Totally.

Yeah.

Speaker 2

this detail where we're we're running these on computers, and so we don't actually have, like, real decimal numbers or real numbers. We have, like, floating point representations of numbers, which are like very it kind of like throws a wrench into the mix. Do you have any consideration of, like, superposition?

Speaker 1

When, like, sort of nonlinearity? Like, you you keep, like, stuff information in the lower bits, But I don't know if that matters. I I really don't.

It's just like a nice thing to think about. Yeah. Yeah.

It is a great question.

Speaker 2

get a sense that, like, a lot of the less important bits are more useful for computation. And maybe the higher order bits are more important for, like, storing data or something like that. But I I'm not sure.

These are the kinds of questions I'm actually hoping to explore over the next few years. Like, I'll skip ahead for a second. So we we have this result that's like maybe the the third sort of like discovery I was alluding to, which is like a way to measure the exact capacity of a language model.

And we get this number if you train a language model on a ton of random data and you measure its rate of memorization. Yeah. Can you open the right curve?

This is sort of the discovery I'm talking about. Like, no matter how you scale the training size, you hit this, like, perfect perfect ish plateau in auto memorization, which we call the model capacity. And the the question I've been stuck on in the back of my mind for a while is, like, how is that actually implemented?

So, like, this is a transformer that is trained for many, many data points and many, many training steps. And so, like, it's almost like, if you have okay. The 10 to the six point on the x axis, the capacity we don't have to actually say the numbers, but it's basically perfectly dividing its computation between all of the data points.

Like, every one of the 10 to the six data points gets, like, a tiny sliver of the model parameters because they're completely independent random strings. So I don't really know if superposition is occurring here. Like, it seemed possible to me that the model would learn, like, these completely independent columns of computation, one per data point.

But it's also possible it's learning some kind of, like, combined thing where it's maybe it learns, like, a load and a store, and it's, like, sort of, like, loading and storing bits using these generic operations. And then in the end, it reconstructs the random string. So even though, like, data is completely independent, the kind of, like, compute is is very similar in terms of, like, predicting random strings.

But, yeah, I guess this is all to say, like, about superposition and everything. I have no idea how the mechanisms are actually implemented inside the models, and that's, like, one thing I'm hoping to learn about the next couple years. It's a reasonable question whether it's meaningful to learn.

Speaker 1

I think there's a lot of things that is, like, nice to know, but maybe not that useful. Latent space lane space alignment is very, very useful. Dataset efficiency, in theory, cool.

But, like, practically, people are just gonna go for the biggest dataset they can. Like like Yeah. Yeah.

The scaling laws are kinda worked out in so far as, like, the relationship of compute and data. Ammono memorization, I I don't know. I think maybe this is a good point to maybe also bring in the idea that Andre has been pushing for the last, like, think year and a bit of the cognitive core.

Like, the what is the dumbest possible model that knows nothing, but is you is smart enough for tool use to do everything else. Right? So you can run it on device and fast inference is open source, whatever.

So Gemma three n is like a really good candidate right now because it's like a four b model that is like claimed to be better than Lama four and g p c 4.

Speaker 2

arenas that shall not be named. This is where things get complicated. Like, I it feels like language models kind of implement things and know things almost in the same way.

And it's, like, really difficult to disentangle, like, whether they're memorizing facts from whether they're, like, learning useful ways to generalize about new stuff. But I I agree this would be really nice. I don't think we have a lot of evidence that we can build a system like this that, like, is really, really good at reasoning, but really dumb about the world.

Like, I don't know if we have the tools.

Speaker 1

Yeah. Maybe, maybe not. I think the existence proof is humans.

Right? People always lean on humans as, like, the existence proof. It's not a great existence proof because I think if you've talked to people about the number of neurons that we have and you make a neuron roughly equivalent to a parameter, we have something like a 100,000,000,000,000 in our brains.

So like and like we consume like 20 watts of energy. Like, that's nothing. Like, we're so much better than than language models.

It's not even funny. And then the the last feature of us is that we are self pruning, which is not something that language models do as well. Oh, like we forget stuff?

No. Like, we are not deeply, densely connected. Like, we like, connections will drop, therefore we're more efficient, you know?

See, I see. Unlike a language model where everything is always connected all the time. Yeah.

Or you're like, you you preset the skip layers or whatever and then that's it. You know, like, it's not it's not really actually anything we've evolved with learning. It's just like something you do based on ablations and, like, guesstimates.

Even if we did want that, I'm not sure if we have, like, the right frameworks or methods for actually building, like, what you're talking about yet. I think the world is much closer to where you're at than where Andre is at. Andre is, like, kind of wishing for a a an optimistic world.

Our conversation with Noam Brown was like, yeah, reasoning is emergent. If you gave the o one harness on top of GPT two, you would get nothing because GPT two didn't know enough. You need a GPT three and GPT four in order to then get o one.

Like, as as GPT four is the base model. Which is like yeah. That's I mean, that's that's reasonable.

The way I put it is like, in order to use tools, you need to like, in order to search Google, you need to to know at least search terms in order to like, then search Google and then learn what you need. And if you don't know what what to search, then like, you might just be too dumb.

Speaker 2

I like the kind of ethos. Like, maybe you could do some kind of free training where whenever the model doesn't know something, it can just Google for it. And that way, you try to encourage it to learn words without or, like, to to guess words correctly without actually storing the information into its weights.

Yeah. It seems like a nice, like, goal at least. Yeah.

Speaker 1

probably or memory and some combination of that. Yeah. It's exciting.

You know? Like, I think, like, if that is the the direction of of where this all ends up, that's great. But, like, people aren't are not doing that.

Instead, we're building, you know, $500,000,000,000 data centers in the middle of Texas. And like, you know, all hail the the the god cluster that just all, you know, eventually wrap around the sun and consume solar energy because that's that's all we need. Do we finish out the universal geometry thing?

Speaker 2

methodological description. So so we had this goal. So so so, yeah, back to the embedding universality.

We started with going from embeddings to text. We know about this platonic representation hypothesis. And maybe I'll skip over the details, but basically, had total inspiration from computer vision in this model from 2017 called CycleGAN, which is among other things, it's a way to map between two different distributions without any underlying notion of, like, which thing should be mapped where.

It's just based on some kind of idea of closeness. So, like, the cool thing about this, if you look at the top left so I guess the the top left is Monet, so impressionist paintings. And this picture on the right is a photograph.

So, like, it's learning this kind of, like, semantic notion of what content goes where just by mapping a distribution of Monet pictures to a distribution of photographs without actually telling it which Monet picture should map to which photograph. It's kind of a subtle point I'm making. It takes a little bit of time to wrap your head around, or maybe, like, go to the middle one, if you don't mind, the zebras and the horses.

So, like, it's clearly learning, like, what an animal is and what legs are and sort of, like, more abstract stuff like what the camera position should be and and what grass is and stuff like that. And it's learning, like, what a horse that looks like a zebra is, which is actually like a complicated semantic concept. Like, we don't have a a dataset that has a horse and then that horse as a zebra.

We just have separate horses and separate zebras. But somehow this this GAN system is able to, like, elicit this sort of mapping property. It's like kind of a magical connection that it learns.

And I'm still, like, in awe that it's possible at all. But we more or less, like, repurpose this system and, like, we we built our own. But, like, this idea, we took it and we applied it to model embeddings where instead of zebras and horses, we have, like, BERT embeddings and GPT embeddings or, like, two completely different models with different architectures.

So I think these are GTR, which is a t five based retrieval model, and GTE, which is based on BERT. So they have different training data, different architectures, different downstream objectives, different embeddings. But yet when we do the CycleGAN in the embedding space, they just perfectly sort of snap to the same place, which is amazing and has some pretty deep implications of, like, the platonic stuff.

Like, maybe the models actually are learning a lot of the same functions or something, and in some semantic way, they're, like, very close. And, yeah, this is a diagram of how our system looks.

Speaker 1

It's weird to me how profound it seems. Again, you you seem you seem, like, deeply impressed by it. And then the other thing is, like, when we talked to the the to Emmanuel from Anthropic who did the circuit tracing and mech mech interpret mechanistic interpretability work, they were, like, excited that, like, the same thing in different languages maps to the same circuits.

And I'm like, what you would expect? Yeah. Yeah.

I don't, like, I I don't know, like, why like, I I I I don't know. I think I feel like this this feels more profound to you than it does to me. I'm like, yeah, obviously.

No. That's that's so fair.

Speaker 2

and we're happy that we're, like, the people that got it to work. Yeah. Exactly.

Yeah. It does it does seem obvious in retrospect. And I think that's, like, constant feedback I've gotten from research from, you know, people will tell you that this seems obvious to them.

But you have to realize that, like, you came from a perspective of no one ever having done this before, and they're coming from a perspective of you telling them it's true.

Speaker 1

maybe obvious to you too, if that makes sense. The way I put it is that we have the intuition, but not the proof. You have the you did the work and you have at least some evidence that it's true.

Whereas we just have intuitions. Right? So part of research is just confirming intuitions.

Speaker 2

The applied part comes from like, okay. Now that you know this for a fact, what do you do with it? Yeah.

Right. I think the details can be really interesting. Like, the details of the proof, like, which models are most similar to one another and to what degree can you get them to align and on which distributions does this property actually emerge.

And, like, that's why reading papers can be fun sometimes is because they kind of answer all those little question.

Speaker 1

Yeah. I would say okay. I pulled up something very current, which is Gemma three n, which launched which sort of was generally available yesterday.

I would say the for me, and you can correct me if I'm wrong, the most immediate implication is mapping adapters to language models. So the the dream is that you have a language model backbone. Let's say this one is like a two b language model backbone.

And then you offload your vision so you only you only load in the vision and para parameters or the the vision adapter when you need vision. You only load in audio. You only only only speech text to speech whenever you need it.

Because these are all separately trained. You're you're just sort of aligning latent spaces, and you can sort of train them separately. And I think, like, this helps to make us more confident in one, it's it's more efficient.

That's a that's a given. Two, it makes helps makes us confident that we can just just add capabilities without taking away or catastrophically forgetting others.

Speaker 2

So they're just sort of like stacking more parameters. Just stackable. So that's very cool.

Yeah.

Speaker 1

stackable. It's like a a fatter version of Laura's that is not really Yeah. That model specific.

I would say Apple and and Google are pursuing this for their on device stuff, is is where is my sense. Is this open source? Gemma?

Yeah. For a given definition of open source, which is like we release the weights of Hugging Face, here you go. Oh, that sounds like open source to me.

Speaker 2

Oh, yeah. I guess it's it's open weights, but not the data.

Speaker 1

Not the data, not the code. Not the code. Yeah.

Right. Yeah. I would say that this is quite soda in terms of efficient models.

Speaker 2

a small l m also from Hugging Face would be also in that in that category. There's not that many people working on very good, very efficient models. Yeah.

This is a a very deeply related question and something that really interests me, which is like, what is the limit of, like, a 100,000,000 parameter model? Like, if you imagine, you know, a hundred years from now when we have maybe our computers are gelatinous blobs and we all communicate through telepathy, will we have 100,000,000 parameter models that are at the level of today's o three pro or whatever? And, like, if so, like, how would that even be the case?

Like, based on scaling laws. Like, do we have special data? Do we come up with, a brilliant new training scheme or some type of magical architecture?

Like, I really don't know. Or maybe we're we really are on the at the plateau already. I don't know.

Speaker 1

as as small, like that's what miss Charles is doing, you know, we've plateaued a little bit in terms of like what we can do to compress things. I have a fun theory that this is where we mix quantum computing with models. Like, we have to change what a parameter means.

We have to search through very high dimensional space and and resolve them much quicker than we can with, like, conventional compute.

Speaker 2

That would be my pie in the sky thing. I said a hundred years. That's very reasonable to me.

Throw quantum at it. Yeah. I probably have to get a second PhD to know what's going on there.

I think that we should establish the definition of small model as being a model that a grad student can inference at reasonable time on a single GPU. Which is probably like 7b, maybe. I don't think 27 is small under any reasonable.

Is it wait. Is it MOE?

Speaker 1

Mistral? I don't think so. I think I'll I think their stuff is default dense.

Don't quote me on that. This is this is coming off of just a lot of pre trained data that is that is potentially collided. Okay.

There was one there's two more papers that we wanted to cover, and then we can we can sort of wrap it. You had an approximating you had a language model training data. I think this is a little bit also newer.

Speaker 2

How does this rank in terms of your your overall work? Yeah. Let's return to the kind of information theory question.

So, yeah, maybe we'll we'll skip over the contextual embeddings in the case of time, but we'll group those papers. Great paper. Hopefully, people start training with that technique.

It's kind of a free lunch. Those questions are all about information and model activations. Like how much can we recover from this given vector?

Or like, what data does this vector represent? Or what computation does this vector represent? And there's really two types of like, if you wanna taxonomize, there's there's two types of, whatever you call it, dense information storage mechanisms.

One of them is activations or embeddings, which we were discussing already. And then the other is weights, which are the things that are used to perform the computation, but not the computation itself. And so we have now two papers in this direction of what is stored in the weights.

The first one is about language model capacity, which is called how much can language models memorize or how much do language models memorize. I never remember which one we settled on. And then the other one is called approximating language model training data from weights.

The first one is, like, I I think has a lot of deep messages about how language models store information and how they work in general. The second thing is, like, a proof of concept of maybe, like, a longer term research project. Let's start with the capacity stuff, if that's good with you.

Do I have the paper for that? I don't I don't know. You know, we can return to the question you asked me, which is something like, why do we care?

Or, like, what is this useful for? And I don't know if I have a good answer for this. I think this is this is somewhat profound.

Like, it's kind of like in in physics, you know, when they try to measure these constants, like gravity. People tried to measure the rate of acceleration of gravity for a long time or, like, those Greek guys, like, back in in the BC era when they were trying to to approximate the radius of the Earth based on shadows. We're trying to take the GPT architecture, like the main one, and just measure how much information it it can store.

And we did this through the the lens of memorization, which I think we can skip over for the podcast, and and we'll just talk about, like, information storage and and weights. Like, these curves to me are are pretty crazy. Again, maybe maybe it's like the sort of discoverer's fat folly or something where I'm like, oh, this didn't exist before, so it seems so cool.

But then you're saying, like, it seems somewhat obvious?

Speaker 1

No. No. No.

Don't don't let me take that away from you. Yeah. No.

No. Like, I I independently was asking how come there's not enough people exploring LLMs from information theory? And then, like, you used to come along and your embeddings works become, like, an information theory exploration.

And I'm, like, suddenly, like, I'm very aligned to, like, exploring this, promoting this, and encouraging more people to to figure it out. Because, like, that's ultimately how we figure out this whole compression issue and that, you know, what what Andre wants, which is like the cognitive core. Right?

The most efficient model for the most capability. Like, that is an information theory question. Totally agree with that.

We could start here.

Speaker 2

that are trained in 32 bit precision, we approximate can store about 3.6 bits of information to maybe 3.9 bits, somewhere in there per parameter.

And, like, why is this? I mean, for some perspective, this is this is quite bad. Like, if you have 32 bits available and you can only use three to four of them, like you're Just store 32, bro.

Yeah. Yeah. Then you'll like, you know, you could build your own AI lab if you can make the these models that much more efficient.

I don't know how they're implementing this mechanism or where the kind of bottlenecks come from or even now that we know this, what it's necessarily useful for. I guess the tools that would be interesting to me are knowing, like, given a dataset, if you could predetermine the exact model size and maybe architectural properties required to get a certain level of performance, that would be really neat. And, like, we don't even know how to do that.

We don't even know what the difference is between doing low retraining, which trains less than 1% of the parameters, and full fine tuning, which trains all the parameters. We don't even really understand the difference there. So I think this is like maybe like a baby step sort of in that direction, but there's a lot of unknown ahead of us.

Okay. Do you think this is a hard limit? Do you think someone can come up with a better algorithm, better architecture, and then sort of just change the slope?

There are two axes here. One is the ability of the model to store data. And I think we can definitely improve that.

I think, like, maybe even if we tested this with the LAMA architecture, like, there's sort of, like, a GPT plus plus architecture. Like, I would guess that can store better data just because the kind of numerical flow is a little bit better. The nonlinearities are maybe, like, a little bit more suitable to training.

Like, that will probably raise the bound a little bit. And then the second axis is that our measurement tools are just not that good. Like, this is you know, me, I'm a grad student.

I'm running all these hyperparameter sweeps and sort of like where we draw conclusions from them. But even that being said, like, there are probably ways to measure this better. And but all that would do is push the number up.

So it's possible, like, there is a way to store five bits per parameter if you have, like, a better optimization technique or if you were a super genius and you could just perfectly set the weights to store the data, then maybe you can do better. And this is just sort of like what we can reach through optimization is this 3.6 bits per parameter.

But I would be happy if someone came along with a much better measurement tool. Like Yeah. This is just sort of, like, the first measurement.

I mean, I would I would guess in the future, like, you know, people will look back and say, like, this is, like, somewhat off in one direction or another for whatever reason. And and that's just how science goes, and I have no problem with it. What we do is we call we we call this the Morse concert three point six.

Right?

Speaker 1

I would never. And then we we set a, like, a, like, a challenge, like a leaderboard of, like, beat this. Right?

And, like, let let people go. Yeah. That assumes that we know the truth constant ahead of time, and we can measure the error rate.

It's not it's doable. You you laid it out here. Yeah.

Yeah. Yeah. Yeah.

That makes sense. One minor doubt I have is like, the goal actually isn't memorization, it's generalization. Right?

Mhmm. The best memorizer model may not be the best generalizer model and you're like, this incentivizing people to to max this number might actually just be fruitless in terms of actual intelligence. Like, you just get the best actual compressor.

Like, you're just gonna get g zipped. That's totally true.

Speaker 2

And there's this pattern in research, you know, time after time. It's like someone poses a question and then people answer it over and over and over again, but it's it's often much more fruitful to just ask a new question. Maybe it just doesn't matter how much GPT models can store, and you should just, like, work on something else.

Speaker 1

We'll figure that out.

Speaker 2

dwell on this this side at all? Yeah. Let's just talk about it real quick.

Definitely not the algorithm itself.

Speaker 1

By the way, what are your tools for doing these kinds of charts and these kinds of diagrams? Like, I I just come kinda curious of behind the scenes on the tools. I think, like, visualization has definitely been some hobby of mine during grad school.

This one, actually, Oscar, my co author made this one.

Speaker 2

but he made it I I think most the last few papers have all been in diagrams Google diagrams. I was using Figma for a while and Illustrator. I think Illustrator actually is is the best tool.

Speaker 1

Oh, did you know the transformers transformers diagram was in Adobe Illustrator?

Speaker 2

Oh, yeah. Yeah. I did know that, actually.

Yeah. Because that's the only way you can get arrows that sort of like curve like that. And they they've good shadows.

Diagrams is like the least robust, but it's the most accessible. And honestly, if you if you're good, you can make pretty good stuff.

Speaker 1

is is nice too if it's not going in a a paper. Yeah. It's just too rough for a paper.

But you need something professional looking, you know? It helps like, if you're gonna publish your work, you need to make it look nice and professional and, like, official. Right?

So this is what it is. Yeah. Yeah.

And I think there's something worthwhile about saying, like, okay.

Speaker 2

all the references perfect and and all the diagrams professional, all the captions are correct. And I think it's, like, important to put that level of detail into That's a little chewy. But let's finish this off.

So okay. We're talking about bits, information theory, what what information sort embeddings. We're talking about language model capacity.

I think a much more quest more practical question is maybe some maybe this is more analogous to the vector database hacking, embedding, threat model we discussed is like, if you have access to a set of model weights, what can you learn about the data? So like you were just mentioning, Gemma three b came out yesterday, and you can download it. And it takes up a certain amount of space on disk, and it was trained on some data, but we have no insight into what the data was.

I mean, it's probably English. It's probably some distribution of web text. I would guess there's a lot of code.

And we seem to have a lot of information about the model. Right? You have this file, and there's, like, many ones and zeros, which means something.

But it's kind of like a very highly compressed version of the training data. But I I I would be extremely surprised if they do any type of, like, private training. Like, there are these mechanisms for doing, like, differentially private language model training or even just anonymization in the pretraining pipeline.

I bet they don't do any of that. They just sort of like train on the data, and then they kind of know that we don't have the right tools to decrypt to the model weights. And so that's like my dream is we can come up with some way of of translating model weights back into text datasets.

And so in the the most recent kind of drop paper drop is that paper approximating language model training data from weights. And it turns out to be a really hard problem. Like, trying to go from model weights to text is really hard, and we do something a lot simpler, which is like well, there's there's two ways we make it simpler.

The first thing is we assume access to two checkpoints, which I think is is probably not the case in Gemma. But in in the case of DeepSeek, if you download the 400,000,000,000 parameter model weights, it's this giant file, and you can actually get two of them. You can get the base model weights and the fine tune model weights.

So the way we put this, we you have this kind of like difference in parameter space telling you what DeepSeek fine tuned on, and it's very controversial. I mean, they're sort of like geopolitical. Definitely at the at the corporation level, they're really interested in the implications of, like, what a deep sea train on.

And they've released this kind of treasure trove of information of what they trained on, which is the actual model weights, but we have no tool for, like, interpreting or kind of decrypting this weight difference. And so we started with something really simple, which is instead of even just trying to, like, regenerate the training data, we take just a web corpus and try to do selection of training data that kind of, like, looks like the true training data and gives us performance that's as close as possible to the true training data. So there's this complicated method, but it's something like you just sort of, like, look at the data point gradient and see if it points in the direction and weight space of the fine tune, and then you take, like, the top dataset.

There are some tricks to it, but it's basically just like gradient based selection based on this weight difference. And it seems to be okay. Like, it can get us pretty good training data.

So I guess if you actually wanted to use this, it would be like your competitor releases a base model and a fine tune, and you're trying to recreate their dataset. So you can take this weight difference and take a giant web dataset. Like, if I was doing this at a company, I'd probably try to scale it up to trillions of tokens and then select the exact data points that try to produce the model.

And and it turns out you you can train a pretty good model with that. We don't get to quite the performance of the original model, but it does seem to be, like, trending in that direction.

Speaker 1

This is, like, very creative. I don't know what the the the the use of it exactly is.

Speaker 2

Yeah. Like, when would you be in this exact situation?

Speaker 1

Decently often for the Open Model Labs. Like, even DeepSeq r r one, like, has released an update. Mistral does it pretty frequently.

Lava does it frequently. It's not impossible. But like, I I think it's like it's like think that I really like the creativity in using quote unquote synthetic checkpoints to do this.

Mhmm. Which is I don't think I've heard this from any other place. So I don't know if you you came up with the idea.

It's like linear interpolation in in weight space. Okay. That's a bunch of the recent work.

I wanted to sort of cap things off with the datasets question. If if that's is is that a good? You can ask me whatever you want.

Well, it's not it's not like a it's not an ask. It's just like, I think this is a very good thesis. I think it's a hot take.

I almost invited you to speak based on just this alone, but it was it was a little bit late to talk. Oh, for the for the conference. Yes.

When I look for conference keynotes, I look for something that has a broad overview that can I put the last few years in perspective, or it's an insight that you can reasonably rely on to, like, last for a while so you can get some mileage out of it? You know? And I think a lot of ideas in AI come and go.

But, like, things that are trend these things are scaling laws, things are trend lines, things are things that are like there's, you know, there's no new ideas in AI, that I pay attention to. So maybe you wanna recap. Like, what's what's the backstory if if there was one?

Yeah. Yeah. Sure.

Speaker 2

and this is a post that I wrote a a few months ago. The highest art form of humanity. Yeah.

Yeah. Yeah. Publishing papers wasn't doing it for me anymore.

And and I moved to Subsec. And this is the name of the post. There are no new ideas in AI, only new datasets.

One one guy pledged me. But then I found out he was, like, my former student from a class I was teaching. So it I don't think it really counts.

It counts. He's a friend. He's our first supporter.

A pledge is a pledge, man. I'll take whatever I can get. So so the underlying thesis is that whenever maybe I'll I'll lay out this framework first.

So there's this this book called The Structure of Scientific Revolutions by Thomas Kuhn that I read near the beginning of my PhD, which suggests that science kind of moves in these cycles where not very often there's something he calls a paradigm shift, which is like a you could think of it as a zero to one innovation where everything changes. And then it's followed by a rapid period of small innovations, a lot of, like, reapplication of previous techniques, pre paradigm shift techniques to the new era. And then things sort of slow down as we wait for a new paradigm shift.

And I was kind of asking myself what's unique to the paradigm shifts that we've seen in in AI. And and by the way, to me, AI and language models are somewhat synonymous at this point, like, at least for the foreseeable future. I'm I'm certain that will change.

But basically, everything that's pushed the boundary to whatever we have now that resembles, like, intelligence has come from language models. And so those breakthroughs came in in a few steps. So I think the idea is also like a meta commentary on the research community because what everyone wants as a researcher is some kind of, like, cute new method that no one has thought of before that just works on the existing data better than the previous methods.

That's like, for whatever reason, like, the kind of most glamorous thing people think you can do as a researcher like Mamba. It's like it's like a transformer, but it's like more efficient and works better. So that's what a good idea looks like.

And I think everyone wants to, like, find something like that. But if you look at what's actually born out in practice, it's never been like that. I think, like, all of the things that I would consider paradigm shifts in the Kunian sense came from a new technique but trained on new data.

And I think the new data is super, super important. So I wrote it as a series of four paradigm shifts. The first is the emergence of deep neural networks with AlexNet, which I think was, like, 2010 to 2012 era where we just started training on ImageNet, which is like a scale no one had ever seen before of millions of images.

And then the second thing was transformers and BERT and this attention is all you need paper twenty seventeen. The first GPT is twenty eighteen, which is web scale pre training. Like, no one had ever done that before.

No one had ever tried to scrape all the text off the Internet and then tokenize it and feed it into models. Like, it's a crazy idea. And I think, like, we should be honest.

I mean, transformers are incredible, and, like, their staying power is never going to cease to amaze me. They're, like, much more optimal than I think anyone ever knew. And I don't know if we'll ever beat them, but the real innovation is web scale pretraining.

And I think, like, we honestly probably could have gotten this with RNNs. I know, like, the scaling laws paper shows that RNNs have worse curves for scaling, but probably people would have been like, I bet you could have built ChatGPT with a very sophisticated RNN. Like, you didn't even need transformers.

What you need is web scale pre training and the third innovation, which is instruction tuning. And we thought it it came with, like, reinforcement learning. But I think the big innovation of instruction tuning is actually the human preference data, which is like gathering positive and negative pairs of what looks good, like, in terms of a chatbot interface.

And actually, it turns out you can do supervised learning on that too. You can do DPO, which is a form of supervised learning. You don't even need the the InstructGPT techniques.

You just need the data. So, like, I'm sort of playing devil's advocate here, but I actually think this is true that, like, if we had the right datasets, we almost could have scaled, twenty fifteen era techniques and gotten something that looks like at least in StructGPT. Reasoning models are are a little different.

Like, they're I I'm not sure if we could have that with RNNs or not. Like, I don't know if I'm in a position to comment on that with certainty, but they do fall into this framework, which is they really did emerge from a new data source. In this case, it's something like a little different.

It's like verification with symbolic systems, like math, calculators, coding environments, unit tests, like things where we can provide numerical feedback to language model outputs. But we we built a way to to learn that and leverage it to get more intelligent systems.

Speaker 1

using yet. That's a really good thesis. I would say that the researchers I talked to would somewhat disagree.

Yeah. Obviously, this is like a hot take type of thing. And, like, you already acknowledge that RNNs don't don't scale to the same extent.

Like, they operate on the slope of the curve. Whereas, you know, I I guess, like, the amount of data or the type of data or the core insight just changes the order of magnitude of the x axis, right, that that we are mostly working on. But, like, both are important.

The way that I think someone put it to me was an improvement on compute or data efficiency is the equivalent of having a whole bunch more data that we, you know, that otherwise would be a lot more expensive to collect. It's likely that, you know, the frontier models right now are just a collection of hundreds of these small little experiments that just stack up. You mentioned mu Muon in your in your post, which seems to be the atom killer.

Speaker 2

vibes are good. Yeah. Yeah.

And the value of building better optimizers is really incredible. Like, it's just a free launch. You can just sort of, like, plug in a slightly better training mechanism, and then you save, like, a ton of compute and a ton of training time.

That's, like, hugely valuable.

Speaker 1

I think this is cool because I think, like, it's puts us in a mode of, like, if you were ever to ask what comes after reasoning, it has to be something on the order of this. And most ideas are not. Most ideas are not.

And so this is cool in a sense of like, it just jolts you out of incremental thinking into like, what really is missing for the next paradigm. And I don't have an answer. Do you have one?

Do you have do you have candidates?

Speaker 2

Oh, I I really haven't even considered that too much. I guess, like, scaling, reasoning You gotta do the autocomplete.

Speaker 1

Per step five.

Speaker 2

you you gotta show us the way now. We can say it's an exercise left to the reader. But, I mean, the reality is, like, predicting the future is too damn hard.

You know? Like, I I don't maybe it'll be obvious to me in hindsight in five years, but sitting here today, I I really can't derive from first principles what the next wave of innovation will come from.

Speaker 1

Yep. I think we have a few years left. Like, each of these phases lasted for a few years.

Reasoning just started last year, kind of. We got some some juice on this one. Cool.

I think that is a broad overview. We've went way over time, but, like, I really enjoyed this. I guess my parting question for you is kind of a meta one.

So I'm not an academic. I'm, like, kinda self taught. I just read a bunch of papers, and, like, I talk to people all day as part of the podcast.

How do I rate in terms of, like, my my questions? As though, like, could I pass as a grad student? Like or, like, what what's my distribution like?

Maybe I was maybe more industry oriented than academics.

Speaker 2

I think you gotta realize that, like, the only person that's an expert in your area as a grad student is you. And even, like, eventually your adviser defers to you for a small set of questions that fall within your very niche expertise. So, like, I think you're clearly, like, a very good generalist and have, like, a huge amount of background on these topics.

And to the point where I would say you're passing the the grad student Turing test. And I think if you went to a talk, like, people would just assume you have some weird research area of your own that they don't understand. You know?

My research area is AI engineering. Like, I'm I'm trying to, like, making it up as I go.

Speaker 1

But, no, this is super helpful. Okay. Well, I I I that's about all we we prepared.

All the best in your search, all the best in your PhD. I assume, like, apparently, the current PhD meta is you just you do a bunch of small papers. You staple them together and like find like an overall theme, you do a defense, that's it.

Like, that's the journey. Which is kinda cool. Like, I I would I would love to do that.

I you know, I'm too old to do it, but like, it's cool. Yeah. Yeah.

It's it's a great thing to do at at any age. Well, it's better to do a substack. Right?

Then you have, like, people, like, subscribing and pledging along the way and, like, getting validation and, like, yeah, that's that's better than a PhD. Substack is cool. That's the title of the episode, like, Substack Better Than PhD.

But no. Thanks. Yeah.

Thanks for your time. This is really great. Where can people find you?

What are you looking for, really? I'm online. You know, you can follow my Substack and Twitter.

Speaker 2

I tweet pretty consistently, and you're putting papers out. I guess, like, the most meaningful thing, to be honest, is to engage with the research and send me an email. If you really care, that that's amazing.

And, like, I love having those kinds of discussions. And you mean, like, when I'm looking for an end job or out of life? Your research direction.

Speaker 1

Like, what interests you over anything else that's like if if there's someone out there looking who has a problem and is looking for someone to help them on it, like, you are the guy for?

Speaker 2

Underscore. Oh, yeah. Hopefully, if you listen this long, like, I think, like, my research is a lot more well connected than some people's Ph.

Research and that it all falls into, like, a very small manifold of, like, all possible problems. And so if you if you wanna work on anything within that space or that's sort of, like, adjacent to the problems that we discussed in terms of, like, language model Maybe not even language model, but model, weight, and activation information. I think anything that can be described as that is very interesting to me, and I would love to talk.

Awesome. Well, we'll put your contact info in the show notes, and thanks for your time. Thank you.

Shared via Hopper