πŸ”¬ Training Transformers to solve 95% failure rate of Cancer Trials β€” Ron Alfa & Daniel Bear, Noetik

Latent Space: The AI Engineer Podcast
20 April 2026 1h 25m
0:00 --:--
Episode Description
Today, we explain this piece of β€œclickbait” from our guest!TL;DR: 95% of cancer treatments fail to pass clinical trials, but it may be a matching problem β€”Β if we better understood what patients have which tumors which will respond to which treatments, success rates improve dramatically and millions of lives can be saved β€”Β with the treatments we ALREADY have.See our full episode dropping today:Why Big Pharma is licensing AI ModelsTolstoy famously wrote, β€˜All healthy cells are alike; each cancer c

Summary

Noetik aims to revolutionize cancer treatment by addressing the 95% clinical trial failure rate, attributing it to poor patient selection rather than drug efficacy. They develop foundation models trained on extensive, multimodal patient data, including spatial transcriptomics and H&E images, to identify therapeutically relevant cancer subtypes and predict drug response. This approach, exemplified by their OctoVC and Tario models, enables pharmaceutical companies like GSK to better match drugs to patients and rescue trials.

Chapters

Noetik's Founding ThesisNoetik was founded on the contrarian thesis that cancer drug failures stem from poor patient selection, not pharmacology, aiming to build models to understand patient biology and match drugs.
Limitations of Preclinical ModelsThe current drug development process relies on 'Frankensteinian' cell lines and animal models that often fail to translate to human patient biology, leading to broad, inefficient clinical trials.
Intentional Data GenerationNoetik emphasizes generating its own high-quality, multimodal patient data, including H&E, protein stains, and spatial transcriptomics, to create datasets specifically designed for training robust foundation models.
Virtual Cell Models & InferenceNoetik's virtual cell models aim to simulate cell biology in a practical context for drug development, using patient-derived data to predict gene expression and drug response from H&E images.
Bridging Mouse and Human BiologyNoetik uses a specialized mouse platform with barcoded genetic perturbations to validate human-trained models and infer human biology from mouse data, addressing the challenge of animal model translation.
Tario: Autoregressive TransformerThe Tario model, an autoregressive transformer, represents a new architecture for Noetik, demonstrating improved scaling behavior for spatial transcriptomics data, especially with longer context lengths.
GSK Licensing Deal & Industry ShiftNoetik secured a significant licensing deal with GSK for its OctoVC foundation model, signaling a shift in biopharma towards broad model licensing for pipeline-wide application rather than bespoke collaborations.
Conviction in Data-Driven AIThe Noetik team, influenced by their Recursion background, emphasizes the conviction required to generate massive, high-quality datasets from first principles, believing it's essential for developing truly effective AI in biology.

Topics

Cancer clinical trialsPatient selectionDrug developmentFoundation models in biologyMultimodal data generationSpatial transcriptomicsH&E stainingVirtual cell modelsSelf-supervised learningBiomarker discoveryMouse models in oncologyGenetic perturbationsTransformer architecturesAI in biopharmaData moatsScaling laws in AI

People

RJ Hanaki (host) Brandon Anderson (host) Ron Alfa (guest) Dan Baer (guest) Tycho Brahe (mentioned) Kepler (mentioned) Newton (mentioned)
Key Concepts (22)
Contrarian Thesis on Cancer Drug Failure β€” Noetik's core belief that 90-95% of cancer drugs fail in trials not due to poor pharmacology or target selection, but because of inadequate patient selection.
Patient Selection Problem β€” The challenge of identifying which specific patients will respond to a given cancer drug, leading to broad, inefficient clinical trials and drug failures.
Reverse Translation β€” A drug discovery approach that starts from human patient data to understand disease biology and identify new therapeutic targets.
Cancer Subtypes β€” The idea that cancer is not a monolithic disease, but comprises many distinct biological subtypes, even within traditionally classified types, which respond differently to treatments.
Limitations of Preclinical Models β€” The issue with using immortalized, 'Frankensteinian' cancer cell lines and animal models in preclinical drug development, as they often do not accurately represent human patient biology or mutations.
Intentional Data Design β€” The philosophy that in biology, data generation must be deliberate and designed with foresight to support specific model training goals, rather than just collecting random data.
Multimodal Data Generation β€” Noetik's strategy of collecting diverse data types from patient tumor samples, including H&E pathology images, protein stains (immunofluorescence), spatial transcriptomics, and DNA genotyping.
H&E Staining β€” A standard pathology stain (hematoxylin and eosin) that creates contrast in tissue, allowing pathologists to identify cellular structures and classify tumors, and serving as a key input for Noetik's models.
Spatial Transcriptomics β€” A technology that measures RNA expression in a spatially resolved pattern within tissue, providing molecular information linked to specific cell locations.
Virtual Cell Models (Noetik's View) β€” Noetik's practical approach to virtual cells, focusing on models that simulate cell biology in a context useful for drug development, such as understanding drug targets or mapping cell-level biology to patient-level responses.
Self-Supervised Learning in Biology β€” Training models to learn basic biology (genes, proteins, cells, tissue) purely from generated data, without bias from existing clinical records or doctor's notes.
Inference from H&E β€” The ability of Noetik's trained models to predict complex biological information, like gene expression and patient response, solely from a standard H&E image, making it a powerful diagnostic tool.
Data Moat β€” Noetik's competitive advantage derived from generating an order of magnitude larger, high-quality, paired multimodal dataset (H&E, protein, spatial transcriptomics) compared to public or academic sources.
World Models β€” Models designed to simulate what will happen if a particular action (e.g., knocking down a gene) is taken, allowing for counterfactual perturbation analysis.
In Vivo Perturbations (Perturb Map) β€” A mouse-based platform developed by Noetik for highly multiplexed genetic knockouts in cancer cells, injected into mice, allowing for validation of human-trained models in a living system.
Barcoding (Genetic) β€” A technology used in Perturb Map where individual gene knockouts are associated with unique protein tags, allowing researchers to identify which gene was perturbed in each tumor within a mouse.
In Silico Humanization of Mouse Models β€” A method developed by Noetik to use their models to infer human biology directly from mouse data, translating mouse transcriptome outputs into human gene forms to bridge the species gap.
Masked Auto Encoding (MAE) β€” A self-supervised learning objective where a model predicts masked-out chunks of data from revealed chunks, similar to BERT, used in earlier Noetik models like OctoVC.
Autoregressive Model (Tario) β€” A type of model (like Tario) that predicts the next token in a sequence, a training task known to scale well with LLMs, and which Noetik is applying to spatial transcriptomics data.
Scaling Behavior with Context Length β€” The observation that larger models, particularly autoregressive ones, show significant performance benefits when trained with longer context lengths (i.e., seeing more tissue area at once) in biological data.
Foundation Model Licensing Deal β€” A business development model where a pharmaceutical company licenses a pre-trained foundation model for broad use across its pipeline, rather than bespoke project-driven collaborations, as exemplified by Noetik's deal with GSK.
Top-Down Approach to Biology Modeling β€” Noetik's strategy of focusing on solving patient-level heterogeneity and drug response prediction first, rather than building up from subcellular mechanistic models.
References (16)
NOEtic company
PDB dataset
ImageNet dataset
BERT model
Octo Virtual Cell model
CRISPR tool
Agenus company
GSK company
Recursion company
Tario model
Keytruda drug
Merck company
GPT model
Bayer company
Isomorphic company
Leash Bio Labs company
Transcript (107 segments)
Speaker 1

So we basically opened the lab. We hired a team. We got all the instruments.

We started sourcing tumor samples. There was no prior here that any of this would be like zero. We just started generating data and, like, sourcing human tumors, processing.

We built this whole processing pipeline to to get the tumors into, like, these arrays and the formats. So you've got, like, these two week runs where you're processing two slides, and and we're just churning data for months. And we couldn't even train them all off.

So we sort of just built all this, and then then, like, let's say, eighteen months later, hey. I wonder, can we train them all off? And then it was not, you know, like, wasn't obvious.

Yeah. There wasn't really, like, anything major to go off of.

Speaker 3

transformers developed for single cell data. There just, like, weren't really datasets out there that people had been able to develop on. We do a lot of, like, custom model building.

Speaker 2

Hi there. I'm RJ Hanaki, and this is Brandon Anderson. We're the cohosts of the Latent Space Science Podcast.

And today, we're really happy to be in the studio with some of the people from NOEtic.

Speaker 1

I'm Ron Alpa, cofounder and CEO of NOEtic, physician scientist by training. My hobbies are making hot takes about AI curing cancer.

Speaker 3

Hi. I'm Dan Baer. I'm VP of AI at NOEtic.

I'm a biologist by training. Did PhD work in neuroscience, and then moved into CompNeuro, computer vision, self supervised learning, and have, you know, been doing AI research at Noetic for the past few years.

Speaker 2

what is Noetic? Why did you found it? What is the difference between Noetic and the other virtual cell Yeah.

Companies?

Speaker 1

Maybe just start with a little bit of a contrarian thesis, which is really the reason for founding Noetic. It's we all know the numbers that ninety percent, ninety five percent of cancer drugs fail in the clinic. Why do they fail?

So our thesis is they fail not because we're bad at pharmacology, not because we're bad at target selection, you're making the drug. We're actually better at that process than we have ever been in the history of drug development. Most of those drugs fail, we'd argue, is because we're bad at selecting which patients those drugs are gonna work in.

And oftentimes, see trials where there is no placebo effect in cancer. Some patients respond to these drugs. And if you have a patient that responds, that tells you something that there's some biology that that's active there, but you have a problem in in patient selection.

And so really that's the thesis behind OIDIQ is can we build models that can fundamentally understand patient biology from the very beginning and help you position molecules in the right patient population.

Speaker 2

So you're actually using the models partly at least to select the patient cohort, not just so it you can imagine working either way. You could design, oh, I think that this molecule will do well because I know something about the patient population, but you could also say, I think that this patient population is the match for this molecule.

Speaker 1

you can use them on both sides of the equation. So you can use them for discovering new targets directly from the patient data, which people often refer to as reverse translation. So starting from humans and then trying to understand which targets to go after, and then you can use that to develop molecules.

But you can also use them directly on patient data if you have, you know, let's say, a phase two or a phase three trial. You can use these models to understand which patients or or what underlying biology of the patients in the trial is a predictor of response. And we've been doing a ton of that recently.

Speaker 2

Are you doing a lot of, like, rescuing

Speaker 1

trials that had a bad effect?

Speaker 4

and understand whether there's underlying biology that would help us design the next trial. We haven't shared any of that yet, but You'll see us too. So cancer is kind of like infamous in that, like, there are many, many different types of cancers.

Whenever it says, like, cure cancer, that is almost a meaningless vacuous statement. So your point is even amongst cancer or you pick a specific type of cancer and then a subtype and a subtype, there's a bunch of different patient populations that each one of them will respond differently to drugs. And your point is you can figure this out right now.

That way, some subpopulation will do well and respond to this drug when you think generally speaking, the rest of the population would not, even though we have historically classified this as like, oh, what type of cancer, what indication or so on. Yeah, that's exactly right.

Speaker 3

even go further and say, like, nobody actually knows what the subtypes are. There are cancers that originate in a certain tissue like the lung that, you know, have been classified into subtypes based on pathologists looking at them for, you know, more than a century. And, you know, those subtypes certainly have some connection to the real, like, carving nature at its joints, like, what are the actual functional subtypes of disease there.

But our thesis is kind of that if you look at the data, a much richer kind of data, so the multimodal data that we're generating in our lab, we're going to see that actually, you know, what people thought was one subtype of lung cancer is really three distinct subtypes of cancer.

Speaker 1

And that is gonna be critical for figuring out which patients should get which drugs. Yeah. Maybe I'll just go back to you like, one of your first questions.

And and, you know, I was saying, like, drugs don't you know, many drugs fail in patients because we don't understand which patients they will work in in oncology. Why do we end up in that situation? So whenever you make a new drug, you do a set of experiments in cell culture, cells in a dish.

Those cells are often cell lines. These cell lines have existed for forty, fifty years, and and they're immortalized. So they have genomes that allow them to persist that have abnormal numbers of chromosomes.

They have gene expression patterns that don't represent any known cell in, like, the human body, really. These are sort of Frankensteinian cells. They're cancer and dry.

They're mostly cancer. And then and so you can do your experiments in in in these cell lines in a dish, or then you can move these into animal models. And in oncology, you often have, you know, sort of a panel of of different animal models with with, you know, different cancer types that you'll test these in.

And we in doing these experiments, we sort of convince ourselves that that some of these cell lines are, let's say, lung cancer cell lines or colon cancer cell lines, and then even that some of them in in the mouse context are colon cancer cell lines and lung cancer. And then we in the mouse, we implant them under the skin in, like, weird places, and we treat the mice with drugs, and we we see how they respond. But, ultimately, there's a big gap because they don't translate to to patient biology most of the time.

So these cancer cell lines, most of them don't even, you know, even if they are derived from a colon cancer, they don't even have the mutations that human colon cancers have in many cases. And so and pharma has done this for, you know, twenty, thirty years where you you develop a drug, you test it against, you know, hundreds of these. It's not an art experiment.

We can you can send this out to any CRO. They'll test your drug against hundreds of of different cancer cell lines, and then you can first sit back and say, okay. Well, which of the 50 colon lines responded to my drug and which of the 50 ovarian cancer line?

And you could try and map that to human biology, but the problem is these cell lines as an abstraction do do not relate in any way to to human, you know, patients. And so what happens is ultimately, no matter what you do preclinical, the the molecule gets in the clinic, the clinical team says, look. We don't really know how to design this trial because none of the data that you've produced gives us any insight on on which patients to run.

So we're gonna we're gonna basically enroll an open label study. So we're gonna enroll all tumors, all patients that are, you know, enrollable in the in this trial, and and we're gonna see where we get signal. Imagine doing that in an early phase trial where, let's say, you have 50 patients, and you're you're trying to do you test different doses, and you don't really know the dose of the drug, and you don't know what the safety margins are, and you're also trying to figure out where is my signal.

And then what if I told you that, let's say, in in just lung cancer, hypothetically, let let's say there's only 10 different subtypes of lung cancer. And you don't even know if it's lung. It could be any.

So, you know, this is what happens. And oftentimes, you get to the end of these early stage trials, and you don't see very many responders as you would expect statistically, and then these molecules get canceled.

Speaker 2

that your Noetic system, you help the pharmaceutical company to characterize. We expect that people with a certain genetic profile or even transcriptomic profile will will respond to this drug.

Speaker 1

from the patient and you say, yes. This is a match or no. Is that the sort of grand vision?

Yeah. I mean, I I would say we are even less biased than that. We are saying, okay.

Well, we want the model to learn, let's say, from lung cancers. We want the model to learn, like, how many different therapeutically relevant subtypes of lung cancers are just from self supervised learning from the data. And those subtypes could be driven by large genetic changes.

They could be driven by, you know, immune changes.

Speaker 3

is learning in the process of training. Yeah. And we do see, you know, different types.

I mean, feel free to contradict this, like, as the actual doctor here, but, like, you know, the the biomarkers that, you know, people have been using are, you know, biased towards simplicity, you know, does the patient have this particular mutation, sometimes, like, stain for this single protein, or, you know, do transfer comics like to to look for a particular gene signature. But like, there's no reason to think that biology or like biology of cancer is that simple that you're gonna capture, you know, most of the meaningful variation with such simple biomarkers. And, you know, most of them, they have like weak correlations with, you know, clinical success.

But the hypothesis really is here like, again, if you were to card nature at its joints and figure out what's really going on is there, you know, these five subtypes that the correlation there between which patients you give a particular drug and whether you have success is much much stronger than if you're forcing yourself to go with these, like, very simple biomarkers.

Speaker 2

You mentioned the lab. You do a lot of data generation in the lab. So why do you think that that versus using existing public repositories or whatever is appropriate?

Speaker 1

Yeah. We generate all our data in in the lab. Everything from sourcing tumor samples themselves to processing them and generating the data.

Maybe another another hot take I have just in AI and bio is you're sort of not at the order of magnitude of data that you are in other spaces of building training models. And so it becomes really hard to brute force these problems just by collecting data. We have a couple pretty good examples of where someone has designed a dataset.

So PDB was designed and has been built over the past fifty years or so, and so it's not an accident that that dataset exists. Someone decided that we are going to design this dataset. We're gonna collect this data over decades and decades, and then with the intuition that potentially this would help solve protein folding down the road, and and and it did.

So it's not just that PDB is a bunch of random data that, you know, has been that people have organized from from the web. I think that in bio, you really need to be intentional about the data that you generate and how you generate it, and have some foresight around, well, what are the models we're we're gonna wanna train, and what are the models gonna need to learn from, from the very beginning? So that's why we've taken taken this approach.

Yeah.

Speaker 3

that, you know, neural networks can do better than other methods on object categorization. ImageNet is at least the the part of it that people were developing models on is 1,200,000 images, very carefully curated. These are high quality images, not like random images from the internet or like multiple data sets cobbled together.

And labeled? Yeah, and labeled. And I think with the data that we're generating, we're around that scale right now.

But, you know, of course, people have gone much much larger in image data sets and language data sets, text data sets, obviously for LLMs. So we think that we need to get the data up to that scale before we can really see the meaningful progress on the algorithm side. The scale of language data.

Yeah. Language is really the only modality where people are seeing these very impressive scaling results. And, you know, part of that has to be just the scale of data that's there and that the models are trained on.

That can't be the only thing because, you know, there's a lot of, like, video data as well.

Speaker 4

like, thousands of hours of video data and, you know, haven't seen kind of the scaling results that you have in language modeling. But having the right scale of data is necessary if not sufficient to, like, really make progress here. Kind of for a contrariate take to that?

Sure. So, I mean, there's this whole concept about the giant frontier of a limbs in jerk of AI. It held, like, certain regions that could be really good at solving some problems and then remarkably stupid at solving nearby problems.

And maybe the arguments with happening is that a lot of these twin tier models are just becoming massively like, everything is becoming in distribution. Like, if everything starts out OD, if you just get more data, it now becomes in distribution. Is it possible that for biological systems, because these are their underlying physical processes here, that you can basically make things more in distribution earlier and that you can actually cover the space?

I kind of have some follow ups with PDP, but maybe I'm just curious at this point.

Speaker 3

Yeah. I mean, I think it's a good question is, like, sort of how much data and what kind of diversity do you need, like, in biology to solve, you Say, like the drug translation problem, like figuring out which drugs are gonna work in which patients. My intuition from working in biology, like, for a while is that we're still pretty far from that, like, you know, we're building data sets that are focused on right now cancer, and, you know, have generated data from thousands of patients in a few major cancer subtypes.

But there's like every other disease, there's healthy tissue, there's even other species, you know, there's a lot of biology to learn, especially if you think about it as we have to learn kind of the spatial and functional patterns of tens of thousands of genes, tens of thousands of proteins, how their spatial arrangement contributes to the function of organs, and so forth. You know, my hunch is that biology is like pretty complex and that we still need to generate a lot more data. But, yeah, I I I don't know.

Yeah. But as a cancer company, do you think you could actually do this hypothetically for cancer? I mean, for at least some, you know, subclass of cancer?

Definitely. Yeah. I I think that we've done experiments that suggest that, you know, if we can generate data from several 100 patients in all of the major cancer indications and some of the less major indications, that that will result in a model that can generalize pretty well to kind of any type of cancer we would throw at it.

Speaker 2

Backing up, what is the data you're collecting? Because it my understanding is you use some pretty specialized instruments and gathering very specific datasets. So how did you come to that that decision about how much data, how much to spend on it, and what types of data?

I'll give a hat tip to my previous employer, Recursion.

Speaker 1

the very beginning, and a lot of what we were doing in the early days was figuring out, like, the things we didn't understand about the datasets and figuring out what the problems would be in the dataset. So batch effects, control side of orient samples on plates, things like that. Flash forward to founding of Noetic, started the company, you know, already with some with some principles around how we should think about building the dataset.

What are some things that we know matter? So for example, over many years, we learned that images are actually a really powerful dataset for machine learning for many reasons. One, they're scaled.

So we can put patient samples on slides, and on a single slide, we can capture many patients worth the buyout. The images themselves are very rich sources of biological information beyond that. Now we have a very information dense modality, and we can decrease the cost of data generation so then we can increase the amount of data generation over the whole dataset.

And that's always been a really big benefit to image based modalities over, let's say, sequencing where every time you run a sequencing run, you're basically your aunt is, you know, a patient's first. That was one one way to think about it. The other was how do we design these datasets so we can control for things that we know are gonna be important, such as batch effects.

So for example, if I have a slide, we do a let's say, a spatial transcriptomics run on that slide. You stay in the slide, do a bunch of, you know, wet lab processing, you put it into a machine, you get data out. If you do that on two different days, there are going to be different variables that impact the data.

That's gonna be a large source of variation in datasets. So you wanna be able to control for things like batch effects. So really, you want you have more patients represented on multiple different slides so you can process them different in different batches.

So you wanna be able to control for things like this so you can go downstream and look at the data and say, okay.

Speaker 2

staining batches? So you're you're actually taking different pay one patient, and you're spreading across multiple slides so that you can get a like a it's sort of a calibration across the slides? Yes.

Our data looks very different than anyone in the space of generating data on histology or digital pathology types of specimens.

Speaker 1

So we we receive a sample. We sample those samples dozens of times to build these arrays, and each array has, you know, hundreds of different patient samples randomized, and every patient is represented on multiple different arrays. And so we're getting a lot of different representations of each patient that we're sending through the data process and pipeline.

And then that lets you downstream be able to answer some of these questions and control for some of these barriers.

Speaker 2

You'd mentioned some terms I just wanna define for people at spatial transcriptomic.

Speaker 1

Yeah. What is that? Yeah.

So what be I mean, this was your first question. So what are the data types? Yeah.

So you just sit back, and this is not my background in terms of spatial. Again, everything we did on your previously was cell biology in a dish. If you just sat back and you say, okay.

I wanna train a foundation model that understands human biology. What does that mean? What will be how would you go after that problem?

And and that was really the starting point for the company. Said, okay. But from first principles, how would we do this?

So you probably want tissue level biology. You wanted to understand tissue. Cells are organized into tissues.

You probably want some modality that is relevant in in clinical use so you can relate clinical data to to what your models are learning. That's why we generate pathology H and E. So that's, you know, what every patient gets a tumor removed, and then they get the stain on H and E, and that's what the pathologist I can't explain what H and E is on Basically, two two different dyes, hematogenous and eosin, and it, you know, really just creates a contrast over the tissue.

So you've probably seen these, like, purplish pathology specimens. So pathologists can look at those, and they can identify different cellular structures, and they've used those to classify tumors based on, you know, the classical classifications of, you know, adenocarcinomas, small cell carcinomas, things like that. But basically cellular structures.

Speaker 4

Okay.

Speaker 1

based on? Based on, yeah, pathology on your classifications. And this is what every basically, every tumor, you know, that gets processed in the hospital will get this HEE state.

And it's how the pathologist typically classifies a tumor from from the first level. So okay. So you want that.

You probably also wanna understand cell types. It's really how to understand cell types from just that stain because it doesn't reveal that much that a human can use to classify cell types at least. So you can say, well, I I wanna know whether there are immune cells and different subtypes of immune cells.

Speaker 4

We wanna have some layer of cell biology. Okay.

Speaker 1

the immune response dictates whether or not, like, it'll be you have an effective treatment or It's like the the immune environment of the tumor will be a core. We know it's a core constituent of of of whether a patient's gonna respond or not. So you wanna know, okay, you wanna give them all this.

So the models are gonna get this tissue level information. There's not enough cell level information in there for them all to learn enough cell biology about different subtypes. So we also wanna present it with some cell level information.

So we use protein stains, so standard neofluorescence. So you basically use antibodies against small set of of cell markers to label your different T cells, B cells, your standard subtypes of cells in the tumor and microburden.

Speaker 2

just to for the those who are familiar, the stain on the antibiotic antibody has a fluorescing protein when you hit it with a certain frequency of light, then it fluoresces so you can tell the antibody bound to a certain protein, and now it has a fluorescing

Speaker 1

guillotine attached to it. Yep. And in terms of the data, so from from the from the tissue layer, you have an RGB image.

From the next layer, you have multichannel image with each channel representing, you know, let's say, one color. And so, for example, certain immune cells are each in in a different channel, so you have this multichannel image. Now okay.

So that's great. So we've got tissue, we've got cells. But if we actually wanna make drugs, we need some some type of molecular information.

We need to tie all of this down to what's happening in the genome. What is the cell doing? What are the mechanistic principles of of the biology?

So then we get spatial transcriptome. So that that's spatially resolvable RNA. So DNA transcribed into RNA, which is translated into proteins.

So we get basically the RNA in a spatially resolved pattern for the same cells that we're seeing all of these other layers. So now you have between a thousand or 19,000 different genes.

Speaker 2

that are spots of where though those RNA are and in which cells. And this this one works a little bit similar to the how we talk about protein where you have a segment of RNA, and then you have a fluorescing protein, and usually there's some sort of combinatorial thing. So you have if you see these four colors in this amplitude, then that means this gene because there's they're right next each other or something like that.

So for the detection method, you're basically binding a probe at each one of those RNAs, and then you're cycling it. And it takes weeks to run one of those assays. So you're cycling the machine.

Speaker 1

and you'll get a signal for each RNA species. Now at this point, you you now have basically this very rich data layer where you have the tissue, you have the cells, and you have the molecular information, and you can use all of that to train the model. And so we you think of it as yeah.

If it's essentially the central dogma, if you will. Yeah. And we also have DNA.

Speaker 2

alterations in these tumors. Right. So you get the stack of images basically that you can train models on with understanding the expression of genes and the proteins that are being expressed at the time that the sample is taken all in the image information, and then you can train your models with that.

Yeah.

Speaker 3

is, like, particularly dense because if you think, let's say, there are 20,000 genes in the genome, Now, you know, we're running assays that are detecting nearly all of them in a single sample. So you can think of one of those data points as an image except instead of being an RGB image that has three color channels.

Speaker 2

Now all of sudden it has like 20,000 color. So it's like a very meaty computer vision problem to try to look at those data and figure out what makes patient a different from patient b, and then go from that to which drug is gonna work and which one. And so you you have a hot take about virtual cell?

Like, I wanna understand how okay. So you, you know, you have this big pile of data that every single sample has a massive dataset with it, and then you have many many samples. So how do you turn that into useful knowledge?

Speaker 1

Maybe just what is a what is a virtual yeah, everyone's always, you know, asking that question. I think there there are really two ways to think about it. You know, one is we want to be able to simulate all the biochemical processes in a cell.

So we want to have this sort of comprehensive foundation model where we understand, you know, if if some signal from outside the cell, interacts with the cell, then here are the millions of intracellular chemical reactions that are gonna happen, and you could sort of predict them, yep, from the model. So that that's one view. I think that's interest it's sort of an interesting intellectual pursuit.

I don't think we have all the modalities of data that you would need to solve that problem. I tend to see the virtual cell problem as something more practical. We're trying to make drugs that work in patients.

So from a virtual cell perspective, really, what we want to do is understand cell biology in in some heuristic that's useful for for making drugs. And the heuristic could be, you know, a way to under understand drug targets or a way to, you know, map your cell level biology up to patient level biology. And so the way we've designed these first virtual cell models is really just to simulate the biology of a cell in some context.

And the biology of that cell being, you know, let's say, the the cell being in some context and the output being, you know, the the transcriptome in that context or, you know, the protein in that in that context. And these types of of, you know, input output relationships allow us to to essentially design experiments.

Speaker 3

and allow you run some simulations in that that regime. Yeah. I mean, I think what most of the things that people are calling, like, virtual cell models right now are focused on single cell gene expression, so transcriptomics data, RNA data, and they're largely geared toward the problem of predicting what's gonna happen to the transcriptome.

So the set of genes expressed when you hit cells with either a small molecule, a drug, or a genetic perturbation. And typically, this is cells grown in vitro, like either cell culture or primary cells, something like that. I think that Genetic perturbation being where I Like.

Speaker 2

and see how that impacts the expression of the de various RNA.

Speaker 3

So and I think my view, and I think Ron shares it too, is that like, maybe of interest in some cases, but the problem we're really trying to solve is predicting what's gonna happen in a patient. And you're just modeling data that comes from a patient is, in my mind, much more likely to translate to what happens when you give a patient a drug than something that's happening in cell culture.

Speaker 2

Is there other clinical data that you're pulling into the model besides the actual so you're calling it context of the cell just the surrounding cells, but it is there other

Speaker 3

this drug caused a bad reaction kind of stuff? Yeah. I mean, we're pulling in data from the entire patient, so not just, you know, the very local neighborhood of the patient.

So far, we haven't done much integration of, you know, like electronic health records or, you know, other information that one could get about the patient. And that's pretty intentional. Like, we really want these models to learn basic biology.

Again, like the central dogma, not just the central dogma, but, you know, the basic biology of genes, proteins, cells, tissue in a self supervised way. So purely from the data that we're generating and not be biased by, you know, what the doctor wrote about that patient. Because, you know, our thesis is kind of that, like, most of the therapeutically predictive and important information is not contained in those very small number of, you know, patients who have been treated with a given drug and whatever the doctors thought was important to write down given the state of knowledge at that time.

So it's much more about trying to discover what's really there in in patient biology than go based on the text that people have written about it. So you have this self supervised model.

Speaker 4

You eat a lot of data. You have essentially some clusters of patients now. How do you translate those clusters of patients to making decisions?

Like, you go to a pharma company and you say, we can repurpose or we can suggest this subtype should be the focus of your phase two trials. Like, what is the process for that? What data do they need to provide to you, and how do you translate your models?

So it depends on what the problem is. I think it's important. So I wanna mail the backup.

Speaker 1

One of the more interesting aspects of these models is they are useful for a broad array of use cases as, you know, as we were talking about from the very beginning. So you as the pharma company could say, okay. Well, I have this molecule, and the target of the molecule is x, and I wanna design my clinical trial.

The molecule has seen zero patients so far. Yeah. All I know is the target, and, you know, some biology around the target.

So we can run simulations using the models and our our cohorts of patients. And let's say, if we were to look at, you know, in lung cancer, we can run simulations around the target and ask, okay, which sets of patients here would this target be be important in across a cohort of, you know, lung cancers and colon cancers and, you know, across all of oncology. And you might see and we see this some sometimes.

You might see that, you know, your target probably don't wanna put it in lung cancer. Maybe you wanna put it in ovarian cancer because it's not really important in lung cancer. Yeah.

What are you simulating here?

Speaker 4

drug is expected to knock down this gene and therefore, it will result that you want to look for clusters where knocking down this gene

Speaker 3

inhibits tumor growth rather than enhancing tumor growth? I mean, that's certainly one one way we could do it. There are other types of simulation where you might just wanna ask, like, if there were immune cell here, like a T cell which is responsible for actually killing tumor cells, what would happen to it or what genes would it express or what proteins would express in this particular patient's tumor microenvironment.

And, you know, that's what we've called like these virtual cell simulations. Like, we have a model called Octo Virtual Cell that does this, and that can give quite powerful answers to the question of, are these drugs gonna work in these patients? Because you might find like, actually, as Ron was saying, the thing that this drug targets is just not important in this particular patient's tumor in that there's not like it's not gonna have any effect on the T cells or the macrophages or some other cell type there.

Then, you know, there's the type of simulation you alluded to where you can ask the model what would happen to this patient's tumor if you were to knock down this particular target gene or its protein product. And you might be looking for cases where the model predicts that removing that gene or that protein is gonna have a large effect, like either increase the immune system function, its ability to fight that tumor, or decrease the tumor's ability to grow, or some other readout that you think is correlated with clinical success. I just wanna call out maybe like the the simplest use case is the one where there's like a company that has a drug and they've given it to some patients, and we know some of those patients responded.

And then it just becomes like a question of, like, has the the space of patients that the model has learned via self supervision tell us that all of the responsive patients are in one of these clusters and not the other nine clusters or something.

Speaker 2

hypothesis that this is the right cluster. So that's the scenario where you would sequence something. What would you collect about those?

So you have a cohort Yeah. Responded and one that didn't. Yeah.

Speaker 3

something Ron mentioned earlier, which is this type of data called HNE. It's a stain, the standard pathology stain that makes these, you know, pinkish and purplish looking images. Right now, what we do is we've built models that are trained on kind of all of the multimodal data we generate.

But then once they're trained at inference time, all they need is an image of H and E. And that could be something that we generate in our lab, or it could just be, you know, a digital image that they have from a trial that was run years ago. And the reason that that is so powerful and flexible is, again, because H and E is kind of like the lingua franca of of pathology and especially oncology.

So almost every patient who's been given a, you know, clinical stage drug is gonna have that.

Speaker 2

and say these H and Es live in this this part of the latent space and these H and Es do not. Yeah. Exactly.

Speaker 3

way we've gone further than that even is given the H and E, they can say, I predict that these genes are expressed at this location in this patient. So not only do we have these clusters, these embeddings that say, you know, all of the responders to this drug are over here, all of the non responders are over there, But we can actually see, okay, for the responders, these are the genes that are expressed much more highly or predicted to be expressed much more highly in the responder cluster versus the non responder cluster. And so that adds a major, like, level of interpretability there because, you know, we can see things like, okay, like, good, the responders are actually expressing the the protein target of this drug.

So we would be worried if that weren't the case, but, you know, we can see it is. On the other hand, we also see that, you know, the biology is very very complicated.

Speaker 2

you know, what is predictive of therapeutic response. Yeah. So I have, like, a million directions that I'm gonna go here.

They H and E, that actually gives you a pathway to a diagnostic then as well. Exactly. Yeah.

Right. Yeah. Yeah.

And so that you you can imagine after the drug hopefully makes it to the market, then a doctor says, oh, you have cancer. I'm very sorry.

Speaker 1

and then we're going to put in the model and it's it says, oh, you know, this one won't work for you, but this one won't. That's right. And you you can so we're we're using the same the same approach for actually, today, we're we're looking at many different mechanisms from different collaborations that we're we have in place.

You know, one of them we've announced with a company called Agenus. These are all different mechanisms. The input is still H and E using you know, and some of the same indications.

So using the H and E, we're asking whether Dara Hay works in in some sets of patients, whether drug b works in other sets of patients. And so you can take that, you know, to its natural progression and say, well, okay. If you can use that same input, just H and E, for, you know, experimental drugs, why not use it also for drugs that are on the market already?

In a sense, the same assay can they can be very predictive across many different cancers and many different potential therapeutics.

Speaker 2

There are model lots of models that take HNEs and go to gene expression out there, open source, whatever. They do, you know, so so. I've read in Twitter, your Twitter feed, and whatever that you feel that you have a data moat.

Right? And so why is Noetic's model better?

Speaker 3

Sure. I mean, I think, you know, the scale of data that we've trained these models on is, like, you know, pretty different from a lot of what's out there. Like, the reality is there's just not that much of this kind of paired H and E plus other data modalities.

Typically, you know, there are some datasets generated by academic labs, others where, you know, they might have maybe, like, a 100 or a few 100 patients worth of data with paired spatial transcriptomics. That might even be an overestimate. In comparison, we're generating these data that are, you know, multiple patients per slide, individual patients distributed across multiple slides.

We've generated now, you know, more than a 100,000,000 cells spatially resolved and spatial transcriptomics. That's all paired with HNE and protein as well. At least in order of magnitude larger than any of the other datasets that we've seen out there.

And I think that makes like a pretty enormous difference. I mean, we've seen with our own models that if you drop down to 40% or 10% of that data used in training, models get a lot worse. And they especially get worse at kind of generalizing to other types of cancer from the ones that they've been trained on.

So I I think that's a big piece of it. I also think that, you know, the algorithmic side of it is important. You know, we've developed custom architectures specifically for training on this multimodal data.

And again, my background is in computer vision, and specifically in self supervised learning there. And so we've tried to develop, you know, self supervised learning approaches for these data that are really adapted for solving this problem of, you know, figuring out what is different in one patient versus another, and then simulating what would happen if you were to, like, knock down a particular gene or protein or something. So this is why we call these world models where we're trying to build models that can simulate what's gonna happen if you if you take a particular action.

I think that's another another big differentiator for these models.

Speaker 4

as well is probably a third one. It's funny because you were just talking about how one of the other strategies people take for this is to, do perturbations on cells and then watch the response. And, you know, your experience plus, like, your strategy here is you can simulate this sort of counterfactual perturbation idea without even having to collect the data to do that.

And you can see this. Well, there's yeah.

Speaker 3

piece that we haven't talked about yet, which is actually we are running perturbation experiments except they're in vivo perturbations, using a platform, based in mouse. We have another platform where we are, it's called perturb map. Ron, if you wanna describe any of it.

But basically, have there's a platform for generating highly multiplexed knockouts of individual genes. So the same kind of like CRISPR knockouts that people are doing for individual cells in vitro, except when we knock out a gene in a cancer cell, that cancer cell gets injected into a mouse. It's barcoded so we know which gene was knocked out, and it's being injected alongside, like, roughly a 100 other cell types with different genes knocked out.

So you end up with mice that have tumors that are barcoded, that have a 100 different genetic perturbations in them. We can actually use that to validate our models and ask our, you know, what the models are predicting in humans via simulation actually borne out when you do these perturbations in a mouse system. Sorry, because there's a lot to go with it there.

Bar barcode. Yeah. So sorry.

Barcoding. This is a technology in which an individual gene is knocked out with, CRISPR, but also this introduces a set of protein tags in that cell that get expressed. It's a combinatorial code.

So gene x might have, you know, proteins a, b, and c. Gene y, when it's knocked out, has proteins d, e, and f, and we can tag those proteins or label them with antibodies so that when we go and look in the mouse, we know exactly which gene was knocked out based on which of those protein tags were expressed.

Speaker 1

encoded on them. Yeah. Exactly.

And, I mean, the the system's designed so everything that we're doing here is tissue level. You could be in vivo of, you know, tumors that came for you in that are in the form of the tumor that are, you know, the whole tissue. And then here and then this mouse system, you have hundreds of tumors in the lungs of a mouse.

And if you look at these images, it's a mouse lung with, like, literally hundreds of tumors in it. And each tumor has a distinct biology that's driven by the biology of the knockout of the gene that's being perturbed, and we can capture basically the the biology of each tumor in a spatially resolved way. So what you can see is, okay.

Well, we have a bunch of tumors in human that we have we have certain tumors in humans, let's say, don't have immune cells in them. And so those tumors are very aggressive, and they don't respond to immune therapies. You can generate those same tumors in this mouse system, and, again, they don't have immune cells in them.

And you can do it genetically, so you can start to map kind of the gene, the causative gene relationships between these different immune or just broadly tumor genotypes or biological profiles, if you will, to to what you see in a human. And then you can treat those mice with drugs, and you see how, you know, hundreds of tumors in a single mouse responds to treatment with one drug, or you can treat many different you know, let's say, 50 different knockouts across a panel of mice with 50 different drugs, and and you can start to build this intersectional pharmacology and, you know, genetic experiment.

Speaker 4

On Twitter and in various places, I've heard you say, noetic is no no cell lines, no war bottles. Maybe you even said that, you know, a few minutes ago. And then then we just said we have mouse ball.

Yes. Yes.

Speaker 2

cell,

Speaker 1

like, to In the lungs. In the lungs, not under the sky. So yes.

So, you know, fundamentally, we think it's really important to build models that are trained on human data, and we are sourcing all these tumor pick tumors to build, you know, human centric models. So that is also that is true. From the very beginning, we have asked this question of, you know, let's say we wanna develop a drug from the very beginning, and let's say the FDA and I know things have changed a little bit with the FDA, but let's say the FDA wants you to have some data in an animal that says your new mechanism works in some animal system.

What do you do? You're kinda stuck because you've now generated arguably the best data that you can in the human system. And then the FDA says, well, cool, but does it work in mouse?

How does it work in the mouse? And then so you have to back into this system that it doesn't translate. And so from the very beginning of the company, this has been, you know, sort of a question.

And so we started, you know, at the same time we started generating the mouth to the human day, we started building this mouse platform with the aim of drawing connectivity between these two systems. And so we focused on a platform. We wanted a platform that, one, allows you to to map a diversity of human tumors because we know that if we just run a mouse model with one tumor, that tumor has no connectivity.

So in the mouse system, we wanna have diversity of tumors, and we wanna see a mapping of diverse tumor biology to the tumor biology that we're seeing in in the human across many different mutations. So we licensed this system and have been building it so you can see many different perturbations that produce a lot of the tumor biologies, plural, that you see in the human. And then we also want to be able to get from this mouse system to biologically relevant, let's say, targets or genes in the human as well.

So one of the fundamental problems in mouse systems is we share many genes with mice, but there are a lot of genes in biological process we don't share with mice as is obvious. And so oftentimes, run into these when you're developing drugs. It's okay.

You have a target, you know, you have some biology that works really well in mice. Maybe that doesn't even exist in humans, or, like, maybe that pathway is, like, useless in humans. So one of the things we've started to develop that we'll share more about soon is a way to use one of these models to essentially infer human biology from the mouse directly.

And so we're in silico humanizing the mouse. So all the outputs in terms of the tran the transcriptome from the mouse are in the form of of human genes. And so when we read out this mouse system, we are reading out in the form of of human or alkenhout.

How do you validate that?

Speaker 4

impressive claim if you can do it, but, man, it's

Speaker 1

seems like a tricky validation task. In my experience, both here here at Noetic and my previous employer, I could say recursion. Recursion.

Like, a lot of a lot of the, you know, a lot of the approaches you're looking for when you're building these types of models is you're trying to ask whether the models are recognizing biology that you know to be true. So for example, in the human context, we know that twelve percent of patients with lung cancer respond to immune checkpoint inhibitors. Do the models recognize those patients?

Can they recover those patients without training? Like, in cold. Yeah.

Yeah. And and we see that. And then when you go look at those patients, we see the underlying features of those patients maps to what we know about those patients in, you know, the client.

In the mouse system, we have control genes. So we ask, if if you look at the mouse tumor embedding space, do the tumors that should be really cold look really cold from the human inverts? Cold in the sense we have, like, they don't have immune cells.

No mice. OAAT. Yeah.

Yeah. And then hot in the sense of, like, clots and immune cells.

Speaker 2

and then, you know, the more of these examples that you know to be true that that work that you see, the more confidence you have. Obviously, when when you're into the regime of something very new, it's it's still uncertain to some extent. So the bridge is sort of the bridge between the mouse and the human is you build a world model on a human, you build the world model on the mouse, and then you say, what are the parallel structures in the two latent spaces?

Is that kind of the intuition here?

Speaker 3

we've trained models on human HME, spatial transcriptomics, etcetera, and then are just inferencing them on mouse H and E, which is easy to generate. And apparently, mouse H and E looks enough like human H and E that the models think is perfectly valid H and E makes predictions about is this like immune hot, like immune infiltrated versus cold versus fibrotic versus some other tumor phenotype. And those predictions are accurate.

So, you know, these are like some of the controls that Ron mentioned. So, you know, we know that in mice and humans and everything, if you knock down tumor cells ability to present antigens to immune cells, you know, those are very cold. Like immune cells are nowhere near those tumors.

And, you know, that's exactly what we see in the mouse and that's exactly what the models, the in silico humanized models predict. And, you know, then there are other examples where again, we're recovering the biology that we expect to see there. And then there are findings that are novel, but also make total biological sense.

For instance, we have done knockouts in the mouse of, let's say, half a dozen genes that are all in the same pathway. So you might predict that knocking down those genes are gonna produce the same phenotype because they're on the same pathway. And that way What is the pathway?

Yeah. So a pathway is like protein a signals to protein b signals to protein c, and, you know, there's like a chain of events that leads to the cell having some behavior, you know, changes in its metabolism, its growth, etcetera. So these are I don't know if you've ever seen these crazy looking protein signaling diagrams that, you know, make you want to stay away from biology.

But, you know, like, you know, people have, you know, worked down a lot, and they know that these two proteins interact physically and signal to each other and so forth.

Speaker 2

binds to this protein and that causes it to upregulate a gene that causes another protein to be formed blah blah blah until Yeah.

Speaker 3

meaning the cell changed the way it looks or the Exactly. Yeah. And so, you know, based on decades of biological literature doing experiments on these, There's a very strong biological prior that if you hit gene A, gene B, gene C, and they're all in the same pathway, you should get similar phenotypes.

I mean, is kind of how Yeah. Like old school genetics was done. And we see that with these in silico humanized mouse models, which is amazing to me as a biologist that you have a model that's trained on human data, then you show it some mouse histology, and it's able to say these five different tumor genotypes all look like they have the same phenotype, and lo and behold, there are, you know, five genes that are in the same pathway.

Speaker 2

So you guys, switching gears a little bit, because we wanna talk about models in on the on Latent Space podcast. You guys recently, there was an interesting blog post, Tario, model. It's a some transformer based model.

Do wanna talk about that? Sure. Yeah.

Speaker 3

new model architecture that we developed post sort of the first virtual cell model, OctoVC, that we developed. So Tario, this model is, just a different transformer architecture. One major difference between it and, you know, our prior models.

I guess, if this is a model podcast, this is getting into, like, the self supervised learning objective. So, you know, for a while, including with OctoVC, we were training models on what's called the the masked auto encoding loss function or objective where you have a piece of data, you chunk it up into small chunks, you mask out some of those chunks and the the training task is the model has to predict the masked out chunks from the revealed chunks, so like BERT. Yeah, exactly like BERT.

What are the chunks? Because this is multimodal and, like, I would imagine the different channels contain wildly different levels of information.

Speaker 4

masking in OctoVC if I'm Yeah. Oh, yeah. So And I was like, that was kinda surprising because when you have, you know, 19,000 channels and maybe some of the channels are fairly, like, most of the signal is fairly sparse Yep.

Speaker 3

or you really risk, like, just throwing baby out with the path. Yeah. What are the chunks?

That totally depends on which modalities we're talking about. So spatial transcriptomics, one chunk or one token might be the level of expression for a particular gene at a particular spatial location. For protein images, multiplex protein images, again, it might be, you know, the image patch for that particular protein at a particular location and so on.

And, you know, for, like, histology images, again, those are usually just patches of the image. So pretty standard, like, vision transformer style. The masking and the maybe surprising result that like, you can and actually need to mask out large amounts of the data to get the model to learn anything interesting.

If you ran the hypothetical where you only mask out like 10% of the the image, you know, maybe more like BERT, for instance, in language modeling, what do the models learned? And, you know, they learn these kind of like boring behaviors, like how to, like, continue an edge a little bit, you know, between two like regions of an object or something. So they can learn that task very well, but they don't end up learning anything about sort of the holistic structure of the image data.

And we found pretty early on at Noetic that the same thing was true with these multimodal like transformers where if you mask out a lot of it, there are actually pretty strong correlations between where protein A is expressed and where protein B is expressed, and forcing the models to learn them is really what gives it this predictive power. And so Karyo though Yeah. Is an is a is a auto aggressive model.

Yeah. Exactly. So, yeah, that was gonna be the the tie in.

So, you know, prior models including OctoVC were of this masked auto encoding style training objective. Tario is an autoregressive model, which if you think about it is kind of a particular choice of masked auto encoding except, you know, instead of randomly masking on front of the data, you're always asking the model to predict the next token in a sequence. We know that this is something that scales very well with LLMs, like training on the next token prediction task.

And with still an open question, how do you get models of other data modalities to scale the way that LLMs have scaled? Tario was not actually our first attempt, one of our subsequent attempts to bring that autoregressive, like, next token prediction task into modeling spatial transcriptomics data. We found that when we use this architecture and this task, we started to see, you know, much better scaling behavior where bigger models and especially at longer context lengths were really outperforming, you know, the smaller models at at shorter context lengths.

Because they can see further an image? Yeah. That's probably a big part of it.

I think like the you know, there's actually a pretty subtle, but very interesting result in that blog post with Tarja, which is that you only really see the benefits of using larger models when you're looking at longer context lengths. And here, longer context really means, again, like you're seeing more tissue at once, More area at once. And I'm not like super deep into the language modelling literature, but I don't know if there's an analogous thing with like language models where like you only see these scaling behaviors at at longer context.

So it could be that we're finding here is that like with patient data, you really do need to incorporate sort of more of the patient spatial context to really get the models to learn these more complicated nonlinear patterns in, you know, the spatial transcriptomics and take advantage of it. Is it possible part of this is because you have some number of low expression genes and that the that the bit behavior is driven entirely by some under better in your modeling of low expression genes? Yeah.

Definitely possible that, like, the more context you have, like, the more likely you are to catch kind of these low expression but highly predictive genes, etcetera. I would guess it's a combination of that and larger area. Like we've done some experiments just like comparing model with the same amount of context but in smaller or larger areas, and there definitely seems to be an advantage to looking at larger regions of tissue as well.

Speaker 2

you did a big deal recently, you got a lot of press, and and I think have the distinction of being one of the only AI for bio tooling companies that is is making money.

Speaker 1

Accidental. No. So could you tell whatever you can disclose about that, we'd love to hear.

Yeah. So we were really excited to announce a deal with GSK where we licensed them Octo VC, which is for virtual cell foundation model. So we announced that back in January.

It's a 50,000,000 deal. It includes an upfront payment, milestones, and then separate than that, it also includes a annual license fee, model licensing fee. You know, I think this was a, you know, attractive deal for both parties, for us and for GSK, because, you know, really, the deal focuses on models that we've trained already on lung cancer, colon cancer, allows us to, you know, provide them with access to the models.

You know, GSK is one of, you know, the top AI teams in biopharma. So, you know, they know how to use these types of capabilities. They can use them for their internal use.

They can also use them to fine tune on their data. So that was a really big sell for GSK as well because, you know, GSK and every pharma is sitting on mountains and mountains of so called translational data. So the types of data that we're training the models on that come from clinical trials, you know, pathology specimens across many different therapeutics.

That, you know, everyone's sitting on a lot of this data, and it's been very hard to unlock. And so all of a sudden, you know, GSK can can use our models both to do simulations and to do therapeutic discovery, but they can also fine tune the models on their data. And in a way, the the model then becomes, you know, sort of GSK's version of the model.

This was super exciting. You know? It was the first you know, at least first announced foundation model licensing deal in the space.

And, you know, frankly, it was one, you know, we we've been trying to do for a long time even before Noetic. You know, I think a lot of companies have been trying to do these types of deals, and it's been I think it's just been historically slow for adoption on the pharma side, and it's been slow to demonstrate, like, a very clear value proposition for different types of of capabilities. And so what's unique about this deal is it looks you know, it doesn't look exactly like a software, you know, licensing framework for, let's say, a small amount of money with number of seats where you're licensed.

Well, it looks like a real business development deal in the industry where there's a very significant multimillion dollar, you know, cash upfront near term payment, but then the substrate of the deal is not a molecule. It's not doing therapeutic discovery work together.

Speaker 2

It the substrate is actually a model, which is what really made this pretty eek. Why do you think there's appetite for this suddenly? And it seems like almost whiplash that Yeah.

It, you know, it seems like only a maybe a year or two ago that Bayer was dying and whatever. And now suddenly, there's this deal, both is getting ton of attention.

Speaker 1

People are AI pill. In some extent, we increase it even more. I mean, maybe not totally, but increasingly more.

People are, you know, in pharma, you know, across the industry are seeing the value of different capabilities. They're able to use some of the open source capabilities, and they're able to demonstrate the value to themselves internally. And if you look at a if you look at a pharma company, you know, these companies are working on dozens and dozens of programs.

And so I you know, my opinions, just frankly, my opinion is that I think pharma increasingly wanna be able to access models, not just for one collaboration where you and I are working together on this one program. They wanna be able to access the technology across the whole pipeline. And so I think that's gonna create sort of a driving force for not just, you know, bespoke project driven licensing, but actual license broad licensing where a pharma can can access the technology in many different therapeutic programs.

Speaker 3

Yeah. And I think also, you know, with the structure prediction models, protein structure prediction, binding prediction models, there is, like, this massive public dataset. There are increasing amounts of data.

People can generate data to augment that. So, you know, there's enough data to the point where people can train very good models, but maybe not just on the data that any one biopharma company has. And I think that the same is true, but even more so for the types of models that we are building, which are, you know, foundation models at the patient biology level where, like, you know, no one company, I mean, these companies may have a lot of data, but it's, you know, scattered, it's siloed, and pulling everything together to, like, train an actual foundation model may not be as easy as it sounds, like, within a single company.

Whereas, we have just said, you know what, we're gonna generate enough data ourselves to actually train a real foundation model. And that's the nice thing about being a startup here is like we can make that bet that like you actually do benefit from generating all of this data in a, you know, uniformized way, like very high quality, etcetera. And then use that to develop and train the models.

And my opinion is that you need to have data at that scale before you can even think about developing models that actually work. It's like you can't do the AI R and D, like, or build the algorithms until you have good enough data set to tell you whether your favorite algorithmic idea is actually working or not.

Speaker 1

idea or someone else's idea about how to build a model, like, actually leading to improvements there. Yeah. I mean, this is a good point.

I mean, so like sometimes people ask me, well, why doesn't GSD just generate your data? So we just started generating data for years. There was no mob.

It was like, how how many years? Like, how Like, two years, maybe? A year and a half at least before we had the first trained models working?

Speaker 3

2024.

Speaker 1

So Yeah. That's like two years after. Yeah.

Through whom they were starting. So we how year or four years of SIL. So this is year four.

And so we basically opened the lab. We hired a team. We got all the instruments.

We started sourcing tumor samples. There was no prior here that any of this would work. Like, zero.

Big crazy peck. Like, I was just going for it. And, like, we just started generating data and, like like, sourcing human tumors, processing.

We built this whole processing pipeline to to get the tumors into, like, these arrays and the formats, and it takes weeks to you know, it takes literally two weeks for a machine to run a couple slides on the spatial transcriptomics. So so you've got, like, these two week runs where you're processing two slides, and and we're just churning data for months. And we couldn't even train up we didn't even have enough data to train a model for, like, at least a year and a half.

And then you're building, like, processing pipelines. You have to align all the data. You've gotta, like, post process it off the machine.

So we sort of just built all this, and then then, like, let's say, eighteen months later, hey, I wonder if this stuff and then it was not like, it wasn't obvious. There wasn't like, oh, we're gonna, like, off the shelf, you know, train this on some, like, open source architecture. You know, we've had we've you know, Dan and the team have done a ton of work.

Yeah. There wasn't really, like, anything major to go off of.

Speaker 3

transformers developed for single cell data, but, like, incorporating spatial data into that was, you know, again, there just like weren't really data sets out there that people had been able to develop on. So we do a lot of, like, custom model building and I enjoy that. I think people enjoy that.

Just say a lot for joining.

Speaker 1

About to build custom model. Yeah. Really unique, innovative, and powerful.

Speaker 2

Who who are you looking for? Like, kind of people?

Speaker 3

research on, again, this kind of alien landscape of data where you really have to figure out what's working from first principles, and obviously the work we do should have very very large impact. So definitely not restricted to people who have a biology background, you know, people who just like tackling very challenging machine learning problems, and are, you know, open to to learning the minimum amount of biology necessary to, like, make progress, I think, you know, would be great candidates.

Speaker 4

Talking to you guys reminds me a lot of the Leash Bio Yeah. Labs, which I know that both of you are part of the Recursion Mafia, you know. I'm not.

Yeah. Well, yeah. Yeah.

Yeah. Yeah. We're we're working at you on the show in the future too.

So yeah. Yeah. We're looking forward to that.

But like, it's it's interesting because both of you seem to have really similar philosophies and that like, you have deep convictions that, like, you're just gonna start collecting data before you know this is going to work. And you are going to just brute force it, go go go, and eventually, it will work. And, you know, you have signs.

I don't know. I think that's really impressive. I wonder, is there something about recursion which is in the water, which has led to this sort of thinking of just like, we're gonna commit to doing things at scale and it may not work at first.

It you have to hit a certain point before it will.

Speaker 1

I mean, we failed a lot at the beginning. Yeah. Give me the abrocurge.

Abrocurge. Yeah. Yeah.

And so you and we had we I said we had to build it from first principles, we really did. And so we spent many years trying to figure out, like, what should the data look like? Ian, myself, we're all involved in kind of platform development, how to design, you know, these datasets, how to design the experiments, iterative cycles over the years seeing, you know, things that did work, things that didn't work.

And so at the end of, you know, coming out of Recursion, I think what a lot of folks there had was, like, an understanding of what are the things we need to think about so that even if I wanna design a different dataset, you know, today, but it's, like, totally different. What are the things that we learn that we had to learn, like, over mistakes over like, not mistakes, but, like, trial and error, basically, over that many months that we would try to insert in our new approach? And so I don't know that every everything that I've predicted at Noetic in terms of, like, how to generate the data set has been important necessarily.

I know that we could start at the very beginning and say, okay. Well, let's make sure we do these 10 things. I know every one of these 10 things was important before.

Let's at least make sure we do these 10 things. I don't know that all 10 things are important for us today, but I would presume that, you know, many of them are, and it lets you sort of leapfrog that process of trial and error a little bit. Certainly, we do have trial and error still.

Speaker 4

you know, three problems, four problems overtell.

Speaker 1

their own data mode, like, do you have any advice or any suggestions about how to be more successful there? I think you sort of need to I mean, you think ahead to, okay, what am I trying to do on the machine learning side? And, like, what is the right data for solving this problem?

I think oftentimes I see, like, a lot of companies are like, okay. Well, I wanna generate x dataset. I'm I'm just gonna generate x dataset, and I'm gonna do machine machine learning on that.

Mhmm. Like, that might not be the right dataset. You might not have designed it the right way.

You know? It doesn't follow that like any dataset is a machine learning dataset, etcetera, yada It doesn't follow that that that dataset is gonna solve the problem you're trying to solve. So and I for me, it's really and even in Fowning and Away, it was, okay, what what problem are we trying to solve?

And then what are the data that are gonna help solve that problem? And rather than like, you know, going from from the, you know, data directly to to try to solve.

Speaker 3

where the technology is and, you know, where it's changing rapidly. So, you know, I finished my PhD in 2016. I did a lot of looking at spatial RNA like via this technique called in situ hybridization, same technique that is like at the base of what we're doing.

I could look at maybe two genes at a time on a single sample and that took me a full week of manual work. And, you know, I came to Noetic like five years later, six years later, and all of a sudden, you know, there are platforms where you can look at a thousand genes or 20,000 genes at once, you know, it's a single machine that can run this assay. It's expensive, but it's just like data beyond the wildest dreams of Dan Bear in 2016.

And that is only improving, like, rapidly. So I think it's important to see what the technology of today, you know, allows and also where it's going in terms of what data to generate. And what what does that pitch look like?

Speaker 2

$50,000,000

Speaker 1

and then I mean, it wasn't 50. It was maybe it was maybe closer to 10. But I if so yeah.

I mean, it isn't. So yeah. So you have to do that if you if I mean, if you're going into a regime where there's no data, yeah, and you wanna do something different, then, I mean, there's no shortcut to it.

Right? You're gonna have to generate the dataset. And so you're not gonna know the answer until it's there.

And that mean, and that's why a lot of companies are not going into that space where where there are no data sense because, you know, I think it it can be challenging to to do that. Yeah.

Speaker 4

try this pattern where they first will we either start with a public open source dataset or they will try a pilot, will they will internally collect a small amount of data and see if something works or something that doesn't. And oftentimes, there's almost like a critical point where below this, you're just not gonna get a new signal. Then you have to have conviction that you need to collect up to a certain point before you start, like, really driving something, like, fundamentally valuable.

Yeah. Yeah. I mean, imagine trying to train a foundation model on not enough data.

Yeah. Yeah. The kind and then then then that's it's sort of your well, least your clinical trial called, right?

GPT two, GPT three, GPT three you know, GPT one, two, and three, like, there was a clear progression there. As each one of them, you could see there was something which work with scale and there was this insight to, oh, we're gonna scale this up. Yeah.

You know, sometimes a biological data, like, the process of collecting lots of data is just very expensive to begin with. You can't just take something off the shelf and expect that you're gonna hit the threshold of, you know, GP three like usefulness. Yeah.

Yeah. So Yeah. Take some conviction.

It definitely takes conviction.

Speaker 3

I think, you know, it also takes sort of like a a scientific belief. Then there's a lot out there like that we just don't know yet and that you're not gonna capture the biology you need to by having, right now, like an agent that reads all of the biological literature. Because again, that's just like a tiny slice of what's out there.

Like, this is I don't know if it's a great analogy or if I'm gonna botch the history here, but like, in astronomy, it was required like Tycho Brahe, like collecting this enormous amount of astronomical data at his observatory, that then was the substrate for Kepler, you know, figuring out the first laws of motion of the the planets and know, that was superseded by like Newton's laws and so forth. But like, I I don't I sometimes don't know how you even get started without like this large repository of really high quality data being with. And, you know, maybe there's, like, a tragedy of the commons problem here of, like, who's gonna generate that data and who's gonna capture the value of it.

But I'm very glad that we're we're taking that bet and, you know, we're seeing it pay off. Yeah.

Speaker 1

but if, you know, hypothetically speaking,

Speaker 4

yeah, how much of PDB do you need to train? I mean, there there was some people I argued that yeah. And then you can get some pretty good models with a pink 11%.

Yeah. Delay. And there are people going back in the nineteen nineties argued that there was the PDB was already complete in the sense of, like, if you had a sufficiently smart algorithm, you could have done a pretty reasonable job of protein folding even back then.

Interesting. So you don't need a lot to get a pretty big boost, but the community was sort of inopinably collecting PDB data for quite some time Yeah.

Speaker 2

being convicted that this was going to lead to solving protein folding.

Speaker 4

Yeah.

Speaker 3

just knowing a protein was very helpful for some useful data set. And we did see we did see a transition from, like, early data. Like, how many samples did we get?

I'm guessing probably on the order of a few 100 before there's like Yeah. There was a there was definitely a moment, like, very soon after I joined where, like, we the dataset just kind of doubled in size overnight because there was like a huge bolus and like the models immediately got a lot better, at that point. And, you know, now we'd run these more controlled experiments of seeing, you know, what happens if you train on 10% of the data versus forty percent versus a hundred percent.

What happens if you hold out all of the pancreatic cancer or all of the breast cancer? And so, you know, we have a much better idea of what kind of diversity and scale we need now. I guess I would say, if we were sticking to cancer, maybe we're not like that far off.

I think, you know, again, if we end up generating a few 100 patients in a bunch of major and, you know, some minor indications which we're, you know, gonna do this year, maybe that's enough to generalize to kind of all cancer. Because there is a lot of shared biology in, you know, cancer and immune cells across different tissues and different, you know, mutations and so forth. But if you think about all of the disease biology that there is for a model to learn, you know, maybe that's like another order of magnitude.

Speaker 4

But even being able to solve all cancer biology would be a pretty impressive Yeah. To cure cancer would be would be great. Oh, if it's all in other biology, I did not say cure cancer.

It was such a different place. But, yeah, at least if you go back and just sort of a like, just take one drug.

Speaker 1

If you could look at one drug mechanism across the whole of oncology,

Speaker 2

that's incredibly powerful. I mean, imagine what Merck has done with Keytruda.

Speaker 1

Merck has run hundreds of trials with Keytruda. Like, it might even be over a thousand trials of Keytruda in different populations to find, you know, all these different indications. Okay.

The subset of ovarian cancers, the subset of lung cancers, the subset of colon cancers. That's all been done, you know, by enrolling trials. Mhmm.

If you can look at that biology from model embeddings and at least have a very well defined starting point for, okay, if I'm gonna run a trial, it doesn't have to be as broad as as it would need to be if I didn't have any answer, then that can be a really powerful tool for, you know, a diversity of mechanisms.

Speaker 3

Yeah. Maybe it's just like last point, like going back to the the virtual cell hot takes. Like, you know, if your goal is to build like an actual mechanistic model of an individual cell, and then build up from one cell to an entire tissue, and then, you know, tissue to patient, and so forth.

Like, you might need a lot more data and a lot more data modalities than, you know, just like gene expression or something like that. But, you know, we're taking much more of like a top down approach of we're trying to first solve the problem of what is determining heterogeneity among actual patients, and which of that variability is predictive of drug response. And my intuition is that you don't need to model the mechanism at the subcellular level necessarily to solve that problem of which patient should get which drug or, you know, which targets are important in which patients.

And I saw a similar debate play out in neuroscience and computational neuroscience where for a long time people were really trying to build these biophysical models of individual neurons and then they were gonna stitch them together into models of, you know, the brain and so forth. And what actually ended up working in, you know, in terms of building computational models of the brain and behavior is this abstraction. You know, we're just gonna treat individual neurons as, you know, linear, nonlinear units, and, you know, put them together in neural networks that are connected by, you know, linear weight matrices and, you know, stack a bunch of layers together and then build neural network models of the brain that abstract away kind of all of the details of biophysically what a neuron is doing.

And, you know, those are now by far the the most predictive models of how a given neuron is gonna respond to real world stimuli in a real brain. And I think that my bet is that the same is gonna be true for these models too, is that like by modeling sort of at the level of functional tissue where you have a bunch of cells interacting in like a disease context that that's gonna get you to the problem of predicting kind of the the patient level behavior much faster than trying to first model a cell and then stitch a bunch of those cells together.

Speaker 2

Yeah. That makes sense to me. It's a good analogy.

Like that. Do you have any call to action for the listeners?

Speaker 1

Yep. I mean, I would say, one, everyone should be excited about biology. You know, sometimes a lot of my hot takes on on x recently are just that I feel like there's a huge amount of enthusiasm in sort of, like, the mainstream tech ecosystem, and, like, people aren't really following a lot of, like, what's happening in the biology space.

But at the same time, like, you're hearing, you know, French of your lab saying we're gonna cure cancer. And, yeah, people should actually look at the the folks working on curing cancer or working on aging or working on areas of biology. These are really exciting, you know, problems.

There are real, like, significant NL problems in the space. One call to action is with love for for people to just, like, be more stoked about learning about applications, machine learning, and, like, biological sciences and, like, solving some of these hard problems because I think these are the problems that are gonna, like, massively impact humanity in, like, the next ten years, and we're just, like, really the very beginning. Like, you know, maybe we're in in in the, like, first inkling of the chat GPT moment for bio, but it's, like, very much just the very beginning.

So with like Catch you, Moi, can I Yeah? Yeah.

Speaker 3

dig in and learn more about the details. I think, you know, a lot of the times it's presented as we have these protein folding models, we have these binding models, you know, we have AI for science agents that are, you know, like reading all of the literature and automating these computational biology workflows. And I think it's important to realize that there are a lot of problems in AI for biology, AI for biochemistry, etcetera.

And some of them, and they're very important. But like solving any one of those is not gonna, like, solve the problem of how do we develop better therapeutics. And, you know, we're focused on, you know, a pretty particular slice of that process, which is again, translating things that we know work well in some patients into actual, like, successful drug trials where we know exactly which patients to give them to.

And that requires building foundation models at a particular level, you know, the patient level. But people should not be under the impression that like this is all gonna be solved immediately because, you know, AI agents like LLMs are gonna just read the literature and figure out what the right drug is. Like, there are a lot more data to generate.

There's a lot more ML problems to solve, and there's the need to translate those methods into actual successful drugs. And there's a lot of different places to contribute. It's a lot to do.

Yeah, dude. Great. Thank you very much.

Here we are.

Shared via Hopper