🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
Episode
85 min
Read time
3 min
Topics
Startups, Leadership, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Patient selection as root cause: 90-95% of cancer drugs fail in clinical trials not because of poor pharmacology or target selection, but because trials enroll the wrong patients. Drugs often work in a subset of patients, but without models to identify that subset upfront, trials run on broad populations, diluting signal and leading to cancellation of molecules that could help specific subgroups.
- ✓Spatial transcriptomics as training data: Noetik generates multimodal tissue data stacking H&E pathology images, multiplex fluorescence protein stains, and spatial transcriptomics capturing up to 20,000 genes per spatial location. Each data point functions like a 20,000-channel image rather than a standard RGB image. Over 100 million spatially-resolved cells have been generated, representing at least one order of magnitude more paired data than any known public dataset.
- ✓H&E as universal inference input: Despite training on expensive multimodal data, Noetik's models run inference using only standard H&E pathology slides at deployment. Because H&E is collected for virtually every cancer patient globally, this allows retrospective analysis of existing trial cohorts — splitting responders from non-responders using images already on file, without requiring new data collection from past participants.
- ✓Autoregressive scaling for spatial biology: Noetik's Tario model applies next-token autoregressive training — the same objective scaling LLMs — to spatial transcriptomics data. Larger models only outperform smaller ones at longer context lengths, meaning the model must observe larger tissue regions simultaneously to capture nonlinear spatial patterns. This mirrors LLM scaling behavior and suggests tissue context length is a key variable for biological foundation model performance.
- ✓In vivo perturbation validation via barcoded mouse tumors: To validate human model predictions without relying on cell lines, Noetik uses a multiplexed CRISPR knockout platform injecting ~100 barcoded cancer cell variants into single mice, producing hundreds of genetically distinct tumors per animal. Human-trained models are then inferenced directly on mouse H&E, and predictions about immune infiltration and tumor phenotype are validated against known pathway biology across multiple gene knockouts simultaneously.
What It Covers
Noetik co-founders Ron Alfa and Daniel Bear explain how 90-95% of cancer drug trial failures stem from poor patient selection rather than bad pharmacology. They describe building multimodal foundation models trained on spatially-resolved human tumor data — combining H&E pathology, multiplex protein imaging, and 20,000-gene spatial transcriptomics — to match drugs to the right patient subpopulations.
Key Questions Answered
- •Patient selection as root cause: 90-95% of cancer drugs fail in clinical trials not because of poor pharmacology or target selection, but because trials enroll the wrong patients. Drugs often work in a subset of patients, but without models to identify that subset upfront, trials run on broad populations, diluting signal and leading to cancellation of molecules that could help specific subgroups.
- •Spatial transcriptomics as training data: Noetik generates multimodal tissue data stacking H&E pathology images, multiplex fluorescence protein stains, and spatial transcriptomics capturing up to 20,000 genes per spatial location. Each data point functions like a 20,000-channel image rather than a standard RGB image. Over 100 million spatially-resolved cells have been generated, representing at least one order of magnitude more paired data than any known public dataset.
- •H&E as universal inference input: Despite training on expensive multimodal data, Noetik's models run inference using only standard H&E pathology slides at deployment. Because H&E is collected for virtually every cancer patient globally, this allows retrospective analysis of existing trial cohorts — splitting responders from non-responders using images already on file, without requiring new data collection from past participants.
- •Autoregressive scaling for spatial biology: Noetik's Tario model applies next-token autoregressive training — the same objective scaling LLMs — to spatial transcriptomics data. Larger models only outperform smaller ones at longer context lengths, meaning the model must observe larger tissue regions simultaneously to capture nonlinear spatial patterns. This mirrors LLM scaling behavior and suggests tissue context length is a key variable for biological foundation model performance.
- •In vivo perturbation validation via barcoded mouse tumors: To validate human model predictions without relying on cell lines, Noetik uses a multiplexed CRISPR knockout platform injecting ~100 barcoded cancer cell variants into single mice, producing hundreds of genetically distinct tumors per animal. Human-trained models are then inferenced directly on mouse H&E, and predictions about immune infiltration and tumor phenotype are validated against known pathway biology across multiple gene knockouts simultaneously.
- •Data generation must precede model development: Noetik spent roughly 18 months generating data before training any functional model. The lesson for AI biotech startups: design datasets around the specific ML problem first, control for batch effects by distributing each patient sample across multiple slides and arrays, and reach a critical data threshold before expecting meaningful model signal. Dropping to 10-40% of training data causes substantial generalization failure, particularly across cancer types not seen during training.
Notable Moment
Noetik ran its lab for approximately 18 months — sourcing human tumors, building processing pipelines, and running two-week spatial transcriptomics machine cycles — before accumulating enough data to train a single model. There was no prior evidence any of it would work. The first functional foundation model, Octo VC, emerged roughly two years after the company launched.
Episode Transcript
So we basically opened the lab. We hired a team. We got all the instruments. We started sourcing tumor samples. There was no prior here that any of this would work. Like, zero. We just started generating data and, like, like, sourcing human tumors, processing. We built this whole processing pipeline to to get the tumors into, like, these arrays and the formats. So you've got, like, these two week runs where you're processing two slides, and and we're just churning data for months. And we couldn't even train a model. So we sort of just built all this. And then then, like, let's say, eighteen months later, hey, I wonder, can we train a model off of and then it was not, you know, like, it wasn't obvious. Yeah. There wasn't really, like, anything major to go off of. I mean, there were, like, transformers developed for single cell data. There just, like, weren't really datasets out there that people have been able to develop on. We do a lot of, like, custom model building. Hi there. I'm RJ Haneke, and this is Brandon Anderson. We're the co hosts of the Latent Space Science podcast. And today, we're really happy to be in the studio with some of the people from NOEIC. I'm Ron Alpa, co founder and CEO of NOEIC, physician scientist by training. My hobbies are making hot takes about AI curing cancer. Hi. I'm Dan Baer. I'm VP of AI at Noetic. I'm a biologist by training, did PhD work in neuroscience, and then moved into comp neuro, computer vision, self supervised learning, and have, you know, been doing AI research at Noetic for the past few years. Maybe we should start with what is Noetic, why did you found it? What is the difference between Noetic and the other virtual cell Yeah. Companies? Maybe just start with a little bit of a contrarian thesis, Ted, which is really the reason for founding Noetic. It's we all know the numbers that ninety percent, ninety five percent of cancer drugs fail in the clinic. Why do they fail? So our thesis is they fail not because we're bad at pharmacology, not because we're bad at target selection, you know, making the drug. We're actually better at that process than we have ever been in the history of drug development. Most of those drugs fail, we'd argue, because we're bad at selecting which patients those drugs are gonna work in. And oftentimes, you see trials where there is no placebo effect in cancer. Some patients respond to these drugs. And if you have a patient that responds, that tells you something that there's some biology that that's active there, but you have a problem in in patient selection. And so really, that's the thesis behind your why they because can we build models that can fundamentally understand patient biology from the very beginning and help you position molecules in the right patient population. So you're actually using the …
Get the full transcript (15,198 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 82-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
Beyond Biotech
Turning cancer cell dependencies into targeted therapies
Jul 10
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Citeline Podcasts
Killing Cancer Loudly: Onchilles Pharma's Neutrophil-Derived Path to Pan-Cancer Therapy
Mar 11
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Noetik
“Noetik's Tario model applies next-token autoregressive training — the same objective scaling LLMs — to spatial transcriptomics data.”
by Noetik
“The first functional foundation model, Octo VC, emerged roughly two years after the company launched.”
company
“Noetik co-founders Ron Alfa and Daniel Bear explain how 90-95% of cancer drug trial failures stem from poor patient selection.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
Beyond Biotech
Jul 10
Turning cancer cell dependencies into targeted therapies
Citeline Podcasts
Mar 11
Killing Cancer Loudly: Onchilles Pharma's Neutrophil-Derived Path to Pan-Cancer Therapy
My First Million
Jan 12
Why Balance Is the Enemy of Greatness | David Senra
The Long Run with Luke Timmerman
Oct 28
Ep188: Art Krieg on Innate Immune System Activators for Cancer
The Long Run with Luke Timmerman
Oct 14
Ep187: Eric Fischer on Creating a New Class of Medicines
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime