🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Episode
89 min
Read time
3 min
Topics
Startups, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
- ✓PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
- ✓Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
- ✓Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
- ✓Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.
What It Covers
Xaira Therapeutics researchers Bo Wang and Ci Chu explain how their X-Cell virtual cell model uses genome-wide CRISPR perturbation data across 25 million cells and 16 cell types to predict gene expression changes after genetic interventions, demonstrating generalization to unseen cell types and primary human T cells where linear baselines fail entirely.
Key Questions Answered
- •Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
- •PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
- •Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
- •Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
- •Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.
- •Phase 3 clinical trial failure rate frames the goal: Roughly 90% of diseases have no cure, and phase 3 clinical trial success rates run as low as 5–10%. Xaira's three-platform strategy — virtual cell models, protein design (from David Baker's UW group), and patient representation learning — targets this failure rate by connecting cellular causal predictions to patient stratification, aiming to select the right patients before large-cohort trials rather than discovering mismatches during them.
Notable Moment
When researchers visualized X-Cell's predictions as heat maps alongside linear baseline predictions and ground truth gene expression changes, the model's output was visually indistinguishable from ground truth while the linear baseline diverged clearly — the first time seven genome-wide perturbation campaigns combined into one training set produced this level of accuracy.
You just read a 3-minute summary of a 86-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
Jul 16 · 101 min
a16z Podcast
Mark Zuckerberg & Priscilla Chan: How AI Will Cure All Disease
Nov 6
More from Latent Space
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Jul 8 · 57 min
a16z Podcast
Mark Zuckerberg & Priscilla Chan: How AI Will Help Cure Disease
Jul 9
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI
Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Nov 6
Mark Zuckerberg & Priscilla Chan: How AI Will Cure All Disease
a16z Podcast
Jul 9
Mark Zuckerberg & Priscilla Chan: How AI Will Help Cure Disease
20VC (20 Minute VC)
Jun 29
20VC: Leo Aschenbrenner's Largest Holding: Inside the $90BN Bloom Energy | Why Electricity, Not AI Models, Will Decide the Winners of the AI Race | Why We Are Not in an AI Capex Bubble | Energy Sovereignty and The Future of Power with KR Sridhar
No Priors: Artificial Intelligence | Technology | Startups
Jun 10
Biohub: The Future of Biology is Open-Source with Co-Founders Mark Zuckerberg, Priscilla Chan, and Head of Science Alex Rives
NVIDIA AI Podcast
Apr 29
How Dassault Systèmes Is Building AI That Understands Physics - Ep. 296
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime