🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Episode
89 min
Read time
3 min
Topics
Startups, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
- ✓PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
- ✓Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
- ✓Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
- ✓Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.
What It Covers
Xaira Therapeutics researchers Bo Wang and Ci Chu explain how their X-Cell virtual cell model uses genome-wide CRISPR perturbation data across 25 million cells and 16 cell types to predict gene expression changes after genetic interventions, demonstrating generalization to unseen cell types and primary human T cells where linear baselines fail entirely.
Key Questions Answered
- •Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
- •PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
- •Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
- •Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
- •Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.
- •Phase 3 clinical trial failure rate frames the goal: Roughly 90% of diseases have no cure, and phase 3 clinical trial success rates run as low as 5–10%. Xaira's three-platform strategy — virtual cell models, protein design (from David Baker's UW group), and patient representation learning — targets this failure rate by connecting cellular causal predictions to patient stratification, aiming to select the right patients before large-cohort trials rather than discovering mismatches during them.
Notable Moment
When researchers visualized X-Cell's predictions as heat maps alongside linear baseline predictions and ground truth gene expression changes, the model's output was visually indistinguishable from ground truth while the linear baseline diverged clearly — the first time seven genome-wide perturbation campaigns combined into one training set produced this level of accuracy.
Episode Transcript
And what really blew my mind away is when I saw the model make prediction, just print out the heat map of the gene expression changes, look at the actual raw data, and line up the linear baseline prediction, the ground truth, and the Excel prediction altogether. It's visually very clear to see that Excel prediction is much more similar to ground truth than the linear baseline. This is a wow moment I was talking about in the beginning. This is the first time that someone can put together not just one perturbisic, but seven genome wide perturbisic campaigns together. Something that, jumped out to us biologists right away is that some of the perturbations are complex universal. Hi. I'm RJ Haneke, CTO of Mero Omics. This is Brandon Anderson who builds RNA therapeutics at Atomic AI, and this is the latent space AI for science podcast. One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether, something is AI for science or something like b to b SAS. We're really happy to have in the studio with us today, Mo Wang and Si Chu from Xeira Therapeutics. At Xeira, they're building with a bunch of other people an AI drug discovery platform. They're using high throughput experimentation system to collect very large data sets and then training AI models that can predict the way that your cells in your body will respond to drugs and therapeutics. Really happy to have you. Big fan of your work. Why don't you two introduce yourselves to the listeners? Hello, everyone. My name is Bowen. I'm SVP and head of biomedical AI at Xero Therapeutic. Joined Xero about eight months ago and before that I was associate professor at the University of Toronto in Canada. And, I'm Sichu. My first name is incredibly difficult to pronounce. Honestly, sweet Mandarin, so I go by Chu, as in Chubak or Pikachu, as your favorite fictional character. I'm the SVP of AI enabled discovery at Xera. I joined about more than two years ago when I was still in stealth mode. And here I lead the high throughput biology group, generating the kind of data that will feed our AI models and also think about their applications. Before this, I, spent about a decade in, at the intersection of AI and, big data and biology. So I previously worked at In Situil, leading, the in vitro discovery platform there. And before that, I was at Verily, which is spun out of Google X. Okay. So you you're at Xara, the, company which is on the Pareto frontier of confusing names and mega rounds. So, Xera is, I think, kinda came out of stealth, like, a few years ago and just really big org kind of out of nothing. So, I'm curious if you can explain a little bit about what is Xyra's mission, what …
Get the full transcript (14,591 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 86-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
a16z Podcast
Mark Zuckerberg & Priscilla Chan: How AI Will Cure All Disease
Nov 6
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
20VC (20 Minute VC)
20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile
Aug 1
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks.”
“X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states.”
by Xaira Therapeutics
“Xaira Therapeutics researchers Bo Wang and Ci Chu explain how their X-Cell virtual cell model uses genome-wide CRISPR perturbation data across 25 million cells and 16 cell types to predict gene expression changes after genetic interventions.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Nov 6
Mark Zuckerberg & Priscilla Chan: How AI Will Cure All Disease
20VC (20 Minute VC)
Aug 1
20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile
a16z Podcast
Jul 9
Mark Zuckerberg & Priscilla Chan: How AI Will Help Cure Disease
20VC (20 Minute VC)
Jun 29
20VC: Leo Aschenbrenner's Largest Holding: Inside the $90BN Bloom Energy | Why Electricity, Not AI Models, Will Decide the Winners of the AI Race | Why We Are Not in an AI Capex Bubble | Energy Sovereignty and The Future of Power with KR Sridhar
No Priors: Artificial Intelligence | Technology | Startups
Jun 10
Biohub: The Future of Biology is Open-Source with Co-Founders Mark Zuckerberg, Priscilla Chan, and Head of Science Alex Rives
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime