Skip to main content
Latent Space

🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

89 min episode · 3 min read
·
Bo Wang,Ci Chu

Episode

89 min

Read time

3 min

Topics

Startups, Fundraising & VC, Design & UX

AI-Generated Summary

Key Takeaways

  • âś“Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
  • âś“PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
  • âś“Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
  • âś“Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
  • âś“Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.

What It Covers

Xaira Therapeutics researchers Bo Wang and Ci Chu explain how their X-Cell virtual cell model uses genome-wide CRISPR perturbation data across 25 million cells and 16 cell types to predict gene expression changes after genetic interventions, demonstrating generalization to unseen cell types and primary human T cells where linear baselines fail entirely.

Key Questions Answered

  • •Causal vs. observational data: Foundation models trained on observational datasets like CellxGene (33M+ cells) consistently fail to outperform linear baselines on perturbation prediction tasks. The core reason is that correlational data supports multiple equally valid causal structures — if genes A, B, and C co-vary, any of them could be the regulator. Only interventional, perturbation-specific training data breaks this ambiguity and enables genuine counterfactual prediction.
  • •PerturbSeq scaling strategy: Xaira's PISCES dataset combines pooled CRISPR-Cas9 knockdown with single-cell RNA sequencing to generate a two-dimensional matrix — every gene perturbed against every gene measured — across 16 cell types. Pooled experimental design eliminates batch effects entirely since all perturbations occur in one vessel. After quality filtering, 25 million cells remain from an initial harvest of hundreds of millions.
  • •Diffusion over autoregressive architecture: X-Cell replaces autoregressive gene-ordering (used in scGPT) with a diffusion language model that iteratively refines predictions from noisy gene expression states. This removes the need to impose an arbitrary gene order on expression matrices. Ablation studies show this switch produces measurable improvement on out-of-distribution generalization tasks, particularly predicting perturbation effects in unseen cell types.
  • •Multi-prior conditioning improves context specificity: X-Cell incorporates five biological prior types during training: LLM-derived gene descriptions (essentially GPT embeddings per gene), protein-protein interaction networks, DepMap cancer essentiality scores, cell morphology data, and scGPT cell-type embeddings. These priors become encoded as learnable parameters, so inference requires no prior input — but providing priors at inference time yields additional accuracy gains in context-specific predictions.
  • •Generalization benchmark design: To rigorously test out-of-distribution performance, Xaira trained X-Cell exclusively on resting T cells, then asked it to predict perturbation effects in activated T cells — a combinatorial prediction problem. The model correctly predicted TCR complex inactivation and identified putative T cell inactivators later confirmed in a separately published primary T cell screen from Alex Marson's lab at UCSF, validating transfer from cell lines to donor-derived primary cells.
  • •Phase 3 clinical trial failure rate frames the goal: Roughly 90% of diseases have no cure, and phase 3 clinical trial success rates run as low as 5–10%. Xaira's three-platform strategy — virtual cell models, protein design (from David Baker's UW group), and patient representation learning — targets this failure rate by connecting cellular causal predictions to patient stratification, aiming to select the right patients before large-cohort trials rather than discovering mismatches during them.

Notable Moment

When researchers visualized X-Cell's predictions as heat maps alongside linear baseline predictions and ground truth gene expression changes, the model's output was visually indistinguishable from ground truth while the linear baseline diverged clearly — the first time seven genome-wide perturbation campaigns combined into one training set produced this level of accuracy.

Know someone who'd find this useful?

You just read a 3-minute summary of a 86-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime