🔬Beyond AlphaFold: How Boltz is Open-Sourcing the Future of Drug Discovery
Episode
81 min
Read time
2 min
Topics
Career Growth, Remote Work, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Training Under Constraints: Boltz-1 was trained only once due to compute limitations, requiring live debugging during training runs. The team stopped training mid-run to fix bugs, then resumed without restarting from scratch. They used Department of Energy cluster resources with two-day training windows followed by week-long queue waits, eventually completing training with help from Genesis compute resources.
- ✓Coevolution as Structural Hints: AlphaFold models decode evolutionary patterns where amino acid positions that mutate together across species indicate spatial proximity in 3D structure. This coevolutionary data acts like database lookup, providing strong priors that guide models to approximate solution spaces before physics-based refinement finds low-energy states. Models struggle without this evolutionary signal on novel proteins.
- ✓Generative Modeling Over Regression: AlphaFold3 shifted from predicting single structures to sampling posterior distributions of possible conformations. This generative approach handles uncertainty better than regression, which averages conflicting predictions into incorrect structures. The architecture uses diffusion models with cubic computational complexity from pairwise operations, requiring fewer parameters but more compute than language models.
- ✓Validation Through Distributed Testing: Boltz coordinated 25 academic and industry labs to test designs across diverse applications, reporting results from 8-10 labs in their paper. For nanobodies targeting 14 novel proteins with no known interactions in training data, they achieved nanomolar binders on two-thirds of targets using just 15 designs per target, demonstrating true generalization beyond training distribution.
- ✓Atomic-Level Sequence Prediction: BoltzGen predicts both protein structure and sequence simultaneously by encoding amino acids through atomic composition. The model receives blank tokens for designed proteins and predicts atomic positions, which implicitly determine amino acid identity since different residues have unique atomic arrangements. This unified supervision signal scales better than separate discrete and continuous objectives.
What It Covers
Gabriela Corso and Jeremy Volven from Boltz explain how they open-sourced protein structure prediction after AlphaFold3 remained proprietary. They trained their model once with limited compute, fixing bugs mid-training, and built BoltzLab to democratize drug discovery through accessible AI tools for designing proteins and small molecules that bind therapeutic targets.
Key Questions Answered
- •Training Under Constraints: Boltz-1 was trained only once due to compute limitations, requiring live debugging during training runs. The team stopped training mid-run to fix bugs, then resumed without restarting from scratch. They used Department of Energy cluster resources with two-day training windows followed by week-long queue waits, eventually completing training with help from Genesis compute resources.
- •Coevolution as Structural Hints: AlphaFold models decode evolutionary patterns where amino acid positions that mutate together across species indicate spatial proximity in 3D structure. This coevolutionary data acts like database lookup, providing strong priors that guide models to approximate solution spaces before physics-based refinement finds low-energy states. Models struggle without this evolutionary signal on novel proteins.
- •Generative Modeling Over Regression: AlphaFold3 shifted from predicting single structures to sampling posterior distributions of possible conformations. This generative approach handles uncertainty better than regression, which averages conflicting predictions into incorrect structures. The architecture uses diffusion models with cubic computational complexity from pairwise operations, requiring fewer parameters but more compute than language models.
- •Validation Through Distributed Testing: Boltz coordinated 25 academic and industry labs to test designs across diverse applications, reporting results from 8-10 labs in their paper. For nanobodies targeting 14 novel proteins with no known interactions in training data, they achieved nanomolar binders on two-thirds of targets using just 15 designs per target, demonstrating true generalization beyond training distribution.
- •Atomic-Level Sequence Prediction: BoltzGen predicts both protein structure and sequence simultaneously by encoding amino acids through atomic composition. The model receives blank tokens for designed proteins and predicts atomic positions, which implicitly determine amino acid identity since different residues have unique atomic arrangements. This unified supervision signal scales better than separate discrete and continuous objectives.
- •Infrastructure Cost Advantage: Running Boltz models on their platform costs significantly less than self-hosting open-source versions. Their small molecule screening pipeline runs 10x faster than open-source implementations through optimization. Platform users can parallelize 100,000 candidate designs across GPU fleets, completing in minutes what would take weeks serially, amortizing compute costs across customers.
Notable Moment
The team revealed their flagship model went through an unrepeatable training curriculum because they could only afford one training run. While the model trained, they discovered and fixed bugs on the fly, stopping and restarting training multiple times without returning to the beginning. This improvised approach somehow produced a working model that matched AlphaFold3 performance despite the chaotic development process.
Episode Transcript
Actually, we only trained the big model once. That's how much compute we had. We could only train it once. And so, like, while the model was training, we were, like, finding bugs left and right. Yeah. A lot of them that I wrote. And, like, I would I remember, like, us, like, sort of, like, you know, doing, like, surgery in the middle, like, stopping the run, making the fix, like, relaunching. And, yeah. We never actually went back to the start. We just, like, kept training it with, like, the bug fixes along the way, which was It's impossible to reproduce now. Yeah. Yeah. Yeah. No. That model is, like, Yeah. Yeah. No. That model is, like, has gone through such a curriculum that, you know, it's learned some weird stuff. But, yeah. Somehow by miracle, I worked out. It's a pleasure to have with us today Gabriela Corso and Jeremy Volven. They are they recently founded Boltz, a company trying to democratize and bring art structure prediction in biology to, you know, the masses. They were both, recent PhD grads from MIT and have been working on all sorts of foundational papers in, like, generative biology. Anyway, pleasure to have you here. Thanks for coming. Thank you. Thank you. I guess we're maybe, what, Six years post AlphaFold two right now, which was like kind of a big moment. Is that right? I think it was at 2021. So, yeah. Going on five years. Five years. Five years. Yeah. Yeah. Yeah. So maybe for the audience, like, let's go back to that moment in time and explain, like, what was this big moment and why was it interesting? Why was everyone so excited? And I think you two were probably quite excited. So why were you personally excited? I would start on kind of why that was interesting kind of, you know, from a scientific standpoint. So what AlphaFold, so maybe first as a kind of introduction for, the ones in the audience and not structure biologists. So the idea of structure biology is that, you know, we want to try to understand how, you know, proteins and other molecules take shape inside our cells and how they interact. And Structural Biology is this beautiful discipline, where we are somehow able to understand this minuscule structure, atomic details using, these incredibly, complex methods like, you know, x-ray crystallography. And, you know, the dream has always been of a computational biology. Can we understand kind of the structures without having to, you know, resolve this crystal, you know, shoot x rays and so on? And so AlphaFold was a real breakthrough in this problem of protein folding, which is trying to understand the structure of a single, protein. And to me, it was exciting across kind of many dimensions. One, I was computer scientist. I was working a lot on machine learning and I saw kind of the impact that kind of the work similar, somewhat similar to …
Get the full transcript (15,179 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 78-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
a16z Podcast
OpenAI Researchers on the Future of Mathematical Reasoning
Sep 8
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Stuff You Should Know
Selects: Blacksmiths? You got that right!
Sep 5
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by DeepMind
“Gabriela Corso and Jeremy Volven from Boltz explain how they open-sourced protein structure prediction after AlphaFold3 remained proprietary.”
- BoltzLabBy guest
by Boltz
“Gabriela Corso and Jeremy Volven from Boltz explain how they open-sourced protein structure prediction after AlphaFold3 remained proprietary...and built BoltzLab to democratize drug discovery through accessible AI tools for designing proteins and small molecules that bind therapeutic targets.”
- Boltz-1By guest
by Boltz
“Boltz-1 was trained only once due to compute limitations, requiring live debugging during training runs. The team stopped training mid-run to fix bugs, then resumed without restarting from scratch.”
- BoltzGenBy guest
by Boltz
“BoltzGen predicts both protein structure and sequence simultaneously by encoding amino acids through atomic composition. The model receives blank tokens for designed proteins and predicts atomic positions, which implicitly determine amino acid identity since different residues have unique atomic arrangements.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Sep 8
OpenAI Researchers on the Future of Mathematical Reasoning
Stuff You Should Know
Sep 5
Selects: Blacksmiths? You got that right!
Deep Questions with Cal Newport
Sep 3
Did OpenAI Create “Secret AI Civilizations”? | Tech Decoded
Lenny's Podcast
Aug 30
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
Practical AI
Aug 28
Building the Foundation for the Agentic AI Era
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime