Proteomics and AI with Peter Cimermančič
Episode
57 min
Read time
2 min
Topics
Startups, Fundraising & VC, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓The 80% Dark Matter Problem: Current mass spectrometry search algorithms leave up to 80% of measured spectra unidentified, meaning most proteomics-based drug discovery operates on only 20-30% of available data. Researchers should treat preprocessing quality — not downstream analytics — as the primary bottleneck limiting biological discovery and target identification accuracy.
- ✓Cloud Migration as Dual Solution: Moving proteomics preprocessing from single workstations to cloud infrastructure solves two problems simultaneously: parallelizing compute across hundreds of nodes accelerates processing speed, while the added compute capacity enables replacement of legacy rule-based algorithms with transformer and vision-based AI models that capture previously ignored signal.
- ✓Absence of Signal as Data: Current scoring algorithms match spectra using barcode-style peak matching, ignoring peak intensities, non-canonical fragmentation patterns, and crucially, the absence of expected peaks. AI models trained without these human-imposed rules learn from missing signal too, producing 70% more peptide identifications on standard human proteomes and 200-300% gains on complex metaproteomics datasets.
- ✓Expanding the Search Space: Standard proteomics searches only consider canonical proteins longer than 50 amino acids in unmodified form. Deliberately expanding searches to include small open reading frames, post-translational modifications, and sequence variants — enabled by a sufficiently accurate scoring model — reveals biologically relevant proteins that canonical pipelines structurally cannot detect.
- ✓Proteomics Covers the Full Drug Discovery Pipeline: Mass spectrometry proteomics applies across every drug discovery stage: affinity purification identifies protein-protein interactions for target discovery; chemoproteomics maps covalent small-molecule binding sites; immunopeptidomics identifies peptides for cancer vaccines; and plasma proteomics predicts patient treatment response and disease outcomes, consistently outperforming other omics modalities in multiomics studies.
What It Covers
Peter Cimermančič, cofounder of Tesserai and former seven-year Verily researcher, explains how AI-powered preprocessing of mass spectrometry proteomics data can recover up to 80% of currently unidentified spectra, unlocking drug targets and biological insights that conventional search algorithms systematically miss.
Key Questions Answered
- •The 80% Dark Matter Problem: Current mass spectrometry search algorithms leave up to 80% of measured spectra unidentified, meaning most proteomics-based drug discovery operates on only 20-30% of available data. Researchers should treat preprocessing quality — not downstream analytics — as the primary bottleneck limiting biological discovery and target identification accuracy.
- •Cloud Migration as Dual Solution: Moving proteomics preprocessing from single workstations to cloud infrastructure solves two problems simultaneously: parallelizing compute across hundreds of nodes accelerates processing speed, while the added compute capacity enables replacement of legacy rule-based algorithms with transformer and vision-based AI models that capture previously ignored signal.
- •Absence of Signal as Data: Current scoring algorithms match spectra using barcode-style peak matching, ignoring peak intensities, non-canonical fragmentation patterns, and crucially, the absence of expected peaks. AI models trained without these human-imposed rules learn from missing signal too, producing 70% more peptide identifications on standard human proteomes and 200-300% gains on complex metaproteomics datasets.
- •Expanding the Search Space: Standard proteomics searches only consider canonical proteins longer than 50 amino acids in unmodified form. Deliberately expanding searches to include small open reading frames, post-translational modifications, and sequence variants — enabled by a sufficiently accurate scoring model — reveals biologically relevant proteins that canonical pipelines structurally cannot detect.
- •Proteomics Covers the Full Drug Discovery Pipeline: Mass spectrometry proteomics applies across every drug discovery stage: affinity purification identifies protein-protein interactions for target discovery; chemoproteomics maps covalent small-molecule binding sites; immunopeptidomics identifies peptides for cancer vaccines; and plasma proteomics predicts patient treatment response and disease outcomes, consistently outperforming other omics modalities in multiomics studies.
Notable Moment
Cimermančič describes how HIV-human protein interaction studies using mass spectrometry were conducted while seeing only 20% of actual interactions — raising the pointed question of how many viable drug targets for infectious disease have been permanently overlooked due to preprocessing limitations rather than biological absence.
Episode Transcript
Welcome to the Axial podcast. Axial is an early stage investment firm based in San Francisco. We partner with great founders and inventors investing in early stage life science companies often when they are no more than an idea. Axial is fanatical about helping the right venture who's compelled to build their own and journey business. Hey, Peter. Thank you for doing this. I'm I'm really excited to record this conversation for a few months now. Maybe to start off with just a brief instruction yourself and what, Tesoro AI does. Well, hi, Josh. Much appreciate this opportunity, to be on the podcast. And, yeah, you've been one of the most helpful, transparent like, the feedback that you're given me on just starting the company was fantastic. So, yeah, much appreciate that as well. Appreciate it. But, yeah, my name is Peter Simmermansek. I'm currently cofounder of Tesserai, but previously, I've been at Verily for seven years. Before that, I've been grad student at UCSF, and my background is I'm I'm from Slovenia. I did my, I I took biochemistry in college. So coming from a different noncomputer science background, but turning into, let's say, an AI, ML, specialist or application scientist. Yeah. I've been spending quite a bit of time on computational biology, applying working at the intersection of AI, ML, and problems, really interesting problems in life sciences. Could you speak to me? What's the premise or what's the problem to sort of AI solving? Yeah. So as mentioned, I've been working in combiospace for some time now for, let's say, almost two decades, and, have seen how data workflows work, operates, and have identified or have seen quite a few challenges that we are having, as as the field more broadly. So maybe just at a very high level, data, goes through multiple stages. It all starts with instruments, single app generating raw data. The next step being data preprocessing where this raw signal is turned into nice tidy data tables, for example, protein expression tables, gene expression tables, mutation statuses, and so on depending on exactly what the instrument does. But once you have these nice, tidy data tables, the next stage is data modeling or analysis. And this is where, let's say, computational biologists come in. They build bunch of analysis tools to identify new biomarkers, to identify new biology. Basically, this is where, let's say, insights are being generated. And, what was really interesting to me is that a lot of our efforts are being spent on this second latter stage, so data analytics or insight generation. And that the first stage, data preprocessing, are very much overlooked. Even in when I was a grad student, for example, the cool kids on the block were always, let's say, students that were analyzing data for new biology insights. It's not those that were building tools to, to improve preprocessing, so to speak. So so so just diving back into, what sort of success we have …
Get the full transcript (9,585 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 54-minute episode.
Get Axial Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Axial Podcast
Modern Computational Tools for Chemistry with Corin Wagen
Mar 23 · 50 min
The SaaS Podcast
Product-Market Fit: From Edtech Vitamin to $100M Painkiller
Feb 19
More from Axial Podcast
Evolutionary Intelligence and Biologics Discovery with Jeremy Agresti
Mar 23 · 51 min
The SaaS Podcast
Product-Market Fit: How Tito Goldstein Found It After 2 Years of Near-Zero Revenue
Jan 29
More from Axial Podcast
We summarize every new episode. Want them in your inbox?
Modern Computational Tools for Chemistry with Corin Wagen
Evolutionary Intelligence and Biologics Discovery with Jeremy Agresti
AI Workflows for Biopharma with Alex Telford
AI Legal Software with Scott Stevenson
Scaling Proteomics with Milad Dagher
Similar Episodes
Related episodes from other podcasts
The SaaS Podcast
Feb 19
Product-Market Fit: From Edtech Vitamin to $100M Painkiller
The SaaS Podcast
Jan 29
Product-Market Fit: How Tito Goldstein Found It After 2 Years of Near-Zero Revenue
Huberman Lab
Oct 6
How to Make Yourself Unbreakable | DJ Shipley
Conversations with Tyler
Oct 1
John Amaechi on Leadership, the NBA, and Being Gay in Professional Sports
Everything Everywhere Daily
Aug 18
The Liberal Arts
Explore Related Topics
This podcast is featured in Best Biotech Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Axial Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Axial Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime