Skip to main content
Axial Podcast

Proteomics and AI with Peter Cimermančič

57 min episode · 2 min read
·
Peter Cimerman

Episode

57 min

Read time

2 min

Topics

Startups, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • The 80% Dark Matter Problem: Current mass spectrometry search algorithms leave up to 80% of measured spectra unidentified, meaning most proteomics-based drug discovery operates on only 20-30% of available data. Researchers should treat preprocessing quality — not downstream analytics — as the primary bottleneck limiting biological discovery and target identification accuracy.
  • Cloud Migration as Dual Solution: Moving proteomics preprocessing from single workstations to cloud infrastructure solves two problems simultaneously: parallelizing compute across hundreds of nodes accelerates processing speed, while the added compute capacity enables replacement of legacy rule-based algorithms with transformer and vision-based AI models that capture previously ignored signal.
  • Absence of Signal as Data: Current scoring algorithms match spectra using barcode-style peak matching, ignoring peak intensities, non-canonical fragmentation patterns, and crucially, the absence of expected peaks. AI models trained without these human-imposed rules learn from missing signal too, producing 70% more peptide identifications on standard human proteomes and 200-300% gains on complex metaproteomics datasets.
  • Expanding the Search Space: Standard proteomics searches only consider canonical proteins longer than 50 amino acids in unmodified form. Deliberately expanding searches to include small open reading frames, post-translational modifications, and sequence variants — enabled by a sufficiently accurate scoring model — reveals biologically relevant proteins that canonical pipelines structurally cannot detect.
  • Proteomics Covers the Full Drug Discovery Pipeline: Mass spectrometry proteomics applies across every drug discovery stage: affinity purification identifies protein-protein interactions for target discovery; chemoproteomics maps covalent small-molecule binding sites; immunopeptidomics identifies peptides for cancer vaccines; and plasma proteomics predicts patient treatment response and disease outcomes, consistently outperforming other omics modalities in multiomics studies.

What It Covers

Peter Cimermančič, cofounder of Tesserai and former seven-year Verily researcher, explains how AI-powered preprocessing of mass spectrometry proteomics data can recover up to 80% of currently unidentified spectra, unlocking drug targets and biological insights that conventional search algorithms systematically miss.

Key Questions Answered

  • The 80% Dark Matter Problem: Current mass spectrometry search algorithms leave up to 80% of measured spectra unidentified, meaning most proteomics-based drug discovery operates on only 20-30% of available data. Researchers should treat preprocessing quality — not downstream analytics — as the primary bottleneck limiting biological discovery and target identification accuracy.
  • Cloud Migration as Dual Solution: Moving proteomics preprocessing from single workstations to cloud infrastructure solves two problems simultaneously: parallelizing compute across hundreds of nodes accelerates processing speed, while the added compute capacity enables replacement of legacy rule-based algorithms with transformer and vision-based AI models that capture previously ignored signal.
  • Absence of Signal as Data: Current scoring algorithms match spectra using barcode-style peak matching, ignoring peak intensities, non-canonical fragmentation patterns, and crucially, the absence of expected peaks. AI models trained without these human-imposed rules learn from missing signal too, producing 70% more peptide identifications on standard human proteomes and 200-300% gains on complex metaproteomics datasets.
  • Expanding the Search Space: Standard proteomics searches only consider canonical proteins longer than 50 amino acids in unmodified form. Deliberately expanding searches to include small open reading frames, post-translational modifications, and sequence variants — enabled by a sufficiently accurate scoring model — reveals biologically relevant proteins that canonical pipelines structurally cannot detect.
  • Proteomics Covers the Full Drug Discovery Pipeline: Mass spectrometry proteomics applies across every drug discovery stage: affinity purification identifies protein-protein interactions for target discovery; chemoproteomics maps covalent small-molecule binding sites; immunopeptidomics identifies peptides for cancer vaccines; and plasma proteomics predicts patient treatment response and disease outcomes, consistently outperforming other omics modalities in multiomics studies.

Notable Moment

Cimermančič describes how HIV-human protein interaction studies using mass spectrometry were conducted while seeing only 20% of actual interactions — raising the pointed question of how many viable drug targets for infectious disease have been permanently overlooked due to preprocessing limitations rather than biological absence.

Know someone who'd find this useful?

Episode Transcript

Welcome to the Axial podcast. Axial is an early stage investment firm based in San Francisco. We partner with great founders and inventors investing in early stage life science companies often when they are no more than an idea. Axial is fanatical about helping the right venture who's compelled to build their own and journey business. Hey, Peter. Thank you for doing this. I'm I'm really excited to record this conversation for a few months now. Maybe to start off with just a brief instruction yourself and what, Tesoro AI does. Well, hi, Josh. Much appreciate this opportunity, to be on the podcast. And, yeah, you've been one of the most helpful, transparent like, the feedback that you're given me on just starting the company was fantastic. So, yeah, much appreciate that as well. Appreciate it. But, yeah, my name is Peter Simmermansek. I'm currently cofounder of Tesserai, but previously, I've been at Verily for seven years. Before that, I've been grad student at UCSF, and my background is I'm I'm from Slovenia. I did my, I I took biochemistry in college. So coming from a different noncomputer science background, but turning into, let's say, an AI, ML, specialist or application scientist. Yeah. I've been spending quite a bit of time on computational biology, applying working at the intersection of AI, ML, and problems, really interesting problems in life sciences. Could you speak to me? What's the premise or what's the problem to sort of AI solving? Yeah. So as mentioned, I've been working in combiospace for some time now for, let's say, almost two decades, and, have seen how data workflows work, operates, and have identified or have seen quite a few challenges that we are having, as as the field more broadly. So maybe just at a very high level, data, goes through multiple stages. It all starts with instruments, single app generating raw data. The next step being data preprocessing where this raw signal is turned into nice tidy data tables, for example, protein expression tables, gene expression tables, mutation statuses, and so on depending on exactly what the instrument does. But once you have these nice, tidy data tables, the next stage is data modeling or analysis. And this is where, let's say, computational biologists come in. They build bunch of analysis tools to identify new biomarkers, to identify new biology. Basically, this is where, let's say, insights are being generated. And, what was really interesting to me is that a lot of our efforts are being spent on this second latter stage, so data analytics or insight generation. And that the first stage, data preprocessing, are very much overlooked. Even in when I was a grad student, for example, the cool kids on the block were always, let's say, students that were analyzing data for new biology insights. It's not those that were building tools to, to improve preprocessing, so to speak. So so so just diving back into, what sort of success we have …

Get the full transcript (9,585 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Axial Podcast transcripts →

You just read a 3-minute summary of a 54-minute episode.

Get Axial Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Axial Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Biotech Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Axial Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Axial Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime