Computational Protein Design with Costas Maranas
Episode
49 min
Read time
2 min
Topics
Productivity, Relationships, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Inverse Protein Design: The core unsolved challenge in computational protein engineering is the inverse folding problem — given a desired protein structure and function, determine which amino acid sequence produces it. Current biophysical force fields like AMBER and CHARMM carry significant uncertainty, meaning even powerful search algorithms succeed only a fraction of the time, far below the 20–40% hit rate that would make wet lab screening practical.
- ✓Negative Data Generation: Training machine learning models on protein function requires balanced datasets of both working and non-working variants, yet most published datasets contain only positive results. Maranas advocates for a moonshot-style initiative: systematically engineer hundreds of enzymes spanning diverse EC classifications across prokaryotic, eukaryotic, and archaeal organisms, generating unbiased positive and negative variant data to properly train protein language models.
- ✓Mathematical Reframing Over Raw Compute: When computational limits block biological design problems, reframing them using established mathematical structures — such as mixed-integer linear programming borrowed from airline scheduling and warehouse logistics — can unlock solutions. Maranas used this approach to design microbial strains requiring up to 10 simultaneous gene knockouts, a result once considered impossible that CRISPR now executes in an afternoon.
- ✓Computational-Experimental Collaboration Protocol: Productive wet lab partnerships require computational researchers to become genuine domain experts in the experimental partner's organism and methods, not vice versa. Maranas estimates it takes multiple back-and-forth cycles — where computational suggestions are rejected, models are updated, and new suggestions are made — before a collaboration becomes reliably productive. Selecting collaborators for personal compatibility, not just scientific overlap, is equally critical.
- ✓Top-Down Genome Streamlining: Rather than building minimal cells from scratch, a pragmatic near-term strategy is stripping 10–20% of dispensable DNA from proven production strains like E. coli or yeast. Removing non-functional genomic segments reduces replication burden and eliminates metabolic pathways that could accidentally activate and divert carbon flux away from the target product, improving both predictability and yield in bioreactor deployments.
What It Covers
Penn State chemical engineering professor Costas Maranas discusses how computational methods — specifically optimization algorithms, biophysical force fields, and emerging transformer models — can engineer proteins, enzymes, and microbial strains to perform functions nature never evolved them to do, and why data quality remains the central bottleneck.
Key Questions Answered
- •Inverse Protein Design: The core unsolved challenge in computational protein engineering is the inverse folding problem — given a desired protein structure and function, determine which amino acid sequence produces it. Current biophysical force fields like AMBER and CHARMM carry significant uncertainty, meaning even powerful search algorithms succeed only a fraction of the time, far below the 20–40% hit rate that would make wet lab screening practical.
- •Negative Data Generation: Training machine learning models on protein function requires balanced datasets of both working and non-working variants, yet most published datasets contain only positive results. Maranas advocates for a moonshot-style initiative: systematically engineer hundreds of enzymes spanning diverse EC classifications across prokaryotic, eukaryotic, and archaeal organisms, generating unbiased positive and negative variant data to properly train protein language models.
- •Mathematical Reframing Over Raw Compute: When computational limits block biological design problems, reframing them using established mathematical structures — such as mixed-integer linear programming borrowed from airline scheduling and warehouse logistics — can unlock solutions. Maranas used this approach to design microbial strains requiring up to 10 simultaneous gene knockouts, a result once considered impossible that CRISPR now executes in an afternoon.
- •Computational-Experimental Collaboration Protocol: Productive wet lab partnerships require computational researchers to become genuine domain experts in the experimental partner's organism and methods, not vice versa. Maranas estimates it takes multiple back-and-forth cycles — where computational suggestions are rejected, models are updated, and new suggestions are made — before a collaboration becomes reliably productive. Selecting collaborators for personal compatibility, not just scientific overlap, is equally critical.
- •Top-Down Genome Streamlining: Rather than building minimal cells from scratch, a pragmatic near-term strategy is stripping 10–20% of dispensable DNA from proven production strains like E. coli or yeast. Removing non-functional genomic segments reduces replication burden and eliminates metabolic pathways that could accidentally activate and divert carbon flux away from the target product, improving both predictability and yield in bioreactor deployments.
Notable Moment
Maranas describes attending an operations research conference in the late 1990s where genome assembly researchers presented to a room of mathematicians who were entirely disengaged. Recognizing that his cross-disciplinary background uniquely positioned him to bridge that gap became the moment he committed to redirecting his entire lab toward computational biology.
Episode Transcript
Welcome to the Axial podcast. Axial is an early stage investment firm based in San Francisco. We partner with great founders and inventors investing in early stage life science companies often when they are no more than an idea. Axial is fanatical about helping the right venture who's compelled to build their own and journey business. Alright. We're we're Yeah. We're live, Costas. I appreciate taking the time to do this. I'm a big fan of your research. And maybe to start off this podcast, you can just introduce yourself and the work you do. Okay. My name is Casas Morales. I'm a professor, in the department of chemical engineering at Penn State. And what we do in our group is we're trying to straddle two worlds, the world of computations and algorithms and the world of biology. So we're trying to apply computational algorithm methods so that we can, we we can train microbes or other organisms to do something different that nature did not, evolve them to do. For example, treat, coax microbes to make your favorite chemical or engineer proteins and antibodies to recognize a new ligand or a new substrate. And so when you have that common thread of computation across all of biology, whether it's engineered a cell, a protein, maybe even just a an ensemble of things, how do you think about, like, choosing proms? Okay. Well, I I tend to read the in the past, before Google search and online libraries, I would go in the library, and I would get lost there. I will look for paper a, and in the process, I will find five or six other papers that were interesting. So that generated a lot of ideas in our group in the late nineties, early two thousands. Nine day nowadays, Google, I mean, sometimes I just, you know, start looking up things that I find interesting. So sometimes serendipity. Sometimes collaborators come up with ideas and questions that, pollinate new ideas. So, unfortunately, they and sometimes it's the students. The students get excited about something, they'll let me know. And if I feel that it's it's within our wheelhouse, then I will encourage it to continue down that path. So so there's no, you know, I think there's no systematic way to you know, that we find new problems. Sometimes with the funding agencies, you know, if they want us to work on organism x, you know, and they give us money, that that's what we will do. Interesting. And you've been running you've been at Penn State for going up three decades now, so congrats. And there's a really cool picture of you during your days at Princeton during grad school. You're in front of a old computer. I can send you the picture I found. And then you look at your picture now, and you're in a suit and stuff. How has the role of, like like, just cheaper compute power? You know? I imagine, like, AWS …
Get the full transcript (8,142 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 46-minute episode.
Get Axial Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Axial Podcast
Modern Computational Tools for Chemistry with Corin Wagen
Mar 23 · 50 min
Latent Space
🔬Why There Is No "AlphaFold for Materials" — AI for Materials Discovery with Heather Kulik
Mar 24
More from Axial Podcast
Evolutionary Intelligence and Biologics Discovery with Jeremy Agresti
Mar 23 · 51 min
10% Happier with Dan Harris
The Science of Journaling: How 15 Minutes of Writing Can Free Up Your Brain, Improve Your Health, and Make You a Better Friend | Dr. James Pennebaker
Aug 14
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“Current biophysical force fields like AMBER and CHARMM carry significant uncertainty, meaning even powerful search algorithms succeed only a fraction of the time.”
“Current biophysical force fields like AMBER and CHARMM carry significant uncertainty, meaning even powerful search algorithms succeed only a fraction of the time.”
More from Axial Podcast
We summarize every new episode. Want them in your inbox?
Modern Computational Tools for Chemistry with Corin Wagen
Evolutionary Intelligence and Biologics Discovery with Jeremy Agresti
AI Workflows for Biopharma with Alex Telford
AI Legal Software with Scott Stevenson
Scaling Proteomics with Milad Dagher
Similar Episodes
Related episodes from other podcasts
Latent Space
Mar 24
🔬Why There Is No "AlphaFold for Materials" — AI for Materials Discovery with Heather Kulik
10% Happier with Dan Harris
Aug 14
The Science of Journaling: How 15 Minutes of Writing Can Free Up Your Brain, Improve Your Health, and Make You a Better Friend | Dr. James Pennebaker
The Joe Rogan Experience
Jul 30
#2533 - Diana Pasulka
Planet Money
Jul 15
Building things and breaking things in China (Summer School World Tour)
How I AI
Jul 6
How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)
Explore Related Topics
This podcast is featured in Best Biotech Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into Axial Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Axial Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime