Modal and Scaling AI Inference with Erik Bernhardsson
Episode
39 min
Read time
2 min
Topics
Productivity, Remote Work, Startups
AI-Generated Summary
Key Takeaways
- ✓Container Cold Start Optimization: Modal achieves sub-second container launches by building custom file systems and container runtimes that cache redundant data between images, since most container data remains unread during execution, enabling rapid GPU deployment without traditional Docker inefficiencies.
- ✓Multi-Tenant Resource Pooling: Aggregating variable AI workloads across shared GPU pools enables 100% effective utilization versus underutilized dedicated resources. Usage-based pricing charges only for active GPU seconds, eliminating capacity planning while pooling bursty demand creates cost efficiency impossible with reserved infrastructure.
- ✓Function-as-Service Programming Model: Developers decorate Python functions to specify GPU types and dependencies, then call them like local code. Modal handles serialization, exception management, and auto-scaling across distributed containers, maintaining sub-second feedback loops similar to front-end development hot reloading.
- ✓Gen AI Inference Characteristics: Stable diffusion and similar models send small text inputs to GPUs that perform trillions of operations before returning small outputs. This compute-intensive, low-IO pattern differs from traditional data processing, making 200-millisecond overhead negligible compared to multi-second inference times.
What It Covers
Erik Bernhardsson discusses Modal's serverless platform for AI workloads, enabling sub-second GPU container deployment through custom infrastructure. He covers multi-tenant architecture, cold start optimization, developer productivity, and Gen AI inference scaling challenges.
Key Questions Answered
- •Container Cold Start Optimization: Modal achieves sub-second container launches by building custom file systems and container runtimes that cache redundant data between images, since most container data remains unread during execution, enabling rapid GPU deployment without traditional Docker inefficiencies.
- •Multi-Tenant Resource Pooling: Aggregating variable AI workloads across shared GPU pools enables 100% effective utilization versus underutilized dedicated resources. Usage-based pricing charges only for active GPU seconds, eliminating capacity planning while pooling bursty demand creates cost efficiency impossible with reserved infrastructure.
- •Function-as-Service Programming Model: Developers decorate Python functions to specify GPU types and dependencies, then call them like local code. Modal handles serialization, exception management, and auto-scaling across distributed containers, maintaining sub-second feedback loops similar to front-end development hot reloading.
- •Gen AI Inference Characteristics: Stable diffusion and similar models send small text inputs to GPUs that perform trillions of operations before returning small outputs. This compute-intensive, low-IO pattern differs from traditional data processing, making 200-millisecond overhead negligible compared to multi-second inference times.
Notable Moment
Bernhardsson rejected a Snowflake job offer in 2012 because he doubted cloud-native databases would succeed, calling it his worst career decision. He now builds Modal on the same multi-tenant cloud principles that made Snowflake successful.
Episode Transcript
Modal is a serverless compute platform that's specifically focused on AI workloads. The company's goal is to enable AI teams to quickly spin up GPU enabled containers and rapidly iterate and auto scale. It was founded by Eric Bernardson, who was previously at Spotify for seven years where he built the music recommendation system and the popular Luigi workflow scheduler. In this episode, Eric joins Sean Falconer to talk about the motivation for founding his company, the market gap in ML and AI tooling, optimizing container cold start, Modal's interface design, and more. This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him. Eric, welcome to the show. Thank you. It's great to be here. Yeah. Thanks so much for being here. So I was diving a little bit into your background preparing for this. And so it seems like you spent a lot of your time working in data throughout your career, which also kinda matches my own experience. You know, you previously were at Spotify for a number of years. You were the CTO of better.com. Now you're the founder and CEO of Modal. You know, were there certain things in these prior roles that led to identifying some sort of need for Modal? Like, what's the story behind essentially going off and deciding to start this company? Yeah. For sure. The answer is yes. And the long story is I was at Spotify for seven years, built a music recommendation system. But as a part of building that, I also realized there's kind of a general gap in the tooling. I ended up building a vector database called Illinois. I know we use it today. And also workflow scheduler called Luigi that very few people use today. But generally, I realized, like, as a part of building all of that stuff at Spotify and also did a lot of other stuff that there's very little tooling in data, AI, machine learning. There's more today, but it still never really felt like much later, like, in 2020, 2021 when I started thinking about building a company. I realized there's kind of a gap in the market for, like, a tool I always wanted to have myself. So that was kind of the genesis of modal. It's almost like building selfishly for what I always wanted to have, you know, throughout my years as recognition system, and also, to some extent, it better as the CTO. So a little bit more like general role. We spent a lot of time thinking about platforms and data and stuff like that too. Yeah. I mean, sometimes people talk about how, you know, the discipline is essentially, like or, like, the tooling and the things that are available for data engineers lag somewhere, you know, five years maybe behind, you know, more app traditional application development. Where would you say that something like ML engineering from an applied sense of …
Get the full transcript (8,908 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 36-minute episode.
Get Software Engineering Daily summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Software Engineering Daily
A Rust Framework to Simplify Distributed Systems
Sep 10 · 50 min
Latent Space
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Jul 8
More from Software Engineering Daily
SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
Sep 8 · 52 min
This Week in Startups
China wants you to cheer for the robots taking your job | E2329
Aug 25
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
- ModalBy guest
“Erik Bernhardsson discusses Modal's serverless platform for AI workloads, enabling sub-second GPU container deployment through custom infrastructure.”
company
“Bernhardsson rejected a Snowflake job offer in 2012 because he doubted cloud-native databases would succeed, calling it his worst career decision.”
More from Software Engineering Daily
We summarize every new episode. Want them in your inbox?
A Rust Framework to Simplify Distributed Systems
SED News: The NVIDIA-Hugging Face Deal, China’s Proxy Economy, the Open Weight Surge
Moving Beyond RAG with Precomputed Context
The Death of Online Anonymity
TypeScript 7 and What Comes Next
Similar Episodes
Related episodes from other podcasts
Latent Space
Jul 8
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
This Week in Startups
Aug 25
China wants you to cheer for the robots taking your job | E2329
No Priors: Artificial Intelligence | Technology | Startups
Aug 13
What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest
a16z Podcast
Jul 21
Why Physical AI Is the Next Frontier | Applied Intuition
Practical AI
Jul 17
The Future of AI Infrastructure with CoreWeave
Explore Related Topics
This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Software Engineering Daily.
Every Monday, we deliver AI summaries of the latest episodes from Software Engineering Daily and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime