No Priors: Artificial Intelligence | Technology | Startups

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

May 1, 2026

42 min episode · 2 min read

Baseten Ceo Tuhin Srivastava

Episode

42 min

Read time

2 min

Topics

Leadership, Artificial Intelligence

AI-Generated Summary

Published May 2, 2026

Key Takeaways

✓Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
✓Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
✓Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
✓Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
✓Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.

What It Covers

Baseten CEO Tuhin Srivastava joins Sarah Guo and Elad to discuss how the AI inference market reached 30x growth in 12 months, why 95% of tokens served run on custom models, how compute scarcity shapes strategy, and what the path from open-source adoption to specialized model deployment looks like.

Key Questions Answered

•Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
•Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
•Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
•Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
•Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.

Notable Moment

Srivastava revealed that Baseten's co-founder's seven-year-old child has learned to ask "is that a P0?" when a pager alert fires at home — illustrating how deeply an always-on operational culture becomes embedded across an infrastructure company scaling at this pace.

Know someone who'd find this useful?

You just read a 3-minute summary of a 39-minute episode.

Get No Priors: Artificial Intelligence | Technology | Startups summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from No Priors: Artificial Intelligence | Technology | Startups

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

This Week in Startups

May 1

Explore Related Topics

👔Leadership 🤖Artificial Intelligence

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into No Priors: Artificial Intelligence | Technology | Startups.

Every Monday, we deliver AI summaries of the latest episodes from No Priors: Artificial Intelligence | Technology | Startups and 192+ other podcasts. Free for up to 3 shows.

Start My Monday Digest

No credit card · Unsubscribe anytime

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

AI-Generated Summary

Key Takeaways

What It Covers

Key Questions Answered

Notable Moment

Keep Reading

SAP: Bringing the ‘Operating System’ of a Company into the AI Era with CTO Philipp Herzig

Can an AI Agent Legally Own a Company? Christian van der Henst's Wild Experiment| E2283

Scaling Global Organizations in the Age of AI with ServiceNow CEO Bill McDermott

Consumer electronics can't keep up with AI

More from No Priors: Artificial Intelligence | Technology | Startups

SAP: Bringing the ‘Operating System’ of a Company into the AI Era with CTO Philipp Herzig

Scaling Global Organizations in the Age of AI with ServiceNow CEO Bill McDermott

The Agentic Economy: How AI Agents Will Transform the Financial System with Circle Co-Founder and CEO Jeremy Allaire

AI for Atoms: How Periodic Labs is Revolutionizing Materials Engineering with Co-Founder Liam Fedus

Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI

Similar Episodes

Can an AI Agent Legally Own a Company? Christian van der Henst's Wild Experiment| E2283

Consumer electronics can't keep up with AI

OpenAI Misses Targets, Codex vs Claude, Elon vs Sam Trial, Big Hyperscaler Beats, Peptide Craze

1977: Ask Farnoosh: How Much Should We Pay for College? Plus: Her Investments Went Missing

The Week AI Grew Up

Explore Related Topics

You're clearly into No Priors: Artificial Intelligence | Technology | Startups.