Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
Episode
42 min
Read time
2 min
Topics
Startups, Leadership, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
- ✓Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
- ✓Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
- ✓Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
- ✓Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.
What It Covers
Baseten CEO Tuhin Srivastava joins Sarah Guo and Elad to discuss how the AI inference market reached 30x growth in 12 months, why 95% of tokens served run on custom models, how compute scarcity shapes strategy, and what the path from open-source adoption to specialized model deployment looks like.
Key Questions Answered
- •Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
- •Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
- •Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
- •Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
- •Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.
Notable Moment
Srivastava revealed that Baseten's co-founder's seven-year-old child has learned to ask "is that a P0?" when a pager alert fires at home — illustrating how deeply an always-on operational culture becomes embedded across an infrastructure company scaling at this pace.
Episode Transcript
Hi, listeners. Today, Elad and I are here with Tuhen Srivastava, the founder and CEO of Base ten, the AI Inference Cloud. We're here to talk about capacity constraints for AI capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source, and perhaps multi chip future, and what 30 x scale in a year looks like. Tuhen, welcome back. Hi. Good to see you. Thanks for having me. Alright. You are in one of the, craziest markets, AI inference. It's very important. There's a lot going on. You guys have grown 30 x over the last year. And I think I can say you're expecting to do more than a billion dollars in revenue this year. Mhmm. What's going on? Tell us about scale. Yeah. No. It's been it's been nuts. I I think what's happened over the last, honestly, at twenty four months, but just kinda keeps getting bigger and bigger is that I think everyone is real realizing that you can put AI everywhere. You have you have all these great options available from closed source to open source models. The open source models have crossed some sort of chasm in terms of their baseline capability, and then I think RL techniques and post training is, for specialized models, has become mainstream enough, and, you know, there's enough examples of it work of it working. The customers are realizing they can, you know, kind of own their inference, more and more. And what that's meant for us is, more, you know, the long tail models coming through, customers in housing a lot of that intelligence themselves. And as the application layer just gets, you know, bigger and bigger and bigger and that's growing, we are we are just someone index on that, and we've been around to be able to collect the demand. There's an existential question in here that I think everybody is, continually asking of, does the independent application layer get to exist at all versus the labs? Like, how do you you have to believe this. Why do you believe it? Yeah. Look. I I I think it'd be it'd be a sad thing if it didn't exist in general, and I think that's, like, my but, you know, sadness is fine. The Sadness is fine. But, like, that that that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons. One is because, you know, I think this idea, that what it what is valuable to a company, is, you know, the the user signal that they can gather, that only they can gather. And to the extent that that is encoded, in a model, I think a lot of their business will, be at risk. But to the to the extent that it is encoded in workflows, that is where they will be able to develop notes. So a good I …
Get the full transcript (8,257 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
Browse all No Priors: Artificial Intelligence | Technology | Startups transcripts →
You just read a 3-minute summary of a 39-minute episode.
Get No Priors: Artificial Intelligence | Technology | Startups summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from No Priors: Artificial Intelligence | Technology | Startups
Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong
Sep 10 · 45 min
Gradient Dissent
The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava
Nov 18
More from No Priors: Artificial Intelligence | Technology | Startups
Redefining Chip Architecture with Arm CEO Rene Haas
Sep 3 · 37 min
The Bike Shed
456: Typescript with Jimmy Thigpen
Feb 25
More from No Priors: Artificial Intelligence | Technology | Startups
We summarize every new episode. Want them in your inbox?
Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong
Redefining Chip Architecture with Arm CEO Rene Haas
Rethinking Legacy Data Infrastructure with Eon Co-Founders Ofir Ehrlich and Gonen Stein
From Restoring Sight to Reimagining the Brain, with Max Hodak
What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest
Similar Episodes
Related episodes from other podcasts
Gradient Dissent
Nov 18
The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava
The Bike Shed
Feb 25
456: Typescript with Jimmy Thigpen
Software Engineering Daily
Sep 1
The Death of Online Anonymity
Invest Like the Best with Patrick O'Shaughnessy
Sep 1
Sarah Guo - What the 250 People Building AI Believe - [Invest Like the Best, EP.489]
The Rich Roll Podcast
Aug 31
Hunter Biden Is Out Of Secrets: A Story of Recovery, Exposure & Amends
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into No Priors: Artificial Intelligence | Technology | Startups.
Every Monday, we deliver AI summaries of the latest episodes from No Priors: Artificial Intelligence | Technology | Startups and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime