Skip to main content
No Priors: Artificial Intelligence | Technology | Startups

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

42 min episode · 2 min read
·
Baseten Ceo Tuhin Srivastava

Episode

42 min

Read time

2 min

Topics

Startups, Leadership, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
  • Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
  • Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
  • Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
  • Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.

What It Covers

Baseten CEO Tuhin Srivastava joins Sarah Guo and Elad to discuss how the AI inference market reached 30x growth in 12 months, why 95% of tokens served run on custom models, how compute scarcity shapes strategy, and what the path from open-source adoption to specialized model deployment looks like.

Key Questions Answered

  • Custom model adoption: 95% of tokens served on Baseten run on customer-modified models, not vanilla open-source weights. Companies customize for both quality and performance through quantization and post-training. Builders should validate product-market fit with best-in-class frontier models first, then invest in specialization once user signal is proven and worth optimizing.
  • Compute scarcity reality: Baseten runs 90 clusters across 18 clouds at mid-90% utilization with a daily 4PM capacity management meeting. Securing 1,024 B200s from a reputable provider now requires a 3-to-5 year contract with 20-30% TCV prepaid upfront. Companies with low cost of capital hold a structural advantage in acquiring GPU supply.
  • Application layer defensibility: Companies like Abridge build moats through proprietary user signal — clinician edits, workflow steps, downstream actions — that frontier labs cannot access. Enterprises should encode differentiation in multi-step workflows and post-train models on those reward signals rather than relying solely on closed-source API access for competitive advantage.
  • Inference-post-training loop: Baseten acquired post-training firm PaaS to close the loop between inference and model improvement. Inference generates data, evals surface reward functions, and post-training improves the model — creating a continuous cycle. Infrastructure companies should treat inference and post-training as paired problems, since quantization decisions and training choices directly affect inference performance.
  • Jevons paradox in AI workloads: As inference costs fall, customers extend agent run times and insert intelligence into more workflow steps rather than reducing usage. Developers consistently respond to lower costs by running longer, more complex tasks. Companies building on AI should plan infrastructure and pricing models around expanding consumption, not stable or declining token volumes.

Notable Moment

Srivastava revealed that Baseten's co-founder's seven-year-old child has learned to ask "is that a P0?" when a pager alert fires at home — illustrating how deeply an always-on operational culture becomes embedded across an infrastructure company scaling at this pace.

Know someone who'd find this useful?

Episode Transcript

Hi, listeners. Today, Elad and I are here with Tuhen Srivastava, the founder and CEO of Base ten, the AI Inference Cloud. We're here to talk about capacity constraints for AI capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source, and perhaps multi chip future, and what 30 x scale in a year looks like. Tuhen, welcome back. Hi. Good to see you. Thanks for having me. Alright. You are in one of the, craziest markets, AI inference. It's very important. There's a lot going on. You guys have grown 30 x over the last year. And I think I can say you're expecting to do more than a billion dollars in revenue this year. Mhmm. What's going on? Tell us about scale. Yeah. No. It's been it's been nuts. I I think what's happened over the last, honestly, at twenty four months, but just kinda keeps getting bigger and bigger is that I think everyone is real realizing that you can put AI everywhere. You have you have all these great options available from closed source to open source models. The open source models have crossed some sort of chasm in terms of their baseline capability, and then I think RL techniques and post training is, for specialized models, has become mainstream enough, and, you know, there's enough examples of it work of it working. The customers are realizing they can, you know, kind of own their inference, more and more. And what that's meant for us is, more, you know, the long tail models coming through, customers in housing a lot of that intelligence themselves. And as the application layer just gets, you know, bigger and bigger and bigger and that's growing, we are we are just someone index on that, and we've been around to be able to collect the demand. There's an existential question in here that I think everybody is, continually asking of, does the independent application layer get to exist at all versus the labs? Like, how do you you have to believe this. Why do you believe it? Yeah. Look. I I I think it'd be it'd be a sad thing if it didn't exist in general, and I think that's, like, my but, you know, sadness is fine. The Sadness is fine. But, like, that that that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons. One is because, you know, I think this idea, that what it what is valuable to a company, is, you know, the the user signal that they can gather, that only they can gather. And to the extent that that is encoded, in a model, I think a lot of their business will, be at risk. But to the to the extent that it is encoded in workflows, that is where they will be able to develop notes. So a good I …

Get the full transcript (8,257 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all No Priors: Artificial Intelligence | Technology | Startups transcripts →

You just read a 3-minute summary of a 39-minute episode.

Get No Priors: Artificial Intelligence | Technology | Startups summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from No Priors: Artificial Intelligence | Technology | Startups

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into No Priors: Artificial Intelligence | Technology | Startups.

Every Monday, we deliver AI summaries of the latest episodes from No Priors: Artificial Intelligence | Technology | Startups and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime