The CEO Behind the Fastest-Growing AI Inference Company | Tuhin Srivastava
Episode
59 min
Read time
2 min
Topics
Productivity, Startups, Leadership
AI-Generated Summary
Key Takeaways
- ✓Company pivoting: Stay lean during market shifts - BaseTen remained 18 people from 2019-2023, enabling rapid pivots when ChatGPT and Stable Diffusion created new opportunities without organizational weight.
- ✓Inference differentiation: Focus on dedicated capacity over shared endpoints - 99% of BaseTen's business serves custom models with dedicated infrastructure, avoiding commoditized shared model serving markets.
- ✓Technical optimization: Modern LLM inference requires both infrastructure scaling across thousands of GPUs and runtime optimization using frameworks like VLLM, TensorRT-LLM, and SGLang for performance improvements.
- ✓Market positioning: Open source adoption follows predictable pattern - companies start with Anthropic/OpenAI, then switch to open source models for cost control, reliability, and data privacy.
What It Covers
BaseTen CEO Tuhin Srivastava explains how his AI inference company pivoted from serving data scientists with small models to becoming fastest-growing inference provider for production applications.
Key Questions Answered
- •Company pivoting: Stay lean during market shifts - BaseTen remained 18 people from 2019-2023, enabling rapid pivots when ChatGPT and Stable Diffusion created new opportunities without organizational weight.
- •Inference differentiation: Focus on dedicated capacity over shared endpoints - 99% of BaseTen's business serves custom models with dedicated infrastructure, avoiding commoditized shared model serving markets.
- •Technical optimization: Modern LLM inference requires both infrastructure scaling across thousands of GPUs and runtime optimization using frameworks like VLLM, TensorRT-LLM, and SGLang for performance improvements.
- •Market positioning: Open source adoption follows predictable pattern - companies start with Anthropic/OpenAI, then switch to open source models for cost control, reliability, and data privacy.
Notable Moment
Srivastava reveals BaseTen killed three of four products in 2022, including an application builder that consumed two dozen employees for 2.5 years of development work.
No transcript yet — request it by email, free
We'll transcribe this episode on request and email you the full transcript and AI summary — usually within a day. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 56-minute episode.
Get Gradient Dissent summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Gradient Dissent
Elon's Former Battery Chief: AI Data Centers Will Make Electricity Cheaper
Aug 18 · 94 min
No Priors: Artificial Intelligence | Technology | Startups
Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
May 1
More from Gradient Dissent
40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks
Aug 3 · 79 min
David Senra
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Sep 9
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“Modern LLM inference requires both infrastructure scaling across thousands of GPUs and runtime optimization using frameworks like VLLM, TensorRT-LLM, and SGLang for performance improvements.”
“Modern LLM inference requires both infrastructure scaling across thousands of GPUs and runtime optimization using frameworks like VLLM, TensorRT-LLM, and SGLang for performance improvements.”
“Modern LLM inference requires both infrastructure scaling across thousands of GPUs and runtime optimization using frameworks like VLLM, TensorRT-LLM, and SGLang for performance improvements.”
More from Gradient Dissent
We summarize every new episode. Want them in your inbox?
Elon's Former Battery Chief: AI Data Centers Will Make Electricity Cheaper
40 Trillion Tokens a Day (Yes, More Than OpenAI) | Lin Qiao, CEO of Fireworks
He's Building an AI That Can't Lie | Dan Klein
He Raised $70M to Cure Every Disease With AI
Uber, Nissan, and Mercedes Chose This Self-Driving Startup | Alex Kendall, Wayve
Similar Episodes
Related episodes from other podcasts
No Priors: Artificial Intelligence | Technology | Startups
May 1
Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
David Senra
Sep 9
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Eye on AI
Jul 28
Video Is About to Stop Being One-Way (and That Changes Everything) | Victor Riparbelli, Synthesia
Masters of Scale
Jul 14
The quiet reinvention of a $42b business, with Canva’s Cameron Adams
Latent Space
Jul 8
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Gradient Dissent.
Every Monday, we deliver AI summaries of the latest episodes from Gradient Dissent and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime