20VC: How to Build Your Own Data Center & Why Every Startup Should Do It | How ElevenLabs Leapfrogged Us: What I Learned | The AI Talent War: How Your Hiring Process Needs to Change with Cliff Weitzman, Speechify
Episode
65 min
Read time
3 min
Topics
Career Growth, Relationships, Startups
AI-Generated Summary
Key Takeaways
- ✓GPU Ownership Economics: Renting an H100 GPU from AWS or GCP costs $3.50–$5 per hour, totaling $35,000–$50,000 annually, versus a $30,000 purchase price. Owning delivers roughly 1.5x cost savings per year, and GPUs remain warrantied for three years while staying usable for inference on older workloads indefinitely. NVIDIA's new buyback agreement with Blackstone, BlackRock, Apollo, and Goldman Sachs underwrites up to 25% of GPU resale value, creating a liquid secondary market that lowers financing risk further.
- ✓Colocated Compute for Model Training: Large-scale AI model training requires colocated GPU clusters with adjacent high-capacity memory — something impossible to replicate by renting spot instances from hyperscalers. Speechify's Simba 3.2 model, ranked first globally for text-to-speech quality at $10 per million characters versus ElevenLabs' $100, was built on owned hardware. Engineers on rented compute self-censor experiments due to cost anxiety; owned hardware removes that friction entirely and accelerates research velocity.
- ✓B2B API Strategy Mistake: Weitzman admits that dismissing ElevenLabs' API-first approach in 2022 was Speechify's single largest strategic error. He assumed text-to-speech APIs would commoditize, missing that the API serves as a wedge — voice cloning, emotional prosody, speech-to-text, and duplex conversation models all layer on top. The lesson: offer a core product at low or no cost to embed in customer stacks, then expand product surface area continuously as the relationship deepens.
- ✓AI Hiring Process Overhaul: Replace traditional coding interviews with two functional tests: build a specified feature and run it through unit tests, then navigate a large open-source codebase, make targeted changes, and identify what breaks. Prioritize Math Olympiad winners, Kaggle award holders, and physics or mathematics graduates over candidates with conventional software engineering backgrounds. Raw technical intelligence and work ethic now matter more than prior coding experience because agents can teach syntax; humans must supply judgment and architectural thinking.
- ✓Production Shipping as Competitive Moat: Weitzman's internal rule: engineers receive zero credit until a feature reaches production with no bugs and real users engage with it. He demonstrates this by screen-sharing live product use during team calls, recording all bugs, and sending the recording to engineers immediately. A 19-year-old engineer resolved 14 product notes overnight between training runs. Speed from hypothesis to user feedback, not benchmark performance, is what separated Speechify from competitors across 770 billion words served.
What It Covers
Speechify CEO Cliff Weitzman covers three interconnected topics with Harry Stebbings: why buying NVIDIA GPUs outright beats renting from hyperscalers by 1.5x annually, how ceding the B2B API market to ElevenLabs was his biggest strategic error, and how AI-era hiring now prioritizes raw mathematical aptitude over traditional software engineering credentials.
Key Questions Answered
- •GPU Ownership Economics: Renting an H100 GPU from AWS or GCP costs $3.50–$5 per hour, totaling $35,000–$50,000 annually, versus a $30,000 purchase price. Owning delivers roughly 1.5x cost savings per year, and GPUs remain warrantied for three years while staying usable for inference on older workloads indefinitely. NVIDIA's new buyback agreement with Blackstone, BlackRock, Apollo, and Goldman Sachs underwrites up to 25% of GPU resale value, creating a liquid secondary market that lowers financing risk further.
- •Colocated Compute for Model Training: Large-scale AI model training requires colocated GPU clusters with adjacent high-capacity memory — something impossible to replicate by renting spot instances from hyperscalers. Speechify's Simba 3.2 model, ranked first globally for text-to-speech quality at $10 per million characters versus ElevenLabs' $100, was built on owned hardware. Engineers on rented compute self-censor experiments due to cost anxiety; owned hardware removes that friction entirely and accelerates research velocity.
- •B2B API Strategy Mistake: Weitzman admits that dismissing ElevenLabs' API-first approach in 2022 was Speechify's single largest strategic error. He assumed text-to-speech APIs would commoditize, missing that the API serves as a wedge — voice cloning, emotional prosody, speech-to-text, and duplex conversation models all layer on top. The lesson: offer a core product at low or no cost to embed in customer stacks, then expand product surface area continuously as the relationship deepens.
- •AI Hiring Process Overhaul: Replace traditional coding interviews with two functional tests: build a specified feature and run it through unit tests, then navigate a large open-source codebase, make targeted changes, and identify what breaks. Prioritize Math Olympiad winners, Kaggle award holders, and physics or mathematics graduates over candidates with conventional software engineering backgrounds. Raw technical intelligence and work ethic now matter more than prior coding experience because agents can teach syntax; humans must supply judgment and architectural thinking.
- •Production Shipping as Competitive Moat: Weitzman's internal rule: engineers receive zero credit until a feature reaches production with no bugs and real users engage with it. He demonstrates this by screen-sharing live product use during team calls, recording all bugs, and sending the recording to engineers immediately. A 19-year-old engineer resolved 14 product notes overnight between training runs. Speed from hypothesis to user feedback, not benchmark performance, is what separated Speechify from competitors across 770 billion words served.
- •Compound Startup Necessity: At a certain scale, staying single-product becomes a strategic liability. Speechify now competes simultaneously in consumer text-to-speech, B2B API, voice agents, speech-to-text, and a Siri competitor — not because it chose to diversify, but because incumbents like Apple repeatedly failed to ship and left gaps. The framework: identify where large players fumble execution, enter with a free or low-cost wedge product, accumulate users, then build adjacent products on top of the established distribution base.
Notable Moment
Weitzman describes sequencing his family member's blood weekly for 15 weeks, running proteomics and RNA analysis on a GPU cluster, and cross-referencing six years of daily symptom data — work no physician had attempted. He now organizes genome-sequencing meetups for others with the same rare disease to find shared epigenetic patterns.
Episode Transcript
100% is on. It's the biggest strategic mistake I made in the history of speech for fire. The best way to lose is not to be in the race. Be in the race. You don't want to be a fat manager who's like a general sitting in the back saying take that hill. You want to be the warrior who runs up with their sword and engages the enemy first. This is 20 VC with me, Harry Stebbings. Now I am fed up of the simple question answer back and forth podcast. Today is a real freaking discussion. Cliff Weitzman, founder and CEO of Speechify, one of the fastest growing text to speech startups in the world on the show where we have a real debate about whether it's right to scale into enterprise from a phenomenal consumer business, what it takes to build an amazing go to market motion when you've already built this amazing consumer business. And then he also tells us some wild freaking stories about spending tens of millions of dollars on NVIDIA GPUs and why so many more companies should be doing that over relying on other providers. This and so much more in the episode today. But before we dive into the show today, today, I wanna tell you about how the first AI law firm, Crosby, helped us close a big sponsor. As you know, some of the biggest companies in the world advertise on 20 VC, my British dulcet tones clearly convert well. Was working to close this big sponsor, and they wanted to get through legal review quite quickly to close the deal. Crosby turned red lines around in three hours and caught major issues that would have caused a serious problems in the future. Crosby combines AI, some of the best engineers in the world from companies like Ramp and Stripe, and some of the best attorneys in the world from top 10 law firms. Customers get the best of both worlds. An elite human attorney reviews every contract, but they move incredibly quickly, returning redlines in under four hours. They help the fastest growing companies like Cognition, Ramp, and Clay close deals in hours, not weeks. Learn more at crosby.ai/20vc. If you wanna redline NDAs, MSAs, DPAs, and any other procurement contracts faster, go to crosby.ai/20vc. It's speed that you can really trust. While Crosby keeps your numbers sharp, OneMind keeps your customer conversation sharper. Our friends over at OneMind have a hot take. The The b to b GTM model we've been using for, well, the last twenty years, it's collapsing. Predictable revenue isn't so predictable, and buyers are just tired of explaining themselves at every handoff between SDRs, AEs, CSMs, and support. You feel it in your board reporting. Your sellers feel it in their coverage. Your buyers feel it as they wait for answers. Well, enter OneMind and their GTM superhumans. HubSpot out on an investment that helped us close an $8,000,000 deal. 8,000,000, baby. That's …
Get the full transcript (14,499 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 62-minute episode.
Get 20VC (20 Minute VC) summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from 20VC (20 Minute VC)
20VC: NVIDIA Crushes Quarter and Buys Hugging Face | OpenAI Cuts Off Cursor | Instinct Hits $2.5BN Valuation and The Race for AI Assistants | Cognition Raises at $46BN, Linear $2.5BN and Clay $7BN
Sep 3 · 77 min
Cognitive Revolution
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
Aug 8
More from 20VC (20 Minute VC)
20VC: The AI Bubble Is Wrong | AI Margins Need to Improve | Revenue Concentration Should be a Concern | Why People Over-Estimate Open Models But Enterprises Still Fear Frontier Models with Aaron Katz, ClickHouse
Aug 31 · 63 min
a16z Podcast
Marc Andreessen on Evaluating Founders and AI's Consumer Surplus
Mar 30
More from 20VC (20 Minute VC)
We summarize every new episode. Want them in your inbox?
20VC: NVIDIA Crushes Quarter and Buys Hugging Face | OpenAI Cuts Off Cursor | Instinct Hits $2.5BN Valuation and The Race for AI Assistants | Cognition Raises at $46BN, Linear $2.5BN and Clay $7BN
20VC: The AI Bubble Is Wrong | AI Margins Need to Improve | Revenue Concentration Should be a Concern | Why People Over-Estimate Open Models But Enterprises Still Fear Frontier Models with Aaron Katz, ClickHouse
20VC: Is Anthropic's Coding Business Worth $2 Trillion? | Should American Enterprises Work With Open-Source Chinese Models? | Why 80–90% of Neo-Labs Die in the Next 18 Months? with Eno Reyes, Co-Founder @ Factory
20VC: NVIDIA Bonanza: Buys Poolside & Invests in Mercor and Perplexity | Anthropic's $30TRN Revenue Assumption & OpenAI Confirms IPO | Why Customer Service, Defence and Robotics are Overinflated
20VC: Inside Sequoia's Investment Committee: Lessons from Don Valentine, Doug Leone and Alfred Lin | How the SpaceX and Citadel Deals Went Down | What Sequoia Specifically Looks for in Founders with Julien Bek
Similar Episodes
Related episodes from other podcasts
Cognitive Revolution
Aug 8
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
a16z Podcast
Mar 30
Marc Andreessen on Evaluating Founders and AI's Consumer Surplus
The Ezra Klein Show
Aug 7
AMA: Peter Thiel, Chris Rufo and the D.S.A.
The Vergecast
Aug 5
Hotline: E Ink, the fediverse, and smart Puka shells
Up First (NPR)
Jul 11
Election Betting on Prediction Markets, Special Education, Breastmilk Storage
Explore Related Topics
This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into 20VC (20 Minute VC).
Every Monday, we deliver AI summaries of the latest episodes from 20VC (20 Minute VC) and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime