Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute
Episode
151 min
Read time
3 min
Topics
Productivity, Investing, Startups
AI-Generated Summary
Key Takeaways
- ✓CapEx-to-compute timeline: Hyperscaler CapEx announcements (Google's $180B, combined big-four $600B) do not represent compute coming online this year. A significant portion funds turbine deposits for 2028–2029, data center construction for 2027, and power purchasing agreements years out. Roughly 20 incremental gigawatts deploy in the US this year. Investors and analysts should model CapEx as a multi-year pipeline, not a current-year capacity signal, to avoid misreading AI infrastructure buildout pace.
- ✓Commitment timing determines compute margins: AI labs that signed five-year compute contracts early (OpenAI with Microsoft, CoreWeave, Oracle) locked in pricing at roughly $1.40/hour for H100s. Anthropic's conservative approach forced it to acquire last-minute capacity at spot rates as high as $2.40/hour for two-to-three year Hopper deals. The margin difference between early commitment and late acquisition is roughly 70%, making compute procurement timing one of the highest-leverage strategic decisions an AI lab makes.
- ✓GPU value appreciates with model capability: Contrary to Michael Burry's thesis that GPU depreciation cycles are two to three years, H100s are worth more today than at launch. GPT-5.4 runs on H100s at higher token throughput than GPT-4 did, while producing higher-quality output. As models improve efficiency and capability simultaneously, older hardware serving newer models extracts more economic value per chip. This inverts standard depreciation logic and supports longer five-year contract structures for cloud providers.
- ✓ASML EUV tools are the hard ceiling on AI compute: ASML produces roughly 70 EUV tools annually, scaling to approximately 100 by 2030. Each gigawatt of AI data center capacity requires 3.5 EUV tools worth of wafer passes across logic and memory. With 700 total EUV tools deployed by decade's end and assuming 25% AI allocation, the realistic ceiling is around 50 gigawatts per year — consistent with Sam Altman's targets but leaving no room for Elon Musk's 100-gigawatt ambitions without crowding out all consumer semiconductor production.
- ✓HBM memory creates a 4x demand destruction multiplier on consumer devices: HBM requires three to four times more DRAM wafer area per bit than standard DRAM. Every byte of HBM capacity allocated to AI effectively destroys four bytes of consumer device capacity. This is already cutting mid-range and low-end smartphone volumes — Xiaomi and Oppo are reportedly halving production. iPhone memory costs are projected to rise by $100–$150 per unit. Investors in consumer electronics should model significant margin compression and volume decline through 2026.
What It Covers
Dylan Patel, CEO of SemiAnalysis, breaks down the three compounding bottlenecks constraining AI compute scaling through 2030: semiconductor manufacturing capacity (logic wafers, HBM memory, EUV tooling), power and data center infrastructure, and capital deployment timing. The conversation quantifies how $600B in hyperscaler CapEx translates to actual gigawatts, why Anthropic undershot compute commitments, and why ASML's 70 machines per year caps the entire AI buildout.
Key Questions Answered
- •CapEx-to-compute timeline: Hyperscaler CapEx announcements (Google's $180B, combined big-four $600B) do not represent compute coming online this year. A significant portion funds turbine deposits for 2028–2029, data center construction for 2027, and power purchasing agreements years out. Roughly 20 incremental gigawatts deploy in the US this year. Investors and analysts should model CapEx as a multi-year pipeline, not a current-year capacity signal, to avoid misreading AI infrastructure buildout pace.
- •Commitment timing determines compute margins: AI labs that signed five-year compute contracts early (OpenAI with Microsoft, CoreWeave, Oracle) locked in pricing at roughly $1.40/hour for H100s. Anthropic's conservative approach forced it to acquire last-minute capacity at spot rates as high as $2.40/hour for two-to-three year Hopper deals. The margin difference between early commitment and late acquisition is roughly 70%, making compute procurement timing one of the highest-leverage strategic decisions an AI lab makes.
- •GPU value appreciates with model capability: Contrary to Michael Burry's thesis that GPU depreciation cycles are two to three years, H100s are worth more today than at launch. GPT-5.4 runs on H100s at higher token throughput than GPT-4 did, while producing higher-quality output. As models improve efficiency and capability simultaneously, older hardware serving newer models extracts more economic value per chip. This inverts standard depreciation logic and supports longer five-year contract structures for cloud providers.
- •ASML EUV tools are the hard ceiling on AI compute: ASML produces roughly 70 EUV tools annually, scaling to approximately 100 by 2030. Each gigawatt of AI data center capacity requires 3.5 EUV tools worth of wafer passes across logic and memory. With 700 total EUV tools deployed by decade's end and assuming 25% AI allocation, the realistic ceiling is around 50 gigawatts per year — consistent with Sam Altman's targets but leaving no room for Elon Musk's 100-gigawatt ambitions without crowding out all consumer semiconductor production.
- •HBM memory creates a 4x demand destruction multiplier on consumer devices: HBM requires three to four times more DRAM wafer area per bit than standard DRAM. Every byte of HBM capacity allocated to AI effectively destroys four bytes of consumer device capacity. This is already cutting mid-range and low-end smartphone volumes — Xiaomi and Oppo are reportedly halving production. iPhone memory costs are projected to rise by $100–$150 per unit. Investors in consumer electronics should model significant margin compression and volume decline through 2026.
- •Inference performance gap between Hopper and Blackwell is 20x, not 3x: Raw FLOP comparisons between GPU generations understate real-world performance differences. Running DeepSeek or Kimi K2.5 on Hopper versus Blackwell yields roughly 20x throughput difference at 100 tokens per second, driven by memory bandwidth, NVLink interconnect speed, and architectural improvements rather than transistor count alone. This means going back to older process nodes (seven nanometer DUV) to bypass EUV constraints would sacrifice far more performance than FLOP-per-dollar calculations suggest.
- •Fast AI timelines favor the US; slow timelines favor China: China currently operates entirely on ASML DUV tools and lacks indigenized EUV capability, though working tools are plausible by 2030 with mass production lagging several years behind. If AI revenue compounds fast enough that US labs reach 10+ gigawatt scale by end of 2026, the economic gap widens before China can close the semiconductor supply chain. If capability timelines extend to 2035, China's vertically integrated domestic supply chain and engineering scale become structural advantages.
Notable Moment
Patel reveals that Carl Zeiss — the German optics supplier whose precision mirrors are the bottleneck inside ASML's EUV machines, which are themselves the bottleneck for all advanced AI chips — has a market capitalization of roughly $2.5 billion. A company worth less than a mid-size tech startup effectively sets the hard ceiling on global AI compute expansion through the end of the decade.
Episode Transcript
Alright. This is the episode of my roommate teaches me semiconductors. It's also the send off for this, this current set. It's yeah. You're you know, after you use it, I'm like, I can't use this again. I gotta get out of here. Those sloppy suckers for dark guys. Okay. Dylan is the, CEO of Semia Analysis. Dylan, the burning question I have for you, if you add up the big four, Amazon, Meta, Google, Microsoft, their combined, forecasted CapEx that you published recently this year is $600,000,000,000. And given, you know, yearly prices of renting that compute, that would be, like, close to 50 gigawatts. Now, obviously, we're not putting on 50 gigawatts this year. So presumably, that's paying for compute that is gonna be coming online over the coming years. So I have a question about what how to think about the timeline around when that CapEx comes online. Similar question for the labs where, you know, OpenAI just announced that they raised a $110,000,000,000. Anthropic just announced they raised $30,000,000,000. And if you look at the compute that they have coming online this year, you you should tell me how much it is. But, like, is it not is it not another four gigawatts total that they'll have this year? It feels like the cost to rent the compute that OpenAI and Anthropic will have this year to, like, sustain their compute spend at, you know, $1,013,000,000,000 dollars a gigawatt. Those individual raises alone are, like, enough to cover their compute spend for the year. And then this is not even including the revenue that they're gonna earn this year. So help me understand, first, when is the time scale at which the big tech CapEx is actually coming online? And two, what are the labs raising all this money for if, like, the the yearly price of a a one gigawatt data center is, like, $13,000,000,000? So when you talk about the CapEx of these hyperscalers, right, on the order of $600,000,000,000 and you look at the across the rest of the supply chain, gets you to on the order of a trillion dollars. A portion of this is, you know, immediately for compute going online this year. Right? The chips and the, the the other parts of CapEx that do get paid this year. But there's a lot of setup CapEx as well. Right? So when we have when we're talking about 20 gigawatts this year in America, roughly Incremental. Incremental added capacity, A portion of this is not spent this year. A portion of that CapEx is actually spent the prior year. And so when you look at, hey. Google's got a $180,000,000,000. Actually, a big chunk of that is spent on turbine deposits for '28 '29. A chunk of that is spent on data center construction for '27. A chunk of that is spent on, you know, power purchasing agreements and down payments and all these other things that they're doing, for further …
Get the full transcript (31,282 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 148-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
AI researchers debate how close we are to recursive self-improvement
Sep 11 · 97 min
Invest Like the Best with Patrick O'Shaughnessy
Dylan Patel - The Infinite Demand for Tokens, Claude Mythos, and Supply Constraints - [Invest Like the Best, EP.468]
Apr 23
More from Dwarkesh Podcast
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1 · 140 min
10% Happier with Dan Harris
The Science Of Building The Life You Want | Arthur Brooks
Sep 7
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
company
“Dylan Patel, CEO of SemiAnalysis, breaks down the three compounding bottlenecks constraining AI compute scaling through 2030.”
“Carl Zeiss — the German optics supplier whose precision mirrors are the bottleneck inside ASML's EUV machines, which are themselves the bottleneck for all advanced AI chips — has a market capitalization of roughly $2.5 billion.”
“AI labs that signed five-year compute contracts early (OpenAI with Microsoft, CoreWeave, Oracle) locked in pricing at roughly $1.40/hour for H100s.”
“ASML EUV tools are the hard ceiling on AI compute. ASML produces roughly 70 EUV tools annually, scaling to approximately 100 by 2030.”
“Anthropic's conservative approach forced it to acquire last-minute capacity at spot rates as high as $2.40/hour for two-to-three year Hopper deals.”
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
AI researchers debate how close we are to recursive self-improvement
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
The rise and fall of agent civilizations
Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
Similar Episodes
Related episodes from other podcasts
Invest Like the Best with Patrick O'Shaughnessy
Apr 23
Dylan Patel - The Infinite Demand for Tokens, Claude Mythos, and Supply Constraints - [Invest Like the Best, EP.468]
10% Happier with Dan Harris
Sep 7
The Science Of Building The Life You Want | Arthur Brooks
20VC (20 Minute VC)
Aug 29
20VC: Is Anthropic's Coding Business Worth $2 Trillion? | Should American Enterprises Work With Open-Source Chinese Models? | Why 80–90% of Neo-Labs Die in the Next 18 Months? with Eno Reyes, Co-Founder @ Factory
Software Engineering Daily
Aug 25
The Gap Between AI Spending and AI Value
Huberman Lab
Aug 6
Essentials: Control Your Brain Chemistry for Focus, Motivation & Well-Being
Explore Related Topics
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime