Why Cerebras CEO Andrew Feldman Built The World's Largest Computer Chip
Episode
51 min
Read time
2 min
Topics
Fundraising & VC, Leadership, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Wafer-Scale Memory Architecture: Cerebras achieves 15x faster inference than GPUs—and up to 1,000x faster on specific workloads—by using fast SRAM instead of slow HBM memory. The tradeoff is lower storage density per square millimeter, solved by building a chip covering an entire silicon wafer, roughly dinner-plate sized, stuffed with high-speed memory.
- ✓Speed Premium Pricing: Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed. Cerebras operates at 15x faster than that tier, suggesting substantial pricing power. Slow tokens cost less to produce on GPUs, but GPU cost-per-token rises sharply as speed requirements increase.
- ✓Supply Chain Differentiation: Cerebras avoids three major AI chip bottlenecks simultaneously: HBM memory shortages, TSMC's constrained CoWoS packaging process, and TSMC's oversubscribed 3nm node. By using 5nm fabrication and on-chip SRAM, Cerebras sidesteps constraints choking NVIDIA and other GPU vendors, leaving data center availability as the primary growth limiter.
- ✓CUDA Moat Erosion: CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes. In training, two of three leading frontier models (Gemini on TPUs, Claude on Trainium) now train without CUDA, representing a 70% market share loss for NVIDIA's software ecosystem compared to three years ago.
- ✓Open vs. Closed Source Economics: Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization. Closed source models outperform open source by roughly 4–5% on quality benchmarks but cost significantly more per token, creating a cost-versus-capability tradeoff enterprises must actively evaluate.
What It Covers
Cerebras CEO Andrew Feldman explains how his company built a chip 58 times larger than any competitor, achieving inference speeds 15 times faster than leading GPUs. The episode covers wafer-scale engineering breakthroughs, inference economics, CUDA's declining relevance, open vs. closed source AI models, and semiconductor supply chain constraints.
Key Questions Answered
- •Wafer-Scale Memory Architecture: Cerebras achieves 15x faster inference than GPUs—and up to 1,000x faster on specific workloads—by using fast SRAM instead of slow HBM memory. The tradeoff is lower storage density per square millimeter, solved by building a chip covering an entire silicon wafer, roughly dinner-plate sized, stuffed with high-speed memory.
- •Speed Premium Pricing: Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed. Cerebras operates at 15x faster than that tier, suggesting substantial pricing power. Slow tokens cost less to produce on GPUs, but GPU cost-per-token rises sharply as speed requirements increase.
- •Supply Chain Differentiation: Cerebras avoids three major AI chip bottlenecks simultaneously: HBM memory shortages, TSMC's constrained CoWoS packaging process, and TSMC's oversubscribed 3nm node. By using 5nm fabrication and on-chip SRAM, Cerebras sidesteps constraints choking NVIDIA and other GPU vendors, leaving data center availability as the primary growth limiter.
- •CUDA Moat Erosion: CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes. In training, two of three leading frontier models (Gemini on TPUs, Claude on Trainium) now train without CUDA, representing a 70% market share loss for NVIDIA's software ecosystem compared to three years ago.
- •Open vs. Closed Source Economics: Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization. Closed source models outperform open source by roughly 4–5% on quality benchmarks but cost significantly more per token, creating a cost-versus-capability tradeoff enterprises must actively evaluate.
Notable Moment
Feldman reveals that despite solving a 75-year-old unsolvable engineering problem and building the world's fastest inference chip, Cerebras' primary growth constraint today is not manufacturing capacity or software—it is simply the availability of powered data center buildings, a limitation expected to persist for at least 15–18 months.
Episode Transcript
Odd lots is brought to you by VanEck. For years, investors basically forgot about real assets, energy, gold, and infrastructure. But look what's driving markets now. Central banks loading up on gold, massive CapEx cycles, currencies doing weird things. These assets are at the center of it. Racks, the VanEck real asset ETF, is an actively managed one stop shop for real assets spanning gold, commodities, natural resource equities, and more. Go to vaneck.com/raaxpod to learn more. Fun disclosures later in this episode. So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now a global workforce of 300,000 can use AI to fill their HR questions, resolving 94% of common questions. Not noise, proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM. You need to make a huge presentation in an hour. Adobe Acrobat uses AI to take all your documents and generate a presentation with a single click. Build slides quickly and streamline the process. Need a last minute pitch deck? Do that with Acrobat. Need to level up your presentation design? Do that with Acrobat. You have 30 plus documents that need to be simplified into a proposal. Do that. Do that. Do that with Acrobat. Learn more at adobe.com slash do that with Acrobat. Bloomberg Audio Studios. Podcasts, radio, news. Hello, and welcome to another episode of the Odd Lots podcast. I'm Jill Weisenthal. And I'm Tracy Alloway. Tracy, I have to say, unfortunately, I don't have AI psychosis. I'm certain of that. Debatable. I'm pretty sir I'm pretty sure I don't have AI psychosis. I do have to say, unfortunately, like, the amount of time now where it's like it feels like AI related questions Mhmm. And there's many of them, are sort of, like, swallowing up the other thoughts that I have in my head, whether it's questions about which model is best and why, and what are the economics of inference, and how much training is pretraining versus posttraining for each model. Like, it's just sort of like this blob that's growing that's taking up more and more of my thoughts. What is your definition of AI psychosis? Because one would argue that maybe thinking about AI literally all the time would be a form of psychosis. Well, let's just say, like, I'm not the type who thinks that, like, I don't, like, think that the AI is a friend, for one thing. I'm not in love with the AI models. I don't think that in collaboration with ChatGPT that I'm stumbling on, unified theory of physics Mhmm. And things like that. So, like But you do spend a lot of time inputting instructions, pressing the button, and seeing what comes out. And seeing what comes …
Get the full transcript (10,003 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 48-minute episode.
Get Odd Lots summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Odd Lots
Jasmine Sun on What the AI Industry Got Wrong About the Public Backlash
Aug 21 · 51 min
20VC (20 Minute VC)
20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman
May 26
More from Odd Lots
Nick Bostrom on What Happens if AI Solves All of Our Problems
Aug 20 · 55 min
No Priors: Artificial Intelligence | Technology | Startups
The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman
May 21
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
“Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization.”
by NVIDIA
“CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes.”
Gear
- Cerebras ChipBy guest
by Cerebras
“Cerebras CEO Andrew Feldman explains how his company built a chip 58 times larger than any competitor, achieving inference speeds 15 times faster than leading GPUs.”
Products
by Anthropic
“Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed.”
More from Odd Lots
We summarize every new episode. Want them in your inbox?
Jasmine Sun on What the AI Industry Got Wrong About the Public Backlash
Nick Bostrom on What Happens if AI Solves All of Our Problems
What the OpenAI-Hugging Face Hack Really Tells Us About AI Danger
A Historic El Niño Is Coming That Could Cost the World Trillions
Trucking Is Booming Again, And Drivers Aren't Happy About It
Similar Episodes
Related episodes from other podcasts
20VC (20 Minute VC)
May 26
20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman
No Priors: Artificial Intelligence | Technology | Startups
May 21
The Story Behind Cerebras’ $63 Billion IPO with Founder and CEO Andrew Feldman
Latent Space
May 21
Giving Agents Computers — Ivan Burazin, Daytona
All-In with Chamath, Jason, Sacks & Friedberg
Jan 23
Coinbase CEO Brian Armstrong Breaks Down the Three Biggest Trends in Crypto + More from Davos!
20VC (20 Minute VC)
Oct 6
20VC: Cerebras CEO on Why Raise $1BN and Delay the IPO | NVIDIA Showing Signs They Are Worried About Growth | Concentration of Value in Mag7: Will the AI Train Come to a Halt | Can the US Supply the Energy for AI with Andrew Feldman
Explore Related Topics
This podcast is featured in Best Finance Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Odd Lots.
Every Monday, we deliver AI summaries of the latest episodes from Odd Lots and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime