Skip to main content
Odd Lots

Why Cerebras CEO Andrew Feldman Built The World's Largest Computer Chip

51 min episode · 2 min read
·

Episode

51 min

Read time

2 min

Topics

Fundraising & VC, Leadership, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Wafer-Scale Memory Architecture: Cerebras achieves 15x faster inference than GPUs—and up to 1,000x faster on specific workloads—by using fast SRAM instead of slow HBM memory. The tradeoff is lower storage density per square millimeter, solved by building a chip covering an entire silicon wafer, roughly dinner-plate sized, stuffed with high-speed memory.
  • Speed Premium Pricing: Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed. Cerebras operates at 15x faster than that tier, suggesting substantial pricing power. Slow tokens cost less to produce on GPUs, but GPU cost-per-token rises sharply as speed requirements increase.
  • Supply Chain Differentiation: Cerebras avoids three major AI chip bottlenecks simultaneously: HBM memory shortages, TSMC's constrained CoWoS packaging process, and TSMC's oversubscribed 3nm node. By using 5nm fabrication and on-chip SRAM, Cerebras sidesteps constraints choking NVIDIA and other GPU vendors, leaving data center availability as the primary growth limiter.
  • CUDA Moat Erosion: CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes. In training, two of three leading frontier models (Gemini on TPUs, Claude on Trainium) now train without CUDA, representing a 70% market share loss for NVIDIA's software ecosystem compared to three years ago.
  • Open vs. Closed Source Economics: Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization. Closed source models outperform open source by roughly 4–5% on quality benchmarks but cost significantly more per token, creating a cost-versus-capability tradeoff enterprises must actively evaluate.

What It Covers

Cerebras CEO Andrew Feldman explains how his company built a chip 58 times larger than any competitor, achieving inference speeds 15 times faster than leading GPUs. The episode covers wafer-scale engineering breakthroughs, inference economics, CUDA's declining relevance, open vs. closed source AI models, and semiconductor supply chain constraints.

Key Questions Answered

  • Wafer-Scale Memory Architecture: Cerebras achieves 15x faster inference than GPUs—and up to 1,000x faster on specific workloads—by using fast SRAM instead of slow HBM memory. The tradeoff is lower storage density per square millimeter, solved by building a chip covering an entire silicon wafer, roughly dinner-plate sized, stuffed with high-speed memory.
  • Speed Premium Pricing: Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed. Cerebras operates at 15x faster than that tier, suggesting substantial pricing power. Slow tokens cost less to produce on GPUs, but GPU cost-per-token rises sharply as speed requirements increase.
  • Supply Chain Differentiation: Cerebras avoids three major AI chip bottlenecks simultaneously: HBM memory shortages, TSMC's constrained CoWoS packaging process, and TSMC's oversubscribed 3nm node. By using 5nm fabrication and on-chip SRAM, Cerebras sidesteps constraints choking NVIDIA and other GPU vendors, leaving data center availability as the primary growth limiter.
  • CUDA Moat Erosion: CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes. In training, two of three leading frontier models (Gemini on TPUs, Claude on Trainium) now train without CUDA, representing a 70% market share loss for NVIDIA's software ecosystem compared to three years ago.
  • Open vs. Closed Source Economics: Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization. Closed source models outperform open source by roughly 4–5% on quality benchmarks but cost significantly more per token, creating a cost-versus-capability tradeoff enterprises must actively evaluate.

Notable Moment

Feldman reveals that despite solving a 75-year-old unsolvable engineering problem and building the world's fastest inference chip, Cerebras' primary growth constraint today is not manufacturing capacity or software—it is simply the availability of powered data center buildings, a limitation expected to persist for at least 15–18 months.

Know someone who'd find this useful?

Episode Transcript

Odd lots is brought to you by VanEck. For years, investors basically forgot about real assets, energy, gold, and infrastructure. But look what's driving markets now. Central banks loading up on gold, massive CapEx cycles, currencies doing weird things. These assets are at the center of it. Racks, the VanEck real asset ETF, is an actively managed one stop shop for real assets spanning gold, commodities, natural resource equities, and more. Go to vaneck.com/raaxpod to learn more. Fun disclosures later in this episode. So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now a global workforce of 300,000 can use AI to fill their HR questions, resolving 94% of common questions. Not noise, proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM. You need to make a huge presentation in an hour. Adobe Acrobat uses AI to take all your documents and generate a presentation with a single click. Build slides quickly and streamline the process. Need a last minute pitch deck? Do that with Acrobat. Need to level up your presentation design? Do that with Acrobat. You have 30 plus documents that need to be simplified into a proposal. Do that. Do that. Do that with Acrobat. Learn more at adobe.com slash do that with Acrobat. Bloomberg Audio Studios. Podcasts, radio, news. Hello, and welcome to another episode of the Odd Lots podcast. I'm Jill Weisenthal. And I'm Tracy Alloway. Tracy, I have to say, unfortunately, I don't have AI psychosis. I'm certain of that. Debatable. I'm pretty sir I'm pretty sure I don't have AI psychosis. I do have to say, unfortunately, like, the amount of time now where it's like it feels like AI related questions Mhmm. And there's many of them, are sort of, like, swallowing up the other thoughts that I have in my head, whether it's questions about which model is best and why, and what are the economics of inference, and how much training is pretraining versus posttraining for each model. Like, it's just sort of like this blob that's growing that's taking up more and more of my thoughts. What is your definition of AI psychosis? Because one would argue that maybe thinking about AI literally all the time would be a form of psychosis. Well, let's just say, like, I'm not the type who thinks that, like, I don't, like, think that the AI is a friend, for one thing. I'm not in love with the AI models. I don't think that in collaboration with ChatGPT that I'm stumbling on, unified theory of physics Mhmm. And things like that. So, like But you do spend a lot of time inputting instructions, pressing the button, and seeing what comes out. And seeing what comes …

Get the full transcript (10,003 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Odd Lots transcripts →

You just read a 3-minute summary of a 48-minute episode.

Get Odd Lots summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • Open source models like Kimi K2 (1 trillion parameters) run on Cerebras today at a cost reflecting only compute and power—not training amortization.
  • by NVIDIA

    CUDA has zero role in inference workloads—migrating a model from GPU to Cerebras requires roughly 10 configuration changes.

Gear

  • by Cerebras

    Cerebras CEO Andrew Feldman explains how his company built a chip 58 times larger than any competitor, achieving inference speeds 15 times faster than leading GPUs.

Products

  • by Anthropic

    Anthropic's 2x-faster inference tier sold out at 6x the standard price, demonstrating that enterprise buyers pay significant premiums for speed.

More from Odd Lots

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Finance Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Odd Lots.

Every Monday, we deliver AI summaries of the latest episodes from Odd Lots and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime