Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480]
Episode
87 min
Read time
3 min
Topics
Investing, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Low-Voltage Inference Architecture: GPUs thermal-throttle because voltage scales quadratically with power — doubling voltage quadruples power draw. Etched runs at under half the voltage of any competing AI chip by redesigning power delivery planes entirely. This unlocks significantly higher flop density without thermal throttling, enabling more compute per watt. Bitcoin miners already proved sub-quarter-voltage operation was physically possible; the question was whether transformer workloads could be restructured to support it.
- ✓Cluster-Scale Memory as Decode Advantage: The correct metric for decode performance is not single-chip memory bandwidth but full cluster memory bandwidth. NVIDIA Blackwell chip-to-chip latency runs approximately 4,000 nanoseconds point-to-point, meaning 8-chip tensor-parallel setups deliver far less than 8x throughput gains. Etched built a fully custom interconnect stack above Layer 2, cutting latency by more than 5x, enabling the entire cluster's SRAM and HBM to function as a unified memory pool for token generation.
- ✓Prefetch Everything Before Silicon Returns: Etched compressed post-silicon bring-up from the industry benchmark of 10 months down to 40 days by completing all parallel work before chips arrived. This included deploying 700 FPGAs running full inference stacks, shipping racks to customer data centers pre-chip for software validation, building thermal mock chips to validate cold plates, and standing up full production lines. Every workstream that did not require physical silicon was finished in advance.
- ✓Project-Based Legend Recruiting: Map every hard technical problem in the target domain, identify who specifically did the zero-to-one work — not who managed it — then pursue those individuals across 20+ conversations over months. Etched recruited Brian Leuler, who built NVIDIA's HGX and DGX rack systems representing the majority of NVIDIA's revenue, by identifying him as one of three people globally who fit the exact profile needed, then converting two of the other candidates into investors.
- ✓Bimodal Talent Philosophy — Legends Plus First-Principles Thinkers: Pair domain legends who know what scaled success looks like with young engineers who have no inherited constraints. Legends prevent billion-dollar mistakes; first-principles thinkers take aggressive risks legends would avoid. Etched pairs figures like Leuler with robotics world-record holders like Sanford, who built a functional cold plate prototype in one week — a task conventional thermal engineers would estimate at months — by simply not knowing it was considered impossible.
What It Covers
Etched founders Gavin Huberti and Rob Walken explain how they built a transformer-specific AI inference chip on the first tape-out attempt, raising $800M with over $1B in customer demand. They detail their low-voltage inference architecture, cluster-scale memory interconnects, and the vertical integration strategy behind their full rack-scale inference product launched in 2023.
Key Questions Answered
- •Low-Voltage Inference Architecture: GPUs thermal-throttle because voltage scales quadratically with power — doubling voltage quadruples power draw. Etched runs at under half the voltage of any competing AI chip by redesigning power delivery planes entirely. This unlocks significantly higher flop density without thermal throttling, enabling more compute per watt. Bitcoin miners already proved sub-quarter-voltage operation was physically possible; the question was whether transformer workloads could be restructured to support it.
- •Cluster-Scale Memory as Decode Advantage: The correct metric for decode performance is not single-chip memory bandwidth but full cluster memory bandwidth. NVIDIA Blackwell chip-to-chip latency runs approximately 4,000 nanoseconds point-to-point, meaning 8-chip tensor-parallel setups deliver far less than 8x throughput gains. Etched built a fully custom interconnect stack above Layer 2, cutting latency by more than 5x, enabling the entire cluster's SRAM and HBM to function as a unified memory pool for token generation.
- •Prefetch Everything Before Silicon Returns: Etched compressed post-silicon bring-up from the industry benchmark of 10 months down to 40 days by completing all parallel work before chips arrived. This included deploying 700 FPGAs running full inference stacks, shipping racks to customer data centers pre-chip for software validation, building thermal mock chips to validate cold plates, and standing up full production lines. Every workstream that did not require physical silicon was finished in advance.
- •Project-Based Legend Recruiting: Map every hard technical problem in the target domain, identify who specifically did the zero-to-one work — not who managed it — then pursue those individuals across 20+ conversations over months. Etched recruited Brian Leuler, who built NVIDIA's HGX and DGX rack systems representing the majority of NVIDIA's revenue, by identifying him as one of three people globally who fit the exact profile needed, then converting two of the other candidates into investors.
- •Bimodal Talent Philosophy — Legends Plus First-Principles Thinkers: Pair domain legends who know what scaled success looks like with young engineers who have no inherited constraints. Legends prevent billion-dollar mistakes; first-principles thinkers take aggressive risks legends would avoid. Etched pairs figures like Leuler with robotics world-record holders like Sanford, who built a functional cold plate prototype in one week — a task conventional thermal engineers would estimate at months — by simply not knowing it was considered impossible.
- •Vertical Integration Bounded by Economies of Scale: Integrate vertically only where doing so adds token capacity or removes a binding constraint — not as a default strategy. Etched builds chips, boards, cold plates, interconnects, and production lines in-house because each was a bottleneck. They do not build data centers because customers are already moving power infrastructure to accommodate Etched hardware. The natural integration boundaries sit at chip fabrication on one end and model architecture on the other, with full-stack ownership between.
Notable Moment
Rob Walken described uploading a pre-diagnosis photo of his back tumor — taken before his stage-four bone cancer diagnosis at age 16 — to GPT-4V, which immediately flagged it as a potential tumor requiring urgent MRI. A process that took six months of medical evaluation in 2015 took seconds in 2023, motivating his decision to build inference infrastructure.
Episode Transcript
Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth. Ramp customers grew revenue 3.2 times faster than the average American business. Visa, Vercel, Cursor, Stripe, Notion, ElevenLab, Shopify, and 70,000 other businesses all run on ramp. Mine does too, and so should yours. Learn more at ramp.com/invest. OpenAI, Cursor, Anthropic, Perplexity, and Vercel all have something in common. They all use Work OS. And here's why. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, RBAC, and audit logs. That's where Work OS comes in. Instead of spending months building these mission critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit workos.com to get started. Felix by Rogo is a personal finance agent that turns a single prompt into finished client ready work using your firm's own templates, context, and standards. Send Felix an email like, take these comments and turn them for me, or update my tracker with the context of these emails. Or run the ability to pay math on this buyer, and Felix sends back finished PowerPoint decks, Excel models, and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo.ai/felix. Hello, and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and wanna go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at colossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guests today are Gavin Huberti and Rob Walken, the founders of Etched. A few years ago when they set out to build a better AI chip than the largest companies in the world, almost everyone I called told me it could not be done. They have since done it, taping out a working chip on their first attempt and becoming the first hardware company founded after ChatGPT to do so. They already have more than a billion dollars of customer demand for their first product and have raised $800,000,000 to …
Get the full transcript (19,386 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
Browse all Invest Like the Best with Patrick O'Shaughnessy transcripts →
You just read a 3-minute summary of a 84-minute episode.
Get Invest Like the Best with Patrick O'Shaughnessy summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Invest Like the Best with Patrick O'Shaughnessy
Eric Vishria - A Decade of Lessons Investing in Software & Hardware - [Invest Like the Best, EP.486]
Aug 11 · 65 min
a16z Podcast
AI for America's Small Businesses | Lassie
Jul 30
More from Invest Like the Best with Patrick O'Shaughnessy
Gavin Baker - AI Market Jitters - [Invest Like the Best, EP.485]
Aug 4 · 65 min
a16z Podcast
Building AI Agents for Enterprise Operations
Jun 1
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
by OpenAI
“Rob Walken described uploading a pre-diagnosis photo of his back tumor — taken before his stage-four bone cancer diagnosis at age 16 — to GPT-4V, which immediately flagged it as a potential tumor requiring urgent MRI.”
Gear
by NVIDIA
“NVIDIA Blackwell chip-to-chip latency runs approximately 4,000 nanoseconds point-to-point, meaning 8-chip tensor-parallel setups deliver far less than 8x throughput gains.”
by NVIDIA
“Etched recruited Brian Leuler, who built NVIDIA's HGX and DGX rack systems representing the majority of NVIDIA's revenue”
by NVIDIA
“Etched recruited Brian Leuler, who built NVIDIA's HGX and DGX rack systems representing the majority of NVIDIA's revenue”
More from Invest Like the Best with Patrick O'Shaughnessy
We summarize every new episode. Want them in your inbox?
Eric Vishria - A Decade of Lessons Investing in Software & Hardware - [Invest Like the Best, EP.486]
Gavin Baker - AI Market Jitters - [Invest Like the Best, EP.485]
Sam Altman - How to Make an Abundant Future - [Invest Like the Best, EP.484]
Matthew Smith — Natural Gas: The Next Bottleneck - [Invest Like the Best, EP.483]
John Kim - How to Raise a Few Billion Dollars - [Invest Like the Best, EP.482]
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Jul 30
AI for America's Small Businesses | Lassie
a16z Podcast
Jun 1
Building AI Agents for Enterprise Operations
Capital Allocators
Mar 9
Katelin Holloway – Human Side of Venture Investing at 776 (EP.490)
Equity
Feb 10
This Sequoia-backed lab thinks the brain is 'the floor, not the ceiling' for AI
The Rework Podcast
Oct 8
Built on Trust
Explore Related Topics
This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Invest Like the Best with Patrick O'Shaughnessy.
Every Monday, we deliver AI summaries of the latest episodes from Invest Like the Best with Patrick O'Shaughnessy and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime