The System Behind Self-Driving: Waymo’s Dmitri Dolgov
Episode
64 min
Read time
3 min
Topics
Investing, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Sensor Fusion Architecture: Waymo uses three complementary sensing modalities — cameras, LiDAR, and radar — each with 360-degree coverage. Rather than switching between sensors, all three feed separate encoders that jointly produce a unified world model. Radar excels in fog and heavy rain where cameras degrade; LiDAR provides high-resolution 3D structure. No real-time cloud dependency exists; all safety-critical inference runs locally onboard the vehicle.
- ✓Foundation Model Distillation Pipeline: Waymo builds one large off-board foundation model, then specializes it into three onboard "teacher" models — the Driver, the Simulator, and the Critic. Each teacher distills a smaller, faster "student" model deployable on the vehicle. This architecture enables closed-loop reinforcement learning fine-tuning, realistic synthetic environment generation, and automated behavioral evaluation without requiring pixel-level simulation throughout the entire training pipeline.
- ✓Full Autonomy vs. Driver Assist — A Qualitative Gap: Dolgov argues that driver-assist systems and full autonomy are fundamentally different engineering problems, not points on a single spectrum. A basic vision-language model fine-tuned on trajectories can handle nominal driving but falls orders of magnitude short of the safety threshold required for driverless operation. Reaching full autonomy requires the Simulator and Critic infrastructure that driver-assist development never demands, making incremental convergence from Level 2 upward practically implausible.
- ✓Generation 6 Hardware Cost Reduction: Waymo's sixth-generation sensor stack costs a fraction of the fifth generation — comparable to a premium ADAS system — through unification and simplification across all three modalities. The driving software stack transfers largely unchanged across hardware generations and vehicle platforms, including the upcoming Hyundai Ioniq deployment. LiDAR, radar, and camera component costs follow predictable downward trends as automotive supply chains mature and manufacturing volumes increase.
- ✓Scaling Signals and City Expansion Velocity: Waymo operates 3,000 vehicles across 11 U.S. cities, generating roughly 4 million fully autonomous miles per week. The company launched riders in four new cities simultaneously in a single day — a milestone that took eight years to achieve from first autonomous passenger operation in Chandler, Arizona in 2020. London and Tokyo deployments are planned for 2025, with the core technology generalizing well to new geographies with targeted data collection and validation work.
What It Covers
Waymo co-CEO Dmitri Dolgov explains the technical architecture behind 500,000 weekly autonomous rides, covering the sensor fusion stack, the foundation model distillation pipeline, why driver-assist systems cannot incrementally evolve into full autonomy, and how Generation 6 hardware cuts costs to levels comparable to premium ADAS systems while enabling accelerated global deployment.
Key Questions Answered
- •Sensor Fusion Architecture: Waymo uses three complementary sensing modalities — cameras, LiDAR, and radar — each with 360-degree coverage. Rather than switching between sensors, all three feed separate encoders that jointly produce a unified world model. Radar excels in fog and heavy rain where cameras degrade; LiDAR provides high-resolution 3D structure. No real-time cloud dependency exists; all safety-critical inference runs locally onboard the vehicle.
- •Foundation Model Distillation Pipeline: Waymo builds one large off-board foundation model, then specializes it into three onboard "teacher" models — the Driver, the Simulator, and the Critic. Each teacher distills a smaller, faster "student" model deployable on the vehicle. This architecture enables closed-loop reinforcement learning fine-tuning, realistic synthetic environment generation, and automated behavioral evaluation without requiring pixel-level simulation throughout the entire training pipeline.
- •Full Autonomy vs. Driver Assist — A Qualitative Gap: Dolgov argues that driver-assist systems and full autonomy are fundamentally different engineering problems, not points on a single spectrum. A basic vision-language model fine-tuned on trajectories can handle nominal driving but falls orders of magnitude short of the safety threshold required for driverless operation. Reaching full autonomy requires the Simulator and Critic infrastructure that driver-assist development never demands, making incremental convergence from Level 2 upward practically implausible.
- •Generation 6 Hardware Cost Reduction: Waymo's sixth-generation sensor stack costs a fraction of the fifth generation — comparable to a premium ADAS system — through unification and simplification across all three modalities. The driving software stack transfers largely unchanged across hardware generations and vehicle platforms, including the upcoming Hyundai Ioniq deployment. LiDAR, radar, and camera component costs follow predictable downward trends as automotive supply chains mature and manufacturing volumes increase.
- •Scaling Signals and City Expansion Velocity: Waymo operates 3,000 vehicles across 11 U.S. cities, generating roughly 4 million fully autonomous miles per week. The company launched riders in four new cities simultaneously in a single day — a milestone that took eight years to achieve from first autonomous passenger operation in Chandler, Arizona in 2020. London and Tokyo deployments are planned for 2025, with the core technology generalizing well to new geographies with targeted data collection and validation work.
- •Emergent AI Behavior as a Capability Signal: A concrete example of emergent model capability occurred when a Waymo vehicle detected a pedestrian obscured behind a bus using peripheral LiDAR returns bouncing beneath the vehicle chassis — a detection method no engineer explicitly programmed. This type of emergent behavior, enabled by intermediate world representations rather than pure pixel-to-trajectory end-to-end models, signals that the foundation model approach produces capabilities that exceed explicit engineering specifications.
Notable Moment
Dolgov describes watching a Waymo vehicle detect a pedestrian hidden entirely behind a bus and respond correctly — then discovering the system had used faint LiDAR reflections bouncing under the bus chassis to infer the person's presence and predict their movement. No engineer designed this behavior; the model derived it independently from training.
Episode Transcript
When you're driving around or being driven around, say, you know, we think about what we're building as a driver. I can imagine building a big model that understands how the physical world works and understands the important properties of what it means to drive, the social aspects of driving, and what it means to be a good driver as opposed to a bad one. I would say that we've clearly moved past the stage of scientific research and kind of deep core technology development to this new phase of accelerated global scaling and deployment. Waymo is now doing nearly half a million fully autonomous rides a week across multiple cities, a shift from long term research to real world scale. In this episode, originally aired on the Cheeky Pint podcast, Waymo co CEO Dmitry Dolgov joins John Collison to break down how they built the system behind it, from the sensor stack and why lidar still matters to the role of simulation and critic models in training the AI. They also get into why driver assist won't naturally evolve into full autonomy, what it takes to scale globally, and how the product itself is changing from custom built vehicles to entirely new economies of ride hailing. Dmitry Dolgaard is co CEO of Waymo. He joined Google's self driving car project in 2009 as one of its first engineers and was repeatedly promoted until he took it over in 2021. Waymo is Google's most successful moonshot and now provides over 500,000 fully autonomous rides each week. Cheers, by the way. Yeah. Cheers. You grew up in Russia. Yes. I grew up in Russia. Yep. Then I I was actually Soviet Union. Right. Right. Exactly. My dad is a physicist. Mhmm. So the Soviet Union started falling apart, and then, you know, he, got I had a position a visiting position in university, in Kyoto University Mhmm. For a year. We moved there as a family, and then he went to Berkeley, and then kinda tied it along. And then I ran out of, you know, I graduated from high school. Yep. I was thinking about the next thing I wanted to do, and I really liked that, that that, technical school in Russia. The Russians are serious about the physics. They are. They are. So I went back to Russia, and I got my bachelor's and master's. What year was this that you went back to Russia? 1994. Okay. So that was kind of almost peak Russian optimism in a sense where It was. It was opening up. It was. Yeah. Yeah. No. I actually remember, talking to my mom about it. And, yeah, of course, my parents grew up in the Soviet Union. They've seen it. Yep. You know? I mean, they were born right before Yeah. The war, and then they saw you know, they lived through some really tough times. And I remember talking to my mom and saying, she she you know, in fact, I got …
Get the full transcript (11,655 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 61-minute episode.
Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from a16z Podcast
Daniel Litt: The Mathematician's Guide to AI
Sep 1 · 63 min
Cognitive Revolution
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
Aug 10
More from a16z Podcast
Gavin Baker: Why AI Demand Is Outrunning Compute Supply
Aug 31 · 75 min
The Rich Roll Podcast
The Most Decorated Winter Olympian Ever: Johannes Klæbo On Training, Mindset & Reinventing His Sport
Aug 10
More from a16z Podcast
We summarize every new episode. Want them in your inbox?
Daniel Litt: The Mathematician's Guide to AI
Gavin Baker: Why AI Demand Is Outrunning Compute Supply
Why a16z Launched the Machine Age Fund | Jen Kha
Why 1,200 AI Agents Started Working Together | Ryan Greenblatt
The Infrastructure Behind the Machine Age
Similar Episodes
Related episodes from other podcasts
Cognitive Revolution
Aug 10
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
The Rich Roll Podcast
Aug 10
The Most Decorated Winter Olympian Ever: Johannes Klæbo On Training, Mindset & Reinventing His Sport
Cognitive Revolution
Jul 4
Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models
Latent Space
May 28
The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray
Cognitive Revolution
Apr 8
Calm AI for Crazy Days: Inside Granola's Design Philosophy, with co-founder Sam Stephenson
Explore Related Topics
This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into a16z Podcast.
Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime