Skip to main content
NVIDIA AI Podcast

Autonomous Driving, Visual AI, and the Road Ahead with Porsche and Voxel51 - Ep. 267

41 min episode · 2 min read
·
Tim Sohne,Brian Moore

Episode

41 min

Read time

2 min

Topics

Fundraising & VC, Design & UX, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Data Quality Over Quantity: Auto labeling using foundation models achieves comparable performance to human annotation at lower cost and higher speed, removing the bottleneck of manually labeling billions of kilometers of driving data for training autonomous systems.
  • Simulation for Edge Cases: Synthetic data generation enables testing scenarios impossible to replicate safely in real world, like helicopter landings on roadways, while generative models like NVIDIA Cosmos improve simulation fidelity to near-video realism for validation.
  • Foundation Model Capabilities: Vision language action models require four competencies for autonomous navigation: semantic understanding (classes, attributes), spatial awareness (object locations), temporal reasoning (past and future states), and physical understanding (forces, vehicle dynamics). Current models excel at semantics but need improvement in other areas.
  • Situated Safety Approach: Future autonomous systems will shift from testing every possible scenario to reasoning-based safety, where models derive actions from basic concepts, explain decisions in natural language, and request driver takeover when encountering operational design domain boundaries.

What It Covers

Porsche's Tim Sohne and Voxel51's Brian Moore explain how autonomous vehicle development shifts from modular systems to end-to-end AI models, requiring massive data curation, synthetic simulation, and foundation models for safe operation.

Key Questions Answered

  • Data Quality Over Quantity: Auto labeling using foundation models achieves comparable performance to human annotation at lower cost and higher speed, removing the bottleneck of manually labeling billions of kilometers of driving data for training autonomous systems.
  • Simulation for Edge Cases: Synthetic data generation enables testing scenarios impossible to replicate safely in real world, like helicopter landings on roadways, while generative models like NVIDIA Cosmos improve simulation fidelity to near-video realism for validation.
  • Foundation Model Capabilities: Vision language action models require four competencies for autonomous navigation: semantic understanding (classes, attributes), spatial awareness (object locations), temporal reasoning (past and future states), and physical understanding (forces, vehicle dynamics). Current models excel at semantics but need improvement in other areas.
  • Situated Safety Approach: Future autonomous systems will shift from testing every possible scenario to reasoning-based safety, where models derive actions from basic concepts, explain decisions in natural language, and request driver takeover when encountering operational design domain boundaries.

Notable Moment

Researchers discovered that autonomous systems trained entirely on automatically labeled data from foundation models can match the performance of systems trained on expensive human-annotated datasets, fundamentally changing the economics and scale of AV development.

Know someone who'd find this useful?

Episode Transcript

Hello, and welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. It's often said that modern cars are computers rolling along on wheels. From performance and safety systems to in vehicle infotainment, computers control and oversee many, many functions in today's vehicles. Autonomous driving systems, of course, are no exception. The quest to build self driving cars depends on compute power and, yes, data, lots of data. Here to delve into the inner workings of autonomous vehicles and the increasingly vital role of data in modern car making, Our Tin Sohne, technical lead for vision language action models at Porsche, and Brian Moore, CEO and cofounder of Voxel fifty one, whose visual AI and computer vision data platform, 51, is used by customers across a range of industries, including, you guessed it, Portia. Tim, Brian, welcome to the NVIDIA AI podcast, and thank you so much for taking the time to join. Thanks for having us. Thank you for having us. So maybe we can start with each of you introducing yourselves and just kind of talking a little bit about what you do at your respective companies and and a little bit about how that relates to autonomous systems. And, of course, we'll get into it. So, Tim, maybe you can start. Right. So I'm Tim. I'm a PhD student and tech lead at Porsche AG, like you said. I'm dealing with vision language action models for autonomous driving. And in my research, I want to turn cars into embodied agents that can understand space, time, and physical properties of the real world so that they are able to act within it and interact with the driver through natural language or through pose, facial expressions, or gesture. Fantastic. And, Brian, tell us a little bit about voxel fifty one. Yeah. So so as you mentioned, I'm Brian, the cofounder and CEO here at Voxel fifty one. First, my background. So I'm a geek, a nerd by background, and a PhD in machine learning from University of Michigan. CodeBlue? Exactly. That is where over ten years ago, I met my cofounder, Jason, who is a faculty at Michigan. We started off doing some consulting work, of course, being located in Ann Arbor, just down the road from Detroit, the Motor City. We had the opportunity and great pleasure to collaborate with a number of automakers, you know, ten plus years ago in early versions of autonomy, where we kind of reached that key insight that in theory, it's all about models and algorithms. In practice, it's all about data and data quality and data strategy, which led us to the opportunity and need, to provide our product 51 to help solve some of those data challenges. Very cool. And we'll get into as I said, as we talk, we'll get into a little bit more about what fifty one is and what voxel fifty one does with customers like Porsche. But maybe, Tim, let's start with you. On …

Get the full transcript (7,278 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all NVIDIA AI Podcast transcripts →

You just read a 3-minute summary of a 38-minute episode.

Get NVIDIA AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by Nvidia

    generative models like NVIDIA Cosmos improve simulation fidelity to near-video realism for validation.

More from NVIDIA AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into NVIDIA AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from NVIDIA AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime