How World Foundation Models Will Advance Physical AI With NVIDIA’s Ming-Yu Liu - Ep. 240
Episode
20 min
Read time
2 min
Topics
Productivity, Relationships, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Three verification applications: World models test policy checkpoints in simulation before physical deployment, initialize policy models with pretrained physics understanding to reduce training data needs, and enable real-time future rollout simulation for decision-making before robots act.
- ✓Cosmos platform components: NVIDIA releases open-weight models free for commercial use, including diffusion-based and autoregressive architectures, video tokenizers for transformer processing, post-training scripts for custom camera configurations, and video curation toolkits leveraging GPU-accelerated libraries for data processing.
- ✓Autoregressive versus diffusion tradeoffs: Autoregressive models predict tokens sequentially like GPT, run faster due to existing optimizations, but struggle with video compression accuracy. Diffusion models generate token sets simultaneously, produce more coherent outputs with better quality, but require different integration approaches.
- ✓Early industry adoption focus: Self-driving car companies and humanoid robot manufacturers including One X, Huobi, Li Auto, and others partner with NVIDIA to apply world models for testing scenarios difficult to replicate physically and validating agent behavior across diverse environments.
What It Covers
NVIDIA's Mingyu Liu explains world foundation models, deep learning systems that simulate physics and predict futures to train and verify physical AI agents like humanoid robots and autonomous vehicles through the new Cosmos development platform.
Key Questions Answered
- •Three verification applications: World models test policy checkpoints in simulation before physical deployment, initialize policy models with pretrained physics understanding to reduce training data needs, and enable real-time future rollout simulation for decision-making before robots act.
- •Cosmos platform components: NVIDIA releases open-weight models free for commercial use, including diffusion-based and autoregressive architectures, video tokenizers for transformer processing, post-training scripts for custom camera configurations, and video curation toolkits leveraging GPU-accelerated libraries for data processing.
- •Autoregressive versus diffusion tradeoffs: Autoregressive models predict tokens sequentially like GPT, run faster due to existing optimizations, but struggle with video compression accuracy. Diffusion models generate token sets simultaneously, produce more coherent outputs with better quality, but require different integration approaches.
- •Early industry adoption focus: Self-driving car companies and humanoid robot manufacturers including One X, Huobi, Li Auto, and others partner with NVIDIA to apply world models for testing scenarios difficult to replicate physically and validating agent behavior across diverse environments.
Notable Moment
Liu compares world models to having a strategic advisor simulate different futures before making decisions, allowing physical AI systems to avoid costly real-world mistakes by testing policies in thousands of virtual kitchens before deploying in actual environments.
Episode Transcript
Hello, and welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. NVIDIA CEO Jensen Huang recently keynoted the CES Consumer Electronics Show conference in Las Vegas, Nevada. Amongst the many exciting announcements Jensen talked about was NVIDIA Cosmos. Cosmos is a development platform for world foundation models, which I think we're all gonna be talking a lot about in the coming months and years. What is a world foundation model? Well, thankfully, we've got an expert here to tell us all about it. Mingyu Liu is vice president of research at NVIDIA. He's also an I triple e fellow, and he's here to tell us all about world foundation models, how they work, what they mean, and why we should care about them going forward. So without further ado, Mingyu, thank you so much for joining the NVIDIA AI podcast, and welcome. It's great to be here. So let's start with the basics, if you would. What is a world foundation model? Sure. So world foundation models are deep learning based space time visual simulator that can help us look into the future. It can simulate physics. It can simulate people's intentions and activities. It's like data science of AI. Imagine many different environments and can simulate the future. So we can make good decision based on the simulation. We can leverage work foundation models, imagination, and simulation capability to help train physical AI agents. We can also leverage this capability to help the agent make good decision during the inference time. We can generate a virtual world based on text prompts, image prompts, video prompts, action prompts, and the layer combinations. So we call it a a world foundation model because it can generate many different worlds and also because it can be customized through different physical AI setups Right. To become a customized world model. Right? So different physical AI have different number of cameras in different locations. So we want the world foundation model to be customizable for different physical AI setups so they can use in their settings. So I wanna ask you kind of how a world model is similar or different to an LLM and other types of models. But I think first, I wanna back up a step and ask you, how is a world model similar or different to a model that generates video? Because my understanding and please correct me where I'm wrong. My understanding is that you can prompt a world model to generate a video, but that video is generated based on the things you were talking about based on, understanding of, you know, physics and other things in the physical world, and it's a different process. So I don't know what the best way is to kind of unpack it for the listeners, but one place to start might be, how does a world model differentiate from an LLM or a generative AI video model? So world model is different to LN, in the sense that …
Get the full transcript (3,263 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 17-minute episode.
Get NVIDIA AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from NVIDIA AI Podcast
Inside Instacart's AI-Powered Smart Shopping Cart | NVIDIA AI Podcast Ep. 302
Jun 24 · 39 min
Software Engineering Daily
Foundation Models for Structured Data
Jun 23
More from NVIDIA AI Podcast
How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301
Jun 10 · 21 min
Eye on AI
#331 Sergey Levine: The Robot Revolution Nobody Is Talking About
Apr 12
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
- CosmosBy guest
by NVIDIA
“NVIDIA releases open-weight models free for commercial use, including diffusion-based and autoregressive architectures, video tokenizers for transformer processing, post-training scripts for custom camera configurations, and video curation toolkits leveraging GPU-accelerated libraries for data processing.”
More from NVIDIA AI Podcast
We summarize every new episode. Want them in your inbox?
Inside Instacart's AI-Powered Smart Shopping Cart | NVIDIA AI Podcast Ep. 302
How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301
Everyone Can Build a Robot: Open Source Embodied AI With Seeed Studio | NVIDIA AI Podcast Ep. 300
Inside AI Tokenomics: How to Profitably Turn Tokens Into Business Value | NVIDIA AI Podcast Ep. 299
Snap’s Secret to Processing 10 Petabytes a Day: GPU-Accelerated Spark | NVIDIA AI Podcast Ep. 298
Similar Episodes
Related episodes from other podcasts
Software Engineering Daily
Jun 23
Foundation Models for Structured Data
Eye on AI
Apr 12
#331 Sergey Levine: The Robot Revolution Nobody Is Talking About
Cognitive Revolution
Jun 3
Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures
Odd Lots
Apr 2
This Is How to Tell if Writing Was Made by AI
Latent Space
Jan 28
🔬 Automating Science: World Models, Scientific Taste, Agent Loops — Andrew White
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into NVIDIA AI Podcast.
Every Monday, we deliver AI summaries of the latest episodes from NVIDIA AI Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime