Skip to main content
Practical AI

Open models and the future of Physical AI with NVIDIA

47 min episode · 2 min read
·
Ming-Yu Liu

Episode

47 min

Read time

2 min

Topics

Productivity, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • ✓Open Models vs. APIs for Physical AI: APIs aggregate requests from multiple users to maximize GPU efficiency, but physical AI devices run batch-size-one inference with no users to aggregate — making on-device open models architecturally necessary. Developers building robots or autonomous vehicles should evaluate non-transformer architectures optimized for real-time, single-instance inference rather than defaulting to LLM-style deployments.
  • ✓World Model Architecture in Cosmos 3: NVIDIA's Cosmos 3 fuses two distinct capabilities — world understanding (video + text in, text out) and world simulation (observation + action in, video out) — into a single omni-model supporting text, audio, video, and action as both inputs and outputs. Developers can select input-output combinations that match their specific physical AI use case.
  • ✓Neural Simulation Accelerates Policy Iteration: Instead of deploying robot or autonomous vehicle policy checkpoints to physical hardware for testing — which is slow and unscalable — developers can use a world model as a neural simulator. The policy interacts with the simulated environment to validate checkpoints rapidly, reserving real-world testing only for the most promising candidates.
  • ✓Sensor Diversity Requires Fine-Tuning Open Models: Physical AI devices vary significantly in sensor configuration — some robots use two cameras, others mount cameras on grippers, and autonomous vehicles use seven to eleven cameras plus LiDAR. Open model weights provide a transferable prior that developers can fine-tune on their specific sensor data, enabling multi-view support and device-specific optimization not achievable through closed APIs.
  • ✓NVIDIA's Open Model Resources Include Training Frameworks and Data: NVIDIA's open model releases for Cosmos, NeMo Tron, and other models include open weights, training frameworks for fine-tuning on custom data, open-source datasets, developer blogs, and GitHub recipe cookbooks. Developers can access these directly on Hugging Face and GitHub to begin building physical AI applications without starting from scratch.

What It Covers

NVIDIA's VP of Cosmos Lab, Ming-Yu Liu, explains why open models are foundational to physical AI development — covering world models, neural simulation, multi-agent robotics systems, and how developers can access NVIDIA's open-source frameworks, training tools, and datasets via Hugging Face and GitHub.

Key Questions Answered

  • •Open Models vs. APIs for Physical AI: APIs aggregate requests from multiple users to maximize GPU efficiency, but physical AI devices run batch-size-one inference with no users to aggregate — making on-device open models architecturally necessary. Developers building robots or autonomous vehicles should evaluate non-transformer architectures optimized for real-time, single-instance inference rather than defaulting to LLM-style deployments.
  • •World Model Architecture in Cosmos 3: NVIDIA's Cosmos 3 fuses two distinct capabilities — world understanding (video + text in, text out) and world simulation (observation + action in, video out) — into a single omni-model supporting text, audio, video, and action as both inputs and outputs. Developers can select input-output combinations that match their specific physical AI use case.
  • •Neural Simulation Accelerates Policy Iteration: Instead of deploying robot or autonomous vehicle policy checkpoints to physical hardware for testing — which is slow and unscalable — developers can use a world model as a neural simulator. The policy interacts with the simulated environment to validate checkpoints rapidly, reserving real-world testing only for the most promising candidates.
  • •Sensor Diversity Requires Fine-Tuning Open Models: Physical AI devices vary significantly in sensor configuration — some robots use two cameras, others mount cameras on grippers, and autonomous vehicles use seven to eleven cameras plus LiDAR. Open model weights provide a transferable prior that developers can fine-tune on their specific sensor data, enabling multi-view support and device-specific optimization not achievable through closed APIs.
  • •NVIDIA's Open Model Resources Include Training Frameworks and Data: NVIDIA's open model releases for Cosmos, NeMo Tron, and other models include open weights, training frameworks for fine-tuning on custom data, open-source datasets, developer blogs, and GitHub recipe cookbooks. Developers can access these directly on Hugging Face and GitHub to begin building physical AI applications without starting from scratch.

Notable Moment

Ming-Yu Liu reframes the long-term vision of physical AI as a multi-agent economy — where high-level LLM planners coordinate factory floor agents, mobile platforms, robotic arms, and self-driving trucks — but explicitly notes this full picture is likely decades away from commercial reality.

Know someone who'd find this useful?

Episode Transcript

Narrator: Welcome to the Practical AI Podcast, where we break down the real world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind the scenes content, and AI insights. You can learn more at practicalai.fm. Now onto the show. Daniel: Welcome to another episode of Practical AI Podcast. This is Daniel Whitenack. I am CEO at Prediction Guard, and I'm joined as always by my cohost, Benson, who is a principal AI and autonomy research engineer. How are you doing, Chris? Chris: Hey. Hey. I'm doing great today, Daniel. Really excited about today's conversation. Daniel: Yeah. Well, I I think in in, in the world that both of us inhabit, we are open to interesting discussions. And, I I think we'll talk about things open and things related to the world in in today's episode because we have with us MingYu Liu, who is vice president of Cosmos Lab at NVIDIA. Welcome to the show, MingYu. Ming-Yu Liu: Hello. Thanks to have me here. Great to be here. Hi, Daniel and Chris. Daniel: Yeah. Yeah. Great great to have you. And one of the things I kind of alluded to is maybe something open, which I know is dear to NVIDIA's heart, which is open models. And I've seen in the news recently in what NVIDIA has posted, obviously, they are promoting open models as a key piece of what is important in the e AI ecosystem. I'm wondering, MingYu, if you could help us understand maybe from your perspective, from NVIDIA's perspective, why in today's, world, is open is the idea of having open models and promoting open models, why is that such a critical thing in the ecosystem? Ming-Yu Liu: Yes. Okay. So when I was a student, I learned a lot about computer vision and machine learning. At that time, there are great researchers, students, open source layer code and model. And looking at layer code and model actually help me understand the concept and help me to innovate. Right? And and this the academy always work like that and has been years, you know, even before GBT, and GBT also have open source model in the past. And that had been, you know, pushed the field forward around innovation. It's just until recently, the model is getting bigger. The amount of resources required to build a model is tremendous amount, and hence, the amount of open model reduced. Right? But open model generally is there for the field, it's been very useful to push the field forward. And so, NVIDIA, won't this continue to happen? I think the world will be in a better place …

Get the full transcript (6,997 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Practical AI transcripts →

You just read a 3-minute summary of a 44-minute episode.

Get Practical AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by NVIDIA

    “NVIDIA's open model releases for Cosmos, NeMo Tron, and other models include open weights, training frameworks for fine-tuning on custom data, open-source datasets, developer blogs, and GitHub recipe cookbooks.”
  • “Developers can access these directly on Hugging Face and GitHub to begin building physical AI applications without starting from scratch.”
  • by NVIDIA

    “World Model Architecture in Cosmos 3: NVIDIA's Cosmos 3 fuses two distinct capabilities — world understanding (video + text in, text out) and world simulation (observation + action in, video out) — into a single omni-model supporting text, audio, video, and action as both inputs and outputs.”
  • “Developers can access these directly on Hugging Face and GitHub to begin building physical AI applications without starting from scratch.”

More from Practical AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Practical AI.

Every Monday, we deliver AI summaries of the latest episodes from Practical AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime