Skip to main content
ML

Ming-Yu Liu

Nvidia's VP of Cosmos Lab**open Models Vs**world Model Architecture in Cosmos 3**neural Simulation Accelerates Policy Iteration**sensor Diversity Requires Fine-tuning Open Models
2episodes
2podcasts

We have 2 summarized appearances for Ming-Yu Liu so far. Browse all podcasts to discover more episodes.

Featured On 2 Podcasts

Top resources Ming-Yu Liu mentions

Books, tools, and gear cited across podcast appearances. Ranked by frequency.

SignalCast may earn commission on purchases via affiliate links on each resource page.

All Appearances

2 episodes
Practical AI

Open models and the future of Physical AI with NVIDIA

Practical AI
47 minVice President of Cosmos Lab at NVIDIA

AI Summary

→ WHAT IT COVERS NVIDIA's VP of Cosmos Lab, Ming-Yu Liu, explains why open models are foundational to physical AI development — covering world models, neural simulation, multi-agent robotics systems, and how developers can access NVIDIA's open-source frameworks, training tools, and datasets via Hugging Face and GitHub. → KEY INSIGHTS - **Open Models vs. APIs for Physical AI:** APIs aggregate requests from multiple users to maximize GPU efficiency, but physical AI devices run batch-size-one inference with no users to aggregate — making on-device open models architecturally necessary. Developers building robots or autonomous vehicles should evaluate non-transformer architectures optimized for real-time, single-instance inference rather than defaulting to LLM-style deployments. - **World Model Architecture in Cosmos 3:** NVIDIA's Cosmos 3 fuses two distinct capabilities — world understanding (video + text in, text out) and world simulation (observation + action in, video out) — into a single omni-model supporting text, audio, video, and action as both inputs and outputs. Developers can select input-output combinations that match their specific physical AI use case. - **Neural Simulation Accelerates Policy Iteration:** Instead of deploying robot or autonomous vehicle policy checkpoints to physical hardware for testing — which is slow and unscalable — developers can use a world model as a neural simulator. The policy interacts with the simulated environment to validate checkpoints rapidly, reserving real-world testing only for the most promising candidates. - **Sensor Diversity Requires Fine-Tuning Open Models:** Physical AI devices vary significantly in sensor configuration — some robots use two cameras, others mount cameras on grippers, and autonomous vehicles use seven to eleven cameras plus LiDAR. Open model weights provide a transferable prior that developers can fine-tune on their specific sensor data, enabling multi-view support and device-specific optimization not achievable through closed APIs. - **NVIDIA's Open Model Resources Include Training Frameworks and Data:** NVIDIA's open model releases for Cosmos, NeMo Tron, and other models include open weights, training frameworks for fine-tuning on custom data, open-source datasets, developer blogs, and GitHub recipe cookbooks. Developers can access these directly on Hugging Face and GitHub to begin building physical AI applications without starting from scratch. → NOTABLE MOMENT Ming-Yu Liu reframes the long-term vision of physical AI as a multi-agent economy — where high-level LLM planners coordinate factory floor agents, mobile platforms, robotic arms, and self-driving trucks — but explicitly notes this full picture is likely decades away from commercial reality. 💼 SPONSORS [{"name": "Prediction Guard", "url": "https://predictionguard.com/practicalai"}, {"name": "Midwest AI Summit", "url": "https://midwestaisummit.com"}] 🏷️ Physical AI, Open Source Models, World Models, Robotics, Neural Simulation

AI Summary

→ WHAT IT COVERS NVIDIA's Mingyu Liu explains world foundation models, deep learning systems that simulate physics and predict futures to train and verify physical AI agents like humanoid robots and autonomous vehicles through the new Cosmos development platform. → KEY INSIGHTS - **Three verification applications:** World models test policy checkpoints in simulation before physical deployment, initialize policy models with pretrained physics understanding to reduce training data needs, and enable real-time future rollout simulation for decision-making before robots act. - **Cosmos platform components:** NVIDIA releases open-weight models free for commercial use, including diffusion-based and autoregressive architectures, video tokenizers for transformer processing, post-training scripts for custom camera configurations, and video curation toolkits leveraging GPU-accelerated libraries for data processing. - **Autoregressive versus diffusion tradeoffs:** Autoregressive models predict tokens sequentially like GPT, run faster due to existing optimizations, but struggle with video compression accuracy. Diffusion models generate token sets simultaneously, produce more coherent outputs with better quality, but require different integration approaches. - **Early industry adoption focus:** Self-driving car companies and humanoid robot manufacturers including One X, Huobi, Li Auto, and others partner with NVIDIA to apply world models for testing scenarios difficult to replicate physically and validating agent behavior across diverse environments. → NOTABLE MOMENT Liu compares world models to having a strategic advisor simulate different futures before making decisions, allowing physical AI systems to avoid costly real-world mistakes by testing policies in thousands of virtual kitchens before deploying in actual environments. 💼 SPONSORS None detected 🏷️ World Foundation Models, Physical AI, Autonomous Vehicles, Humanoid Robotics

Explore More

Never miss Ming-Yu Liu's insights

Subscribe to get AI-powered summaries of Ming-Yu Liu's podcast appearances delivered to your inbox weekly.

Start Free Today

No credit card required • Free tier available