World Models, Robotics, and the Future of 3D AI
Episode
23 min
Read time
2 min
Topics
Startups, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓World Models vs. Language Models: World models represent a separate model category from LLMs, grounded in visual and physical understanding rather than discrete text tokens. Just as LLMs became horizontal platforms powering countless applications, world models are designed to serve industries spanning entertainment, construction, robotics, and VR through a single generalized architecture.
- ✓Real-to-Sim Robotics Pipeline: Atlas enables a five-minute onboarding workflow for robots: photograph a specific environment with a phone, upload images to Atlas for reconstruction, describe the target task in natural language, generate a simulation, then RL fine-tune a pretrained robotics foundation model. This compresses environment-specific robot adaptation from weeks to minutes.
- ✓Spatial Grounding Solves Video Drift: Standard video generation models lose coherence beyond short durations. Atlas addresses this by anchoring reference images in three-dimensional space as fixed coordinates, then allowing precise camera path control frame-by-frame. This directorial approach enables coherent generations exceeding one minute, using reference images as spatial breadcrumbs along the trajectory.
- ✓3D as Input Control Surface, Not Just Output: Three-dimensional representation serves two distinct roles in Atlas. As output, it produces Gaussian splat scenes compatible with existing game engines and VFX pipelines. As input, it provides camera control as a native modality, letting creators physically steer camera movement with precision that text prompts like "pan left" cannot replicate.
- ✓Gaussian Splats Remain Relevant for Client-Side Rendering: Despite Atlas supporting direct two-dimensional frame generation, explicit 3D representations like Gaussian splats retain value for mobile and VR devices where server-side streaming is cost-prohibitive. Meeting existing VFX, architecture, and gaming workflows where they already operate accelerates adoption faster than requiring fully AI-native pipelines.
What It Covers
WorldLabs cofounder Justin Johnson joins the a16z podcast to explain Atlas, a multimodal world model capable of generating, reconstructing, and simulating three-dimensional environments. The conversation covers applications across robotics, VFX, gaming, and VR, and positions world models as a horizontal AI platform distinct from language models.
Key Questions Answered
- •World Models vs. Language Models: World models represent a separate model category from LLMs, grounded in visual and physical understanding rather than discrete text tokens. Just as LLMs became horizontal platforms powering countless applications, world models are designed to serve industries spanning entertainment, construction, robotics, and VR through a single generalized architecture.
- •Real-to-Sim Robotics Pipeline: Atlas enables a five-minute onboarding workflow for robots: photograph a specific environment with a phone, upload images to Atlas for reconstruction, describe the target task in natural language, generate a simulation, then RL fine-tune a pretrained robotics foundation model. This compresses environment-specific robot adaptation from weeks to minutes.
- •Spatial Grounding Solves Video Drift: Standard video generation models lose coherence beyond short durations. Atlas addresses this by anchoring reference images in three-dimensional space as fixed coordinates, then allowing precise camera path control frame-by-frame. This directorial approach enables coherent generations exceeding one minute, using reference images as spatial breadcrumbs along the trajectory.
- •3D as Input Control Surface, Not Just Output: Three-dimensional representation serves two distinct roles in Atlas. As output, it produces Gaussian splat scenes compatible with existing game engines and VFX pipelines. As input, it provides camera control as a native modality, letting creators physically steer camera movement with precision that text prompts like "pan left" cannot replicate.
- •Gaussian Splats Remain Relevant for Client-Side Rendering: Despite Atlas supporting direct two-dimensional frame generation, explicit 3D representations like Gaussian splats retain value for mobile and VR devices where server-side streaming is cost-prohibitive. Meeting existing VFX, architecture, and gaming workflows where they already operate accelerates adoption faster than requiring fully AI-native pipelines.
Notable Moment
Johnson described a bullet-time shot demonstration where three synchronized iPhones captured a strawberry dropping into oat milk. Atlas reconstructed the splash moment from only three input views and generated a freeze-frame fly-around sequence — a result the team described as exceeding their own expectations for sparse-view reconstruction.
Episode Transcript
Language models are these general horizontal engines for processing streams of discrete text, discrete tokens. Those have had tons of applications from, you know, everything that we know and love today. And our thesis is that there exists another category of model called world models that should be based in visual understanding, should be based in physical understanding that can be used to generate, simulate, reconstruct worlds. Mhmm. And if we can build these with the right generality, they should be applicable to tons of different industries. Right? From entertainment to VR to construction to robotics. We live our lives in this physical built space all around us, We and need models to help us with these things as well. Language models gave AI a way to work with words. What happens when models can understand and simulate the physical world? WorldLabs cofounder Justin Johnson joins Theo Jaffe and Sofia Puccini on MTS to discuss Atlas and the broader idea behind world models. They explore how models could generate and reconstruct environments, what separates them from video generation, and applications across gaming, VFX, and robotics, including turning a few photos of a real space into a simulation for training a robot. We're live with Justin Johnson, who is a cofounder of World Labs, which is a startup that builds spatial intelligence products. And just today, they announced Atlas, the world's first multimodal world model that generates image and video frames with pixel perfect camera control and reconstructs them in three d, model the world, move the camera, and simulate space and time. So congrats on the launch. Like, the launch video looks really cool. Mhmm. Welcome to MTS. Yeah. Thanks for having me. Absolutely. So tell us a little bit more about, like, what this product is specifically, what you can do with it, what people will use it for. Yeah. So this is not a product launch. This is a this is a model announcement. So the products are coming later. This is the base model. It's gonna be used to power our future products from WorldLabs. But at its heart, it's a world model. It does basically three different kinds of things, generation, reconstruction, and simulation. Right? Generation is you want to generate new worlds that don't exist. Right? Like, I can start from a text prompt. I can start from an image prompt, then I can direct the camera and sort of fly through that world and whatever. And and have have the model generate a world that never existed before. So that obviously has tons of applications across many things in creativity, right, from VFX to gaming to all the kinds of use cases that people are using there. The second major capability of Atlas is reconstruction. Right? There, sometimes I don't want to generate a new world. Sometimes I have an existing space in in the real world that I like to model and bring into virtual reality or for the virtual world in some …
Get the full transcript (5,319 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 20-minute episode.
Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from a16z Podcast
Why Companies Are Becoming a Series of Loops | Anish Acharya on Lenny’s Podcast
Sep 12 · 78 min
How I Built This
Advice Line with Kip Tindell of The Container Store
Sep 10
More from a16z Podcast
What It Takes to Build a Startup | Andrew Chen & Matt Perault
Sep 11 · 38 min
The Ezra Klein Show
What ‘Hyperpolitics’ Explains About This Era
Sep 8
More from a16z Podcast
We summarize every new episode. Want them in your inbox?
Why Companies Are Becoming a Series of Loops | Anish Acharya on Lenny’s Podcast
What It Takes to Build a Startup | Andrew Chen & Matt Perault
How AI Is Rewriting the Power Law of Venture Capital
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
OpenAI Researchers on the Future of Mathematical Reasoning
Similar Episodes
Related episodes from other podcasts
How I Built This
Sep 10
Advice Line with Kip Tindell of The Container Store
The Ezra Klein Show
Sep 8
What ‘Hyperpolitics’ Explains About This Era
How I Built This
Sep 3
Advice Line with Ben Goodwin of Olipop
The Rich Roll Podcast
Aug 31
Hunter Biden Is Out Of Secrets: A Story of Recovery, Exposure & Amends
10% Happier with Dan Harris
Aug 31
The Science of Habit Change, Why GLP-1s are Disrupting It, and How To Use It To Have Better Conversations | Charles Duhigg
Explore Related Topics
This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into a16z Podcast.
Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime