Skip to main content
a16z Podcast

World Models, Robotics, and the Future of 3D AI

23 min episode · 2 min read
·
Justin Johnson

Episode

23 min

Read time

2 min

Topics

Startups, Fundraising & VC, Design & UX

AI-Generated Summary

Key Takeaways

  • World Models vs. Language Models: World models represent a separate model category from LLMs, grounded in visual and physical understanding rather than discrete text tokens. Just as LLMs became horizontal platforms powering countless applications, world models are designed to serve industries spanning entertainment, construction, robotics, and VR through a single generalized architecture.
  • Real-to-Sim Robotics Pipeline: Atlas enables a five-minute onboarding workflow for robots: photograph a specific environment with a phone, upload images to Atlas for reconstruction, describe the target task in natural language, generate a simulation, then RL fine-tune a pretrained robotics foundation model. This compresses environment-specific robot adaptation from weeks to minutes.
  • Spatial Grounding Solves Video Drift: Standard video generation models lose coherence beyond short durations. Atlas addresses this by anchoring reference images in three-dimensional space as fixed coordinates, then allowing precise camera path control frame-by-frame. This directorial approach enables coherent generations exceeding one minute, using reference images as spatial breadcrumbs along the trajectory.
  • 3D as Input Control Surface, Not Just Output: Three-dimensional representation serves two distinct roles in Atlas. As output, it produces Gaussian splat scenes compatible with existing game engines and VFX pipelines. As input, it provides camera control as a native modality, letting creators physically steer camera movement with precision that text prompts like "pan left" cannot replicate.
  • Gaussian Splats Remain Relevant for Client-Side Rendering: Despite Atlas supporting direct two-dimensional frame generation, explicit 3D representations like Gaussian splats retain value for mobile and VR devices where server-side streaming is cost-prohibitive. Meeting existing VFX, architecture, and gaming workflows where they already operate accelerates adoption faster than requiring fully AI-native pipelines.

What It Covers

WorldLabs cofounder Justin Johnson joins the a16z podcast to explain Atlas, a multimodal world model capable of generating, reconstructing, and simulating three-dimensional environments. The conversation covers applications across robotics, VFX, gaming, and VR, and positions world models as a horizontal AI platform distinct from language models.

Key Questions Answered

  • World Models vs. Language Models: World models represent a separate model category from LLMs, grounded in visual and physical understanding rather than discrete text tokens. Just as LLMs became horizontal platforms powering countless applications, world models are designed to serve industries spanning entertainment, construction, robotics, and VR through a single generalized architecture.
  • Real-to-Sim Robotics Pipeline: Atlas enables a five-minute onboarding workflow for robots: photograph a specific environment with a phone, upload images to Atlas for reconstruction, describe the target task in natural language, generate a simulation, then RL fine-tune a pretrained robotics foundation model. This compresses environment-specific robot adaptation from weeks to minutes.
  • Spatial Grounding Solves Video Drift: Standard video generation models lose coherence beyond short durations. Atlas addresses this by anchoring reference images in three-dimensional space as fixed coordinates, then allowing precise camera path control frame-by-frame. This directorial approach enables coherent generations exceeding one minute, using reference images as spatial breadcrumbs along the trajectory.
  • 3D as Input Control Surface, Not Just Output: Three-dimensional representation serves two distinct roles in Atlas. As output, it produces Gaussian splat scenes compatible with existing game engines and VFX pipelines. As input, it provides camera control as a native modality, letting creators physically steer camera movement with precision that text prompts like "pan left" cannot replicate.
  • Gaussian Splats Remain Relevant for Client-Side Rendering: Despite Atlas supporting direct two-dimensional frame generation, explicit 3D representations like Gaussian splats retain value for mobile and VR devices where server-side streaming is cost-prohibitive. Meeting existing VFX, architecture, and gaming workflows where they already operate accelerates adoption faster than requiring fully AI-native pipelines.

Notable Moment

Johnson described a bullet-time shot demonstration where three synchronized iPhones captured a strawberry dropping into oat milk. Atlas reconstructed the splash moment from only three input views and generated a freeze-frame fly-around sequence — a result the team described as exceeding their own expectations for sparse-view reconstruction.

Know someone who'd find this useful?

Episode Transcript

Language models are these general horizontal engines for processing streams of discrete text, discrete tokens. Those have had tons of applications from, you know, everything that we know and love today. And our thesis is that there exists another category of model called world models that should be based in visual understanding, should be based in physical understanding that can be used to generate, simulate, reconstruct worlds. Mhmm. And if we can build these with the right generality, they should be applicable to tons of different industries. Right? From entertainment to VR to construction to robotics. We live our lives in this physical built space all around us, We and need models to help us with these things as well. Language models gave AI a way to work with words. What happens when models can understand and simulate the physical world? WorldLabs cofounder Justin Johnson joins Theo Jaffe and Sofia Puccini on MTS to discuss Atlas and the broader idea behind world models. They explore how models could generate and reconstruct environments, what separates them from video generation, and applications across gaming, VFX, and robotics, including turning a few photos of a real space into a simulation for training a robot. We're live with Justin Johnson, who is a cofounder of World Labs, which is a startup that builds spatial intelligence products. And just today, they announced Atlas, the world's first multimodal world model that generates image and video frames with pixel perfect camera control and reconstructs them in three d, model the world, move the camera, and simulate space and time. So congrats on the launch. Like, the launch video looks really cool. Mhmm. Welcome to MTS. Yeah. Thanks for having me. Absolutely. So tell us a little bit more about, like, what this product is specifically, what you can do with it, what people will use it for. Yeah. So this is not a product launch. This is a this is a model announcement. So the products are coming later. This is the base model. It's gonna be used to power our future products from WorldLabs. But at its heart, it's a world model. It does basically three different kinds of things, generation, reconstruction, and simulation. Right? Generation is you want to generate new worlds that don't exist. Right? Like, I can start from a text prompt. I can start from an image prompt, then I can direct the camera and sort of fly through that world and whatever. And and have have the model generate a world that never existed before. So that obviously has tons of applications across many things in creativity, right, from VFX to gaming to all the kinds of use cases that people are using there. The second major capability of Atlas is reconstruction. Right? There, sometimes I don't want to generate a new world. Sometimes I have an existing space in in the real world that I like to model and bring into virtual reality or for the virtual world in some …

Get the full transcript (5,319 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 20-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime