Skip to main content
a16z Podcast

Fei Fei Li: The Race to Build World Models For AI

44 min episode · 2 min read
·
Fei Fei Li

Episode

44 min

Read time

2 min

Topics

Startups, Artificial Intelligence, Software Development

AI-Generated Summary

Key Takeaways

  • New View Prediction as Base Primitive: Atlas operates on new view prediction rather than next-frame prediction, treating every input image as spatially grounded with an associated 3D camera pose. This enables precise reconstruction and generation within a single unified architecture — a capability no prior model achieved at the pre-training stage, making it architecturally distinct from existing video diffusion models.
  • 50–100x Sparse Reconstruction Reduction: Traditional dense 3D reconstruction requires 200–300 images of a space to avoid holes. Atlas reduces that requirement to as few as three inputs by combining generative fill with reconstruction, allowing users to reconstruct entire environments from casual smartphone footage or existing internet imagery that was previously considered unusable for 3D capture.
  • Bullet Time Video with Three Cameras: Replicating the Matrix-style frozen-time rotating camera shot previously required hundreds of synchronized cameras on a green screen studio set. Atlas achieves the same output using three iPhones mounted on tripods, with no calibration or controlled environment, enabling cinematic multi-perspective freeze-frame video for independent creators without production infrastructure.
  • Scaling Law Conviction Drove Architecture Decisions: The team validated Atlas by training progressively larger models and observing consistent quality improvements at each step, confirming the scaling hypothesis applies to spatial intelligence. The released model was constrained by a launch deadline rather than architectural limits, meaning current performance represents an early point on the scaling curve with compute as the primary bottleneck.
  • Real-to-Sim Pipeline for Robotics Data: World Labs acquired a robotics company (formerly Synnex) whose real-to-simulation pipeline previously relied on dense reconstruction — the same laborious multi-hundred-image process. Atlas replaces that bottleneck, enabling faster environment digitization and domain randomization (varying object sizes, positions, colors) to generate the training data diversity robotic policies require without exhaustive physical recapture.

What It Covers

World Labs cofounders Fei-Fei Li, Justin Johnson, and Ben Mildenhall present Atlas, a spatial AI model built on "new view prediction" — a novel primitive analogous to next-token prediction in LLMs — that unifies 3D reconstruction and generation, reducing capture requirements from hundreds of images to as few as three.

Key Questions Answered

  • New View Prediction as Base Primitive: Atlas operates on new view prediction rather than next-frame prediction, treating every input image as spatially grounded with an associated 3D camera pose. This enables precise reconstruction and generation within a single unified architecture — a capability no prior model achieved at the pre-training stage, making it architecturally distinct from existing video diffusion models.
  • 50–100x Sparse Reconstruction Reduction: Traditional dense 3D reconstruction requires 200–300 images of a space to avoid holes. Atlas reduces that requirement to as few as three inputs by combining generative fill with reconstruction, allowing users to reconstruct entire environments from casual smartphone footage or existing internet imagery that was previously considered unusable for 3D capture.
  • Bullet Time Video with Three Cameras: Replicating the Matrix-style frozen-time rotating camera shot previously required hundreds of synchronized cameras on a green screen studio set. Atlas achieves the same output using three iPhones mounted on tripods, with no calibration or controlled environment, enabling cinematic multi-perspective freeze-frame video for independent creators without production infrastructure.
  • Scaling Law Conviction Drove Architecture Decisions: The team validated Atlas by training progressively larger models and observing consistent quality improvements at each step, confirming the scaling hypothesis applies to spatial intelligence. The released model was constrained by a launch deadline rather than architectural limits, meaning current performance represents an early point on the scaling curve with compute as the primary bottleneck.
  • Real-to-Sim Pipeline for Robotics Data: World Labs acquired a robotics company (formerly Synnex) whose real-to-simulation pipeline previously relied on dense reconstruction — the same laborious multi-hundred-image process. Atlas replaces that bottleneck, enabling faster environment digitization and domain randomization (varying object sizes, positions, colors) to generate the training data diversity robotic policies require without exhaustive physical recapture.

Notable Moment

A turning point came during early model development when a smaller pre-Atlas checkpoint spontaneously generated a camera path flying underneath a garden table — a scene element never explicitly captured. The team recognized within seconds this emergent spatial reasoning validated their entire architectural thesis and committed fully to the approach.

Know someone who'd find this useful?

Episode Transcript

On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. You know LLMs are built on NEXT token prediction. We've seen video models as being built on frame prediction. Atlas is really new view prediction. This is the real place where AI can actually unlock a ton of value for people and their process. We're saying, like, 50, a 100 x reduction. There was a famous shot in the first Matrix movie where Neo is, like, falling down. Exactly. They had hundreds of cameras viewing that angle on a green screen. On Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. No one has ever seen this result. When you set out to do this, did you know it was gonna work? I was pretty sure. Each time we made the model bigger and each time we trained it for longer, it got significantly better. Does that mean we're gonna get four d video? Do I go walk around? Language models are built around predicting the next token. What happens when a model instead learns to predict the next view of the world? In this episode, Martin Casado sits down with World Labs cofounders Fei Fei Li, Justin Johnson, and Ben Mildenhall to discuss Atlas and the broader challenge of building AI that can reason about physical space. They explain how Atlas combines generation and three d reconstruction, why the team chose NewView prediction as its underlying primitive, and what they learned trying to scale an approach that hadn't been tested before. They also get into what's still missing, including richer dynamics and interaction, and how world models could eventually connect simulation with robotics and planning. Underlying it all is a bigger hypothesis. Could new view prediction play a similar role for spatial intelligence that next token prediction has played for language? So big day yesterday, you launched a new frontier model, which has got amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So maybe, Justin, you wanna talk about what was launched yesterday, why it's significant? Yeah. So Atlas is our new next generation world model. It has three basic things. It can generate, reconstruct, and simulate the world. So within that, there's a couple different major capabilities. It has really good camera condition generation. So you can input an image together with a camera trajectory and steer the model and have it generate video frames, mark any perspective you want. It's really good at sparse three d reconstruction. You can input one or multiple up to a 100 frames that are views of the real world and use those to reconstruct the real world, and that reconstruction can take the case either of novel …

Get the full transcript (9,319 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 41-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime