Skip to main content
a16z Podcast

The Next Frontier of AI Video Is Control

39 min episode · 2 min read
·
Gorka Mirdsevan,Batuhan Taskaya,Jennifer Lee

Episode

39 min

Read time

2 min

Topics

Productivity, Startups, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Post-training efficiency stack: Combining RL-based step reduction (50 steps down to 20), kernel optimization, and hardware utilization improvements (40% to 80% MFU) compounds into order-of-magnitude gains. H3 Max Turbo generates five-second video in 1.5 seconds at half the cost, making further speed optimization less valuable than pursuing quality and controllability improvements.
  • Real-time video memory architecture: H3 Max Director maintains scene coherence by attending to compressed representations of the prior two minutes of generated video, then layering an evolving system prompt for context beyond that window. This enables continuous streams up to 60 minutes where characters, environments, and camera positions remain consistent across scenes.
  • Blender-plus-AI workflow for 100% controllability: VFX professionals render low-resolution scene layouts in Blender, then pass that video as a reference input to H3 Max. This combination gives creators near-complete control over composition and motion, and GPT-integrated Blender scene generation further automates the pipeline for studio-level production work.
  • Structured camera control via JSON input: Fal's post-training infrastructure conditions H3 Max to accept precise camera trajectory descriptions as structured JSON, specifying position and angle at each timestamp. The model treats this as the sole source of truth, eliminating camera hallucination and enabling 3D scene reconstruction from a single input reference image.
  • Hollywood is Fal's fastest-growing segment: Studio adoption was near zero one year ago and now dominates Fal's generative media conference attendance. Studios need targeted point solutions — video extension, camera adjustment, lighting changes — not full generation from scratch. Fal addresses legal barriers by offering US-hosted models including Stable Diffusion and supporting custom IP unlocking for studio-owned content.

What It Covers

A16z's Jennifer Lee speaks with Fal co-founders Gorka Mirdsevan and Batuan Tashkaya about H3 Max, a post-trained open-source video model achieving 35x speed improvements over base Minimax H3, enabling real-time generation, two-minute scene memory, and professional camera and lighting controls for Hollywood workflows.

Key Questions Answered

  • Post-training efficiency stack: Combining RL-based step reduction (50 steps down to 20), kernel optimization, and hardware utilization improvements (40% to 80% MFU) compounds into order-of-magnitude gains. H3 Max Turbo generates five-second video in 1.5 seconds at half the cost, making further speed optimization less valuable than pursuing quality and controllability improvements.
  • Real-time video memory architecture: H3 Max Director maintains scene coherence by attending to compressed representations of the prior two minutes of generated video, then layering an evolving system prompt for context beyond that window. This enables continuous streams up to 60 minutes where characters, environments, and camera positions remain consistent across scenes.
  • Blender-plus-AI workflow for 100% controllability: VFX professionals render low-resolution scene layouts in Blender, then pass that video as a reference input to H3 Max. This combination gives creators near-complete control over composition and motion, and GPT-integrated Blender scene generation further automates the pipeline for studio-level production work.
  • Structured camera control via JSON input: Fal's post-training infrastructure conditions H3 Max to accept precise camera trajectory descriptions as structured JSON, specifying position and angle at each timestamp. The model treats this as the sole source of truth, eliminating camera hallucination and enabling 3D scene reconstruction from a single input reference image.
  • Hollywood is Fal's fastest-growing segment: Studio adoption was near zero one year ago and now dominates Fal's generative media conference attendance. Studios need targeted point solutions — video extension, camera adjustment, lighting changes — not full generation from scratch. Fal addresses legal barriers by offering US-hosted models including Stable Diffusion and supporting custom IP unlocking for studio-owned content.

Notable Moment

After H3 Max launched on a Saturday, a Fal engineer began streaming continuous AI-generated video from his personal laptop on Twitch using prompt tricks to maintain narrative coherence — an unplanned demonstration that the model had crossed the real-time generation threshold without any internal coordination or preparation.

Know someone who'd find this useful?

Episode Transcript

Generative media is along with the coding agent market, what we call is token market fit. Everyone's waiting for a large consumer moment in AI. I believe H3Max makes it possible. Were you surprised by the speed up and the gain you could get from post training this model? We have a version called HD Max Turbo that's public that can generate, like, a five second video in, like, one point five seconds. From a cost standpoint, it's also, like, 2x less. People starting creating these beautiful scenes using an LLM model, GPT Astra, in Blender, and all of a sudden, it unlocked the whole new workflow for Hollywood and professional people. We have been very, very focused towards speed, performance, quality, and now we have a really good pace model. The next month or two is going be fully focused on. What happens when AI video becomes fast enough to generate in real time? A16z general partner Jennifer Lee sits down with Pfal co founder Gorka Mirdsevan and head of engineering Batwan Tashkaya to discuss H3 Max and the rapidly changing generative video stack. They unpack how post training and systems made video generation significantly faster, opening up new experiences where video can run continuously, remember previous scenes, and respond to direction as it plays. But speed is only part of the story. They also discuss the push toward greater control over camera angles, lighting, characters, and motion, and why those tools could make generative video more useful for professional creative workflows. Welcome, Gorkum Botouin, to our podcast again. We did the last one last year. This is long overdue, and we have such an exciting model to talk about, which is False H3 Max. The day when it came out, I was calling it, it's really in the league of its own. Like, it's so funny to see the benchmarks where you have the dot of this model on the far left or far right. And then everything else is on the other And that graph is actually log scale. So it's actually further, but we had to fit it in. We had to do log scale. That is hilarious. The time portion, the quality is not. Yeah. For sure. The internet noticed for sure. There are so many viral tweets about it. Like people really played around with this model. Maybe just give us the backstory of what inspired you to post train this open weight model from Minimax and how did you get the quality and speed to artists? First of all, the Minimax H3 model is the first truly open source, very capable, latest generation video model out there. So even though we work with some of the other model labs to run inference for them, we never had this capability, like had the right to add this capability on top of it. So when Minimax came up with their very capable open source model that is truly last generation, can take references, like very …

Get the full transcript (6,875 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 36-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • VFX professionals render low-resolution scene layouts in Blender, then pass that video as a reference input to H3 Max. This combination gives creators near-complete control over composition and motion, and GPT-integrated Blender scene generation further automates the pipeline for studio-level production work.

Products

  • H3 MaxBy guest

    by Fal

    A16z's Jennifer Lee speaks with Fal co-founders Gorka Mirdsevan and Batuan Tashkaya about H3 Max, a post-trained open-source video model achieving 35x speed improvements over base Minimax H3, enabling real-time generation, two-minute scene memory, and professional camera and lighting controls for Hollywood workflows.
  • H3 Max TurboBy guest

    by Fal

    H3 Max Turbo generates five-second video in 1.5 seconds at half the cost, making further speed optimization less valuable than pursuing quality and controllability improvements.
  • by Fal

    H3 Max Director maintains scene coherence by attending to compressed representations of the prior two minutes of generated video, then layering an evolving system prompt for context beyond that window.
  • H3 Max, a post-trained open-source video model achieving 35x speed improvements over base Minimax H3, enabling real-time generation, two-minute scene memory, and professional camera and lighting controls for Hollywood workflows.
  • Fal addresses legal barriers by offering US-hosted models including Stable Diffusion and supporting custom IP unlocking for studio-owned content.

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime