Gemini Omni: Clone yourself with AI in under 15 minutes
Episode
20 min
Read time
2 min
Topics
Investing, Design & UX, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Avatar capture speed: Google Flow's mobile QR code scanning process captures a usable facial avatar in under two minutes, requiring only frontal and side-profile head turns. The system automatically pulls background details from the scan environment — posters, books, wall color — and incorporates them into generated scenes without additional prompting.
- ✓AI as creative director: Rather than jumping straight to video generation, prompting Flow to build a storyboard first produces a structured seven-scene shot list with specific camera directions, lighting notes, and character blocking. This intermediate step prevents generic output and gives non-video-literate creators a professional production framework before a single frame renders.
- ✓Dual-version rendering: Flow automatically generates two versions of every video clip simultaneously, mirroring Veo 2 behavior. Reviewing both versions per scene and selecting the stronger take before editing meaningfully improves final output quality without additional generation cost or time investment.
- ✓Character consistency limitations: At current capability, the avatar matches the source face roughly 50% of the time across scenes. Hair length, background color, shelf contents, and lighting shift between clips. Mitigation strategy: use consistent background descriptors in every scene prompt and supply multiple reference images to the Omni model to tighten character coherence.
- ✓Browser-native timeline editing: Flow includes a built-in video editor accessible directly in the browser, eliminating the need for external software. Stitching seven AI-generated scenes into a finished one-minute video takes approximately five minutes by dragging clips into the storyboard-specified sequence and selecting preferred takes per scene.
What It Covers
Host Claire documents a live experiment using Google Flow and the Gemini Omni video model to build a one-minute AI avatar hype video for her podcast. Starting with zero tool knowledge, she completes the full workflow — avatar creation, storyboard generation, video rendering, and timeline editing — in under fifteen minutes.
Key Questions Answered
- •Avatar capture speed: Google Flow's mobile QR code scanning process captures a usable facial avatar in under two minutes, requiring only frontal and side-profile head turns. The system automatically pulls background details from the scan environment — posters, books, wall color — and incorporates them into generated scenes without additional prompting.
- •AI as creative director: Rather than jumping straight to video generation, prompting Flow to build a storyboard first produces a structured seven-scene shot list with specific camera directions, lighting notes, and character blocking. This intermediate step prevents generic output and gives non-video-literate creators a professional production framework before a single frame renders.
- •Dual-version rendering: Flow automatically generates two versions of every video clip simultaneously, mirroring Veo 2 behavior. Reviewing both versions per scene and selecting the stronger take before editing meaningfully improves final output quality without additional generation cost or time investment.
- •Character consistency limitations: At current capability, the avatar matches the source face roughly 50% of the time across scenes. Hair length, background color, shelf contents, and lighting shift between clips. Mitigation strategy: use consistent background descriptors in every scene prompt and supply multiple reference images to the Omni model to tighten character coherence.
- •Browser-native timeline editing: Flow includes a built-in video editor accessible directly in the browser, eliminating the need for external software. Stitching seven AI-generated scenes into a finished one-minute video takes approximately five minutes by dragging clips into the storyboard-specified sequence and selecting preferred takes per scene.
Notable Moment
When the avatar video rendered, it accurately reproduced a specific NVIDIA product visible only in the background of Claire's avatar scan photos — a detail she had not mentioned in any prompt. The model extracted and placed environmental context from the original capture without instruction.
Episode Transcript
Today, I am doing a very strange episode where I'm gonna create a video avatar of myself and, in about fifteen minutes, get to a full minute long video starring none other than your favorite podcast host, Clairo. Let's get to it. This episode is brought to you by Merge. Building an AI product is one thing. The hard part is everything around it, connecting to the tools your team and customers rely on, letting agents take action with the right permissions, and all the infrastructure underneath. Merge is the infrastructure layer for production AI. It connects to thousands of tools, gives agents secure ways to act inside them, and optimizes model routing and spend without you building or owning any of it. OpenAI, Dropbox, and Ramp already use Merge to move fast and build AI right. Visit merge.dev/howiai to start building for free. This episode of How I AI is going to be an adventure because I'm gonna be honest. I'm not a 100% sure this is going to work. I'm gonna return to a product I covered very briefly a couple weeks ago called Google Flow and the new Gemini Omni video generation model. And I'm gonna try really hard to create an AI avatar of myself that we can animate or, I guess, cinematically create using AI. So this is Google Flow, and one of the features of Google Flow and the Omni model is you are supposed to be able to create an avatar of yourself. Now we tried this the day it came out. It did not work, but we're gonna give it another call it try and see if we can get a full featured avatar of myself that then we can go and build consistent character videos off of. So I'm gonna select up here. I'm gonna create an avatar. We're gonna click get started. I'm gonna scan this QR code. I have my phone here. I've done this before, so hopefully it'll be fast. Okay. I'm gonna put the mic away just for one second. K. I'm gonna allow access to my camera, and we're just gonna take some photos. Okay. Ready? Start. 178149202522. Okay. Now it's having me turn my head. So I turned my head that way, gave me a check mark. Turn my head the other way, and it's giving me a check mark. And and it says we're done. Now it said we were done last time we tried this. So we're gonna see it's gonna take a couple minutes, and then we will come back and see if I can actually use this avatar of myself. Okay. So look at this beauty. There's this fish eye lens version of me that is now an avatar. So I supposedly can use this, and let's use it to create a hype video for the How I A I podcast. So I'm gonna go in here and say, help me create a storyboard for a hype video. For the How …
Get the full transcript (3,291 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 17-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)
Aug 31 · 46 min
The Vergecast
We react to Google I/O 2026: The Vergecast Livestream
May 19
More from How I AI
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
Aug 24 · 44 min
The AI Breakdown
The Most Useful New AI Features and Tools to Try
Aug 28
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
by Atlassian
“SPONSORS [{"name": "Jira Product Discovery", "url": "https://atlassian.com/howiai"}”
- Gemini OmniRecommended
by Google
“Host Claire documents a live experiment using Google Flow and the Gemini Omni video model to build a one-minute AI avatar hype video for her podcast.”
- Google FlowRecommended
by Google
“Host Claire documents a live experiment using Google Flow and the Gemini Omni video model to build a one-minute AI avatar hype video for her podcast.”
“Flow automatically generates two versions of every video clip simultaneously, mirroring Veo 2 behavior.”
Gear
by NVIDIA
“When the avatar video rendered, it accurately reproduced a specific NVIDIA product visible only in the background of Claire's avatar scan photos.”
More from How I AI
We summarize every new episode. Want them in your inbox?
How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Similar Episodes
Related episodes from other podcasts
The Vergecast
May 19
We react to Google I/O 2026: The Vergecast Livestream
The AI Breakdown
Aug 28
The Most Useful New AI Features and Tools to Try
The Vergecast
Aug 12
Pixel 11 and Pixel Watch 5: Our first impressions | The Vergecast Livestream
All-In with Chamath, Jason, Sacks & Friedberg
Jul 24
The Fight Over Open Source AI, Anthropic's $1.5B Payout, NYC Socialists: Evictions = Violence?
The Vergecast
Jun 5
This is your laptop... on AI
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime