The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More
Episode
59 min
Read time
2 min
Topics
Productivity, Relationships, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Model-Harness Co-training: Google now trains Gemini models in direct partnership with its "anti-gravity" agent harness, meaning the model is optimized for tool-calling loops, orchestration, and agentic workflows from the start. Teams building on top inherit this infrastructure without rebuilding it, cutting iteration cycles from months to days and enabling faster product launches across all Google surfaces.
- ✓Flash-First Strategy: Google deliberately prioritizes cost-adjusted performance over raw capability maximization. Gemini 3.5 Flash benchmarks at roughly 280 tokens per second on Artificial Analysis, runs three times faster than comparable large models, and costs significantly less. For consumer-scale products like Search and the Gemini app, latency improvements outperform quality gains in live experiments when users are unwilling to wait.
- ✓"Model Eats the Scaffolding" Cycle: Every 12–18 months, the surrounding scaffolding that developers build around AI models gets absorbed into the model itself. Product teams should avoid having every team independently rebuild agentic infrastructure from scratch. Standardizing on a shared harness layer reduces redundant engineering, accelerates deployment, and surfaces model failure modes faster through centralized feedback loops.
- ✓Context Window Economics: Context windows have plateaued near one million tokens because serving costs become prohibitive at scale — a single one-million-token request can cost several dollars. Google's direction shifts toward smart context compaction: selectively retrieving relevant information rather than expanding raw window size, which keeps latency and cost manageable while effectively giving models access to much larger information pools.
- ✓Recursive Self-Improvement — Practical, Not Singular: Google uses Gemini internally to improve Gemini, including running ablations, submitting code changes, and generating research reports autonomously. However, humans remain in the driver's seat on large pre-training runs due to the high compute cost of misdirection. The framing is deep human-AI collaboration rather than autonomous AI-led model development, with humans focused on strategic interpretation of results.
What It Covers
Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy across its billions-of-users product surface.
Key Questions Answered
- •Model-Harness Co-training: Google now trains Gemini models in direct partnership with its "anti-gravity" agent harness, meaning the model is optimized for tool-calling loops, orchestration, and agentic workflows from the start. Teams building on top inherit this infrastructure without rebuilding it, cutting iteration cycles from months to days and enabling faster product launches across all Google surfaces.
- •Flash-First Strategy: Google deliberately prioritizes cost-adjusted performance over raw capability maximization. Gemini 3.5 Flash benchmarks at roughly 280 tokens per second on Artificial Analysis, runs three times faster than comparable large models, and costs significantly less. For consumer-scale products like Search and the Gemini app, latency improvements outperform quality gains in live experiments when users are unwilling to wait.
- •"Model Eats the Scaffolding" Cycle: Every 12–18 months, the surrounding scaffolding that developers build around AI models gets absorbed into the model itself. Product teams should avoid having every team independently rebuild agentic infrastructure from scratch. Standardizing on a shared harness layer reduces redundant engineering, accelerates deployment, and surfaces model failure modes faster through centralized feedback loops.
- •Context Window Economics: Context windows have plateaued near one million tokens because serving costs become prohibitive at scale — a single one-million-token request can cost several dollars. Google's direction shifts toward smart context compaction: selectively retrieving relevant information rather than expanding raw window size, which keeps latency and cost manageable while effectively giving models access to much larger information pools.
- •Recursive Self-Improvement — Practical, Not Singular: Google uses Gemini internally to improve Gemini, including running ablations, submitting code changes, and generating research reports autonomously. However, humans remain in the driver's seat on large pre-training runs due to the high compute cost of misdirection. The framing is deep human-AI collaboration rather than autonomous AI-led model development, with humans focused on strategic interpretation of results.
Notable Moment
A Google safety and alignment researcher ran a full suite of model ablations from her phone while sitting in a hot tub, receiving a complete report within an hour. The anecdote illustrates how AI-assisted research workflows have already compressed tasks that previously required days of engineering time into single sessions.
Episode Transcript
Hello, and welcome back to the Cognitive Revolution. Today, after some 340 episodes, I am very excited to share the first episode that I've ever recorded in person with fan favorite Logan Kilpatrick, member of technical staff at Google DeepMind, and Tulsi Doshi, senior director and head of product for Gemini models. The occasion for this conversation is Google's annual IO event, where they're launching the new Gemini 3.5 flash model, all sorts of agent infrastructure and AI product integrations, and plenty more. We've recorded on Friday, May 15, just a couple days before the event. And while many at Google, including my brother Craig, who's giving a keynote on Wednesday, were working overtime to polish their demos and presentations, the overall vibe, at least compared to the rest of the AI space, was one of relatively relaxed confidence. And why not? From twenty twenty four to twenty five, Google grew annual revenue by $50,000,000,000, as much as Anthropic is pulling in today. And they still have 25% of all global compute, the deepest pool of research talent anywhere, and the most comprehensive AI portfolio of any company with top tier positions not just in language models, but also self driving cars, medical and life sciences, and robotics. So after discussing the headline launches that they're announcing this week, which also include a new video generation model called Omni, which they hope will create a nano banana moment for video, a new and improved and more agent focused anti gravity, and a product called Spark, which will bring more agentic functionality to the main consumer Gemini app, I really wanted to take a step back and dig in on Google's overall AI strategy and philosophy. We discussed their decision to lead with the Flash model and more generally to emphasize the cost adjusted performance Pareto frontier, whereas Anthropic and OpenAI are clearly much more focused on competing to have the single most capable model in absolute terms. We talk about how DeepMind is no longer shipping models in isolation and leaving it up to product teams to figure out how to use them, but instead now providing a robust agent harness, which should help elevate and standardize AI experiences across Google's vast product surface. We get into the weeds on questions like why context windows seem to have mostly stopped growing, why Gemini model's knowledge cutoff is now more than a year ago, and whatever happened to that diffusion model line of work. Perhaps most importantly, we discuss how the team at Google relates to the AI's they're creating, how they're thinking about things like model psychology and welfare, and their views on recursive self improvement, which as you'll hear, is definitely a part of their plan, but not something that they seem to be so singularly focused on as other AI leaders. Overall, I think this is a great window into the thinking that underlies Google's AI research and product development, which has clearly sustained the company's historic run …
Get the full transcript (12,014 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 56-minute episode.
Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Cognitive Revolution
AI in the AM — Weekly Highlights: Relaunch Week (Aug 17–20, 2026)
Aug 22 · 153 min
Invest Like the Best with Patrick O'Shaughnessy
Ben Thompson on Big Tech, China, and the AI Boom Running Out of Money - [Invest Like the Best, EP.487]
Aug 18
More from Cognitive Revolution
Let There Be Germicidal Light: This $500 Fixture Could Stop the Next Pandemic, from Complex Systems
Aug 16 · 85 min
Masters of Scale
Pioneers of AI: Why AI conversations are always one-sided
Aug 15
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Google DeepMind
“Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy”
by Google DeepMind
“Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy”
More from Cognitive Revolution
We summarize every new episode. Want them in your inbox?
AI in the AM — Weekly Highlights: Relaunch Week (Aug 17–20, 2026)
Let There Be Germicidal Light: This $500 Fixture Could Stop the Next Pandemic, from Complex Systems
Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses
Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent
Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...
Similar Episodes
Related episodes from other podcasts
Invest Like the Best with Patrick O'Shaughnessy
Aug 18
Ben Thompson on Big Tech, China, and the AI Boom Running Out of Money - [Invest Like the Best, EP.487]
Masters of Scale
Aug 15
Pioneers of AI: Why AI conversations are always one-sided
The Prof G Pod
Aug 13
Why We’re Having Less Sex — and Why It Matters (ft. Dr. Debra Soh)
The Vergecast
Aug 12
Pixel 11 and Pixel Watch 5: Our first impressions | The Vergecast Livestream
The Vergecast
Aug 7
What's behind the Google AI shakeup
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into Cognitive Revolution.
Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime