Skip to main content
Cognitive Revolution

The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More

59 min episode · 2 min read
·
Logan Kilpatrick,Tulsee Doshi

Episode

59 min

Read time

2 min

Topics

Productivity, Relationships, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Model-Harness Co-training: Google now trains Gemini models in direct partnership with its "anti-gravity" agent harness, meaning the model is optimized for tool-calling loops, orchestration, and agentic workflows from the start. Teams building on top inherit this infrastructure without rebuilding it, cutting iteration cycles from months to days and enabling faster product launches across all Google surfaces.
  • Flash-First Strategy: Google deliberately prioritizes cost-adjusted performance over raw capability maximization. Gemini 3.5 Flash benchmarks at roughly 280 tokens per second on Artificial Analysis, runs three times faster than comparable large models, and costs significantly less. For consumer-scale products like Search and the Gemini app, latency improvements outperform quality gains in live experiments when users are unwilling to wait.
  • "Model Eats the Scaffolding" Cycle: Every 12–18 months, the surrounding scaffolding that developers build around AI models gets absorbed into the model itself. Product teams should avoid having every team independently rebuild agentic infrastructure from scratch. Standardizing on a shared harness layer reduces redundant engineering, accelerates deployment, and surfaces model failure modes faster through centralized feedback loops.
  • Context Window Economics: Context windows have plateaued near one million tokens because serving costs become prohibitive at scale — a single one-million-token request can cost several dollars. Google's direction shifts toward smart context compaction: selectively retrieving relevant information rather than expanding raw window size, which keeps latency and cost manageable while effectively giving models access to much larger information pools.
  • Recursive Self-Improvement — Practical, Not Singular: Google uses Gemini internally to improve Gemini, including running ablations, submitting code changes, and generating research reports autonomously. However, humans remain in the driver's seat on large pre-training runs due to the high compute cost of misdirection. The framing is deep human-AI collaboration rather than autonomous AI-led model development, with humans focused on strategic interpretation of results.

What It Covers

Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy across its billions-of-users product surface.

Key Questions Answered

  • Model-Harness Co-training: Google now trains Gemini models in direct partnership with its "anti-gravity" agent harness, meaning the model is optimized for tool-calling loops, orchestration, and agentic workflows from the start. Teams building on top inherit this infrastructure without rebuilding it, cutting iteration cycles from months to days and enabling faster product launches across all Google surfaces.
  • Flash-First Strategy: Google deliberately prioritizes cost-adjusted performance over raw capability maximization. Gemini 3.5 Flash benchmarks at roughly 280 tokens per second on Artificial Analysis, runs three times faster than comparable large models, and costs significantly less. For consumer-scale products like Search and the Gemini app, latency improvements outperform quality gains in live experiments when users are unwilling to wait.
  • "Model Eats the Scaffolding" Cycle: Every 12–18 months, the surrounding scaffolding that developers build around AI models gets absorbed into the model itself. Product teams should avoid having every team independently rebuild agentic infrastructure from scratch. Standardizing on a shared harness layer reduces redundant engineering, accelerates deployment, and surfaces model failure modes faster through centralized feedback loops.
  • Context Window Economics: Context windows have plateaued near one million tokens because serving costs become prohibitive at scale — a single one-million-token request can cost several dollars. Google's direction shifts toward smart context compaction: selectively retrieving relevant information rather than expanding raw window size, which keeps latency and cost manageable while effectively giving models access to much larger information pools.
  • Recursive Self-Improvement — Practical, Not Singular: Google uses Gemini internally to improve Gemini, including running ablations, submitting code changes, and generating research reports autonomously. However, humans remain in the driver's seat on large pre-training runs due to the high compute cost of misdirection. The framing is deep human-AI collaboration rather than autonomous AI-led model development, with humans focused on strategic interpretation of results.

Notable Moment

A Google safety and alignment researcher ran a full suite of model ablations from her phone while sitting in a hot tub, receiving a complete report within an hour. The anecdote illustrates how AI-assisted research workflows have already compressed tasks that previously required days of engineering time into single sessions.

Know someone who'd find this useful?

Episode Transcript

Hello, and welcome back to the Cognitive Revolution. Today, after some 340 episodes, I am very excited to share the first episode that I've ever recorded in person with fan favorite Logan Kilpatrick, member of technical staff at Google DeepMind, and Tulsi Doshi, senior director and head of product for Gemini models. The occasion for this conversation is Google's annual IO event, where they're launching the new Gemini 3.5 flash model, all sorts of agent infrastructure and AI product integrations, and plenty more. We've recorded on Friday, May 15, just a couple days before the event. And while many at Google, including my brother Craig, who's giving a keynote on Wednesday, were working overtime to polish their demos and presentations, the overall vibe, at least compared to the rest of the AI space, was one of relatively relaxed confidence. And why not? From twenty twenty four to twenty five, Google grew annual revenue by $50,000,000,000, as much as Anthropic is pulling in today. And they still have 25% of all global compute, the deepest pool of research talent anywhere, and the most comprehensive AI portfolio of any company with top tier positions not just in language models, but also self driving cars, medical and life sciences, and robotics. So after discussing the headline launches that they're announcing this week, which also include a new video generation model called Omni, which they hope will create a nano banana moment for video, a new and improved and more agent focused anti gravity, and a product called Spark, which will bring more agentic functionality to the main consumer Gemini app, I really wanted to take a step back and dig in on Google's overall AI strategy and philosophy. We discussed their decision to lead with the Flash model and more generally to emphasize the cost adjusted performance Pareto frontier, whereas Anthropic and OpenAI are clearly much more focused on competing to have the single most capable model in absolute terms. We talk about how DeepMind is no longer shipping models in isolation and leaving it up to product teams to figure out how to use them, but instead now providing a robust agent harness, which should help elevate and standardize AI experiences across Google's vast product surface. We get into the weeds on questions like why context windows seem to have mostly stopped growing, why Gemini model's knowledge cutoff is now more than a year ago, and whatever happened to that diffusion model line of work. Perhaps most importantly, we discuss how the team at Google relates to the AI's they're creating, how they're thinking about things like model psychology and welfare, and their views on recursive self improvement, which as you'll hear, is definitely a part of their plan, but not something that they seem to be so singularly focused on as other AI leaders. Overall, I think this is a great window into the thinking that underlies Google's AI research and product development, which has clearly sustained the company's historic run …

Get the full transcript (12,014 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Cognitive Revolution transcripts →

You just read a 3-minute summary of a 56-minute episode.

Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by Brave

    SPONSORS: Brave Search API
  • by Anthropic

    SPONSORS: Anthropic / Claude
  • by Google DeepMind

    Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy
  • by Sequence

    SPONSORS: Sequence
  • by Google DeepMind

    Google DeepMind's Logan Kilpatrick and Tulsee Doshi join host Neil Savage at Google HQ ahead of Google IO to discuss Gemini 3.5 Flash, the Omni video generation model, the Agent harness infrastructure, recursive self-improvement, context window limits, and Google's overall AI product strategy
  • by RoboFlow

    SPONSORS: RoboFlow

More from Cognitive Revolution

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into Cognitive Revolution.

Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime