Skip to main content
Cognitive Revolution

One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids

92 min episode · 3 min read
·
Kirtana Gopalakrishnan

Episode

92 min

Read time

3 min

Topics

Relationships, Fundraising & VC, Design & UX

AI-Generated Summary

Key Takeaways

  • ✓Robotics Development Stage: Gopalakrishnan places the entire robotics field at roughly a GPT-2 level of maturity, citing two specific gaps: few-shot learning does not yet generalize reliably across diverse task types, and models remain highly sensitive to which robot body they were trained on. Unlike LLMs that run identically across any hardware, robotics models lose significant capability when transferred to a new embodiment, making cross-embodiment generalization the field's defining unsolved problem.
  • ✓Gemini Robotics ER2 Architecture: The ER2 model functions as a System 2 reasoning layer built on Gemini 3.5 Flash, available via API with a 128K context window translating to roughly three minutes of dense robot episode memory. Developers can define affordances as tools — grab-at, place-at, point-to — and ER2 orchestrates these via multimodal prompting. The VLA (Vision-Language-Action) model then translates those high-level instructions into joint-space robot actions, with the two models running in a compounding error chain.
  • ✓Simulation-to-Real Gap: Locomotion tasks like bipedal running train well in simulation because flat-surface contact dynamics are mathematically predictable. Manipulation tasks — especially involving deformable objects like cloth, eggs, or trash bags — break simulation fidelity because friction, soft-body deformation, and contact timing are difficult to model accurately. Rigid-body pick-and-place sits in the middle: simulation works reasonably well, making it close to deployment-ready, while dexterous manipulation of soft objects remains a primary research frontier.
  • ✓Data Mixture Strategy: No single data modality solves robotics at scale. Teleoperation data is precise but not scalable and becomes less useful as robot hardware evolves. UMI-style wearable data offers better scale with sensor-accurate action labels but requires specialized hardware. Egocentric human video scales cheaply but introduces noise from body-size variation and imprecise end-effector estimation. The practical path forward is a weighted mixture of all three, with the optimal ratio shifting as new hardware and simulation tools emerge.
  • ✓Whole-Body Control Milestone: Gemini Robotics 2 controls a full humanoid from fingertips to feet in a closed loop, representing a capability that did not exist in any public demo roughly 14 months prior. With as few as 200 task demonstrations, the on-device model can learn new tasks across multiple robot bodies by leveraging a pre-trained cross-embodiment foundation. The primary remaining challenge is orchestration latency: when ER2 and the VLA run sequentially, errors compound across multi-step tasks, reducing end-to-end reliability.

What It Covers

Google DeepMind staff research scientist Keerthana Gopalakrishnan discusses Gemini Robotics 2's three-model architecture — including the publicly available ER2 embodied reasoning API — alongside cross-embodiment generalization challenges, hardware progress on multi-finger dexterity, simulation-to-real gaps, and why robotics remains in its GPT-2 era despite accelerating capability gains.

Key Questions Answered

  • •Robotics Development Stage: Gopalakrishnan places the entire robotics field at roughly a GPT-2 level of maturity, citing two specific gaps: few-shot learning does not yet generalize reliably across diverse task types, and models remain highly sensitive to which robot body they were trained on. Unlike LLMs that run identically across any hardware, robotics models lose significant capability when transferred to a new embodiment, making cross-embodiment generalization the field's defining unsolved problem.
  • •Gemini Robotics ER2 Architecture: The ER2 model functions as a System 2 reasoning layer built on Gemini 3.5 Flash, available via API with a 128K context window translating to roughly three minutes of dense robot episode memory. Developers can define affordances as tools — grab-at, place-at, point-to — and ER2 orchestrates these via multimodal prompting. The VLA (Vision-Language-Action) model then translates those high-level instructions into joint-space robot actions, with the two models running in a compounding error chain.
  • •Simulation-to-Real Gap: Locomotion tasks like bipedal running train well in simulation because flat-surface contact dynamics are mathematically predictable. Manipulation tasks — especially involving deformable objects like cloth, eggs, or trash bags — break simulation fidelity because friction, soft-body deformation, and contact timing are difficult to model accurately. Rigid-body pick-and-place sits in the middle: simulation works reasonably well, making it close to deployment-ready, while dexterous manipulation of soft objects remains a primary research frontier.
  • •Data Mixture Strategy: No single data modality solves robotics at scale. Teleoperation data is precise but not scalable and becomes less useful as robot hardware evolves. UMI-style wearable data offers better scale with sensor-accurate action labels but requires specialized hardware. Egocentric human video scales cheaply but introduces noise from body-size variation and imprecise end-effector estimation. The practical path forward is a weighted mixture of all three, with the optimal ratio shifting as new hardware and simulation tools emerge.
  • •Whole-Body Control Milestone: Gemini Robotics 2 controls a full humanoid from fingertips to feet in a closed loop, representing a capability that did not exist in any public demo roughly 14 months prior. With as few as 200 task demonstrations, the on-device model can learn new tasks across multiple robot bodies by leveraging a pre-trained cross-embodiment foundation. The primary remaining challenge is orchestration latency: when ER2 and the VLA run sequentially, errors compound across multi-step tasks, reducing end-to-end reliability.
  • •Hardware Dexterity Progress: Multi-finger robot hands have advanced to where they now match or exceed the dexterity ceiling that gripper-based robots previously defined. The Woojin hand approximates the grip strength of a 10-year-old child, while the Schunk hand can lift approximately 20 kilograms and has demonstrated jar-opening capability. The remaining hardware gaps are reliability, repeatability under sustained use, and cost reduction — not raw capability. Tactile sensing and soft-robotics glove interfaces remain active research areas without settled solutions.
  • •Safety as Capability Design: Gopalakrishnan frames safety not as a constraint on capability but as a required capability layer spanning three distinct levels: mechanical operational safety (preventing falls and collisions), behavioral guardrail compliance (not overriding human instructions), and sensor-failure response (detecting obstructed vision or unexpected contact and requesting human intervention rather than continuing blindly). Humanoid form factors face a higher public expectation bar than arm robots — visible failures are judged more harshly because the human-like appearance raises implicit competence expectations.

Notable Moment

During filming for Gemini Robotics 2, a production crew unfamiliar with robots began calling "action" to the humanoid as if directing a human actor. Separately, when asked which object was hardest to handle during a packing task, the robot identified a videotape as its favorite item — prompting laughter from the research team watching from the back of the room.

Know someone who'd find this useful?

Episode Transcript

Hello, and welcome back to the Cognitive Revolution. Today, I'm excited to welcome Kirtana Gopalakrishnan, staff research scientist at Google DeepMind and research lead for Gemini Robotics, back for her fourth annual appearance on the show. One of the biggest questions in AI today is how soon will general purpose robots become broadly useful. AI is already affecting the world in major ways, even in its purely digital form. But the most sci fi forecasts for the AI future predict that robotics will soon hit key tipping points. AI 2027, for example, predicts that humanoid robots will become useful sometime in mid twenty twenty seven, and that by 2028, a whole robot economy could take shape, where robots build more robot factories, which in turn produce more robots, leading to unprecedented exponential economic growth but also creating serious risk of AI takeover. So how does that vision line up with the reality of robotics research today? We begin with a discussion of the extremely viral Robot Olympics held in China this summer, as I was really curious to get Kirtan's take on what it means that humanoid robots can now run faster than the fastest humans. Her point of view, as you'll hear, inverts the usual US China dichotomy in AI. She says that while the videos are obviously impressive, skills like running on a flat track are actually relatively easy to train in simulation. And more to the point, foot speed is not really a limiting factor in the utility that robots can provide. As such, her team at Google is more focused on practical value. From there, we get into Gemini Robotics two, a suite of three models that Google released this summer. Gemini Robotics e r two for embodied reasoning, which Kirtana describes as a system two for robotics control, is based on Gemini Flash and is available via the API, allowing developers to define the affordances available to the models as tools just like we do with digital agents. I played around with it and found it remarkably accessible. The other two models, Gemini Robotics two and Gemini Robotics on device two, translate these higher level tool calls down to the level of robot action and are now capable of controlling the whole robot from fingertips to feet on a wide range of form factors, but are currently available only to trusted testers. Having understood how these models work, we again zoom out, and I ask Kirtana how she understands progress in robotics overall. On the one hand, like so many other AI researchers recently, she says that she has been surprised by the pace of progress. But at the same time, she still feels that robotics remains in its GPT two era. We have seen a number of recent demos of robots learning new tasks from just a few or even just a single human demonstration. But Kierfen argues that the range of tasks you can teach this way and the generalization profile of robotics …

Get the full transcript (15,784 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Cognitive Revolution transcripts →

You just read a 3-minute summary of a 89-minute episode.

Get Cognitive Revolution summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • by Google DeepMind

    “Google DeepMind staff research scientist Keerthana Gopalakrishnan discusses Gemini Robotics 2's three-model architecture — including the publicly available ER2 embodied reasoning API”
  • by Google DeepMind

    “The ER2 model functions as a System 2 reasoning layer built on Gemini 3.5 Flash, available via API with a 128K context window translating to roughly three minutes of dense robot episode memory.”
  • by Google DeepMind

    “The ER2 model functions as a System 2 reasoning layer built on Gemini 3.5 Flash, available via API with a 128K context window”

Gear

  • by Woojin

    “The Woojin hand approximates the grip strength of a 10-year-old child, while the Schunk hand can lift approximately 20 kilograms and has demonstrated jar-opening capability.”
  • by Schunk

    “The Woojin hand approximates the grip strength of a 10-year-old child, while the Schunk hand can lift approximately 20 kilograms and has demonstrated jar-opening capability.”

More from Cognitive Revolution

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into Cognitive Revolution.

Every Monday, we deliver AI summaries of the latest episodes from Cognitive Revolution and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime