Skip to main content
KG

Kirtana Gopalakrishnan

Google Deepmind Staff Research Scientist Keerthana**robotics Development Stage**gemini Robotics Er2 Architecture**simulation-to-real Gap**data Mixture Strategy
1episode
1podcast

We have 1 summarized appearance for Kirtana Gopalakrishnan so far. Browse all podcasts to discover more episodes.

Featured On 1 Podcast

Top resources Kirtana Gopalakrishnan mentions

Books, tools, and gear cited across podcast appearances. Ranked by frequency.

SignalCast may earn commission on purchases via affiliate links on each resource page.

All Appearances

1 episode
Cognitive Revolution

One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids

Cognitive Revolution
92 minStaff Research Scientist at Google DeepMind, Research Lead for Gemini Robotics

AI Summary

→ WHAT IT COVERS Google DeepMind staff research scientist Keerthana Gopalakrishnan discusses Gemini Robotics 2's three-model architecture — including the publicly available ER2 embodied reasoning API — alongside cross-embodiment generalization challenges, hardware progress on multi-finger dexterity, simulation-to-real gaps, and why robotics remains in its GPT-2 era despite accelerating capability gains. → KEY INSIGHTS - **Robotics Development Stage:** Gopalakrishnan places the entire robotics field at roughly a GPT-2 level of maturity, citing two specific gaps: few-shot learning does not yet generalize reliably across diverse task types, and models remain highly sensitive to which robot body they were trained on. Unlike LLMs that run identically across any hardware, robotics models lose significant capability when transferred to a new embodiment, making cross-embodiment generalization the field's defining unsolved problem. - **Gemini Robotics ER2 Architecture:** The ER2 model functions as a System 2 reasoning layer built on Gemini 3.5 Flash, available via API with a 128K context window translating to roughly three minutes of dense robot episode memory. Developers can define affordances as tools — grab-at, place-at, point-to — and ER2 orchestrates these via multimodal prompting. The VLA (Vision-Language-Action) model then translates those high-level instructions into joint-space robot actions, with the two models running in a compounding error chain. - **Simulation-to-Real Gap:** Locomotion tasks like bipedal running train well in simulation because flat-surface contact dynamics are mathematically predictable. Manipulation tasks — especially involving deformable objects like cloth, eggs, or trash bags — break simulation fidelity because friction, soft-body deformation, and contact timing are difficult to model accurately. Rigid-body pick-and-place sits in the middle: simulation works reasonably well, making it close to deployment-ready, while dexterous manipulation of soft objects remains a primary research frontier. - **Data Mixture Strategy:** No single data modality solves robotics at scale. Teleoperation data is precise but not scalable and becomes less useful as robot hardware evolves. UMI-style wearable data offers better scale with sensor-accurate action labels but requires specialized hardware. Egocentric human video scales cheaply but introduces noise from body-size variation and imprecise end-effector estimation. The practical path forward is a weighted mixture of all three, with the optimal ratio shifting as new hardware and simulation tools emerge. - **Whole-Body Control Milestone:** Gemini Robotics 2 controls a full humanoid from fingertips to feet in a closed loop, representing a capability that did not exist in any public demo roughly 14 months prior. With as few as 200 task demonstrations, the on-device model can learn new tasks across multiple robot bodies by leveraging a pre-trained cross-embodiment foundation. The primary remaining challenge is orchestration latency: when ER2 and the VLA run sequentially, errors compound across multi-step tasks, reducing end-to-end reliability. - **Hardware Dexterity Progress:** Multi-finger robot hands have advanced to where they now match or exceed the dexterity ceiling that gripper-based robots previously defined. The Woojin hand approximates the grip strength of a 10-year-old child, while the Schunk hand can lift approximately 20 kilograms and has demonstrated jar-opening capability. The remaining hardware gaps are reliability, repeatability under sustained use, and cost reduction — not raw capability. Tactile sensing and soft-robotics glove interfaces remain active research areas without settled solutions. - **Safety as Capability Design:** Gopalakrishnan frames safety not as a constraint on capability but as a required capability layer spanning three distinct levels: mechanical operational safety (preventing falls and collisions), behavioral guardrail compliance (not overriding human instructions), and sensor-failure response (detecting obstructed vision or unexpected contact and requesting human intervention rather than continuing blindly). Humanoid form factors face a higher public expectation bar than arm robots — visible failures are judged more harshly because the human-like appearance raises implicit competence expectations. → NOTABLE MOMENT During filming for Gemini Robotics 2, a production crew unfamiliar with robots began calling "action" to the humanoid as if directing a human actor. Separately, when asked which object was hardest to handle during a packing task, the robot identified a videotape as its favorite item — prompting laughter from the research team watching from the back of the room. 💼 SPONSORS [{"name": "Athena", "url": "https://athena.com/cognitive"}, {"name": "Parallel", "url": "https://parallel.ai/tcr"}, {"name": "Deepgram", "url": "https://deepgram.com"}, {"name": "OutSystems", "url": "https://outsystems.com/tcr"}, {"name": "Anthropic", "url": "https://claude.ai/tcr"}] 🏷️ Robotics Foundation Models, Cross-Embodiment Generalization, Gemini Robotics API, Simulation-to-Real Transfer, Humanoid Dexterity, Robot Safety Alignment, Vision-Language-Action Models

Explore More

Never miss Kirtana Gopalakrishnan's insights

Subscribe to get AI-powered summaries of Kirtana Gopalakrishnan's podcast appearances delivered to your inbox weekly.

Start Free Today

No credit card required • Free tier available