Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Episode
92 min
Read time
3 min
Topics
Career Growth, Productivity, Startups
AI-Generated Summary
Key Takeaways
- ✓On-Policy vs Off-Policy RL: On-policy learning means models generate their own outputs, receive rewards based on those generations, then train on their own trajectories rather than imitating others' successful paths. This approach proves more generalizable than supervised fine-tuning because models learn from first principles by making mistakes and receiving feedback, similar to how humans learn through trial and error rather than pure imitation of expert demonstrations.
- ✓IMO Gold Architecture Decision: DeepMind made the bold choice to completely abandon AlphaProof's specialized symbolic system and use end-to-end Gemini models for IMO 2024. The training process took approximately one week with four co-captains across London, Mountain View, and Singapore coordinating in different time zones. This live competition format with problems released on different days created more adrenaline than standard benchmark optimization since gold medal thresholds depend on human participant performance.
- ✓Reasoning as Post-Training RL: The technical definition of reasoning in modern LLMs centers on using reinforcement learning during post-training to elicit better thinking capabilities. This involves training models to improve with extended thinking time, whether through discrete token chain-of-thought or latent space representations. The field has moved beyond pure architecture innovation toward RL as the primary modeling toolset for capability improvements at the frontier.
- ✓Data Efficiency Gap: Humans demonstrate eight orders of magnitude better data efficiency than current models - a two-year-old child shows more capability than LLMs after seeing vastly less data. The solution likely involves spending more compute per token during training rather than just scaling dataset size. This represents a fundamental research direction as the field approaches data constraints, with the bug potentially residing in backpropagation, architecture, or learning algorithms.
- ✓Transformer Architecture Persistence: Self-attention will likely remain central to AGI systems despite eight years since the original paper. Attempts to remove or simplify attention consistently fail unless at least one attention layer remains. The architecture serves as the interface between learning algorithms and tokens, with any paradigm shift requiring changes to the entire learning system including backpropagation, not just architectural modifications.
What It Covers
Yi Tay, co-captain of Google DeepMind's IMO Gold effort and leader of the Singapore Gemini team, discusses the decision to abandon AlphaProof's symbolic system for end-to-end Gemini models, on-policy reinforcement learning philosophy, the one-week training process for IMO models, and why data efficiency and world models represent the next frontier in AI research beyond pure scaling.
Key Questions Answered
- •On-Policy vs Off-Policy RL: On-policy learning means models generate their own outputs, receive rewards based on those generations, then train on their own trajectories rather than imitating others' successful paths. This approach proves more generalizable than supervised fine-tuning because models learn from first principles by making mistakes and receiving feedback, similar to how humans learn through trial and error rather than pure imitation of expert demonstrations.
- •IMO Gold Architecture Decision: DeepMind made the bold choice to completely abandon AlphaProof's specialized symbolic system and use end-to-end Gemini models for IMO 2024. The training process took approximately one week with four co-captains across London, Mountain View, and Singapore coordinating in different time zones. This live competition format with problems released on different days created more adrenaline than standard benchmark optimization since gold medal thresholds depend on human participant performance.
- •Reasoning as Post-Training RL: The technical definition of reasoning in modern LLMs centers on using reinforcement learning during post-training to elicit better thinking capabilities. This involves training models to improve with extended thinking time, whether through discrete token chain-of-thought or latent space representations. The field has moved beyond pure architecture innovation toward RL as the primary modeling toolset for capability improvements at the frontier.
- •Data Efficiency Gap: Humans demonstrate eight orders of magnitude better data efficiency than current models - a two-year-old child shows more capability than LLMs after seeing vastly less data. The solution likely involves spending more compute per token during training rather than just scaling dataset size. This represents a fundamental research direction as the field approaches data constraints, with the bug potentially residing in backpropagation, architecture, or learning algorithms.
- •Transformer Architecture Persistence: Self-attention will likely remain central to AGI systems despite eight years since the original paper. Attempts to remove or simplify attention consistently fail unless at least one attention layer remains. The architecture serves as the interface between learning algorithms and tokens, with any paradigm shift requiring changes to the entire learning system including backpropagation, not just architectural modifications.
- •AI Coding Productivity Shift: AI coding tools crossed an emergent threshold in 2024 where researchers can paste bugs directly into systems like Anthropic's Claude without examining them, receive fixes, and relaunch jobs with high success rates. This represents "vibe training" beyond basic code generation - the model investigates and solves problems the researcher doesn't fully understand. Time saved per bug can reach full workdays, creating passive productivity buffs across entire teams.
- •Geographic Research Strategy: Singapore's advantage for frontier AI research lies in being far enough from Bay Area culture for mental space while maintaining connectivity. The timezone enables 24-hour job coverage when coordinating with London and Mountain View teams. Talent density matters more than team size, with hiring focused on exceptional RL research track records or competitive programming achievements rather than rapid scaling.
Notable Moment
Tay reveals he knew nothing about the IMO competition itself and only trained the model checkpoint used for the live event. Team members flew to Australia to receive problems as they released, running inference in real-time while others gathered in London for a hackathon-style coordination. The gold medal threshold depended on human participant scores, creating uncertainty until final verification by the IMO committee.
Episode Transcript
The thing that I find the most useful about, like, these models in general is, like, when I have this big spreadsheet of a lot of results and I just need plots of it. I think models can quite go to the screenshot and make a plot of this. I hate making this method like stuff about it's so annoying. There were so many moments this year where AI suddenly crossed that, like, the emergent thing. Like, the AI coding is one of them like we just discussed. I think, like, Nanobana also got to the point where I usually, like, you make these images, it's just like a little for fun. It just throw your friend or something like that. But like Nanon Mile actually really got so good. Welcome back. Yeah. How are you? Yeah. I'm good. I'm good. Great to be back. It's been one one and a half years. Yeah. It's been one and a half. Feels like a long time. So last time we talked, you were at Rekha. Yeah. And then you joined GDM again, working for Quark again. Yeah. And more recently, you've started GDM Singapore. Yeah. Is it GDM Singapore or Gemini Singapore? I don't know if you've named the team. Oh, I I I think we have a Gemini team in Singapore. Yeah. Team in Singapore. It's called Reasoning and AGI. Yeah. Reasoning and AGI. Is it important to have AGI in the name? It was like a white thing that we put AGI in. Yeah. I think that, like, one of the reason why we work on this model is that we wanna get to AGI. And I just with the white thing that we added the AGI to the job posting. Yeah. But there's no, like, formal name of the team yet, but it's basically the Gemini team Singapore. Yeah. I mean, I think people are, like, trying to triangulate Amazon as an AGI team, and you guys have AGI team. And then, let's say Meta now has a superintelligence team. What are people signaling when they choose these names for their teams? Do they have, oh, we have a pen, or is it just vibes? You try to official hot takes on the no. You have officially AGI in your job title? No. It's not a thing, ma'am. It's it's not a thing. Yeah. Yeah. It's just you know, we we just want to signal the north star if we're building this model, let's go to AGI. Yeah. Yeah. No. I I wasn't really fishing politics. Okay. So you rejoined GDM. Yeah. And I think last time when you talked about I listened back to the whole thing. It was amazing episode last time. You were talking about how it's like externally. Word in brain and came out, and now you're back in GDM. Yeah. I wonder what's your just your general reflections just plugging back into the Google infrastructure. Oh, yeah. So I guess coming back, it's …
Get the full transcript (19,823 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 89-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
The Vergecast
What's behind the Google AI shakeup
Aug 7
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The AI Breakdown
Google’s AI Leadership Shakeup: Disaster or Exactly What It Needs?
Aug 6
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
- ClaudeRecommended
by Anthropic
“AI coding tools crossed an emergent threshold in 2024 where researchers can paste bugs directly into systems like Anthropic's Claude without examining them, receive fixes, and relaunch jobs with high success rates.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
The Vergecast
Aug 7
What's behind the Google AI shakeup
The AI Breakdown
Aug 6
Google’s AI Leadership Shakeup: Disaster or Exactly What It Needs?
Accidental Tech Podcast
Jun 9
695: The Crystal Pepsi of Aqua
Cognitive Revolution
May 20
The Model Eats the Scaffolding: DeepMind's Logan Kilpatrick & Tulsee Doshi on 3.5 Flash, Omni & More
Software Engineering Daily
Mar 12
DeepMind’s RAG System with Animesh Chatterji and Ivan Solovyev
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime