[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Read time
2 min
Topics
Remote Work, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓RL Generalization Limits: Reinforcement learning applied to language models excels within training distribution but generalizes poorly beyond it. The strategy requires bringing economically useful tasks into distribution rather than expecting broad generalization, fundamentally changing how products must be designed around model capabilities.
- ✓Reasoning Team Scale: OpenAI's o1 model started with roughly fifty to one hundred contributors, expanding to three hundred people for o3 as safety, evaluation, and product teams joined. The progression from prototype to product requires exponentially more organizational resources than initial research breakthroughs suggest.
- ✓Continuous Learning Advantage: Cursor ships policy updates every two hours for tab autocomplete by co-locating product and ML teams. This rapid iteration cycle proves impossible at larger organizations where product and research groups operate separately, creating competitive advantage through organizational structure rather than pure technical capability.
- ✓Context Over Code: Automating programming jobs requires capturing years of accumulated knowledge—hyperparameter sweep results, architectural decisions, team conversations—not just writing code. Products must ingest this full context (Slack messages, Datadog traces, design documents) to enable models to replicate senior engineer decision-making processes effectively.
What It Covers
Ashvin Nair discusses his transition from OpenAI's reasoning team to Cursor, covering the development of o1/o3 models, reinforcement learning's role in achieving IMO gold performance, and Cursor's approach to co-designing products with models.
Key Questions Answered
- •RL Generalization Limits: Reinforcement learning applied to language models excels within training distribution but generalizes poorly beyond it. The strategy requires bringing economically useful tasks into distribution rather than expecting broad generalization, fundamentally changing how products must be designed around model capabilities.
- •Reasoning Team Scale: OpenAI's o1 model started with roughly fifty to one hundred contributors, expanding to three hundred people for o3 as safety, evaluation, and product teams joined. The progression from prototype to product requires exponentially more organizational resources than initial research breakthroughs suggest.
- •Continuous Learning Advantage: Cursor ships policy updates every two hours for tab autocomplete by co-locating product and ML teams. This rapid iteration cycle proves impossible at larger organizations where product and research groups operate separately, creating competitive advantage through organizational structure rather than pure technical capability.
- •Context Over Code: Automating programming jobs requires capturing years of accumulated knowledge—hyperparameter sweep results, architectural decisions, team conversations—not just writing code. Products must ingest this full context (Slack messages, Datadog traces, design documents) to enable models to replicate senior engineer decision-making processes effectively.
Notable Moment
Nair attended a forecasting conference where AI researchers predicted twenty percent math exam performance by 2027, while OpenAI already had internal models exceeding those benchmarks. The same forecasters simultaneously predicted Dyson spheres by 2035, revealing systematic miscalibration in short-term pessimism and long-term optimism.
Episode Transcript
Okay. We're here at NeurIPS. We're recording a special land space coverage of the folks at NeurIPS. And we're here with Astrid from CURSA. Welcome. Hi. Yeah. Thanks for having me. So I guess the like Astrid from CURSA is like a new identity. I didn't even know if I should say that. Because you only joined CROSSA for three months. Mhmm. Before that, you're opening. I worked in 01/2003. Mhmm. Before that, Berkeley, PhD in RL. Just but focus on robotics. Robotics. Yeah. Is it weird searching for robotics in language models? Okay. This is kind of interesting because a lot of people have been kind of doing this, Are you OpenAI? Yeah. You got a robotics. So actually, I I actually was, at OpenAI in 2017 also, where I was at on robotics. Yeah. I was interning, like, right before my PhD where I worked on robotics there. 2017. Is that, like, Jovan and the Aztum band over there? He was all famously opening as first intern. Oh, really? Okay. Then he might have been before. But, yeah, there was like a like 15 interns. It was a very different company. It was just like, robotics, DOTA, and like 15 interns that summer all having like pretty exciting individual products. Like, yeah, that set of interns if you look over there now is kinda cool. Yeah. But yeah, Anyone from that class that, like, you would shout out? Like, there's just, like, a lot of cool papers that came out, like, Leo Pinto, and I was at, NYU. Yeah. And Faiman. The, the person who leads, reasoning at XAI, I forgot his name. Eric. Well, he then No. Eric. Yeah. I forgot his name. But, he he he worked on, like, K FAC and stuff. I think, yeah. The vision dude, Greg? Not Greg, but yeah. But, yeah, it was it was, like, an exciting time to be there. But, yeah, I think I think robotics is a pretty good fit for LMs because, like, the switch ends up being pretty, like, you know, you kinda do similar things. Like, you wanna look at a lot of data. It's, like, kind of, hard to get it's, like, stuff get hard to get stuff working in robotics world. I think, you know, it kinda builds, like, very greedy people who, like, look at data a lot, that kind of thing. So, yeah. For whatever reason, I think, like, that transfer is, like, yeah. It's happening a lot, and I think it makes a lot of sense. One of my newest highlights so far, had dinner with it's like a small group dinner with Lex Freeman, yes, yesterday. And Lex used to be in robotics. Mhmm. And he was like, my assessment of robotics people robotics people are the best to talk to Ed and NeurIPS Mhmm. Because they're most grounded, he says, Because they don't have a choice. They work with the real world. So they look at data. …
Get the full transcript (9,563 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
The AI Breakdown
What to Use the Latest AI Tools For
Sep 11
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The AI Breakdown
How to Navigate the Next Wave of AI Competition
Aug 31
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Products
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
The AI Breakdown
Sep 11
What to Use the Latest AI Tools For
The AI Breakdown
Aug 31
How to Navigate the Next Wave of AI Competition
This Week in Startups
Aug 26
Bill Gates foresees massive AI job loss: these VCs disagree | E2330
The AI Breakdown
May 19
9 Codex Tips From the Codex Team
How I AI
Apr 27
From a $6.90 newsletter to $3M API: How a non-coder built Memelord | Jason Levin
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime