[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Episode
45 min
Read time
2 min
Topics
Investing, Fundraising & VC, Leadership
AI-Generated Summary
Key Takeaways
- ✓Reasoning Model Development Scale: OpenAI's o1 reasoning model started with approximately 12 core people but expanded to 50-100 contributors for the initial release and eventually 300 people for o3. The breakthrough came in 2023 when RL applied to smaller pretrained models produced surprisingly accurate reasoning traces on math problems, demonstrating capabilities unachievable through additional pretraining alone, leading to full-scale investment.
- ✓RL Generalization Limitations: Reinforcement learning for language models excels at dominating training distributions but generalizes poorly beyond them. The solution requires bringing economically useful tasks into the training distribution rather than expecting broad generalization. This means products must capture complete user context including code repositories, terminal access, conversation history, and workflow data to enable effective RL training on real-world tasks.
- ✓Robotics Market Timing: Language model agents represent a trillion-dollar market opportunity before robotics reaches even ten billion dollars in value. Current AI robotics sits at the GPT-1 to GPT-2 development stage, showing hints of generalization but lacking reliable out-of-distribution performance. The technology requires demonstrable value creation before unit economics can work, including maintenance costs and reliability thresholds for commercial deployment.
- ✓Continual Learning Gap: Models trained on trillions of tokens should theoretically handle millions of deployment tokens without capacity constraints, yet they repeatedly make identical mistakes within and across contexts. The field needs breakthroughs in continual learning that enable models to permanently learn from single experiences, similar to humans avoiding hot stoves after one touch, rather than requiring explicit data curation and filtering.
- ✓Product-Model Co-Design: Cursor's 20-25 person ML team ships competitive models by tightly integrating product and model development. Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference. Internal tooling enables SSH sessions into user environments for direct data inspection, and policy updates occur every two hours, impossible at larger organizations with separated product and research teams.
What It Covers
Ashvin Nair, former OpenAI reasoning team member now at Cursor, discusses the transition from robotics to language models, the development of OpenAI's o1/o3 reasoning models with a 300-person team, achieving IMO/IOI gold medals, and Cursor's approach to co-designing products with models through rapid RL iteration cycles every two hours.
Key Questions Answered
- •Reasoning Model Development Scale: OpenAI's o1 reasoning model started with approximately 12 core people but expanded to 50-100 contributors for the initial release and eventually 300 people for o3. The breakthrough came in 2023 when RL applied to smaller pretrained models produced surprisingly accurate reasoning traces on math problems, demonstrating capabilities unachievable through additional pretraining alone, leading to full-scale investment.
- •RL Generalization Limitations: Reinforcement learning for language models excels at dominating training distributions but generalizes poorly beyond them. The solution requires bringing economically useful tasks into the training distribution rather than expecting broad generalization. This means products must capture complete user context including code repositories, terminal access, conversation history, and workflow data to enable effective RL training on real-world tasks.
- •Robotics Market Timing: Language model agents represent a trillion-dollar market opportunity before robotics reaches even ten billion dollars in value. Current AI robotics sits at the GPT-1 to GPT-2 development stage, showing hints of generalization but lacking reliable out-of-distribution performance. The technology requires demonstrable value creation before unit economics can work, including maintenance costs and reliability thresholds for commercial deployment.
- •Continual Learning Gap: Models trained on trillions of tokens should theoretically handle millions of deployment tokens without capacity constraints, yet they repeatedly make identical mistakes within and across contexts. The field needs breakthroughs in continual learning that enable models to permanently learn from single experiences, similar to humans avoiding hot stoves after one touch, rather than requiring explicit data curation and filtering.
- •Product-Model Co-Design: Cursor's 20-25 person ML team ships competitive models by tightly integrating product and model development. Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference. Internal tooling enables SSH sessions into user environments for direct data inspection, and policy updates occur every two hours, impossible at larger organizations with separated product and research teams.
Notable Moment
Nair reveals that when attending The Curve conference before o1's release, attendees predicted 20% performance on math benchmarks by 2027, yet OpenAI already had models exceeding those estimates internally. The same forecasters predicting Dyson spheres by 2035 were simultaneously underestimating near-term capabilities by multiple years, demonstrating systematic miscalibration in AI progress predictions.
Episode Transcript
Okay. We're here at NeurIPS. We're recording a special land space coverage of the folks at NeurIPS. And we're here with Astrid from Cursa. Welcome. Hi. Yeah. Thanks for having me. So I guess the like Astrid from Cursa is like a new identity. I didn't even know if I should say that because you only joined Corsa for three months. Before that, you're opening your work in 01/2003. Mhmm. Before that, Berkeley, PhD in RL. Just but focus on robotics. Robotics. Yeah. Is it weird searching for robotics to language models? Okay. This is kind of interesting because a lot of people have been kind of doing this, Are you opening AI? Yeah. You got robotics. So actually, I I actually was, at OpenAI in 2017 also, working on robotics. Yeah. I was interning, like, right before my PhD where I worked on robotics there. 2017. Is that, like, Jovan and, as Jovan over there? Or he, he was all famously opening as first intern. Oh, really? Okay. Then he might have been before. But, yeah, there was, like, like, 15 interns. It was a very different company. It was just, like, robotics, DOTA, and, like, 15 interns that summer all having, like, pretty exciting individual products. Like, yeah, that set of interns, if you look over there now, is kinda cool. Yeah. But, yeah, you went to anyone from that class that, like, you would shout out? Like, there's just, like, a lot of cool papers that came out. Like, Lo, Pinto, and I was at, NYU. Yeah. And David. The, the person who leads, reasoning at XAI, I forgot his name. Eric. Well, he then. No. Eric. But, yeah, I forgot his name. But, he he he worked on, like, K FAC and stuff, I think. Yeah. Vision dude, Greg? Not Greg, but yeah. But, yeah, it was it was, like, an exciting time to be there. But, yeah, I think I think robotics is a pretty good fit for LNs because, like, the switch ends up being pretty, like, you know, you kinda do similar things. Like, you wanna look at a lot of data. It's, like, kind of, hard to get, like, stuff get hard to get stuff working in robotics world. I think, you know, it kind of builds, like, very gritty people who, like, look at data a lot, that kind of thing. So, yeah. For whatever reason, I think, like, that transfer is, like, yeah. It's happening a lot, and I think it makes a lot of sense. One of my newest highlights so far, had dinner with it's like a small group dinner with Lex Freeman, yesterday. And Lex used to be in robotics. Mhmm. And he was like, my assessment of robotics people robotics people are the best to talk to at NeurIPS Mhmm. Because they're most grounded, he says, Because they don't have a choice. They work with the real world, such as looking data. And then, like, the …
Get the full transcript (9,551 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 42-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
Making Sense
#420 — Countdown to Superintelligence
Jun 12
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The Ezra Klein Show
The A.I.s Are Already Out of Control
Aug 18
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
- Cursor ComposerBy guest
by Cursor
“Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference.”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
Making Sense
Jun 12
#420 — Countdown to Superintelligence
The Ezra Klein Show
Aug 18
The A.I.s Are Already Out of Control
Lenny's Podcast
Aug 9
The playbook for building high talent density teams | Adam Ward, Head of Talent at Cursor
Cognitive Revolution
Jun 20
Dean Ball, on Joining OpenAI: New Power Centers, Frontier AI Policy, & Main Character Energy
The AI Breakdown
May 19
9 Codex Tips From the Codex Team
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime