Skip to main content
Latent Space

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

45 min episode · 2 min read
·
Ashvin Nair

Episode

45 min

Read time

2 min

Topics

Investing, Fundraising & VC, Leadership

AI-Generated Summary

Key Takeaways

  • Reasoning Model Development Scale: OpenAI's o1 reasoning model started with approximately 12 core people but expanded to 50-100 contributors for the initial release and eventually 300 people for o3. The breakthrough came in 2023 when RL applied to smaller pretrained models produced surprisingly accurate reasoning traces on math problems, demonstrating capabilities unachievable through additional pretraining alone, leading to full-scale investment.
  • RL Generalization Limitations: Reinforcement learning for language models excels at dominating training distributions but generalizes poorly beyond them. The solution requires bringing economically useful tasks into the training distribution rather than expecting broad generalization. This means products must capture complete user context including code repositories, terminal access, conversation history, and workflow data to enable effective RL training on real-world tasks.
  • Robotics Market Timing: Language model agents represent a trillion-dollar market opportunity before robotics reaches even ten billion dollars in value. Current AI robotics sits at the GPT-1 to GPT-2 development stage, showing hints of generalization but lacking reliable out-of-distribution performance. The technology requires demonstrable value creation before unit economics can work, including maintenance costs and reliability thresholds for commercial deployment.
  • Continual Learning Gap: Models trained on trillions of tokens should theoretically handle millions of deployment tokens without capacity constraints, yet they repeatedly make identical mistakes within and across contexts. The field needs breakthroughs in continual learning that enable models to permanently learn from single experiences, similar to humans avoiding hot stoves after one touch, rather than requiring explicit data curation and filtering.
  • Product-Model Co-Design: Cursor's 20-25 person ML team ships competitive models by tightly integrating product and model development. Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference. Internal tooling enables SSH sessions into user environments for direct data inspection, and policy updates occur every two hours, impossible at larger organizations with separated product and research teams.

What It Covers

Ashvin Nair, former OpenAI reasoning team member now at Cursor, discusses the transition from robotics to language models, the development of OpenAI's o1/o3 reasoning models with a 300-person team, achieving IMO/IOI gold medals, and Cursor's approach to co-designing products with models through rapid RL iteration cycles every two hours.

Key Questions Answered

  • Reasoning Model Development Scale: OpenAI's o1 reasoning model started with approximately 12 core people but expanded to 50-100 contributors for the initial release and eventually 300 people for o3. The breakthrough came in 2023 when RL applied to smaller pretrained models produced surprisingly accurate reasoning traces on math problems, demonstrating capabilities unachievable through additional pretraining alone, leading to full-scale investment.
  • RL Generalization Limitations: Reinforcement learning for language models excels at dominating training distributions but generalizes poorly beyond them. The solution requires bringing economically useful tasks into the training distribution rather than expecting broad generalization. This means products must capture complete user context including code repositories, terminal access, conversation history, and workflow data to enable effective RL training on real-world tasks.
  • Robotics Market Timing: Language model agents represent a trillion-dollar market opportunity before robotics reaches even ten billion dollars in value. Current AI robotics sits at the GPT-1 to GPT-2 development stage, showing hints of generalization but lacking reliable out-of-distribution performance. The technology requires demonstrable value creation before unit economics can work, including maintenance costs and reliability thresholds for commercial deployment.
  • Continual Learning Gap: Models trained on trillions of tokens should theoretically handle millions of deployment tokens without capacity constraints, yet they repeatedly make identical mistakes within and across contexts. The field needs breakthroughs in continual learning that enable models to permanently learn from single experiences, similar to humans avoiding hot stoves after one touch, rather than requiring explicit data curation and filtering.
  • Product-Model Co-Design: Cursor's 20-25 person ML team ships competitive models by tightly integrating product and model development. Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference. Internal tooling enables SSH sessions into user environments for direct data inspection, and policy updates occur every two hours, impossible at larger organizations with separated product and research teams.

Notable Moment

Nair reveals that when attending The Curve conference before o1's release, attendees predicted 20% performance on math benchmarks by 2027, yet OpenAI already had models exceeding those estimates internally. The same forecasters predicting Dyson spheres by 2035 were simultaneously underestimating near-term capabilities by multiple years, demonstrating systematic miscalibration in AI progress predictions.

Know someone who'd find this useful?

Episode Transcript

Okay. We're here at NeurIPS. We're recording a special land space coverage of the folks at NeurIPS. And we're here with Astrid from Cursa. Welcome. Hi. Yeah. Thanks for having me. So I guess the like Astrid from Cursa is like a new identity. I didn't even know if I should say that because you only joined Corsa for three months. Before that, you're opening your work in 01/2003. Mhmm. Before that, Berkeley, PhD in RL. Just but focus on robotics. Robotics. Yeah. Is it weird searching for robotics to language models? Okay. This is kind of interesting because a lot of people have been kind of doing this, Are you opening AI? Yeah. You got robotics. So actually, I I actually was, at OpenAI in 2017 also, working on robotics. Yeah. I was interning, like, right before my PhD where I worked on robotics there. 2017. Is that, like, Jovan and, as Jovan over there? Or he, he was all famously opening as first intern. Oh, really? Okay. Then he might have been before. But, yeah, there was, like, like, 15 interns. It was a very different company. It was just, like, robotics, DOTA, and, like, 15 interns that summer all having, like, pretty exciting individual products. Like, yeah, that set of interns, if you look over there now, is kinda cool. Yeah. But, yeah, you went to anyone from that class that, like, you would shout out? Like, there's just, like, a lot of cool papers that came out. Like, Lo, Pinto, and I was at, NYU. Yeah. And David. The, the person who leads, reasoning at XAI, I forgot his name. Eric. Well, he then. No. Eric. But, yeah, I forgot his name. But, he he he worked on, like, K FAC and stuff, I think. Yeah. Vision dude, Greg? Not Greg, but yeah. But, yeah, it was it was, like, an exciting time to be there. But, yeah, I think I think robotics is a pretty good fit for LNs because, like, the switch ends up being pretty, like, you know, you kinda do similar things. Like, you wanna look at a lot of data. It's, like, kind of, hard to get, like, stuff get hard to get stuff working in robotics world. I think, you know, it kind of builds, like, very gritty people who, like, look at data a lot, that kind of thing. So, yeah. For whatever reason, I think, like, that transfer is, like, yeah. It's happening a lot, and I think it makes a lot of sense. One of my newest highlights so far, had dinner with it's like a small group dinner with Lex Freeman, yesterday. And Lex used to be in robotics. Mhmm. And he was like, my assessment of robotics people robotics people are the best to talk to at NeurIPS Mhmm. Because they're most grounded, he says, Because they don't have a choice. They work with the real world, such as looking data. And then, like, the …

Get the full transcript (9,551 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

You just read a 3-minute summary of a 42-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • by Cursor

    Their Composer model balances intelligence with speed to keep programmers in flow state, avoiding context-switching from slow inference.

Products

  • O1By guest

    by OpenAI

    OpenAI's o1 reasoning model started with approximately 12 core people but expanded to 50-100 contributors for the initial release
  • o3By guest

    by OpenAI

    the development of OpenAI's o1/o3 reasoning models with a 300-person team, achieving IMO/IOI gold medals

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime