Skip to main content
Latent Space

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

Read time

2 min

Topics

Remote Work, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • RL Generalization Limits: Reinforcement learning applied to language models excels within training distribution but generalizes poorly beyond it. The strategy requires bringing economically useful tasks into distribution rather than expecting broad generalization, fundamentally changing how products must be designed around model capabilities.
  • Reasoning Team Scale: OpenAI's o1 model started with roughly fifty to one hundred contributors, expanding to three hundred people for o3 as safety, evaluation, and product teams joined. The progression from prototype to product requires exponentially more organizational resources than initial research breakthroughs suggest.
  • Continuous Learning Advantage: Cursor ships policy updates every two hours for tab autocomplete by co-locating product and ML teams. This rapid iteration cycle proves impossible at larger organizations where product and research groups operate separately, creating competitive advantage through organizational structure rather than pure technical capability.
  • Context Over Code: Automating programming jobs requires capturing years of accumulated knowledge—hyperparameter sweep results, architectural decisions, team conversations—not just writing code. Products must ingest this full context (Slack messages, Datadog traces, design documents) to enable models to replicate senior engineer decision-making processes effectively.

What It Covers

Ashvin Nair discusses his transition from OpenAI's reasoning team to Cursor, covering the development of o1/o3 models, reinforcement learning's role in achieving IMO gold performance, and Cursor's approach to co-designing products with models.

Key Questions Answered

  • RL Generalization Limits: Reinforcement learning applied to language models excels within training distribution but generalizes poorly beyond it. The strategy requires bringing economically useful tasks into distribution rather than expecting broad generalization, fundamentally changing how products must be designed around model capabilities.
  • Reasoning Team Scale: OpenAI's o1 model started with roughly fifty to one hundred contributors, expanding to three hundred people for o3 as safety, evaluation, and product teams joined. The progression from prototype to product requires exponentially more organizational resources than initial research breakthroughs suggest.
  • Continuous Learning Advantage: Cursor ships policy updates every two hours for tab autocomplete by co-locating product and ML teams. This rapid iteration cycle proves impossible at larger organizations where product and research groups operate separately, creating competitive advantage through organizational structure rather than pure technical capability.
  • Context Over Code: Automating programming jobs requires capturing years of accumulated knowledge—hyperparameter sweep results, architectural decisions, team conversations—not just writing code. Products must ingest this full context (Slack messages, Datadog traces, design documents) to enable models to replicate senior engineer decision-making processes effectively.

Notable Moment

Nair attended a forecasting conference where AI researchers predicted twenty percent math exam performance by 2027, while OpenAI already had internal models exceeding those benchmarks. The same forecasters simultaneously predicted Dyson spheres by 2035, revealing systematic miscalibration in short-term pessimism and long-term optimism.

Know someone who'd find this useful?

Episode Transcript

Okay. We're here at NeurIPS. We're recording a special land space coverage of the folks at NeurIPS. And we're here with Astrid from CURSA. Welcome. Hi. Yeah. Thanks for having me. So I guess the like Astrid from CURSA is like a new identity. I didn't even know if I should say that. Because you only joined CROSSA for three months. Mhmm. Before that, you're opening. I worked in 01/2003. Mhmm. Before that, Berkeley, PhD in RL. Just but focus on robotics. Robotics. Yeah. Is it weird searching for robotics in language models? Okay. This is kind of interesting because a lot of people have been kind of doing this, Are you OpenAI? Yeah. You got a robotics. So actually, I I actually was, at OpenAI in 2017 also, where I was at on robotics. Yeah. I was interning, like, right before my PhD where I worked on robotics there. 2017. Is that, like, Jovan and the Aztum band over there? He was all famously opening as first intern. Oh, really? Okay. Then he might have been before. But, yeah, there was like a like 15 interns. It was a very different company. It was just like, robotics, DOTA, and like 15 interns that summer all having like pretty exciting individual products. Like, yeah, that set of interns if you look over there now is kinda cool. Yeah. But yeah, Anyone from that class that, like, you would shout out? Like, there's just, like, a lot of cool papers that came out, like, Leo Pinto, and I was at, NYU. Yeah. And Faiman. The, the person who leads, reasoning at XAI, I forgot his name. Eric. Well, he then No. Eric. Yeah. I forgot his name. But, he he he worked on, like, K FAC and stuff. I think, yeah. The vision dude, Greg? Not Greg, but yeah. But, yeah, it was it was, like, an exciting time to be there. But, yeah, I think I think robotics is a pretty good fit for LMs because, like, the switch ends up being pretty, like, you know, you kinda do similar things. Like, you wanna look at a lot of data. It's, like, kind of, hard to get it's, like, stuff get hard to get stuff working in robotics world. I think, you know, it kinda builds, like, very greedy people who, like, look at data a lot, that kind of thing. So, yeah. For whatever reason, I think, like, that transfer is, like, yeah. It's happening a lot, and I think it makes a lot of sense. One of my newest highlights so far, had dinner with it's like a small group dinner with Lex Freeman, yes, yesterday. And Lex used to be in robotics. Mhmm. And he was like, my assessment of robotics people robotics people are the best to talk to Ed and NeurIPS Mhmm. Because they're most grounded, he says, Because they don't have a choice. They work with the real world. So they look at data. …

Get the full transcript (9,563 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Products

  • by OpenAI

    OpenAI's o1 model started with roughly fifty to one hundred contributors, expanding to three hundred people for o3 as safety, evaluation, and product teams joined.
  • Ashvin Nair discusses his transition from OpenAI's reasoning team to Cursor, covering the development of o1/o3 models, reinforcement learning's role in achieving IMO gold performance, and Cursor's approach to co-designing products with models.
  • by OpenAI

    Ashvin Nair discusses his transition from OpenAI's reasoning team to Cursor, covering the development of o1/o3 models, reinforcement learning's role in achieving IMO gold performance, and Cursor's approach to co-designing products with models.

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime