Skip to main content
Dwarkesh Podcast

An audio version of my blog post, Thoughts on AI progress (Dec 2025)

12 min episode · 2 min read

Episode

12 min

Read time

2 min

Topics

Fundraising & VC, Sales & Revenue, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • RL Training Paradox: Labs spend billions having PhDs create training examples for specific tasks like Excel or web browsing, suggesting models cannot learn on-the-job like humans who adapt without rehearsing every software tool beforehand.
  • Deployment Reality Check: If models truly matched human capability, they would generate trillions in annual revenue matching global knowledge worker wages, but current figures fall orders of magnitude short, revealing capability gaps despite benchmark improvements.
  • Continual Learning Timeline: Achieving human-level on-the-job learning may require five to ten years beyond initial continual learning releases, similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length.

What It Covers

Dwarkesh Patel examines contradictions between short AGI timelines and current reinforcement learning approaches, arguing that models lack human-like on-the-job learning capabilities essential for broad automation.

Key Questions Answered

  • RL Training Paradox: Labs spend billions having PhDs create training examples for specific tasks like Excel or web browsing, suggesting models cannot learn on-the-job like humans who adapt without rehearsing every software tool beforehand.
  • Deployment Reality Check: If models truly matched human capability, they would generate trillions in annual revenue matching global knowledge worker wages, but current figures fall orders of magnitude short, revealing capability gaps despite benchmark improvements.
  • Continual Learning Timeline: Achieving human-level on-the-job learning may require five to ten years beyond initial continual learning releases, similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length.

Notable Moment

A biologist describes identifying macrophages in slides as requiring judgment an AI researcher dismissed as solved, illustrating how real jobs demand context-specific skills that resist pre-baked training pipelines.

Know someone who'd find this useful?

Episode Transcript

I'm confused why some people have super short timelines, yet at the same time, are bullish on scaling up reinforcement learning atop LLMs. If we're actually close to a human like learner, then this whole approach of training on verifiable outcomes is doomed. Now, currently, the labs are trying to bake in a bunch of skills into these models through mid training. There's an entire supply chain of companies that are building RL environments, which teach the model how to navigate a web browser or use Excel to build financial models. Now either these models will soon learn on the job in a self directed way, which will make all this free making pointless, or they won't, which means that AGI is not imminent. Humans don't have to go through the special training phase or they need to rehearse every single piece of software that they might ever need to use on the job. Baron Millage made an interesting point about this in a recent blog post he wrote. He writes, quote, when we see frontier models improving at various benchmarks, we should think not just about the increased scale and the clever ML research ideas, but the billions of dollars that are paid to PhDs, MDs and other experts to write questions and provide example answers and reasoning targeting these precise capabilities. You can see this tension most vividly in robotics. In some fundamental science, robotics is an algorithms problem, not a hardware or a data problem. With very little training, a human can learn how to teleoperate current hardware to do useful work. So if you actually had a human like learner, robotics would be, in large part, a solved problem. But the fact that we don't have such a learner makes it necessary to go out into a thousand different homes and practice a million times on how to pick up dishes or fold laundry. Now, one current argument I've heard from the people who think we're gonna have a takeoff within the next five years is that we have to do all this kludgy RL in service of building a superhuman AI researcher. And then the million copies of this automated iliac can go figure out how to solve robust and efficient learning from experience. This just gives me the vibes of that old joke, we're losing money on every sale, but we'll make it up in volume. Somehow, this automated researcher is gonna figure out the algorithm for AGI, which is a problem that humans have been banging their head against for the better half of a century, while not having the basic learning capabilities that children have. I find it super implausible. Besides, even if that's what you believe, it doesn't describe how the labs are approaching reinforcement learning from verifiable reward. You don't need to pre bake in a consultant skill at crafting PowerPoint slides in order to automate Ilia. So clearly, the lab's actions hint at a worldview where these models will …

Get the full transcript (2,407 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Dwarkesh Podcast transcripts →

You just read a 3-minute summary of a 9-minute episode.

Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by OpenAI

    similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length

other

More from Dwarkesh Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Dwarkesh Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime