An audio version of my blog post, Thoughts on AI progress (Dec 2025)
Episode
12 min
Read time
2 min
Topics
Fundraising & VC, Sales & Revenue, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓RL Training Paradox: Labs spend billions having PhDs create training examples for specific tasks like Excel or web browsing, suggesting models cannot learn on-the-job like humans who adapt without rehearsing every software tool beforehand.
- ✓Deployment Reality Check: If models truly matched human capability, they would generate trillions in annual revenue matching global knowledge worker wages, but current figures fall orders of magnitude short, revealing capability gaps despite benchmark improvements.
- ✓Continual Learning Timeline: Achieving human-level on-the-job learning may require five to ten years beyond initial continual learning releases, similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length.
What It Covers
Dwarkesh Patel examines contradictions between short AGI timelines and current reinforcement learning approaches, arguing that models lack human-like on-the-job learning capabilities essential for broad automation.
Key Questions Answered
- •RL Training Paradox: Labs spend billions having PhDs create training examples for specific tasks like Excel or web browsing, suggesting models cannot learn on-the-job like humans who adapt without rehearsing every software tool beforehand.
- •Deployment Reality Check: If models truly matched human capability, they would generate trillions in annual revenue matching global knowledge worker wages, but current figures fall orders of magnitude short, revealing capability gaps despite benchmark improvements.
- •Continual Learning Timeline: Achieving human-level on-the-job learning may require five to ten years beyond initial continual learning releases, similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length.
Notable Moment
A biologist describes identifying macrophages in slides as requiring judgment an AI researcher dismissed as solved, illustrating how real jobs demand context-specific skills that resist pre-baked training pipelines.
Episode Transcript
I'm confused why some people have super short timelines, yet at the same time, are bullish on scaling up reinforcement learning atop LLMs. If we're actually close to a human like learner, then this whole approach of training on verifiable outcomes is doomed. Now, currently, the labs are trying to bake in a bunch of skills into these models through mid training. There's an entire supply chain of companies that are building RL environments, which teach the model how to navigate a web browser or use Excel to build financial models. Now either these models will soon learn on the job in a self directed way, which will make all this free making pointless, or they won't, which means that AGI is not imminent. Humans don't have to go through the special training phase or they need to rehearse every single piece of software that they might ever need to use on the job. Baron Millage made an interesting point about this in a recent blog post he wrote. He writes, quote, when we see frontier models improving at various benchmarks, we should think not just about the increased scale and the clever ML research ideas, but the billions of dollars that are paid to PhDs, MDs and other experts to write questions and provide example answers and reasoning targeting these precise capabilities. You can see this tension most vividly in robotics. In some fundamental science, robotics is an algorithms problem, not a hardware or a data problem. With very little training, a human can learn how to teleoperate current hardware to do useful work. So if you actually had a human like learner, robotics would be, in large part, a solved problem. But the fact that we don't have such a learner makes it necessary to go out into a thousand different homes and practice a million times on how to pick up dishes or fold laundry. Now, one current argument I've heard from the people who think we're gonna have a takeoff within the next five years is that we have to do all this kludgy RL in service of building a superhuman AI researcher. And then the million copies of this automated iliac can go figure out how to solve robust and efficient learning from experience. This just gives me the vibes of that old joke, we're losing money on every sale, but we'll make it up in volume. Somehow, this automated researcher is gonna figure out the algorithm for AGI, which is a problem that humans have been banging their head against for the better half of a century, while not having the basic learning capabilities that children have. I find it super implausible. Besides, even if that's what you believe, it doesn't describe how the labs are approaching reinforcement learning from verifiable reward. You don't need to pre bake in a consultant skill at crafting PowerPoint slides in order to automate Ilia. So clearly, the lab's actions hint at a worldview where these models will …
Get the full transcript (2,407 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 9-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1 · 140 min
Hard Fork
‘Hard Fork’ Live, Part 3: Differing Visions of an A.I. Future
Jun 19
More from Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31 · 24 min
Eye on AI
#335 Sriram Raghavan: Why IBM Is Betting Everything on Small AI Models
Apr 19
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by OpenAI
“similar to how GPT-3 demonstrated in-context learning in 2020 but improvements continue today across comprehension and context length”
other
by Dwarkesh Patel
“An audio version of my blog post, Thoughts on AI progress (Dec 2025)”
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
The rise and fall of agent civilizations
Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
8 Predictions for the Era of Continual Learning
Similar Episodes
Related episodes from other podcasts
Hard Fork
Jun 19
‘Hard Fork’ Live, Part 3: Differing Visions of an A.I. Future
Eye on AI
Apr 19
#335 Sriram Raghavan: Why IBM Is Betting Everything on Small AI Models
Cognitive Revolution
Apr 15
Welcome to AI in the AM: RL for EE, Oversight w/out Nationalization, & the first AI-Run Retail Store
Cognitive Revolution
Apr 1
Success without Dignity? Nathan finds Hope Amidst Chaos, from The Intelligence Horizon Podcast
Machine Learning Street Talk
Dec 24
"I Desperately Want To Live In The Matrix" - Dr. Mike Israetel
Explore Related Topics
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime