The next big breakthrough will be AIs learning on the job
Episode
19 min
Read time
2 min
Topics
Startups, Artificial Intelligence, Software Development
AI-Generated Summary
Key Takeaways
- ✓RLVR Generalization Limits: Reinforcement learning on verifiable, containerized environments works for coding and math but cannot train skills requiring real-world feedback loops — like winning court cases or building a business — because rollouts take months and cannot be parallelized or replayed from identical starting states.
- ✓Grindability Requirement: A domain being verifiable is insufficient for rapid AI progress; it must also be "grindable" — runnable as thousands of parallel, deterministic, replayable simulations. Computer use lags behind coding precisely because cloning real websites like Amazon at scale remains prohibitively labor-intensive today.
- ✓On-Policy Self-Distillation (OPSD): Rather than sparse RL rewards or full transcript replay via supervised fine-tuning, OPSD trains the base model to match per-token predictions of a context-rich "veteran" model, producing targeted weight updates that consolidate session learning without overwriting existing knowledge — superior density versus naive RL.
- ✓"Dreaming" as a Fourth Scaling Axis: Beyond pretraining, RL, and inference-time compute, models could spend compute generating their own RL environments simulating a specific user's real-world context, then train against them before deployment — analogous to EfficientZero's internal simulation strategy but applied to open-ended professional tasks.
What It Covers
Dwarkesh Patel argues that AI's next capability leap requires on-the-job continual learning, explaining why current RLVR training hits hard limits and how techniques like on-policy self-distillation and "dreaming" could unlock genuine AGI-level generalization by 2027–2028.
Key Questions Answered
- •RLVR Generalization Limits: Reinforcement learning on verifiable, containerized environments works for coding and math but cannot train skills requiring real-world feedback loops — like winning court cases or building a business — because rollouts take months and cannot be parallelized or replayed from identical starting states.
- •Grindability Requirement: A domain being verifiable is insufficient for rapid AI progress; it must also be "grindable" — runnable as thousands of parallel, deterministic, replayable simulations. Computer use lags behind coding precisely because cloning real websites like Amazon at scale remains prohibitively labor-intensive today.
- •On-Policy Self-Distillation (OPSD): Rather than sparse RL rewards or full transcript replay via supervised fine-tuning, OPSD trains the base model to match per-token predictions of a context-rich "veteran" model, producing targeted weight updates that consolidate session learning without overwriting existing knowledge — superior density versus naive RL.
- •"Dreaming" as a Fourth Scaling Axis: Beyond pretraining, RL, and inference-time compute, models could spend compute generating their own RL environments simulating a specific user's real-world context, then train against them before deployment — analogous to EfficientZero's internal simulation strategy but applied to open-ended professional tasks.
Notable Moment
Roughly 30–50% of a lab's compute goes to inference, yet none of that compute currently improves the model — meaning the most valuable real-world learning signal is being generated and then completely discarded every single session.
Episode Transcript
So here's the big research bet that all the labs are making. They think that if we train AIs to accomplish millions of verifiable tasks across thousands of diverse RL environments, then we will have basically built AGI. Because this kind of training will have created a kind of problem solving agent. The kind of thing they can make progress on open ended tasks for weeks on end in the face of errors and mistakes and ambiguity. And the people who are optimistic about this vision will say that all these things that we talk about as the fundamental deficits in the current training paradigm, for example, the data inefficiency of these models or the fact that they lack continual learning, These things can just be steamrolled if you just scale training more. And the same way that all the fundamental research problems in natural language processing collapsed when you just threw enough compute into LLMs. So in the previous essay, I talked about how these models are one one millionth the sample efficient as humans. And the people who are in favor of the current training paradigm will say, look, that might be true, but this is only true during training. And training is this one time cost that is amortized across billions of sessions that a model will experience. And what really matters is how smart and general and sample efficient the model is during a session. And this has clearly been improving as we've been doing more RL training. AI agents are able to solve more and more ambitious problems over longer and longer time spans. Anybody who has used these models for coding knows that. Similarly, people would say, look, continue learning, this capability I keep harping about where the model's weights get updated based on what it's learning from deployment may simply not be necessary. Because if in context learning gets so good across longer and longer time horizons, that you don't need to distill back everything the model is learning on the job into the weights. People often say that their employees are not net productive until six months or more of them working on the job. So clearly, online learning is necessary for competence. But But what if you could just fit those six months into the context window? There's been tons of architectural innovations that dramatically increase the amount of information or the amount of context that a transformer can store. And why not think with a couple more years of progress, you might have what feels like infinitely large context windows. Okay. So before we discuss this research a bit further, I wanna step back and I wanna ask a completely tangential question, which I find actually very interesting and confusing about the nature of current AI progress. Why has progress on computer use been so much slower than other domains? Computer use is so clearly verifiable. You could ask a question like, did the desired Etsy item I ordered get …
Get the full transcript (3,886 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 16-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
8 Predictions for the Era of Continual Learning
Aug 7 · 8 min
Machine Learning Street Talk
Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)
May 21
More from Dwarkesh Podcast
Why smarter AI models could drive up compute prices 10x
Aug 3 · 11 min
Cognitive Revolution
It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast
Apr 11
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“SPONSORS [Mercury, https://mercury.com]”
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
8 Predictions for the Era of Continual Learning
Why smarter AI models could drive up compute prices 10x
Adam Brown – A deep but accessible introduction to general relativity
Grant Sanderson – AI and the future of math
The data black hole at the center of AI
Similar Episodes
Related episodes from other podcasts
Machine Learning Street Talk
May 21
Intelligence is collective, not artificial — Prof. Michael I. Jordan (UC Berkeley / Inria)
Cognitive Revolution
Apr 11
It's Crunch Time: Ajeya Cotra on RSI & AI-Powered AI Safety Work, from the 80,000 Hours Podcast
Deep Questions with Cal Newport
Aug 6
Does OpenAI’s Astra Mean AGI Has Arrived? | AI Reality Check
TED Radio Hour
Jul 24
Rethinking addiction in the age of Ozempic
The AI Breakdown
Jun 19
Your Company Doesn’t Need an AI Strategy
Explore Related Topics
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime