Richard Sutton – Father of RL thinks LLMs are a dead end
Episode
66 min
Read time
2 min
Topics
Design & UX, Artificial Intelligence, Software Development
AI-Generated Summary
Key Takeaways
- ✓Experiential Learning vs Imitation: Reinforcement learning enables agents to learn from direct experience through action-sensation-reward cycles, building testable world models with ground truth feedback. LLMs only mimic human responses without predictions about actual world consequences or ability to adjust based on unexpected outcomes, lacking the fundamental learning mechanism animals use.
- ✓Four Component Agent Architecture: Effective AI agents require four distinct parts: a policy determining actions in situations, a value function using TD learning to assess progress, perception constructing state representation, and a transition model predicting world consequences. This transition model learns richly from all sensations, not just rewards, enabling continual adaptation.
- ✓Generalization Through Architecture: Deep learning systems fail at generalization because gradient descent solves training problems without ensuring good transfer to new states. Current systems only generalize well when researchers manually sculpt representations. Catastrophic interference when training on new data demonstrates fundamentally poor generalization, requiring new automated techniques to promote positive transfer across states.
- ✓Digital Intelligence Succession: Four inevitable factors drive AI succession: no unified global governance exists to coordinate development, researchers will eventually solve intelligence, capabilities will exceed human level, and intelligent systems naturally accumulate resources over time. This represents a major universal transition from biological replication to designed entities that understand and modify their own intelligence.
- ✓Cultural Evolution in AI Systems: When AI agents gain sufficient compute, they face a critical choice between self-improvement or spawning copies to learn diverse topics and reintegrate knowledge. The key challenge becomes cybersecurity against corruption, as incorporating external knowledge from spawned copies could introduce hidden goals or viruses that fundamentally alter the original agent's thinking and objectives.
What It Covers
Richard Sutton, Turing Award winner and reinforcement learning pioneer, argues that large language models represent a dead end for AI progress because they lack goals, cannot learn from experience, and fundamentally mimic human behavior rather than understand the world.
Key Questions Answered
- •Experiential Learning vs Imitation: Reinforcement learning enables agents to learn from direct experience through action-sensation-reward cycles, building testable world models with ground truth feedback. LLMs only mimic human responses without predictions about actual world consequences or ability to adjust based on unexpected outcomes, lacking the fundamental learning mechanism animals use.
- •Four Component Agent Architecture: Effective AI agents require four distinct parts: a policy determining actions in situations, a value function using TD learning to assess progress, perception constructing state representation, and a transition model predicting world consequences. This transition model learns richly from all sensations, not just rewards, enabling continual adaptation.
- •Generalization Through Architecture: Deep learning systems fail at generalization because gradient descent solves training problems without ensuring good transfer to new states. Current systems only generalize well when researchers manually sculpt representations. Catastrophic interference when training on new data demonstrates fundamentally poor generalization, requiring new automated techniques to promote positive transfer across states.
- •Digital Intelligence Succession: Four inevitable factors drive AI succession: no unified global governance exists to coordinate development, researchers will eventually solve intelligence, capabilities will exceed human level, and intelligent systems naturally accumulate resources over time. This represents a major universal transition from biological replication to designed entities that understand and modify their own intelligence.
- •Cultural Evolution in AI Systems: When AI agents gain sufficient compute, they face a critical choice between self-improvement or spawning copies to learn diverse topics and reintegrate knowledge. The key challenge becomes cybersecurity against corruption, as incorporating external knowledge from spawned copies could introduce hidden goals or viruses that fundamentally alter the original agent's thinking and objectives.
Notable Moment
Sutton challenges the assumption that children learn through imitation, arguing infants primarily engage in trial and error exploration by waving hands and moving eyes without targets or examples. He contends supervised learning does not occur in nature, with squirrels mastering their environment without formal instruction or imitation processes.
Episode Transcript
Today, I'm chatting with Richard Sutton who is one of the founding fathers of reinforcement learning and inventor of many of the main techniques used there like TD learning and policy gradient methods. And for that, he received this year's Turing award, which, if you don't know, is basically the Nobel Prize for computer science. Richard, congratulations. Thank you, Duwakish. And, thanks for coming on the podcast. It's my pleasure. Okay. So first question, my audience and I are familiar with the LLM way of thinking about AI. Conceptually, what are we missing in terms of thinking about AI from the RL perspective? Well, yes. I think it's really quite a different point of view, and it's it can easily get separated and lose the ability to talk to each other. Mhmm. And, yeah, large language models have become such a big thing. Generative AI in general, a big thing. And our field is subject to the bandwagons and fashions. So we lose we lose track of the, basic basic things. Because I consider reinforcement learning to be basic AI. And what is intelligence? The problem is, is to understand your world. And, reinforcement learning is about understanding your world, whereas large language models are about mimicking people, doing what people say you should do. They're not about figuring out what to do. I guess you would think that to emulate the trillions of tokens in the corpus of Internet text, you would have to build a world model. In fact, these models do seem to have very robust world models, and they they're the best, world models we've made to date in AI. Right? So what what what do you think that that's missing? I would disagree with most of the things you just said. Great. Just to mimic the though what people say is not really to build a model of the world at all, I don't think. You know, you're mimicking things that have a a model of the world, the people. But I don't wanna approach the question in an adversarial way. But but I would I would question the idea that they, they have a world model. So a world model would enable you to predict what would happen. Right. They they have they have the ability to predict what a person would say. They don't have the ability to predict what will happen. What we want, I think, to quote Alan Turing, what we want is a machine that can learn from experience. Right. Where experience is the things that actually happened in your life. You do things. You see what happens. And, that's what you learn from. Yeah. The large language models learn from something else. They learn from here's a situation, and here's what a person did. And, implicitly, the suggestion is you should do what the person did. Right. I guess maybe the the crux, and I'm curious if you disagree with this, is some people will say, okay. So this …
Get the full transcript (11,476 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 63-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1 · 140 min
a16z Podcast
Before Blockchains, There Was State Machine Replication
Jul 13
More from Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31 · 24 min
Deep Questions with Cal Newport
AI Reality Check: Are LLMs a Dead End?
Mar 26
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
The rise and fall of agent civilizations
Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
8 Predictions for the Era of Continual Learning
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Jul 13
Before Blockchains, There Was State Machine Replication
Deep Questions with Cal Newport
Mar 26
AI Reality Check: Are LLMs a Dead End?
20VC (20 Minute VC)
Dec 1
20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing
Hard Fork
Jun 20
Trump Is Selling a Phone + The Start-Up Trying to Automate Every Job + Allison Williams Talks ‘M3GAN 2.0’
Cognitive Revolution
Aug 28
AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?
Explore Related Topics
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime