Skip to main content
Dwarkesh Podcast

Richard Sutton – Father of RL thinks LLMs are a dead end

66 min episode · 2 min read

Episode

66 min

Read time

2 min

Topics

Design & UX, Artificial Intelligence, Software Development

AI-Generated Summary

Key Takeaways

  • Experiential Learning vs Imitation: Reinforcement learning enables agents to learn from direct experience through action-sensation-reward cycles, building testable world models with ground truth feedback. LLMs only mimic human responses without predictions about actual world consequences or ability to adjust based on unexpected outcomes, lacking the fundamental learning mechanism animals use.
  • Four Component Agent Architecture: Effective AI agents require four distinct parts: a policy determining actions in situations, a value function using TD learning to assess progress, perception constructing state representation, and a transition model predicting world consequences. This transition model learns richly from all sensations, not just rewards, enabling continual adaptation.
  • Generalization Through Architecture: Deep learning systems fail at generalization because gradient descent solves training problems without ensuring good transfer to new states. Current systems only generalize well when researchers manually sculpt representations. Catastrophic interference when training on new data demonstrates fundamentally poor generalization, requiring new automated techniques to promote positive transfer across states.
  • Digital Intelligence Succession: Four inevitable factors drive AI succession: no unified global governance exists to coordinate development, researchers will eventually solve intelligence, capabilities will exceed human level, and intelligent systems naturally accumulate resources over time. This represents a major universal transition from biological replication to designed entities that understand and modify their own intelligence.
  • Cultural Evolution in AI Systems: When AI agents gain sufficient compute, they face a critical choice between self-improvement or spawning copies to learn diverse topics and reintegrate knowledge. The key challenge becomes cybersecurity against corruption, as incorporating external knowledge from spawned copies could introduce hidden goals or viruses that fundamentally alter the original agent's thinking and objectives.

What It Covers

Richard Sutton, Turing Award winner and reinforcement learning pioneer, argues that large language models represent a dead end for AI progress because they lack goals, cannot learn from experience, and fundamentally mimic human behavior rather than understand the world.

Key Questions Answered

  • Experiential Learning vs Imitation: Reinforcement learning enables agents to learn from direct experience through action-sensation-reward cycles, building testable world models with ground truth feedback. LLMs only mimic human responses without predictions about actual world consequences or ability to adjust based on unexpected outcomes, lacking the fundamental learning mechanism animals use.
  • Four Component Agent Architecture: Effective AI agents require four distinct parts: a policy determining actions in situations, a value function using TD learning to assess progress, perception constructing state representation, and a transition model predicting world consequences. This transition model learns richly from all sensations, not just rewards, enabling continual adaptation.
  • Generalization Through Architecture: Deep learning systems fail at generalization because gradient descent solves training problems without ensuring good transfer to new states. Current systems only generalize well when researchers manually sculpt representations. Catastrophic interference when training on new data demonstrates fundamentally poor generalization, requiring new automated techniques to promote positive transfer across states.
  • Digital Intelligence Succession: Four inevitable factors drive AI succession: no unified global governance exists to coordinate development, researchers will eventually solve intelligence, capabilities will exceed human level, and intelligent systems naturally accumulate resources over time. This represents a major universal transition from biological replication to designed entities that understand and modify their own intelligence.
  • Cultural Evolution in AI Systems: When AI agents gain sufficient compute, they face a critical choice between self-improvement or spawning copies to learn diverse topics and reintegrate knowledge. The key challenge becomes cybersecurity against corruption, as incorporating external knowledge from spawned copies could introduce hidden goals or viruses that fundamentally alter the original agent's thinking and objectives.

Notable Moment

Sutton challenges the assumption that children learn through imitation, arguing infants primarily engage in trial and error exploration by waving hands and moving eyes without targets or examples. He contends supervised learning does not occur in nature, with squirrels mastering their environment without formal instruction or imitation processes.

Know someone who'd find this useful?

Episode Transcript

Today, I'm chatting with Richard Sutton who is one of the founding fathers of reinforcement learning and inventor of many of the main techniques used there like TD learning and policy gradient methods. And for that, he received this year's Turing award, which, if you don't know, is basically the Nobel Prize for computer science. Richard, congratulations. Thank you, Duwakish. And, thanks for coming on the podcast. It's my pleasure. Okay. So first question, my audience and I are familiar with the LLM way of thinking about AI. Conceptually, what are we missing in terms of thinking about AI from the RL perspective? Well, yes. I think it's really quite a different point of view, and it's it can easily get separated and lose the ability to talk to each other. Mhmm. And, yeah, large language models have become such a big thing. Generative AI in general, a big thing. And our field is subject to the bandwagons and fashions. So we lose we lose track of the, basic basic things. Because I consider reinforcement learning to be basic AI. And what is intelligence? The problem is, is to understand your world. And, reinforcement learning is about understanding your world, whereas large language models are about mimicking people, doing what people say you should do. They're not about figuring out what to do. I guess you would think that to emulate the trillions of tokens in the corpus of Internet text, you would have to build a world model. In fact, these models do seem to have very robust world models, and they they're the best, world models we've made to date in AI. Right? So what what what do you think that that's missing? I would disagree with most of the things you just said. Great. Just to mimic the though what people say is not really to build a model of the world at all, I don't think. You know, you're mimicking things that have a a model of the world, the people. But I don't wanna approach the question in an adversarial way. But but I would I would question the idea that they, they have a world model. So a world model would enable you to predict what would happen. Right. They they have they have the ability to predict what a person would say. They don't have the ability to predict what will happen. What we want, I think, to quote Alan Turing, what we want is a machine that can learn from experience. Right. Where experience is the things that actually happened in your life. You do things. You see what happens. And, that's what you learn from. Yeah. The large language models learn from something else. They learn from here's a situation, and here's what a person did. And, implicitly, the suggestion is you should do what the person did. Right. I guess maybe the the crux, and I'm curious if you disagree with this, is some people will say, okay. So this …

Get the full transcript (11,476 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Dwarkesh Podcast transcripts →

You just read a 3-minute summary of a 63-minute episode.

Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Dwarkesh Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Dwarkesh Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime