Skip to main content
Dwarkesh Podcast

Some thoughts on the Sutton interview

11 min episode · 2 min read
·
Some Thoughts

Episode

11 min

Read time

2 min

Topics

Productivity, Startups, Software Development

AI-Generated Summary

Key Takeaways

  • Compute efficiency critique: LLMs spend most compute during deployment without learning anything, only learning during training on tens of thousands of years of human experience data inefficiently.
  • Imitation learning as foundation: Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.
  • Continual learning gap: Current LLMs learn approximately one bit per episode of tens of thousands of tokens during RL, while animals extract maximum signal continuously from environmental observations.

What It Covers

Dwarkesh reflects on Richard Sutton's perspective that current LLMs waste compute during deployment without learning, requiring new architectures for continual learning and true intelligence.

Key Questions Answered

  • Compute efficiency critique: LLMs spend most compute during deployment without learning anything, only learning during training on tens of thousands of years of human experience data inefficiently.
  • Imitation learning as foundation: Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.
  • Continual learning gap: Current LLMs learn approximately one bit per episode of tens of thousands of tokens during RL, while animals extract maximum signal continuously from environmental observations.

Notable Moment

Dwarkesh compares pretraining data to fossil fuels as non-renewable but essential intermediaries, arguing civilization needed them to reach solar panels despite not being the final solution.

Know someone who'd find this useful?

Episode Transcript

Boy, do you guys have a lot of thoughts about this on an interview? I've been thinking about it myself, and I think I have a much better understanding now of Sutton's perspective than I did during the interview itself. So I wanted to reflect on how I understand his worldview now. And, Richard, apologies if there are still any errors or misunderstandings. It's been very productive to learn from your thoughts. Okay. So here's my understanding of the steel man of Richard's position. Obviously, he wrote the same essay, The Bitter Lesson. And what is this essay about? Well, it's not saying that you just wanna throw away as much compute as you possibly can. The bitter lesson says that you want to come up with techniques which most effectively and scalably leverage compute. Most of the compute that's spent on an LLM is used in running it during deployment, and yet it's not learning anything during this entire period. It's only learning during this special phase that we call training. And so this is obviously not an effective use of compute. And what's even worse is that this training period by itself is highly inefficient because these models are usually trained on the equivalent of tens of thousands of years of human experience. And what's more, during this training phase, all of their learning is coming straight from human data. Now this is an obvious point in the case of free training data, but it's even kind of true for the RLVR that we do with these LLMs. These RL environments are human furnished playgrounds to teach LLMs the specific skills that we have prescribed for them. The agent is in no substantial way learning from organic and self directed engagement with the world. Having to learn only from human data, which is an inelastic and hard to skill resource, is not a scalable way to use compute. Furthermore, what these LLMs learn from training is not a true world model, which would tell you how the environment changes in response to different actions that you take. Rather, they're building a model of what a human would say next. And this leads them to rely on human derived concepts. A way to think about this would be, suppose you trained an LLM on all the data up to the year 1900. That LLM probably wouldn't be able to come up with relativity from scratch. And maybe here's a more fundamental reason to think this whole paradigm will eventually be superseded. LLMs aren't capable of learning on the job, so we'll need some new architecture to enable this kind of continual learning. And once we do have this architecture, we won't need a special trading face. The agents will just be able to learn on the fly like all humans and, in fact, like all animals are able to do. And this new paradigm will render our current approach with LLMs and their special training phase that's super sample …

Get the full transcript (2,116 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Dwarkesh Podcast transcripts →

You just read a 3-minute summary of a 8-minute episode.

Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

other

  • by DeepMind

    Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.
  • by DeepMind

    Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.

More from Dwarkesh Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Dwarkesh Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime