Some thoughts on the Sutton interview
Episode
11 min
Read time
2 min
Topics
Productivity, Startups, Software Development
AI-Generated Summary
Key Takeaways
- ✓Compute efficiency critique: LLMs spend most compute during deployment without learning anything, only learning during training on tens of thousands of years of human experience data inefficiently.
- ✓Imitation learning as foundation: Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.
- ✓Continual learning gap: Current LLMs learn approximately one bit per episode of tens of thousands of tokens during RL, while animals extract maximum signal continuously from environmental observations.
What It Covers
Dwarkesh reflects on Richard Sutton's perspective that current LLMs waste compute during deployment without learning, requiring new architectures for continual learning and true intelligence.
Key Questions Answered
- •Compute efficiency critique: LLMs spend most compute during deployment without learning anything, only learning during training on tens of thousands of years of human experience data inefficiently.
- •Imitation learning as foundation: Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.
- •Continual learning gap: Current LLMs learn approximately one bit per episode of tens of thousands of tokens during RL, while animals extract maximum signal continuously from environmental observations.
Notable Moment
Dwarkesh compares pretraining data to fossil fuels as non-renewable but essential intermediaries, arguing civilization needed them to reach solar panels despite not being the final solution.
Episode Transcript
Boy, do you guys have a lot of thoughts about this on an interview? I've been thinking about it myself, and I think I have a much better understanding now of Sutton's perspective than I did during the interview itself. So I wanted to reflect on how I understand his worldview now. And, Richard, apologies if there are still any errors or misunderstandings. It's been very productive to learn from your thoughts. Okay. So here's my understanding of the steel man of Richard's position. Obviously, he wrote the same essay, The Bitter Lesson. And what is this essay about? Well, it's not saying that you just wanna throw away as much compute as you possibly can. The bitter lesson says that you want to come up with techniques which most effectively and scalably leverage compute. Most of the compute that's spent on an LLM is used in running it during deployment, and yet it's not learning anything during this entire period. It's only learning during this special phase that we call training. And so this is obviously not an effective use of compute. And what's even worse is that this training period by itself is highly inefficient because these models are usually trained on the equivalent of tens of thousands of years of human experience. And what's more, during this training phase, all of their learning is coming straight from human data. Now this is an obvious point in the case of free training data, but it's even kind of true for the RLVR that we do with these LLMs. These RL environments are human furnished playgrounds to teach LLMs the specific skills that we have prescribed for them. The agent is in no substantial way learning from organic and self directed engagement with the world. Having to learn only from human data, which is an inelastic and hard to skill resource, is not a scalable way to use compute. Furthermore, what these LLMs learn from training is not a true world model, which would tell you how the environment changes in response to different actions that you take. Rather, they're building a model of what a human would say next. And this leads them to rely on human derived concepts. A way to think about this would be, suppose you trained an LLM on all the data up to the year 1900. That LLM probably wouldn't be able to come up with relativity from scratch. And maybe here's a more fundamental reason to think this whole paradigm will eventually be superseded. LLMs aren't capable of learning on the job, so we'll need some new architecture to enable this kind of continual learning. And once we do have this architecture, we won't need a special trading face. The agents will just be able to learn on the fly like all humans and, in fact, like all animals are able to do. And this new paradigm will render our current approach with LLMs and their special training phase that's super sample …
Get the full transcript (2,116 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 8-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1 · 140 min
Invest Like the Best with Patrick O'Shaughnessy
Dan Loeb - Lessons from 30 Years of Investing - [Invest Like the Best, EP.475]
May 28
More from Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31 · 24 min
Cognitive Revolution
All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
May 24
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
other
by DeepMind
“Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.”
by DeepMind
“Pretrained LLMs serve as essential priors for reinforcement learning, similar to how AlphaGo used human games before AlphaZero bootstrapped from scratch to superhuman performance.”
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
The rise and fall of agent civilizations
Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
8 Predictions for the Era of Continual Learning
Similar Episodes
Related episodes from other podcasts
Invest Like the Best with Patrick O'Shaughnessy
May 28
Dan Loeb - Lessons from 30 Years of Investing - [Invest Like the Best, EP.475]
Cognitive Revolution
May 24
All Compute Is Food: Palisade's Jeffrey Ladish on AI Shutdown Resistance, Self-Replication & Ecology
Invest Like the Best with Patrick O'Shaughnessy
Apr 28
Paul Tudor Jones - Lessons From 50 Years in Markets - [Invest Like the Best, EP.469]
a16z Podcast
Apr 3
Marc Andreessen on AI Winters and Agent Breakthroughs
a16z Podcast
Mar 17
What's Missing Between LLMs and AGI - Vishal Misra & Martin Casado
Explore Related Topics
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime