Andrej Karpathy — AGI is still a decade away
Episode
145 min
Read time
2 min
Topics
Investing, Fundraising & VC, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Reinforcement Learning Limitations: Current RL assigns credit uniformly across entire solution trajectories based on final outcomes, upweighting every token including wrong paths if the answer is correct. This creates high-variance estimators that waste computational work by sucking supervision through a straw from single reward signals across minutes of rollout.
- ✓Model Collapse Problem: LLMs generate outputs from collapsed distributions with low entropy, producing only three jokes when prompted repeatedly. Training on synthetic self-generated data causes further collapse because models cannot maintain the diversity and entropy needed for robust learning, similar to how humans become more rigid with age and repetitive thought patterns.
- ✓Cognitive Core Size: Optimal intelligence cores could operate at one billion parameters or less, compared to current trillion-parameter models. Most model size stores memorized internet facts rather than cognitive algorithms. Future systems should separate knowledge retrieval from reasoning capabilities, with models knowing what they don't know and looking up information as needed.
- ✓Pre-training Data Quality: Average pre-training examples are garbage, not Wall Street Journal articles but stock tickers and internet slop. Compression ratios show LLAMA 3 stores 0.07 bits per token across 15 trillion tokens, while context windows store 320 kilobytes per token, a 35 million fold difference explaining why in-context learning feels more intelligent than pre-trained knowledge.
- ✓Automation Progression Path: Call center work will automate before radiology because tasks are repetitive, ten-minute interactions with closed databases rather than messy multi-surface jobs. Expect autonomy sliders where AIs handle 80% of volume while delegating 20% to human supervisors managing teams of five AI agents, not instant full replacement across knowledge work.
What It Covers
Andrej Karpathy explains why AGI development will take a decade, not a year, discussing current limitations in continual learning, reinforcement learning's fundamental flaws, model collapse issues, and why coding automation succeeds while other knowledge work automation struggles despite similar text-based interfaces.
Key Questions Answered
- •Reinforcement Learning Limitations: Current RL assigns credit uniformly across entire solution trajectories based on final outcomes, upweighting every token including wrong paths if the answer is correct. This creates high-variance estimators that waste computational work by sucking supervision through a straw from single reward signals across minutes of rollout.
- •Model Collapse Problem: LLMs generate outputs from collapsed distributions with low entropy, producing only three jokes when prompted repeatedly. Training on synthetic self-generated data causes further collapse because models cannot maintain the diversity and entropy needed for robust learning, similar to how humans become more rigid with age and repetitive thought patterns.
- •Cognitive Core Size: Optimal intelligence cores could operate at one billion parameters or less, compared to current trillion-parameter models. Most model size stores memorized internet facts rather than cognitive algorithms. Future systems should separate knowledge retrieval from reasoning capabilities, with models knowing what they don't know and looking up information as needed.
- •Pre-training Data Quality: Average pre-training examples are garbage, not Wall Street Journal articles but stock tickers and internet slop. Compression ratios show LLAMA 3 stores 0.07 bits per token across 15 trillion tokens, while context windows store 320 kilobytes per token, a 35 million fold difference explaining why in-context learning feels more intelligent than pre-trained knowledge.
- •Automation Progression Path: Call center work will automate before radiology because tasks are repetitive, ten-minute interactions with closed databases rather than messy multi-surface jobs. Expect autonomy sliders where AIs handle 80% of volume while delegating 20% to human supervisors managing teams of five AI agents, not instant full replacement across knowledge work.
Notable Moment
Karpathy reveals that during Nanochat development, coding assistants constantly misunderstood his custom implementations, trying to force deprecated APIs and production boilerplate. The models kept assuming he used standard PyTorch containers when he wrote custom gradient synchronization, demonstrating how current AI struggles with novel code patterns outside typical internet examples despite appearing capable.
Episode Transcript
Today, I'm speaking with Andre Karpathy. Andre, why do you say that this will be the decade of agents and not the year of agents? Mhmm. Well, first of all, thank you for, having me here. I'm, excited to be here. So the quote that you've just mentioned, it's the decade of agents. That's actually a reaction to an existing preexisting quote, I should say, where I think a lot of some of the labs I'm not actually sure who said this, but they were alluding to this being the year of agents, with respect to LLMs and, how they were gonna evolve. And I think, I was triggered by that because I feel like there's some over predictions going on in the industry. And, in my mind, this is really a lot more accurately described as the decade of agents. Yeah. And we have some very early agents that are actually, like, extremely impressive and that I use daily, you know, Claude and Codex and so on. But I still feel like there's, so much work to be done. And so I think my, like, my reaction is, like, we'll be working with these things for a decade. They're gonna get better, and, it's gonna be wonderful. But I think I was just reacting to the timelines, I suppose, of the of the, implication. Man, what do you think will take a decade to accomplish? What are the bottlenecks? Well, actually make it work. So in my mind, I mean, when you're talking about an agent, I guess, or what the labs have in mind and what maybe I have in mind as well is it's, you should think of it almost like an employee or like an intern that you would hire to work with you. So, for example, you work with some employees here. When would you prefer to have an agent like Cloud or Codex, do that work? Like, currently, of course, they can't. What would it take for them to be able to do that? Why don't you do it today? And the reason you don't do it today is because they just don't work. So, like, they don't have enough intelligence, they're not multimodal enough, they can't do computer use and all this kind of stuff, and they don't do a lot of the things that you've alluded to earlier. You know, they don't have continual learning. You can't just tell them something and they'll remember it. And they're just cognitively lacking, and it's just not working. And I just think that it will take about a decade to work through all of those issues. Interesting. So, as a professional podcaster and a a viewer of AI from afar, It's easy to identify for me, like, oh, here's what's lacking. Continual learning is lacking or multimodality is lacking. But I don't really have a good, way of trying to put a timeline on it. Mhmm. Like, if somebody's like, how long will continual learning …
Get the full transcript (32,086 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 142-minute episode.
Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Dwarkesh Podcast
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
Sep 1 · 140 min
Beyond Biotech
This nonprofit is building the ecosystem to cure epidermolysis bullosa
Jul 24
More from Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31 · 24 min
Software Engineering Daily
Agentic DevOps at AWS
Jul 16
More from Dwarkesh Podcast
We summarize every new episode. Want them in your inbox?
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
The rise and fall of agent civilizations
Dylan Patel – Anthropic & OpenAI will have most of the world’s compute by 2028
Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032
8 Predictions for the Era of Continual Learning
Similar Episodes
Related episodes from other podcasts
Beyond Biotech
Jul 24
This nonprofit is building the ecosystem to cure epidermolysis bullosa
Software Engineering Daily
Jul 16
Agentic DevOps at AWS
The Jordan Harbinger Show
Jun 11
1342: Jacob Ward | How AI Turns Convenience Into Control
Masters of Scale
Jun 2
The race no one can win: AI’s anti-human crisis, with Aza Raskin
The AI Breakdown
May 21
Anthropic Just Reset AI Expectations
Explore Related Topics
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Dwarkesh Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime