#299 Jacob Buckman: Why the Future of AI Won't Be Built on Transformers
Episode
57 min
Read time
2 min
Topics
Startups, Artificial Intelligence, Software Development
AI-Generated Summary
Key Takeaways
- ✓Power Retention Architecture: Combines recurrent neural networks with attention mechanisms through state space models, allowing independent adjustment of state size from parameter count. This enables linear scaling costs instead of quadratic growth as context windows expand.
- ✓Metamorphosis Retraining Process: Existing transformer models like LLAMA can convert to Power Retention in six hours using dozens of GPUs by swapping attention calls for power retention, preserving original performance while gaining linear-cost inference and unlimited context capabilities.
- ✓Context vs Weight Updates: Future AI systems should inject new knowledge through context state updates rather than weight fine-tuning. This eliminates catastrophic forgetting issues since context-based learning mirrors human experience accumulation rather than evolutionary weight changes through gradient descent.
- ✓Butler vs Consultant Dynamic: Current transformers force chat resets due to expensive state growth, creating consultant-like interactions. Power Retention enables persistent state across all user interactions, creating butler-like AI that accumulates complete user history and preferences for better responses.
What It Covers
Jacob Buckman explains Power Retention, a new AI architecture that solves transformer scaling limitations through linear-cost context windows, enabling models to process unlimited context without quadratic compute costs or performance degradation.
Key Questions Answered
- •Power Retention Architecture: Combines recurrent neural networks with attention mechanisms through state space models, allowing independent adjustment of state size from parameter count. This enables linear scaling costs instead of quadratic growth as context windows expand.
- •Metamorphosis Retraining Process: Existing transformer models like LLAMA can convert to Power Retention in six hours using dozens of GPUs by swapping attention calls for power retention, preserving original performance while gaining linear-cost inference and unlimited context capabilities.
- •Context vs Weight Updates: Future AI systems should inject new knowledge through context state updates rather than weight fine-tuning. This eliminates catastrophic forgetting issues since context-based learning mirrors human experience accumulation rather than evolutionary weight changes through gradient descent.
- •Butler vs Consultant Dynamic: Current transformers force chat resets due to expensive state growth, creating consultant-like interactions. Power Retention enables persistent state across all user interactions, creating butler-like AI that accumulates complete user history and preferences for better responses.
Notable Moment
Buckman reveals that advertised long-context models use sparse or windowed attention rather than true transformers, processing only small context subsets. This industry-wide practice creates performance degradation that users mistake for inherent limitations rather than architectural compromises.
No transcript yet — request it by email, free
We'll transcribe this episode on request and email you the full transcript and AI summary — usually within a day. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 54-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
The Reason 30 Years of Cybersecurity Has Failed - and What Actually Fixes It | Trent Telford, Qanapi
Sep 10 · 55 min
The TWIML AI Podcast
Recurrence and Attention for Long-Context Transformers with Jacob Buckman - #750
Oct 7
More from Eye on AI
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Sep 8 · 54 min
Latent Space
Building Snipd: The AI Podcast App for Learning
Mar 14
More from Eye on AI
We summarize every new episode. Want them in your inbox?
The Reason 30 Years of Cybersecurity Has Failed - and What Actually Fixes It | Trent Telford, Qanapi
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
In 5 to 10 Years, Using Weapons Without AI Will Be Considered Unethical | Yaroslav Azhnyuk, The Fourth Law
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
Similar Episodes
Related episodes from other podcasts
The TWIML AI Podcast
Oct 7
Recurrence and Attention for Long-Context Transformers with Jacob Buckman - #750
Latent Space
Mar 14
Building Snipd: The AI Podcast App for Learning
Deep Questions with Cal Newport
Sep 10
How Worrisome is GPT-6’s “Stealth Thinking”? | Tech Decoded
a16z Podcast
Sep 7
Can Open Source Keep AI Power From Concentrating?
Odd Lots
Sep 4
Why Laser Beams Are the Hottest New Tech in Defense
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime