Recurrence and Attention for Long-Context Transformers with Jacob Buckman - #750
Episode
57 min
Read time
2 min
Topics
Productivity, Startups, Software Development
AI-Generated Summary
Key Takeaways
- ✓State Size Balance: Transformers have states 100,000x larger than LSTMs at long context, while RNNs have states too small. Optimal architectures balance weight FLOPS and state FLOPS within one order of magnitude for compute-efficient training and inference.
- ✓Chunked Algorithm: Power retention uses dual computation forms—recurrent for sequential processing and attention for parallel processing. Breaking sequences into GPU-optimized chunks provides linear cost scaling while maintaining full hardware saturation, achieving best of both approaches without mathematical tradeoffs.
- ✓Model Metamorphosis: Converting existing transformer models to power retention requires only two hours of retraining on 128 H100s. StarCoder 3B recovered full 30% HumanEval performance after this brief metamorphosis period, making adoption practical without pretraining from scratch.
- ✓Vidrial CUDA Framework: Custom CUDA framework enables 20% speedups over Flash Attention on non-standard problem shapes by separating static and dynamic computation. JIT compilation sweeps different configurations to find optimal tile sizes and memory patterns for specific hardware and sequence lengths.
What It Covers
Jacob Buckman explains power retention architecture for transformers, combining recurrence and attention to achieve linear scaling for long context processing while maintaining computational efficiency through balanced weight-state FLOP ratios and chunked algorithms.
Key Questions Answered
- •State Size Balance: Transformers have states 100,000x larger than LSTMs at long context, while RNNs have states too small. Optimal architectures balance weight FLOPS and state FLOPS within one order of magnitude for compute-efficient training and inference.
- •Chunked Algorithm: Power retention uses dual computation forms—recurrent for sequential processing and attention for parallel processing. Breaking sequences into GPU-optimized chunks provides linear cost scaling while maintaining full hardware saturation, achieving best of both approaches without mathematical tradeoffs.
- •Model Metamorphosis: Converting existing transformer models to power retention requires only two hours of retraining on 128 H100s. StarCoder 3B recovered full 30% HumanEval performance after this brief metamorphosis period, making adoption practical without pretraining from scratch.
- •Vidrial CUDA Framework: Custom CUDA framework enables 20% speedups over Flash Attention on non-standard problem shapes by separating static and dynamic computation. JIT compilation sweeps different configurations to find optimal tile sizes and memory patterns for specific hardware and sequence lengths.
Notable Moment
Buckman reveals that typical window attention models plateau in their ability to use context far earlier than their advertised effective context length, which is calculated as depth times window size, demonstrating they fail to leverage most available tokens.
Episode Transcript
All axes of scale are important. I don't wanna say that state is the only important one at all. What what you really want is a is a architecture that's balanced, right? Just like I was saying before, it's you know, in this field, it's all about balance. That's how you know. When a model is beautiful and balanced and elegant, that's how you know it's gonna be a good compute optimal model. And the problem with having a state that's really big or really small fundamentally is that it's imbalanced. So, a way to think about this is Alright, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Jacob Buckman. Jacob is cofounder and CEO of Manifest AI. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Jacob, welcome to the podcast. Thanks so much for having me, Sam. I'm excited to meet you and dig into our conversation. We're going to be talking about achieving long context with transformers and the new power retention architecture that you recently published. To get us started, I'd love to have you share a little bit about your background. Yeah. Absolutely. So I've been working on deep learning for better part of a decade now. I started maybe in 2016, and I was an undergrad at Carnegie Mellon. I did my master's thesis on deep learning for autoregressive language modeling, basically, just because I thought it was was cool. It was exciting. And then, you know, here we are ten years later, and it's apparently the coolest and most exciting thing in the world, which was definitely not something I anticipated when I started getting into it. But, pretty much just Turns out. As it turns out. Yeah. And so I worked for Google Brain for a while, got lots of exposure to different areas of deep learning, some GAN stuff, some general adversarial networks, a little bit of adversarial examples, worked on deep reinforcement learning, and then I eventually went to go get a PhD in mostly deep reinforcement learning at, Mila in in Montreal. So, yeah, for long, I feel like the the writing was on the wall that just actually pushing the current paradigm was gonna take us to unbelievable places. And I started thinking together with my, cofounder, Carlos, how far is it gonna take us? Right? Like, people at this point were speculating, maybe scale will take us all the way to AGI. And, you know, they had a point. Right? And they had a pretty good argument. And one of the things that we started thinking about is, is it gonna take us all the way? Do we see any technical bottlenecks that will actually prevent us from getting there? And the conclusion that we reached was that although several of the major axes of scale are handled well, you know, we know …
Get the full transcript (10,455 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 54-minute episode.
Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from The TWIML AI Podcast
World Models and the Future of Spatial AI with Justin Johnson - #775
Sep 1 · 66 min
Eye on AI
#299 Jacob Buckman: Why the Future of AI Won't Be Built on Transformers
Nov 9
More from The TWIML AI Podcast
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Aug 25 · 55 min
Latent Space
Building Snipd: The AI Podcast App for Learning
Mar 14
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
“Custom CUDA framework enables 20% speedups over Flash Attention on non-standard problem shapes by separating static and dynamic computation.”
Gear
Products
“StarCoder 3B recovered full 30% HumanEval performance after this brief metamorphosis period, making adoption practical without pretraining from scratch.”
More from The TWIML AI Podcast
We summarize every new episode. Want them in your inbox?
World Models and the Future of Spatial AI with Justin Johnson - #775
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
How AI Learns to Smell with Alex Wiltschko - #771
Similar Episodes
Related episodes from other podcasts
Eye on AI
Nov 9
#299 Jacob Buckman: Why the Future of AI Won't Be Built on Transformers
Latent Space
Mar 14
Building Snipd: The AI Podcast App for Learning
Odd Lots
Sep 4
Why Laser Beams Are the Hottest New Tech in Defense
Software Engineering Daily
Aug 18
How LLMs Are Reshaping Recommendation Systems
This Week in Startups
Aug 17
Bittensor creator Const on Affine, dTAO, "mining reasoning," and more | E2326
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into The TWIML AI Podcast.
Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime