He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
Episode
72 min
Read time
3 min
Topics
Investing, Leadership, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Architecture Lock-in: Transformers dominate not because alternatives are worse, but because switching costs are prohibitive. Competing architectures must be "crushingly better" — not marginally better — to displace an established system with mature tooling, fine-tuning pipelines, and inference infrastructure. The same pattern occurred when transformers displaced RNNs: the accuracy jump was so large that researchers had no choice but to migrate.
- ✓Continuous Thought Machine Design: The CTM introduces three architectural novelties: an internal sequential "thought" dimension that applies compute across discrete steps, neuron-level models (NLMs) that treat each neuron as a small MLP processing a history of activations rather than a single ReLU, and synchronization representations that measure dot-product correlations between neuron activation time series to encode richer, temporally-aware state.
- ✓Native Adaptive Computation: Training the CTM on ImageNet with a dual-loss — minimizing cross-entropy at both the lowest-loss step and the highest-certainty step — causes easy examples to resolve in one or two steps while hard examples use the full 50-step budget. This adaptive behavior emerges without explicit computation-penalty terms, unlike Alex Graves' Adaptive Computation Time paper, which required carefully tuned auxiliary losses.
- ✓Calibration as Architecture Signal: After standard training, the CTM produced near-perfect probability calibration on classification tasks — meaning a 90% confidence prediction was correct roughly 90% of the time. Most neural networks trained to convergence become poorly calibrated and require post-hoc correction. The CTM's emergent calibration suggests the synchronization-based representation aligns model uncertainty with actual error rates more naturally.
- ✓SudokuBench Reasoning Gap: Sakana AI released SudokuBench, a dataset of handcrafted variant Sudoku puzzles with unique natural-language rule sets, sourced from thousands of hours of Cracking the Cryptic YouTube videos providing detailed human reasoning traces. Current top models solve only the simplest puzzles at around 15% accuracy. GPT-4 shows improvement but cannot find the "break-in" insight each puzzle requires, exposing a fundamental gap in sequential deductive reasoning.
What It Covers
Llion Jones, co-inventor of the transformer, and Sakana AI researcher Luke Darlow discuss the Continuous Thought Machine (CTM), a spotlight paper at NeurIPS 2025. They examine why AI research is trapped in a transformer-centric local minimum, how biological neuron synchronization inspired a new recurrent architecture, and why research freedom produces better science than commercial pressure.
Key Questions Answered
- •Architecture Lock-in: Transformers dominate not because alternatives are worse, but because switching costs are prohibitive. Competing architectures must be "crushingly better" — not marginally better — to displace an established system with mature tooling, fine-tuning pipelines, and inference infrastructure. The same pattern occurred when transformers displaced RNNs: the accuracy jump was so large that researchers had no choice but to migrate.
- •Continuous Thought Machine Design: The CTM introduces three architectural novelties: an internal sequential "thought" dimension that applies compute across discrete steps, neuron-level models (NLMs) that treat each neuron as a small MLP processing a history of activations rather than a single ReLU, and synchronization representations that measure dot-product correlations between neuron activation time series to encode richer, temporally-aware state.
- •Native Adaptive Computation: Training the CTM on ImageNet with a dual-loss — minimizing cross-entropy at both the lowest-loss step and the highest-certainty step — causes easy examples to resolve in one or two steps while hard examples use the full 50-step budget. This adaptive behavior emerges without explicit computation-penalty terms, unlike Alex Graves' Adaptive Computation Time paper, which required carefully tuned auxiliary losses.
- •Calibration as Architecture Signal: After standard training, the CTM produced near-perfect probability calibration on classification tasks — meaning a 90% confidence prediction was correct roughly 90% of the time. Most neural networks trained to convergence become poorly calibrated and require post-hoc correction. The CTM's emergent calibration suggests the synchronization-based representation aligns model uncertainty with actual error rates more naturally.
- •SudokuBench Reasoning Gap: Sakana AI released SudokuBench, a dataset of handcrafted variant Sudoku puzzles with unique natural-language rule sets, sourced from thousands of hours of Cracking the Cryptic YouTube videos providing detailed human reasoning traces. Current top models solve only the simplest puzzles at around 15% accuracy. GPT-4 shows improvement but cannot find the "break-in" insight each puzzle requires, exposing a fundamental gap in sequential deductive reasoning.
- •Research Freedom as Competitive Strategy: Jones argues that protecting researcher autonomy is a primary leadership responsibility at Sakana AI. Commercial pressure — investor return expectations, product deadlines, publication quotas — systematically narrows the solution space researchers explore. The CTM itself emerged from eight months of unconstrained exploration with no predetermined goal, producing emergent behaviors like backtracking maze navigation and leapfrog path-solving under constrained compute budgets.
Notable Moment
During training, the CTM spontaneously developed two distinct maze-solving strategies depending on available compute steps. With sufficient time, it traced paths sequentially. When steps were constrained, it instead leapfrogged ahead, traced segments backward, then jumped forward again — an algorithm the researchers never designed or anticipated, emerging purely from architectural constraints.
Episode Transcript
This episode is brought to you by White Claw Surge. Nice choice hitting up this podcast. No surprises. You're all about diving into tastes everyone in the room can enjoy, just like White Claw Surge. It's for celebrating those moments when connections have been made and the night's just begun. With bold flavors and 8% alcohol by volume, unleash the night. Unleash White Claw Surge. Please drink responsibly. Hard seltzer with flavors, 8% alcohol by volume. White Claw Seltzer Works, Chicago, Illinois. It's crunch time at work, and you need to bring wings to your workday. Visit redbull.com/gettingitdone and answer a couple questions about your work style to get a Spotify customized playlist tuned to your productivity. Plus, score a can of Red Bull on us while you go from to do to done. And remember, Red Bull gives you wings. Supplies are limited. Terms apply. Visit the website for more information. Despite the fact that I was involved in inventing the transformer, luckily, no one's been working on them as long as I have rice with maybe the exception of the other seven seven authors. So I actually made the decision, earlier this year that I'm gonna drastically reduce the amounts of of research that I'm doing specifically on the transformer because of the feeling that I have that it's it's an oversaturated space. Right? It's not that there's no more interesting things to be done with them, and I'm gonna make use of the opportunity to do something different. Right? To actually turn up the amount of exploration that I'm doing in my research. We just released the continuous thought machine. It's a spotlight at Europe's 2025 this year. You should care about it because it has native adaptive compute. It's a new way of building a recurrent model that uses higher level concepts for neurons and a synchronization as a representation that lets us solve problems in ways that seem more human by being biologically and nature inspired. The atmosphere in AI research was actually quite different back during the transformer, years, because it doesn't feel like something similar could actually happen right now because of the reduced amount of freedom that we have. Right? The Transformers was very, very bottom up. Right? It's not that somebody had this grand plan that came down from on high that this is what we should be working on. It was a bunch of people talking over lunch, thinking about what the current problems are and how to solve them, and having the freedom to have, you know, literally months to dedicate to just trying this idea and having this this, new architecture fall out. We've spent hundreds of millions of dollars. The biggest sort of evolution based search is probably in the tens of thousands. We have all this compute. What happens? What happens if you scale up these search algorithms? And I'm sure you will find something interesting, you know, when someone eventually does buy that bullet and …
Get the full transcript (12,960 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 69-minute episode.
Get Machine Learning Street Talk summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Machine Learning Street Talk
How Researchers Test AI for Hidden Goals — Apollo Research
Jul 31 · 78 min
Equity
This Sequoia-backed lab thinks the brain is 'the floor, not the ceiling' for AI
Feb 10
More from Machine Learning Street Talk
Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)
Jul 13 · 55 min
The Jordan Harbinger Show
1366: Gut Health | Skeptical Sunday
Aug 9
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Books
by Alex Graves
“This adaptive behavior emerges without explicit computation-penalty terms, unlike Alex Graves' Adaptive Computation Time paper, which required carefully tuned auxiliary losses.”
Products
- SudokuBenchBy guest
by Sakana AI
“Sakana AI released SudokuBench, a dataset of handcrafted variant Sudoku puzzles with unique natural-language rule sets, sourced from thousands of hours of Cracking the Cryptic YouTube videos.”
other
- Continuous Thought Machine (CTM)By guest
by Sakana AI
“Llion Jones, co-inventor of the transformer, and Sakana AI researcher Luke Darlow discuss the Continuous Thought Machine (CTM), a spotlight paper at NeurIPS 2025.”
podcast
“SudokuBench, a dataset of handcrafted variant Sudoku puzzles with unique natural-language rule sets, sourced from thousands of hours of Cracking the Cryptic YouTube videos providing detailed human reasoning traces.”
More from Machine Learning Street Talk
We summarize every new episode. Want them in your inbox?
How Researchers Test AI for Hidden Goals — Apollo Research
Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)
The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
The Thermodynamic AI Computing Chip - Thomas Ahle
He won a Nobel here for AlphaFold. Then he left. - John Jumper
Similar Episodes
Related episodes from other podcasts
Equity
Feb 10
This Sequoia-backed lab thinks the brain is 'the floor, not the ceiling' for AI
The Jordan Harbinger Show
Aug 9
1366: Gut Health | Skeptical Sunday
No Priors: Artificial Intelligence | Technology | Startups
Aug 6
Chasing Trillion-Dollar Companies, Founder Ambition, Token Budgets, and Regulatory Capture with Sarah & Elad
Planet Money
Jul 24
Piles of cash and a town of solutions in Kenya, Nigeria (Summer School)
Latent Space
Jul 21
🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Machine Learning Street Talk.
Every Monday, we deliver AI summaries of the latest episodes from Machine Learning Street Talk and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime