Making deep learning perform real algorithms with Category Theory (Andrew Dudzik, Petar Velichkovich, Taco Cohen, Bruno Gavranović, Paul Lessard)
Episode
43 min
Read time
2 min
Topics
Fundraising & VC, Design & UX, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Algorithmic Failure in LLMs: Large language models perform hundreds of billions of multiplications to generate single tokens yet cannot reliably multiply small numbers together, revealing misalignment between training methods and downstream reasoning tasks requiring correctness guarantees.
- ✓Beyond Geometric Deep Learning: Group theory handles spatial symmetries but fails for non-invertible computations like Dijkstra's algorithm where multiple different graphs compress to identical outputs, requiring category theory's broader framework to model information-destroying transformations in algorithmic reasoning.
- ✓Two-Category Weight Tying: Two-morphisms in categorical frameworks formalize when weight sharing is mathematically valid across network layers, enabling provable correctness for parameter sharing beyond simple copying, creating systematic architecture design principles rather than ad-hoc engineering choices.
- ✓Carry Mechanism Challenge: Implementing arithmetic carries in continuous gradient-based systems requires modeling state changes rather than states themselves, a fundamental operation from CPU design that current graph neural networks struggle to represent, potentially solvable through geometric structures like Hopf fibrations.
What It Covers
Category theory provides a mathematical framework for designing neural networks that can reliably execute algorithms like addition and multiplication, addressing fundamental limitations in current large language models and deep learning architectures.
Key Questions Answered
- •Algorithmic Failure in LLMs: Large language models perform hundreds of billions of multiplications to generate single tokens yet cannot reliably multiply small numbers together, revealing misalignment between training methods and downstream reasoning tasks requiring correctness guarantees.
- •Beyond Geometric Deep Learning: Group theory handles spatial symmetries but fails for non-invertible computations like Dijkstra's algorithm where multiple different graphs compress to identical outputs, requiring category theory's broader framework to model information-destroying transformations in algorithmic reasoning.
- •Two-Category Weight Tying: Two-morphisms in categorical frameworks formalize when weight sharing is mathematically valid across network layers, enabling provable correctness for parameter sharing beyond simple copying, creating systematic architecture design principles rather than ad-hoc engineering choices.
- •Carry Mechanism Challenge: Implementing arithmetic carries in continuous gradient-based systems requires modeling state changes rather than states themselves, a fundamental operation from CPU design that current graph neural networks struggle to represent, potentially solvable through geometric structures like Hopf fibrations.
Notable Moment
The discussion reveals that asking neural networks to both translate messy real-world scenarios into structured representations and robustly execute algorithms on fixed computational budgets creates an impossible burden, suggesting future systems need explicit separation between world understanding and algorithmic reasoning components.
Episode Transcript
Language models cannot do addition. Not really. I keep seeing claims that they can, and every time I see this claim, I go again to chat g p t and so on and check, and they can't. And, what they can do is learn patterns, which work a lot of the time, but you can always trip them up by doing something like, okay. So if you ask chat g p t, what is a bunch of eights plus a bunch of ones with a two at the end? It will get the correct answer because it will recognize the trick. It'll say, ah, that's just one and a bunch of zeros. It'll know that you're trying to trick it. But if now you change one of the eights to a seven, now it has to actually know what it's doing. It has to actually sort of walk up, hit the seven, and stop sort of propagating zeros and it simply fails. It, it either chokes and makes up some nonsense or it says it's one with a bunch of zeros anyway, like, it definitely can't add in the in the the basic way that that we know how to do algorithmically that humans learn. And so, like, really teasing apart on a very basic, level, like new Newton's three laws of motion, has it encapsulated it, whether that's VO or Genie? Have these models encapsulated the physics of that a 100% accurately? And right now, they're not. They're kind of approximations, and they look, realistic when you just casually look at them. They're not, accurate enough yet to rely on for, say, robotics. Just because we can achieve some level of moving the needle by hooking up a really potent tool to a language model doesn't mean that we shouldn't think about what would the next generation of these models look like and how can we make them intrinsically better. Because even if you have the best tool in the world, that is not going to save you if you cannot predict the right inputs for that tool. Even some of the current frontier models, they will, as you probably know, perform hundreds of billions of multiplications just to produce a single token of output, yet they cannot reliably multiply even relatively small numbers together without failing. Right? So this is something that, to me hints at a great misalignment between what we are training these systems to do and how we're building them and what we might want to use them for downstream, especially if we're doing reasoning or if we're doing science. But it seems like, let's say an LLM, if you teach it properly, you can just teach it to do, let's say addition up to some failure of its memory just like with humans. We might forget a digit or we forget to carry over or so, you know, we do the algorithm wrong sometimes with some probability. Up to that it can learn this …
Get the full transcript (7,911 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 40-minute episode.
Get Machine Learning Street Talk summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Machine Learning Street Talk
AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
Sep 8 · 89 min
Sean Carroll's Mindscape
343 | Tom Griffiths on The Laws of Thought
Feb 9
More from Machine Learning Street Talk
Designing How AI Grows — Tom McGrath
Sep 2 · 100 min
Odd Lots
Why Soccer Analytics Works Like Volatility Arbitrage Trading
Jul 16
More from Machine Learning Street Talk
We summarize every new episode. Want them in your inbox?
AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
Designing How AI Grows — Tom McGrath
Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov
Every Exponential Ends — Silicon Valley Forgot — Adam Becker
AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart
Similar Episodes
Related episodes from other podcasts
Sean Carroll's Mindscape
Feb 9
343 | Tom Griffiths on The Laws of Thought
Odd Lots
Jul 16
Why Soccer Analytics Works Like Volatility Arbitrage Trading
The TWIML AI Podcast
Jul 27
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
My First Million
Apr 22
25% Of My Portfolio Is One Overvalued Stock, Here's Why
Radiolab
Dec 12
The Alien in the Room
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Machine Learning Street Talk.
Every Monday, we deliver AI summaries of the latest episodes from Machine Learning Street Talk and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime