#326 Zuzanna Stamirowska: Inside Pathway's AI Systems That Work with Live, Real-Time Data
Episode
67 min
Read time
2 min
Topics
Productivity, Health & Wellness, Relationships
AI-Generated Summary
Key Takeaways
- ✓Memory at inference, not just training: BDH updates its fast weights continuously during inference, meaning the model retains context across a session rather than resetting each time. This directly addresses the core LLM limitation where every session starts from scratch — analogous to an employee who never accumulates experience beyond their first day on the job.
- ✓State lives on edges, not nodes: Unlike transformers where knowledge encodes in node weights, BDH stores state on synaptic edges (fast weights) while nodes function purely as computations. This duality, drawn from quantum physics principles, allows individual synapses to represent specific concepts — the paper demonstrates a single synapse activating consistently for the concept of currency.
- ✓Sparse local dynamics reduce compute: BDH uses a graph topology where neurons connect only to relevant neighbors, not all-to-all as in transformers. Attention scales linearly with neuron count n rather than quadratically. At inference, only a small local subgraph activates per step, meaning a model with a massive state accesses only a fraction of it — potentially cutting reasoning compute by 10x per output token.
- ✓Model merging via graph concatenation: Two separately trained BDH graphs can be joined along the shared neuron dimension n, then fine-tuned together to form cross-domain connections. A paper experiment merges two single-language models, producing coherent mixed-language output. This composability enables domain fusion — combining, for example, a finance-trained and a law-trained model into one integrated reasoning system.
- ✓Target use cases: small data and long-horizon reasoning: BDH's first commercial applications focus on high-value, data-scarce domains such as nuclear engineering documentation and healthcare claims resolution. The architecture's interpretability advantage — visible synapse activation patterns — supports regulated industries requiring explainability. AWS customers gain access through a Pathway-NVIDIA-AWS partnership announced at AWS re:Invent in December 2024.
What It Covers
Pathway cofounder Zuzanna Stamirowska presents the Dragon Hatchling (BDH) architecture, a post-transformer neural network modeled on brain-like graph dynamics. The system stores state on edges rather than nodes, enables persistent memory at inference time, and trains comparably to GPT-2 while targeting enterprise reasoning tasks requiring small data and long-horizon coherence.
Key Questions Answered
- •Memory at inference, not just training: BDH updates its fast weights continuously during inference, meaning the model retains context across a session rather than resetting each time. This directly addresses the core LLM limitation where every session starts from scratch — analogous to an employee who never accumulates experience beyond their first day on the job.
- •State lives on edges, not nodes: Unlike transformers where knowledge encodes in node weights, BDH stores state on synaptic edges (fast weights) while nodes function purely as computations. This duality, drawn from quantum physics principles, allows individual synapses to represent specific concepts — the paper demonstrates a single synapse activating consistently for the concept of currency.
- •Sparse local dynamics reduce compute: BDH uses a graph topology where neurons connect only to relevant neighbors, not all-to-all as in transformers. Attention scales linearly with neuron count n rather than quadratically. At inference, only a small local subgraph activates per step, meaning a model with a massive state accesses only a fraction of it — potentially cutting reasoning compute by 10x per output token.
- •Model merging via graph concatenation: Two separately trained BDH graphs can be joined along the shared neuron dimension n, then fine-tuned together to form cross-domain connections. A paper experiment merges two single-language models, producing coherent mixed-language output. This composability enables domain fusion — combining, for example, a finance-trained and a law-trained model into one integrated reasoning system.
- •Target use cases: small data and long-horizon reasoning: BDH's first commercial applications focus on high-value, data-scarce domains such as nuclear engineering documentation and healthcare claims resolution. The architecture's interpretability advantage — visible synapse activation patterns — supports regulated industries requiring explainability. AWS customers gain access through a Pathway-NVIDIA-AWS partnership announced at AWS re:Invent in December 2024.
- •Scale-free network structure emerges organically: During training, BDH's connection degree distribution converges toward a scale-free topology without any top-down design. This structure, common in organic complex systems like the internet and social networks, provides resilience, efficient communicability, and self-similar properties across network sizes — properties the team argues will support generalization and reasoning beyond what fixed transformer architectures can achieve.
Notable Moment
Stamirowska describes an experiment where two BDH models, each trained on a different language, were merged by simply concatenating their graphs. Without additional joint training, the combined model began producing output that mixed vocabulary from both languages in a coherent way — demonstrating that architectural composability is a structural property, not an engineering workaround.
Episode Transcript
So it's a bit like having an intern that you hire on the first day of her job, and she may be brilliant, but she stays an intern forever on the first day of her job. Doesn't get more context, doesn't get better with experience, doesn't get better over time. Do you think there will be creativity in this kind of a network? We'll be able to think beyond the fixed parameters of an LLM and come up with new ideas? This episode is brought to you by Tastytrade. On ION AI, we talk a lot about how artificial intelligence is changing how people analyze information, spot patterns, and make more informed decisions. Markets are no different. The edge increasingly comes from having the right tools, the right data, and the ability to understand risk clearly. That's one of the reasons I like what Tastytrade is building. With Tastytrade, you can trade stocks, options, futures, and crypto all in one platform with low commissions, including zero commissions on stocks and crypto so you keep more of what you earn. The platform is packed with advanced charting tools, back testing, strategy selection, and risk analysis tools that help you think in probabilities rather than guesses. They've also introduced an AI powered search feature that can help you discover symbols aligned with your interests, which is a smart way to explore markets more intentionally. For active traders, there are tools like active trader mode, one click trading, and smart order tracking. And if you're still learning, Tastytrade offers dozens of free educational courses plus live support from their trade desk reps during trading hours. If you're serious about trading in a world increasingly shaped by technology, check out Tastytrade. Visit tastytrade.com to start your trading journey today. I'm going to myself. Tastytrade Inc is a registered broker dealer and member of FINLA, NFA, and SIPC. I mean, I've done some reading. I'm not sure I entirely understand everything yet, but I've been interested in self improving AI me along with everybody else and continue learning and those things. And post transformer architectures, I've spoken to a few people. Carl Friston on I don't know if you followed his work, but he's working with a startup called Versus. And then there's a company called Manifest AI in New York City. I don't know if they're at all related to what you're doing, but in any any case, I'm I'm interested. And you guys are looking at this specifically for, robotics applications or no? Not necessarily. So first of all, Craig, thank you so much for having me. It's a pleasure to be to be here. And, I mean, you also had a chance to to talk to wonderful guests. So we actually do follow the work of the guys at manifest AI. So it was it was actually great to hear them on your podcast as well. We don't necessarily focus on robotics. In fact, what we focus on and the part of, …
Get the full transcript (9,664 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 64-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Sep 8 · 54 min
This Week in Startups
Bittensor creator Const on Affine, dTAO, "mining reasoning," and more | E2326
Aug 17
More from Eye on AI
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
Sep 3 · 38 min
Cognitive Revolution
Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
Jul 30
More from Eye on AI
We summarize every new episode. Want them in your inbox?
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
In 5 to 10 Years, Using Weapons Without AI Will Be Considered Unethical | Yaroslav Azhnyuk, The Fourth Law
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Similar Episodes
Related episodes from other podcasts
This Week in Startups
Aug 17
Bittensor creator Const on Affine, dTAO, "mining reasoning," and more | E2326
Cognitive Revolution
Jul 30
Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard
The TWIML AI Podcast
Jul 27
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
a16z Podcast
Jul 21
Why Physical AI Is the Next Frontier | Applied Intuition
a16z Podcast
Jul 15
Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime