Skip to main content
Dwarkesh Podcast

Adam Marblestone – AI is missing something fundamental about the brain

109 min episode · 2 min read

Episode

109 min

Read time

2 min

Topics

Startups, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Evolution's Loss Functions: The brain uses thousands of specific, genetically-encoded cost functions that activate at different developmental stages, not simple objectives like next-token prediction. Evolution compressed learning curricula into reward signals by encoding innate heuristics (spider detection, social status cues) that the cortex learns to predict, enabling generalization without explicit supervision.
  • Omnidirectional Inference vs Amortized Prediction: The cortex can predict any subset of variables from any other subset, unlike LLMs that only predict forward. This allows filling in blanks bidirectionally, predicting vision from audio or inferring causes from effects. Energy-based models may capture this better than current transformer architectures that amortize inference into feedforward passes.
  • Steering Subsystem Architecture: The hypothalamus and brainstem contain thousands of specialized cell types encoding innate behaviors, far more than cortical regions. These subcortical areas have their own primitive sensory systems (like superior colliculus for face detection) that provide reward signals the cortex learns to predict, solving how abstract concepts trigger instinctive responses.
  • Connectome Economics: Current electron microscopy costs billions per mouse brain, but optical approaches from E11 Bio could reduce costs to tens of millions. A molecularly-annotated connectome showing cell types and synapse properties across multiple mammal species would cost low billions total, trivial compared to AI training budgets approaching trillions.
  • Behavior Cloning with Neural Data: Training AI to predict both task labels and human brain activity patterns as auxiliary loss functions could improve generalization by matching how brains represent information. This regularization approach requires scaling portable brain scanning technology, currently a bottleneck compared to GPU availability for standard supervised learning.

What It Covers

Adam Marblestone explains why AI lacks fundamental brain mechanisms: evolution-encoded loss functions, omnidirectional inference, and a steering subsystem that creates specific reward signals. He argues neuroscience needs technological scaling to answer how brains achieve sample-efficient learning.

Key Questions Answered

  • Evolution's Loss Functions: The brain uses thousands of specific, genetically-encoded cost functions that activate at different developmental stages, not simple objectives like next-token prediction. Evolution compressed learning curricula into reward signals by encoding innate heuristics (spider detection, social status cues) that the cortex learns to predict, enabling generalization without explicit supervision.
  • Omnidirectional Inference vs Amortized Prediction: The cortex can predict any subset of variables from any other subset, unlike LLMs that only predict forward. This allows filling in blanks bidirectionally, predicting vision from audio or inferring causes from effects. Energy-based models may capture this better than current transformer architectures that amortize inference into feedforward passes.
  • Steering Subsystem Architecture: The hypothalamus and brainstem contain thousands of specialized cell types encoding innate behaviors, far more than cortical regions. These subcortical areas have their own primitive sensory systems (like superior colliculus for face detection) that provide reward signals the cortex learns to predict, solving how abstract concepts trigger instinctive responses.
  • Connectome Economics: Current electron microscopy costs billions per mouse brain, but optical approaches from E11 Bio could reduce costs to tens of millions. A molecularly-annotated connectome showing cell types and synapse properties across multiple mammal species would cost low billions total, trivial compared to AI training budgets approaching trillions.
  • Behavior Cloning with Neural Data: Training AI to predict both task labels and human brain activity patterns as auxiliary loss functions could improve generalization by matching how brains represent information. This regularization approach requires scaling portable brain scanning technology, currently a bottleneck compared to GPU availability for standard supervised learning.

Notable Moment

Marblestone describes how the cortex learns that even hearing the word spider activates fear responses by predicting innate flinch reflexes in the hypothalamus. This thought assessor mechanism allows evolution to wire complex social instincts to learned concepts without knowing what future humans would encounter.

Know someone who'd find this useful?

Episode Transcript

The big million dollar question that I have that, I've been trying to get the answer to through all these interviews with AI researchers, how does the brain do it? Right? Like, we're throwing way more data at these LLMs, and they still have a small fraction of the total capabilities that a human does. So what's going on? Yeah. I mean, this might be the quadrillion dollar question or something like that. It's it's it's arguable you can make an argument this is the most important, you know, question in science. I don't claim to know the answer. I I also don't really think that the answer will necessarily come even from a lot of smart people thinking about it as much as they are. My my overall, like, meta level take is that we have to empower the field of neuroscience to just make neuroscience a a more powerful, field, technologically and otherwise, to actually be able to crack a question like this. But maybe the the way that we would think about this now with, like, modern AI, neural nets, deep learning, is that there's sort of these these certain key components of that. There's the architecture. There's maybe hyperparameters of the architecture. How many layers do you have or sort of properties of that architecture? There is the learning algorithm itself. How do you train it? You know, back prop, gradient descent. Is it something else? There is how is it initialized? Okay. So if we take the learning part of the system, it still may have some initialization of of the weights. And then there are also cost functions. There's, like, what is it being trained to do? Yeah. What's the reward signal? What are the loss functions, supervision signals? My personal hunch within that framework is that the the field has neglected, the role of this very specific loss functions, very specific cost functions. Machine learning tends to, like, mathematically simple loss functions. Right? Predict the next token, you know, cross entropy. The you know, these these these these, simple kind of computer scientist loss functions. I think evolution may have built a lot of complexity into the loss functions. Actually, many different loss functions were different areas turned on at different stages of development. A lot of Python code, basically, generating, a specific curriculum for what different parts of the brain need to learn. Because evolution has seen many times what was successful and unsuccessful, and evolution could encode the knowledge of of the learning curriculum. So so in the in the machine learning framework, maybe we can come back and we could talk about, yeah, where do the loss functions of the brain come from? Can that can loss different loss functions lead to different efficiency of learning? You know, people say, like, the cortex has got the universal human learning algorithm, the special science that humans have. What's up with that? A huge question, and we don't know. I've seen models …

Get the full transcript (21,593 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Dwarkesh Podcast transcripts →

You just read a 3-minute summary of a 106-minute episode.

Get Dwarkesh Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

company

  • optical approaches from E11 Bio could reduce costs to tens of millions

More from Dwarkesh Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Dwarkesh Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Dwarkesh Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime