Skip to main content
Latent Space

The First Mechanistic Interpretability Frontier Lab — Myra Deng & Mark Bissell of Goodfire AI

68 min episode · 2 min read
·
Mark Bissell,Myra Deng

Episode

68 min

Read time

2 min

Topics

Health & Wellness, Relationships, Investing

AI-Generated Summary

Key Takeaways

  • Production Interpretability at Scale: Goodfire deploys real-time steering on trillion-parameter models like Qwen QwQ, demonstrating surgical control over specific behaviors through feature manipulation. Their forked SG-Lang codebase enables live activation steering during inference, showing interpretability techniques can scale beyond toy models to frontier systems requiring full H100 nodes for deployment.
  • SAE Limitations in Practice: Sparse autoencoders underperform raw activation probes for detecting harmful behaviors, hallucinations, and PII in production scenarios. While SAEs excel with noisy synthetic datasets requiring generalization, supervised probes trained directly on activations achieve better downstream performance metrics when clean labeled data exists, revealing unsupervised methods have specific optimal use cases.
  • Steering-Prompting Equivalence: Research from Ekdeep Singh establishes formal mathematical equivalence between activation steering and in-context learning. The framework predicts exact steering magnitudes needed to replicate prompting effects, including jailbreaks through many-shot examples. This enables converting between inference-time interventions and understanding their interchangeable nature for model control.
  • Post-Training Surgical Edits: Interpretability enables targeted removal of unintended behaviors like political bias or reward hacking without full retraining. Goodfire positions this as moving beyond crude reinforcement learning that only provides reward signals, toward expert feedback that surgically modifies specific model representations. The approach addresses issues like GPT-4o's sycophancy problems through precise internal adjustments.
  • Healthcare Biomarker Discovery: Partnership with Mayo Clinic, AHRQ Institute, and Prima Menta uses interpretability on genomics foundation models to identify novel Alzheimer's disease biomarkers. The technique extracts superhuman knowledge from narrow AI systems trained on medical imaging and genomic data, demonstrating interpretability as a scientific discovery tool beyond debugging or safety applications.

What It Covers

Goodfire AI announces $150M Series B at $1.25B valuation as the first mechanistic interpretability frontier lab applying research to production. Mark Bissell and Myra Deng discuss using sparse autoencoders and probes to understand model internals, enable surgical steering of behaviors, and solve real-world problems from PII detection at Rakuten to Alzheimer's biomarker discovery.

Key Questions Answered

  • Production Interpretability at Scale: Goodfire deploys real-time steering on trillion-parameter models like Qwen QwQ, demonstrating surgical control over specific behaviors through feature manipulation. Their forked SG-Lang codebase enables live activation steering during inference, showing interpretability techniques can scale beyond toy models to frontier systems requiring full H100 nodes for deployment.
  • SAE Limitations in Practice: Sparse autoencoders underperform raw activation probes for detecting harmful behaviors, hallucinations, and PII in production scenarios. While SAEs excel with noisy synthetic datasets requiring generalization, supervised probes trained directly on activations achieve better downstream performance metrics when clean labeled data exists, revealing unsupervised methods have specific optimal use cases.
  • Steering-Prompting Equivalence: Research from Ekdeep Singh establishes formal mathematical equivalence between activation steering and in-context learning. The framework predicts exact steering magnitudes needed to replicate prompting effects, including jailbreaks through many-shot examples. This enables converting between inference-time interventions and understanding their interchangeable nature for model control.
  • Post-Training Surgical Edits: Interpretability enables targeted removal of unintended behaviors like political bias or reward hacking without full retraining. Goodfire positions this as moving beyond crude reinforcement learning that only provides reward signals, toward expert feedback that surgically modifies specific model representations. The approach addresses issues like GPT-4o's sycophancy problems through precise internal adjustments.
  • Healthcare Biomarker Discovery: Partnership with Mayo Clinic, AHRQ Institute, and Prima Menta uses interpretability on genomics foundation models to identify novel Alzheimer's disease biomarkers. The technique extracts superhuman knowledge from narrow AI systems trained on medical imaging and genomic data, demonstrating interpretability as a scientific discovery tool beyond debugging or safety applications.
  • Rakuten PII Detection System: Deployed token-level PII classification using probes on language model activations processes all user queries daily. The system handles synthetic-to-real transfer learning, multilingual requirements across English and Japanese, and precise scrubbing without routing private data to downstream providers. This demonstrates production-grade interpretability solving compliance problems traditional guardrail models cannot address efficiently.

Notable Moment

The team revealed live demonstration limitations expose engineering challenges at scale. Their hastily assembled demo of steering a trillion-parameter model required custom infrastructure and proved fragile behind the scenes, highlighting how production interpretability demands solving both novel research problems and significant systems engineering hurdles that academic toy models never encounter.

Know someone who'd find this useful?

Episode Transcript

So welcome to the Lanespace five. We're back in the studio with our special MecanTurf co host, Vibhu. Welcome. Mochi. Mochi is your co host. And Mochi, the mechanistic interpretability doggo. We have with us Mark and Myra from Goodfire. Welcome. Thanks for having us on. Maybe we can sort of introduce Goodfire and then and then introduce you guys. How do you introduce Goodfire today? Yeah. It's a great question. So Goodfire, we like to say, is an AI research lab that focuses on using interpretability to understand, learn from, and design AI models. And we really believe that interpretability will unlock the new generation, next frontier of safe and powerful AI models. That's our description right now. And, excited to dive more into the work we're doing to make that happen. Yeah. And there's there's always, like, the official description. Is there an, like, unofficial one that sort of resonates more with a different audience? Well, being an AI research lab that's focused on interpretability, there's obviously a lot of people have a lot that they think about when they think of interpretability. And I think we have a pretty broad definition of what that means and the types of places that can be applied, and in particular, applying it in production scenarios, in high stakes industries, and really taking it sort of from the research world into the real world, which, you know, it's an it's a new field, so that hasn't been done all that much, and we're excited about actually seeing that sort of put into put into practice. Yeah. I would say it it's it wasn't too long ago that Topic was, like, still putting out, like, toy models or supervis position and that kind of stuff. And I wouldn't have pegged it to be this far along. When you and I talked to NeurIPS, you were talking a little bit about your production use cases and your customers. And then not to bury the lead, today we're also announcing the fundraise, your series b, a 150,000,000 million dollars at a 1.25 b valuation. Congrats, you're a unicorn. Thank you. Yeah. No. Things move fast. We were talking to you in December and Yeah. Already some big updates since then. Let's dive, I guess, into a bit of your backgrounds as well. Mark, you you were at Palantir, worked in health stuff, which is, really interesting because, the Goodfire has some interesting, like, health use cases. I don't know how related they are in practice. Yeah. Not not super related, but, I don't know. It was helpful context to know what it's like just to work with, health systems and generally in that domain. Yeah. And, Myra, you were at Two Sigma, which actually I was also at, Two Sigma Oh, really? Back in the day. Wow. Nice. Did we overlap at all? No. No. This is this is when I was the briefly a software engineer before I became a sort of developer …

Get the full transcript (13,214 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

You just read a 3-minute summary of a 65-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • Goodfire deploys real-time steering on trillion-parameter models like Qwen QwQ, demonstrating surgical control over specific behaviors through feature manipulation.
  • Their forked SG-Lang codebase enables live activation steering during inference, showing interpretability techniques can scale beyond toy models to frontier systems requiring full H100 nodes for deployment.

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime