Skip to main content
a16z Podcast

Why Medical AI Needs a Referee | Protege's Engy Ziedan

35 min episode · 2 min read
·
Engy Ziedan

Episode

35 min

Read time

2 min

Topics

Health & Wellness, Investing, Startups

AI-Generated Summary

Key Takeaways

  • Benchmark Gap: AI models scoring 92% on medical licensing exams perform at only 45% on real-world clinical tasks. Passing thousands of Q&A pairs reviewed by physicians does not validate performance in high-stakes scenarios like spinal surgery or oncology pathology, where outcome history and task-specific accuracy matter far more than general knowledge scores.
  • Subtle Misalignment Over Catastrophic Failure: Catastrophic AI failures are narrowly definable and easier to catch. The harder problem is subtle misalignment — for example, an insurance AI optimizing claim denials while a hospital AI maximizes revenue recovery, leaving patients with zero agency and no independent monitor flagging the conflict between both systems.
  • Benchmark Contamination: To produce valid evaluations, Protege sources "net new" patient data — whole slide images and records never previously scanned or entered into any model's training set. Rankings shift based on prompt wording and answer ordering alone, meaning any benchmark built on previously trained data risks measuring memorization rather than clinical reasoning.
  • Independent Arbiter Model: Protege hosts competing vertical AI builders within identical clinical subnodes — such as ambient scribing or oncology pathology — running unified evaluations without revealing participant identities until completion. Deficiencies revealed in evaluations are directly matched to specific training data Protege can supply, creating a closed loop between evaluation and model improvement.
  • Real-Time Monitoring Requirement: Static annual or retrospective quality reports, modeled on government value-based purchasing programs, are inadequate for AI systems that update continuously. A live monitoring layer is needed to detect when an AI agent nudges clinicians toward profit-maximizing decisions, flags early discharge recommendations tied to staff scheduling, or exhibits drift as patient populations and clinical guidelines evolve.

What It Covers

Engy Ziedan, cofounder and chief scientific officer of Protege, explains why medical AI benchmarks fail to predict real-world clinical performance, how subtle model misalignment poses greater risks than catastrophic failures, and why an independent third-party referee is needed to continuously evaluate competing healthcare AI systems.

Key Questions Answered

  • Benchmark Gap: AI models scoring 92% on medical licensing exams perform at only 45% on real-world clinical tasks. Passing thousands of Q&A pairs reviewed by physicians does not validate performance in high-stakes scenarios like spinal surgery or oncology pathology, where outcome history and task-specific accuracy matter far more than general knowledge scores.
  • Subtle Misalignment Over Catastrophic Failure: Catastrophic AI failures are narrowly definable and easier to catch. The harder problem is subtle misalignment — for example, an insurance AI optimizing claim denials while a hospital AI maximizes revenue recovery, leaving patients with zero agency and no independent monitor flagging the conflict between both systems.
  • Benchmark Contamination: To produce valid evaluations, Protege sources "net new" patient data — whole slide images and records never previously scanned or entered into any model's training set. Rankings shift based on prompt wording and answer ordering alone, meaning any benchmark built on previously trained data risks measuring memorization rather than clinical reasoning.
  • Independent Arbiter Model: Protege hosts competing vertical AI builders within identical clinical subnodes — such as ambient scribing or oncology pathology — running unified evaluations without revealing participant identities until completion. Deficiencies revealed in evaluations are directly matched to specific training data Protege can supply, creating a closed loop between evaluation and model improvement.
  • Real-Time Monitoring Requirement: Static annual or retrospective quality reports, modeled on government value-based purchasing programs, are inadequate for AI systems that update continuously. A live monitoring layer is needed to detect when an AI agent nudges clinicians toward profit-maximizing decisions, flags early discharge recommendations tied to staff scheduling, or exhibits drift as patient populations and clinical guidelines evolve.

Notable Moment

Ziedan describes how ambient documentation tools may actually reduce human bias: traditional clinical notes sometimes contained coded language suggesting a patient was untrustworthy, leading to misdiagnosis. AI-transcribed SOAP notes strip that language out, potentially prompting more rigorous clinical investigation than a biased human-written record would have triggered.

Know someone who'd find this useful?

Episode Transcript

Hundreds of millions of people ask chatty peachy questions about their health. Who, if any, is making sure that the answers that are spit out is safe and correct? Models are going to be inhibited in their usefulness by the training data available for them. There is no one that's looking beyond the iceberg of, like, catastrophic failures and misalignment. What is the importance of evals in this industry? No one ever asks, like, what is the value of Uber? Like, show me the eval. But in today's AI market, there is a need for the pricing to be accurate. And without understanding really what is valuable and what's value less technology, this technology does not have a marginal cost of zero. And so it became our mission to provide safe and aligned data that would make AI useful. The right decision may also be very different a year from now versus what it looks like today. As we move into the future A medical AI model can ace thousands of test questions and still fail at a job we actually need it to do. In this episode, Daisy Wolf and Edith Steinman sit down with Protege cofounder and chief scientific officer, N. G. Zden, to unpack why healthcare AI needs a much better way to measure performance. They discuss the gap between benchmarks and real world clinical tasks, why subtle bias and misalignment may be harder to catch than catastrophic failures, and what happens when models evolve faster than the healthcare system can evaluate them. Engie also makes the case for an independent referee, one that can continuously test how AI behaves in real clinical settings, compare competing models, and identify where they actually need to improve. Welcome back to the a16z podcast. I'm Daisy Wolf, partner on a16z's bio and health team, joined by Eva Steinman, investor on our bio and health team as well. Today, we are talking with N. G. Zden, co founder and chief scientific officer of Protege and a healthcare economist, assistant professor at Indiana University whose work has been featured in the New York Times and cited by the CDC. Today, we are gonna dig into why medical AI has a measurement problem, why acing benchmarks does not make a model ready for the hospital, and how Protege is building the referee. Angie, welcome to the podcast. Thanks for having me. Angie, let's start with your background. How did you first meet the Protege team? What were you doing at the time and how has it evolved since then? Yep. I met Bobby when I was two years out of my PhD, an assistant professor at Tulane in my lonely office, and a pandemic had just hit. And I decided that I was going to write papers really fast. So I wanted data basically from two days ago, and that was impossible to obtain at the time. In his past life, he used to work for a data facilitation company, and he offered …

Get the full transcript (6,225 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 32-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime