Hippocratic AI's Munjal Shah on How AI Agents Are Expanding Healthcare Capacity - Ep. 262
Episode
21 min
Read time
2 min
Topics
Productivity, Health & Wellness, Sales & Revenue
AI-Generated Summary
Key Takeaways
- ✓Constellation Architecture: Hippocratic runs 22 models simultaneously per conversation—one 400B parameter model handles dialogue while 19 supervising models check safety in real-time, plus two deep-thinking models perform 30-60 second verification checks, requiring 128 NVIDIA H100 GPUs just to load into RAM before supporting multiple conversations.
- ✓Output Testing Protocol: Rather than validating training data, Hippocratic hired 6,000 licensed US clinicians to conduct 309,000 test calls, marking every error before deployment. This use-case-specific testing approach costs double-digit millions but ensures safety by verifying actual outputs, not architectural assumptions or training sources.
- ✓Inference Latency Requirements: Voice-based healthcare agents need 1.5-2 second end-to-end response times, requiring optimization for latency rather than cost-per-token or throughput. This differs fundamentally from text-based search applications where 20-30 second delays remain acceptable for deeper reasoning capabilities.
- ✓Agent App Store Model: Clinicians submit custom scripts based on specialized expertise, receive validation and safety testing from Hippocratic, then earn revenue share when their agents deploy. A concussion clinic nurse can scale 20 years of knowledge to millions of patients nationwide within four minutes of prompt creation.
What It Covers
Munjal Shah explains how Hippocratic AI deploys safety-focused healthcare agents that have completed 1.85 million patient calls, achieving 8.95/10 satisfaction ratings while addressing clinical staffing shortages through inference-optimized architecture and rigorous output testing protocols.
Key Questions Answered
- •Constellation Architecture: Hippocratic runs 22 models simultaneously per conversation—one 400B parameter model handles dialogue while 19 supervising models check safety in real-time, plus two deep-thinking models perform 30-60 second verification checks, requiring 128 NVIDIA H100 GPUs just to load into RAM before supporting multiple conversations.
- •Output Testing Protocol: Rather than validating training data, Hippocratic hired 6,000 licensed US clinicians to conduct 309,000 test calls, marking every error before deployment. This use-case-specific testing approach costs double-digit millions but ensures safety by verifying actual outputs, not architectural assumptions or training sources.
- •Inference Latency Requirements: Voice-based healthcare agents need 1.5-2 second end-to-end response times, requiring optimization for latency rather than cost-per-token or throughput. This differs fundamentally from text-based search applications where 20-30 second delays remain acceptable for deeper reasoning capabilities.
- •Agent App Store Model: Clinicians submit custom scripts based on specialized expertise, receive validation and safety testing from Hippocratic, then earn revenue share when their agents deploy. A concussion clinic nurse can scale 20 years of knowledge to millions of patients nationwide within four minutes of prompt creation.
Notable Moment
Shah reveals that 30% of patients initially resist AI healthcare, but after agents explain human callback delays and demonstrate empathetic listening, only 15% ultimately refuse. Patients appreciate undivided attention—something increasingly rare in modern interactions—leading to sustained engagement within 30-60 seconds.
Episode Transcript
Hello, and welcome to the NVIDIA AI podcast from GTC twenty twenty five in San Jose, California. I'm Noah Kravitz, and I'm here with Munjal Shah, cofounder and CEO of Hippocratic AI, a start up building a safety focused LLM, large language model, for health care. Hippocratic recently launched a health care AI agent app store and announced their series b funding round and is at the forefront of the AI powered health care transformation that's happening all around us. I'm excited to talk about AI and the big idea behind the Hippocratic AI, the era of health care abundance, with you, Munjal. So welcome, and thanks so much for taking the time to join the AI podcast. Oh, I'm so excited to be here. Thank you for having me. So let's start at the, with the basics. What is Hippocratic AI? Well, you know, as you mentioned, we're a safety focused large language model, focused on health care, and we really used it to build AI clinicians. So we have agents that operate, and reach out to patients and call them on the phone and say, you know, check-in on them post surgery, and let's look at your incision site. And then is it getting infected? And, you know, do you have enough of your medications, and do you need refills? So it's really an agent that talks to patients and, delivers care. So when was the company founded? We started it about two years ago now. Yeah. And you're in production now with agents interacting with patients? Yeah. We've done, at as of the end of this month, we will have done about 1,850,000 calls Wow. To patients all over the country. What's the reaction been like from the patients to, getting health care from an AI agent? You know, people always ask that question. They're always like, well, what are the patients there? Yeah. Yeah. The average I'll give it to you in numbers, and I'll give it to you in anecdotes. Okay. The average patient rating is an 8.95 out of 10. It's pretty good. Yeah. And second, it's I think there's about 30% who are like, I don't wanna talk to AI. Right. With a little bit of rebuttal, the AI goes, look, I I can really help you. And I don't know when that human's gonna call you back because they don't call you back a lot of times. Right. Will you talk to me? It turns out to about fifteen percent will ultimately leave, but the other eighty five percent will talk to it. Right. And within, like, thirty seconds, sixty seconds, when they realize this is not your grandfather's IVR Uh-huh. Like, this truly can understand you and talk to you and is empathetic. They just talk away. People just talk. Yeah. Yeah. I mean, and think about it in this day and age, like, who really listens to every word you say? Like, no one ever. And I think that now …
Get the full transcript (4,469 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 18-minute episode.
Get NVIDIA AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from NVIDIA AI Podcast
Inside Instacart's AI-Powered Smart Shopping Cart | NVIDIA AI Podcast Ep. 302
Jun 24 · 39 min
David Senra
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Sep 9
More from NVIDIA AI Podcast
How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301
Jun 10 · 21 min
Odd Lots
'The Assassin' Fahmi Quadir on How to Survive as a Short-Seller
May 22
More from NVIDIA AI Podcast
We summarize every new episode. Want them in your inbox?
Inside Instacart's AI-Powered Smart Shopping Cart | NVIDIA AI Podcast Ep. 302
How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301
Everyone Can Build a Robot: Open Source Embodied AI With Seeed Studio | NVIDIA AI Podcast Ep. 300
Inside AI Tokenomics: How to Profitably Turn Tokens Into Business Value | NVIDIA AI Podcast Ep. 299
Snap’s Secret to Processing 10 Petabytes a Day: GPU-Accelerated Spark | NVIDIA AI Podcast Ep. 298
Similar Episodes
Related episodes from other podcasts
David Senra
Sep 9
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Odd Lots
May 22
'The Assassin' Fahmi Quadir on How to Survive as a Short-Seller
In Good Company with Nicolai Tangen
Jan 23
HIGHLIGHTS: Mala Gaonkar - Founder of SurgoCap Partners
In Good Company with Nicolai Tangen
Jan 21
Mala Gaonkar: Building SurgoCap, Identifying Great Businesses and Learning from Mistakes
Eye on AI
Aug 24
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into NVIDIA AI Podcast.
Every Monday, we deliver AI summaries of the latest episodes from NVIDIA AI Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime