#316 Robbie Goldfarb: Why the Future of AI Depends on Better Judgment
Episode
63 min
Read time
3 min
Topics
Health & Wellness, Relationships, Investing
AI-Generated Summary
Key Takeaways
- ✓Expert thought process extraction: Form.ai maps how experts reason through complex questions by asking them to verbalize their approach rather than just label data. For political questions, experts might say they would cross-reference reliable sources first, creating a chain of reasoning. Form.ai builds agentic systems that mirror these thought process graphs, testing generalizability across scenarios before deploying judges at scale.
- ✓Consequence mapping methodology: Instead of traditional good-or-bad labeling, Form.ai asks experts to predict outcomes of AI conversations—what emotions users would feel, what actions they would take, what family members would say. This experimental approach provides richer data for training judges and reveals the reasoning behind expert evaluations, particularly valuable for sensitive domains like mental health where clinical nuance matters.
- ✓Four-tier evaluation framework for political content: Form.ai assesses political AI responses across bias (which breaks into multiple subcategories), factuality, source selection (credibility, balance, accurate attribution), and tone-language (avoiding inflammatory phrasing that mirrors user anger). Different combinations of topic, user intent, and evaluation dimension require separate judges, creating a matrix of specialized evaluators rather than one-size-fits-all assessment.
- ✓Trust gap blocking AI adoption: KPMG research shows 80-plus percent of people feel optimistic about AI improving their lives, yet only 40-something percent trust it. This delta represents the critical barrier to realizing AI potential. Form.ai addresses this through transparent expert networks published on their website, allowing users to see exactly who shaped model training rather than relying on internal engineering teams or scaled labelers.
- ✓Dangerous expectation of omniscient AI: ChatGPT established a problematic norm where users expect authoritative answers to any question without context-gathering dialogue. Real expertise requires back-and-forth—doctors do not prescribe after one statement. AI models need awareness of when they lack sufficient context to respond responsibly, but this conflicts with engagement metrics since users migrate to models providing immediate answers over those asking clarifying questions.
What It Covers
Robbie Goldfarb, founder of Form.ai, explains how his company scales expert judgment to evaluate and improve AI systems in contentious domains like healthcare and politics. Form.ai builds transparent networks of credible experts—including Fareed Zakaria and Neil Ferguson—then creates AI judges that capture their reasoning processes to assess models for bias, accuracy, and clinical nuance.
Key Questions Answered
- •Expert thought process extraction: Form.ai maps how experts reason through complex questions by asking them to verbalize their approach rather than just label data. For political questions, experts might say they would cross-reference reliable sources first, creating a chain of reasoning. Form.ai builds agentic systems that mirror these thought process graphs, testing generalizability across scenarios before deploying judges at scale.
- •Consequence mapping methodology: Instead of traditional good-or-bad labeling, Form.ai asks experts to predict outcomes of AI conversations—what emotions users would feel, what actions they would take, what family members would say. This experimental approach provides richer data for training judges and reveals the reasoning behind expert evaluations, particularly valuable for sensitive domains like mental health where clinical nuance matters.
- •Four-tier evaluation framework for political content: Form.ai assesses political AI responses across bias (which breaks into multiple subcategories), factuality, source selection (credibility, balance, accurate attribution), and tone-language (avoiding inflammatory phrasing that mirrors user anger). Different combinations of topic, user intent, and evaluation dimension require separate judges, creating a matrix of specialized evaluators rather than one-size-fits-all assessment.
- •Trust gap blocking AI adoption: KPMG research shows 80-plus percent of people feel optimistic about AI improving their lives, yet only 40-something percent trust it. This delta represents the critical barrier to realizing AI potential. Form.ai addresses this through transparent expert networks published on their website, allowing users to see exactly who shaped model training rather than relying on internal engineering teams or scaled labelers.
- •Dangerous expectation of omniscient AI: ChatGPT established a problematic norm where users expect authoritative answers to any question without context-gathering dialogue. Real expertise requires back-and-forth—doctors do not prescribe after one statement. AI models need awareness of when they lack sufficient context to respond responsibly, but this conflicts with engagement metrics since users migrate to models providing immediate answers over those asking clarifying questions.
- •Mental health scale reveals urgent need: OpenAI reported over one million weekly conversations where users demonstrated suicidal intent, illustrating the massive scale at which people turn to AI for mental health support. Current models lack clinical nuance—the National Eating Disorders Association shut down their Tessa chatbot after it produced adverse effects on users. Form.ai partners with Cleveland Clinic and Mount Sinai to embed medical expertise into health-related AI evaluations.
Notable Moment
Goldfarb describes how a podcaster insisted GPT-3 proved mind-body separation based on controversial Arizona experiments, demonstrating how question phrasing biases LLM responses. The same model gave opposite answers depending on how the question was framed, revealing that probability distributions masquerading as truth engines create dangerous sycophancy problems when users treat outputs as authoritative rather than probabilistic.
Episode Transcript
But our expectation is that I can ask it anything, and it's going to give me an answer. And it's gonna give me an an authoritative answer that I can feel like I can trust. And I think that's a really scary expectation to have set. There was, OpenAI had an announcement a week or two ago where they said that every week they're seeing it was it was, like, over a million conversations where the user demonstrated suicidal intent. That just gives you a sense of where we're headed or the health of information people have, avoiding echo chambers, avoiding misinformation. That's so important to building a healthy society, and AI can go either way on this one. Okay. Robbie, can you start by introducing yourself to listeners and give your background? You you have an interesting background. And, and then give us the, tell us what form.ai is and how you came to found it. Yeah. For sure. So my background is in software engineering and product management. I've spent most of my career at the intersection of AI and trust and safety. I've done quite a bit of work in in education technology. I eventually went to Facebook where I worked on news and misinformation. I was actually I was working on misinformation through the COVID pandemic, through the US twenty twenty elections, through a really interesting period in time, that I think opened my eyes to the potential impact of the technology can have on people, both in a good and a bad way, and maybe more importantly, our ability to shape what those outcomes are. And so that really set me down the path that ultimately, got me working on on forum. I I went though from Facebook to Instagram, where I worked, on similar issues. I also did quite a bit of work on youth safety there. Mhmm. And then most recently was, in Meta's AI Lab, focused on trust and safety more broadly. And then probably about a year ago, I was I was thinking deeply about these issues and the concept of trust in AI. And then, a good friend, Chris Strouhart, who's actually the he's the head of product for Gemini at Google now. I was talking to him about it, and he said, you know, you gotta talk to Campbell. So we connected and and cofounded, Form dot ai together. And and describe what Form dot ai is and how it works. Yeah. So, to put it simply, we scale the world's smartest minds to help evaluate and improve AI systems, particularly in tricky or contentious spaces like health care or politics and news. But to understand what we're doing, our philosophy and our belief is that there is a shift that needs to happen in how AI is being developed, and that experts, credible experts, need to play a bigger role in that process. It's not to say they aren't playing any role in it today, but we believe …
Get the full transcript (9,419 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 60-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Sep 8 · 54 min
David Senra
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Sep 9
More from Eye on AI
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
Sep 3 · 38 min
David Senra
Building Defense Technologies to Protect Democracies | Torsten Reil, Helsing
Aug 26
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Products
company
“Form.ai partners with Cleveland Clinic and Mount Sinai to embed medical expertise into health-related AI evaluations.”
“KPMG research shows 80-plus percent of people feel optimistic about AI improving their lives, yet only 40-something percent trust it.”
“Form.ai partners with Cleveland Clinic and Mount Sinai to embed medical expertise into health-related AI evaluations.”
“Robbie Goldfarb, founder of Form.ai, explains how his company scales expert judgment to evaluate and improve AI systems in contentious domains like healthcare and politics.”
“OpenAI reported over one million weekly conversations where users demonstrated suicidal intent, illustrating the massive scale at which people turn to AI for mental health support.”
“The National Eating Disorders Association shut down their Tessa chatbot after it produced adverse effects on users.”
More from Eye on AI
We summarize every new episode. Want them in your inbox?
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
In 5 to 10 Years, Using Weapons Without AI Will Be Considered Unethical | Yaroslav Azhnyuk, The Fourth Law
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Similar Episodes
Related episodes from other podcasts
David Senra
Sep 9
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
David Senra
Aug 26
Building Defense Technologies to Protect Democracies | Torsten Reil, Helsing
a16z Podcast
Jul 17
Amjad Masad on Going Direct, Building Replit, and the Future of Software
Masters of Scale
Jul 14
The quiet reinvention of a $42b business, with Canva’s Cameron Adams
20VC (20 Minute VC)
Jun 29
20VC: Leo Aschenbrenner's Largest Holding: Inside the $90BN Bloom Energy | Why Electricity, Not AI Models, Will Decide the Winners of the AI Race | Why We Are Not in an AI Capex Bubble | Energy Sovereignty and The Future of Power with KR Sridhar
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime