Skip to main content
The Life Science Rundown

Getting Data Governance for Regulatory Submissions Right Before AI Gets it Wrong with Cary Smithson

28 min episode · 2 min read
·
Cary Smithson

Episode

28 min

Read time

2 min

Topics

Health & Wellness, Artificial Intelligence, Software Development

AI-Generated Summary

Key Takeaways

  • Scope-First Governance: Avoid attempting to govern all data simultaneously. Identify the highest-pain critical data elements (CDEs) — such as product master, substance, dosage form, and manufacturing site data — within a narrow scope first. Demonstrate measurable value quickly to secure ongoing budget, then scale systematically across clinical, regulatory, quality, and manufacturing domains.
  • Business-Owned Stewardship Model: Assign data ownership to business domain experts, not IT. Establish a formal data governance council with named stewards per domain who hold accountability for data exchange, field changes, standards updates, and approval workflows. Tie steward compliance to performance metrics like submission cycle time and right-first-time rates to drive adoption.
  • AI Reliability Depends on Governed Inputs: AI models trained on nonstandard or low-quality data produce erroneous, noncompliant, or unexplainable outputs. A midsize biopharma implementing IDMP-aligned governance achieved 30–50% faster affiliate submissions, 25% fewer health authority queries, and successfully deployed AI-assisted CMC authoring — all within twelve months of launching their data governance program.
  • Regulatory Framework Crosswalk: Maintain a living crosswalk mapping enterprise master data to regulatory data models including RIM and eCTD, with full traceability from source systems to submission artifacts. For IDMP and SPORE, harmonize product and substance attributes across systems. For PQCMC, build CMC data models reflecting manufacturing process parameters, control strategies, and analytical methods fed directly from governed sources.
  • Governance as Persistent Program, Not Project: Embed data governance into existing SOPs, change control processes, system validation (SDLC), and training workflows rather than treating it as a standalone initiative. Monitor health via KPIs including data quality scores, lineage completeness, issue remediation SLA adherence, and submission right-first-time rates, with a continuous improvement loop managed by the governance council.

What It Covers

Cary Smithson of Leap Ahead Solutions outlines how life science companies must establish structured data governance frameworks to meet regulatory submission standards like IDMP, PQCMC, and HL7 FHIR, while building the data foundation required for safe, compliant AI adoption across R&D, regulatory, and quality functions.

Key Questions Answered

  • Scope-First Governance: Avoid attempting to govern all data simultaneously. Identify the highest-pain critical data elements (CDEs) — such as product master, substance, dosage form, and manufacturing site data — within a narrow scope first. Demonstrate measurable value quickly to secure ongoing budget, then scale systematically across clinical, regulatory, quality, and manufacturing domains.
  • Business-Owned Stewardship Model: Assign data ownership to business domain experts, not IT. Establish a formal data governance council with named stewards per domain who hold accountability for data exchange, field changes, standards updates, and approval workflows. Tie steward compliance to performance metrics like submission cycle time and right-first-time rates to drive adoption.
  • AI Reliability Depends on Governed Inputs: AI models trained on nonstandard or low-quality data produce erroneous, noncompliant, or unexplainable outputs. A midsize biopharma implementing IDMP-aligned governance achieved 30–50% faster affiliate submissions, 25% fewer health authority queries, and successfully deployed AI-assisted CMC authoring — all within twelve months of launching their data governance program.
  • Regulatory Framework Crosswalk: Maintain a living crosswalk mapping enterprise master data to regulatory data models including RIM and eCTD, with full traceability from source systems to submission artifacts. For IDMP and SPORE, harmonize product and substance attributes across systems. For PQCMC, build CMC data models reflecting manufacturing process parameters, control strategies, and analytical methods fed directly from governed sources.
  • Governance as Persistent Program, Not Project: Embed data governance into existing SOPs, change control processes, system validation (SDLC), and training workflows rather than treating it as a standalone initiative. Monitor health via KPIs including data quality scores, lineage completeness, issue remediation SLA adherence, and submission right-first-time rates, with a continuous improvement loop managed by the governance council.

Notable Moment

Smithson highlights a counterintuitive risk with AI: failures may go entirely undetected. When AI models ingest poor-quality or nonstandard data, they can generate misleading results that appear valid, making invisible errors potentially more dangerous than obvious ones in regulated submission environments.

Know someone who'd find this useful?

Episode Transcript

Hello, everyone. Welcome to the Life Science Rundown, the podcast where we discuss, regulatory complexities facing the life science industry and explore innovative ways to overcome those challenges. I am your host, Nicholas Catman, president and CEO of the FDA group. Before we get started, here's a quick word about who we are. The FDA group is a technology and consulting firm that helps life science companies in the areas of regulatory submissions, audit projects, mock inspection, staff augmentation, and remediation. So if you ever find yourself in need, just head over to the feagroup.com to check us out and get in touch. So today, I am speaking with Carrie Smithson. Hey, Carrie. How are you doing today? Fine. How are you? I'm doing excellent. Thank you. Thank you very much for joining. And, before we begin, would you kindly introduce yourself? Sure. Yeah. Hi, everyone. I'm Carrie Smithson. I'm managing partner and owner of Leap Ahead Solutions, and I have many years in life sciences industry helping companies, you know, automate their regulatory submissions from the days when they were doing paper, trial master files, and all the quality document management, doing a lot around regulatory compliance. And then as they move to more structured data, which we'll talk about a little bit more, streamlining their regulatory information management data, more recently adding AI to their business processes to further automate what they're doing, and then better govern their data and AI across r and d and quality. Okay. Excellent. Thanks, Nicholas. So today, we're gonna talk about how can life science companies ensure their data is well governed and well managed to meet structured regulatory submission requirements, accelerate regulatory approve regulatory approvals and interoperability, and fully leverage analytics and AI to improve business performance? So that's a mouthful, but I think we're gonna be able to break that down. So let's start at the top. So why is data governance becoming increasingly critical in life sciences? Sure. Well, I see three main reasons. So first, regulatory expectations are becoming more structured. You know, submission data is no longer just documents. It's structured attributes, controlled vocabularies, and traceable lineage. And, poor governance can lead to inconsistencies and rework, and queries from the health authorities that just extend the process to getting your products approved. Then there's also interoperability across the value chain. So all your various functions from early stages of r and d all the way through getting your regulatory approval and then manufacturing and quality, or need consistent and high quality master data, and that's not possible without governance. And then thirdly, AI and analytics, depend on trustworthy data. So without clear definitions, controls, provenance, AI models produce unreliable or noncompliant outputs, and governance enables explainability and validation, which are essential in the regulated context of which we operate. Okay. So would you say there's urgency around this? And if so, what are the trends that are driving that? Sure. Yeah. I'd say it's, you know, it's …

Get the full transcript (4,331 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The Life Science Rundown transcripts →

You just read a 3-minute summary of a 25-minute episode.

Get The Life Science Rundown summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from The Life Science Rundown

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Science Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The Life Science Rundown.

Every Monday, we deliver AI summaries of the latest episodes from The Life Science Rundown and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime