Skip to main content
NVIDIA AI Podcast

Building AI Factories: How Red Hat and NVIDIA Turn Enterprise Data Into Intelligence - Ep. 293

38 min episode · 2 min read
·
Chris Wright,Justin Boitano

Episode

38 min

Read time

2 min

Topics

Career Growth, Productivity, Investing

AI-Generated Summary

Key Takeaways

  • Five-Layer AI Factory Stack: Structure enterprise AI investment across five distinct layers: data center power and cooling, rack-scale GPU infrastructure, software orchestration (Kubernetes/Linux), model delivery, and agentic applications. Enterprises that skip layers or treat these independently create fragmented environments with high failure rates. Addressing all five layers systematically before scaling reduces costly rework.
  • Hybrid Model Architecture for Cost Reduction: Use frontier models only for the planning stage of agentic workflows, while open models handle search and summarization tasks on-premises. NVIDIA's newer blueprints demonstrate a 30x cost reduction using this split architecture, making enterprise-scale agentic search economically viable for knowledge workers across large organizations.
  • Dev-to-Prod Separation as Non-Negotiable: Separate AI development environments from production data until agents pass functional verification, QA, and penetration testing. Agents must inherit role-based access controls matching the requesting user's existing permissions before promotion to production. This mirrors decades of software engineering best practice and prevents unauthorized data exposure at scale.
  • Evals as a Core Infrastructure Component: Build evaluation frameworks into the AI factory stack from day one, not as an afterthought. Evals measure output quality against defined business outcomes and guide iterative refinement of prompting, data sourcing, and problem scoping. Without continuous evals, enterprises cannot distinguish genuine productivity gains from superficially impressive but low-value demonstrations.
  • Treat Agents as Digital Employees with Least-Privilege Access: As agents operate more autonomously, scope their system permissions the same way contractors receive access — starting with minimum necessary privileges and requiring explicit approval to expand. Connecting agents to business systems often reveals existing user permissions are already over-scoped, making a security team audit a recommended first step.

What It Covers

Red Hat CTO Chris Wright and NVIDIA VP Justin Boitano outline how enterprises build AI factories — five-layer technology stacks converting raw data into business intelligence — covering infrastructure sizing, agentic deployment, security guardrails, and a practical first-ninety-day roadmap toward full-scale AI transformation.

Key Questions Answered

  • Five-Layer AI Factory Stack: Structure enterprise AI investment across five distinct layers: data center power and cooling, rack-scale GPU infrastructure, software orchestration (Kubernetes/Linux), model delivery, and agentic applications. Enterprises that skip layers or treat these independently create fragmented environments with high failure rates. Addressing all five layers systematically before scaling reduces costly rework.
  • Hybrid Model Architecture for Cost Reduction: Use frontier models only for the planning stage of agentic workflows, while open models handle search and summarization tasks on-premises. NVIDIA's newer blueprints demonstrate a 30x cost reduction using this split architecture, making enterprise-scale agentic search economically viable for knowledge workers across large organizations.
  • Dev-to-Prod Separation as Non-Negotiable: Separate AI development environments from production data until agents pass functional verification, QA, and penetration testing. Agents must inherit role-based access controls matching the requesting user's existing permissions before promotion to production. This mirrors decades of software engineering best practice and prevents unauthorized data exposure at scale.
  • Evals as a Core Infrastructure Component: Build evaluation frameworks into the AI factory stack from day one, not as an afterthought. Evals measure output quality against defined business outcomes and guide iterative refinement of prompting, data sourcing, and problem scoping. Without continuous evals, enterprises cannot distinguish genuine productivity gains from superficially impressive but low-value demonstrations.
  • Treat Agents as Digital Employees with Least-Privilege Access: As agents operate more autonomously, scope their system permissions the same way contractors receive access — starting with minimum necessary privileges and requiring explicit approval to expand. Connecting agents to business systems often reveals existing user permissions are already over-scoped, making a security team audit a recommended first step.

Notable Moment

Wright cautions that automating existing enterprise processes without redesigning them first simply produces faster versions of flawed workflows. The real transformation, he argues, involves completely redefining how work gets structured around agents — a shift he places not decades away, but measurable in quarters.

Know someone who'd find this useful?

Episode Transcript

Welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. My guests today are Red Hat's Chris Wright and NVIDIA's Justin Boitano, and we're talking AI factories. Why should enterprises build AI factories, and how can they do so with confidence in building AI factories that they can trust? By way of introductions, and I'll keep it brief because both of these guys' work speaks for itself really. Chris Wright is chief technology officer and senior vice president of global engineering at Red Hat, and Justin Boitano is vice president and general manager of enterprise computing at NVIDIA. Gentlemen, welcome to the NVIDIA AI podcast. Thank you so much for taking the time to join us. Thanks for having us. Thanks for having me, Noah. Let's get right into it. And, Justin, I'll I'll start with you, but always, both of you guys feel free to jump in, you know, as the spirit moves you, so to speak, as we go. But, Justin, why don't we start with you? Can you talk a little bit about well, maybe first give kind of a working definition of what we mean, what you mean when we talk about an AI factory, and then get into kind of at a high level, why would an enterprise be interested? Tangible benefits that an enterprise can expect to see from an AI factory? Sure, Noah. Yeah. I you know, and I think it's important to understand kind of the context of where we are, as an industry. And, you know, building digital intelligence to power the productivity of organizations is going to be as critical, in this decade as, you know, energy in running our companies. This is the next industrial revolution, and companies are always asking us, you know, how do we build, these factories, that basically take data in and then produce the intelligence that, you know, helps them run their their businesses more efficiently. And, so as we as we talk about, like, what is an AI factory, you know, we think of them as really kind of five layers of technology that need to come together. At the base layer, you know, you gotta make sure that you have got the the data centers with power to bring into these factories. You've gotta have, you know, chips, is the is the easy way to talk about it. But we're at this point of building rack scale infrastructure that's, you know, six chips with extreme codesign to to build the best token efficiency from the power available to you. Mhmm. The next layer, you you typically wanna have the software infrastructure to orchestrate everything, and then you wanna have models, that run that intelligence, and then ultimately, the apps and the agents on top. And so, what what every business needs to do though is is take this intelligence, and build, you know, use case specific business outcomes that help them, drive innovation, build products faster, and ultimately, you know, grow …

Get the full transcript (6,811 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all NVIDIA AI Podcast transcripts →

You just read a 3-minute summary of a 35-minute episode.

Get NVIDIA AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • software orchestration (Kubernetes/Linux)
  • software orchestration (Kubernetes/Linux)

More from NVIDIA AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into NVIDIA AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from NVIDIA AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime