Skip to main content
Eye on AI

6 in 10 Enterprises Can't Find the Root Cause When Their AI Workloads Fail | Paul Appleby, Virtana

44 min episode · 2 min read
·
Paul Appleby

Episode

44 min

Read time

2 min

Topics

Health & Wellness, Investing, Startups

AI-Generated Summary

Key Takeaways

  • AI Workload Failure Rate: 6 in 10 enterprises lack automated root cause identification across AI infrastructure domains when workloads fail. Without knowing causality, remediation becomes guesswork, leaving GPU compute sitting idle and wasting capital on infrastructure that cannot deliver consistent throughput or meet production-grade reliability requirements at scale.
  • Governance Gap Risk: Enterprise AI infrastructure investment is outpacing governance and controls deployment. IT operators report increasing risk exposure while business executives push for faster AI adoption. Companies should audit whether existing governance frameworks applied to legacy infrastructure have been replicated in new AI data center environments before scaling further.
  • GPU Utilization Economics: Falling token costs do not reduce total AI spend — they accelerate token consumption as more agents deploy. Enterprises should measure GPU utilization rates alongside token costs, since throttled workloads and idle GPUs represent direct ROI loss on hardware investments that can reach hundreds of millions of dollars.
  • Observability Architecture Requirement: Effective AI factory management requires capturing telemetry across every stack layer simultaneously — data pipelines, AI orchestration, compute, network, and storage — correlating 20,000-plus metrics at sub-second intervals. Fragmented monitoring across four to six separate tools prevents real-time causality identification and blocks autonomous remediation from functioning reliably.
  • Executive-Practitioner Disconnect: Study data shows a measurable gap between c-suite AI confidence and IT operator risk assessment. Senior IT operations leaders now report infrastructure resilience metrics weekly to CEOs rather than annually. Enterprises should establish shared ROI metrics combining infrastructure cost data with business outcomes before committing to further AI scaling.

What It Covers

Virtana CEO Paul Appleby presents findings from the AI Factory Reality Check study, revealing that 6 in 10 enterprises cannot automatically identify root causes when AI workloads fail, while governance investment lags behind infrastructure spending across banking, healthcare, airlines, and retail sectors deploying on-premises AI data centers.

Key Questions Answered

  • AI Workload Failure Rate: 6 in 10 enterprises lack automated root cause identification across AI infrastructure domains when workloads fail. Without knowing causality, remediation becomes guesswork, leaving GPU compute sitting idle and wasting capital on infrastructure that cannot deliver consistent throughput or meet production-grade reliability requirements at scale.
  • Governance Gap Risk: Enterprise AI infrastructure investment is outpacing governance and controls deployment. IT operators report increasing risk exposure while business executives push for faster AI adoption. Companies should audit whether existing governance frameworks applied to legacy infrastructure have been replicated in new AI data center environments before scaling further.
  • GPU Utilization Economics: Falling token costs do not reduce total AI spend — they accelerate token consumption as more agents deploy. Enterprises should measure GPU utilization rates alongside token costs, since throttled workloads and idle GPUs represent direct ROI loss on hardware investments that can reach hundreds of millions of dollars.
  • Observability Architecture Requirement: Effective AI factory management requires capturing telemetry across every stack layer simultaneously — data pipelines, AI orchestration, compute, network, and storage — correlating 20,000-plus metrics at sub-second intervals. Fragmented monitoring across four to six separate tools prevents real-time causality identification and blocks autonomous remediation from functioning reliably.
  • Executive-Practitioner Disconnect: Study data shows a measurable gap between c-suite AI confidence and IT operator risk assessment. Senior IT operations leaders now report infrastructure resilience metrics weekly to CEOs rather than annually. Enterprises should establish shared ROI metrics combining infrastructure cost data with business outcomes before committing to further AI scaling.

Notable Moment

Appleby describes visiting a large US enterprise where the senior VP of IT operations shifted from annual to weekly CEO reporting on infrastructure resilience — a detail that illustrates how AI infrastructure risk has escalated from a back-office concern to a board-level priority within just a few years.

Know someone who'd find this useful?

Episode Transcript

Six and ten enterprises cannot automatically identify the root cause across AI infrastructure domains when AI workloads fail. Why root cause is so much harder in an AI factory than in traditional operational efficiency, ergo, cost. And business impact. You know, it might not just be a cost dynamic. The kind of gold rush to be in the AI investment cycle is overriding some of the good governance principles that companies have. We'll see the equilibrium come back. So let's start by having you introduce yourself to listeners. Thanks, Craig. Yeah. I'm Paul Appleby, the CEO of Vertana, and we're a company that exists really for for one really important purpose where, you know, companies across many industries, whether it's banking, telecommunications, health care, retail, airlines, whatever, are so reliant on their technology. In fact, you know, a lot of chief risk officers say that the single biggest point of catastrophic risk of failure for a business now is their technology and infrastructure. Vitan is in the world of trying to protect that. We live in this this world called observability, and that's all about business resilience, and operational efficiency. How do we make sure those services that are critical to your business and your customers stay performant and available? And even more relevant in this era of accelerated AI adoption, of course. Yeah. And how long have you been with Vertana, and and what was your background before coming? I've been here now for a couple of years. Inherited, an amazing business that, in the past, in fact, under, its prior name of Virtual Instruments, was run by John Thompson, the the former chairman of Microsoft. So the company's been around for some time and has, an amazing history with its core technology. But I came to the company two years ago, with a charter from the board to really lean into this this world of, you know, the broad digitization of services and the scale out adoption of AI. Prior to that, I'd I've been working in technology for years, sometimes in startups because I love, you know, that whole idea of the scale up. But sometimes in much larger companies, you know, trend, through through phases of transformation and growth like Salesforce, where, where I was for a number of years, and also, companies like Elasticsearch, more recently as their president. So I've got a a long background in in enterprise software and enterprise technology, and growth and and scale of businesses. Yeah. In Verintana, we were talking before we started recording is an observability platform Yes. For for, large scale operations or tech technology, operations, not necessarily AI. I mean, it it existed before the the the generative AI boom. Is that right? Yeah. I I think it's a a couple of really important things to say. I mean, this class of software, Craig's been around for a long time. For as long as as, you know, technology's been supporting critical business services, there's been a …

Get the full transcript (6,296 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Eye on AI transcripts →

You just read a 3-minute summary of a 41-minute episode.

Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

other

  • by Virtana

    Virtana CEO Paul Appleby presents findings from the AI Factory Reality Check study, revealing that 6 in 10 enterprises cannot automatically identify root causes when AI workloads fail

More from Eye on AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Health & Longevity Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Eye on AI.

Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime