AI Agents Fixing Your IT Before You Even Know Something Broke | Erhan Giral & Ryan Manning, BMC Helix
Episode
59 min
Read time
2 min
Topics
Fundraising & VC, Leadership, Sales & Revenue
AI-Generated Summary
Key Takeaways
- ✓Anomaly Detection Pipeline: BMC Helix processes continuous telemetry streams — machine logs, time-series data, alarms — through proprietary ML models to filter noise before passing condensed, causally-described findings to LLMs. This two-stage approach prevents overwhelming generative models with raw data while still enabling precise root cause analysis down to specific network queues, file systems, or brokers.
- ✓Mixture-of-Experts Fine-Tuning: Rather than using generic foundation models, Helix fine-tunes open-weight models using domain-specific datasets and reward functions, capturing training as discrete "experts" with gate activation values. When a mainframe-specific query arrives, only the relevant expert weights activate — keeping models parameter-efficient enough to run on hardware as accessible as NVIDIA RTX 6000 GPUs on-premise.
- ✓Incident Fingerprinting to Combat Automation Bias: Helix builds causal traces of every incident — recording machine signals, human actions, and failure sequences — to identify recurring patterns. When a new issue matches a prior fingerprint, the system flags it with high confidence for automated remediation. This addresses automation bias by grounding approvals in verified historical precedent rather than blind trust.
- ✓Human-in-the-Loop Remediation Workflow: Diagnostic and read-only actions execute automatically, while any configuration changes require human approval. The system tracks whether engineers follow or deviate from AI-generated plans, using deviations as training signals. This flywheel continuously improves recommendations without requiring engineers to manually document their expertise or write runbooks.
- ✓Savings Range 25–50% with Proof-of-Concept Sales Cycle: Enterprise deployments consistently show 25–50% operational savings, but deals always require a proof-of-concept in a pre-production environment — often alongside competing tools. Organizations evaluate over three-to-five year planning horizons, meaning roadmap alignment matters as much as current capability when competing against ServiceNow and Atlassian in the Fortune 2000 market.
What It Covers
Erhan Giral from BMC Helix explains how agentic AI is transforming IT service management by autonomously detecting anomalies, performing root cause analysis, and generating remediation plans for Fortune 2000 enterprises — shifting IT operations from reactive firefighting to automated, self-improving systems that learn from every resolved incident.
Key Questions Answered
- •Anomaly Detection Pipeline: BMC Helix processes continuous telemetry streams — machine logs, time-series data, alarms — through proprietary ML models to filter noise before passing condensed, causally-described findings to LLMs. This two-stage approach prevents overwhelming generative models with raw data while still enabling precise root cause analysis down to specific network queues, file systems, or brokers.
- •Mixture-of-Experts Fine-Tuning: Rather than using generic foundation models, Helix fine-tunes open-weight models using domain-specific datasets and reward functions, capturing training as discrete "experts" with gate activation values. When a mainframe-specific query arrives, only the relevant expert weights activate — keeping models parameter-efficient enough to run on hardware as accessible as NVIDIA RTX 6000 GPUs on-premise.
- •Incident Fingerprinting to Combat Automation Bias: Helix builds causal traces of every incident — recording machine signals, human actions, and failure sequences — to identify recurring patterns. When a new issue matches a prior fingerprint, the system flags it with high confidence for automated remediation. This addresses automation bias by grounding approvals in verified historical precedent rather than blind trust.
- •Human-in-the-Loop Remediation Workflow: Diagnostic and read-only actions execute automatically, while any configuration changes require human approval. The system tracks whether engineers follow or deviate from AI-generated plans, using deviations as training signals. This flywheel continuously improves recommendations without requiring engineers to manually document their expertise or write runbooks.
- •Savings Range 25–50% with Proof-of-Concept Sales Cycle: Enterprise deployments consistently show 25–50% operational savings, but deals always require a proof-of-concept in a pre-production environment — often alongside competing tools. Organizations evaluate over three-to-five year planning horizons, meaning roadmap alignment matters as much as current capability when competing against ServiceNow and Atlassian in the Fortune 2000 market.
Notable Moment
Giral describes building "gym" environments — simulated data centers where a chaos agent deliberately breaks systems while a separate remediation agent attempts fixes, logging every outcome 24/7. This lights-out training loop mirrors how software engineering was automated, but applied directly to IT operations.
Episode Transcript
What's that gonna do to IT stuff? It sounds like it's gonna get increasingly automated. That space is changing pretty quickly, and that workflow is changing agentic AI pretty dramatically. Right now, AI, if you think about it, AI is only learning through someone else's description. Right? People will start exposing these models to the real world more and more so that they can create they can gain that first person view of things. Once we detect something, we always ask the data why why why until we get to a point where we can take an actionable step to mitigate or perhaps even remedy that issue. Why why don't you introduce yourself to listeners? My name is Arhan. I run the, AI office for BMC Helix now. Work with Ryan on applying AI to various different problems in service ops space. So service ops is essentially you can think of that as, any digital business these days is really a conglomeration of, IT IT services you can imagine. And service ops is essentially an attempt to automate operational activities around these services as much as possible, so that we look at enterprises as combination of software, infrastructure, and human beings operating on these, on these assets. And, we provide various, products and now nowadays, agents, to automate different workflows around the around these services. My personal background is in, monitoring space. I've I've done a, early work on monitoring applications and infrastructure and then combinations of them, which means, you know, you have to process a lot of machine data. You have to reason be able to reason about high throughput, you know, streams of data that these data centers and these machines are constantly emitting. So, yeah, that's what I do in Helix. We're gonna talk about, Helix and and, IT service management. Can you begin, one of you, by talking about how the landscape has changed with the introduction of AI agents? What the process for service management was before and and then what it's migrated to today. Yeah. And and, you know, just thinking back in the old days, it was Jira. Right? You would you would open up a Jira ticket. It is, it is does your solution work alongside Jira, replace it? Just for people that aren't that familiar with. Yeah. Right. And and and, you guys, we'll give the background of BMC first. You were talking about that, before we started. Yeah. Yeah. So, I mean, service management the discipline is, is complicated. The the explanation of what they do is pretty simple. It's the front door, to IT. So if, you know, I'm a new employee trying to get my laptop provisioned or if my if I'm having an issue, with an application, I might go to a portal, send an email, call the service desk, and try to get that remediated. And, that that that space is changing pretty quickly and that that workflow is changing, with agentic AI pretty dramatically. …
Get the full transcript (8,588 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 56-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala
Jul 29 · 46 min
NVIDIA AI Podcast
NVIDIA’s Jacob Liberman on the Power of Agentic AI in the Enterprise - Ep. 250
Apr 2
More from Eye on AI
"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala
Jul 21 · 46 min
Cognitive Revolution
1000 Designs a Day: Neural Concept's Thomas von Tschammer on AI-Native Engineering
Jul 1
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Gear
by NVIDIA
“keeping models parameter-efficient enough to run on hardware as accessible as NVIDIA RTX 6000 GPUs on-premise.”
company
by BMC
“Erhan Giral from BMC Helix explains how agentic AI is transforming IT service management by autonomously detecting anomalies, performing root cause analysis, and generating remediation plans for Fortune 2000 enterprises.”
“Organizations evaluate over three-to-five year planning horizons, meaning roadmap alignment matters as much as current capability when competing against ServiceNow and Atlassian in the Fortune 2000 market.”
“Organizations evaluate over three-to-five year planning horizons, meaning roadmap alignment matters as much as current capability when competing against ServiceNow and Atlassian in the Fortune 2000 market.”
More from Eye on AI
We summarize every new episode. Want them in your inbox?
"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala
"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala
6 in 10 Enterprises Can't Find the Root Cause When Their AI Workloads Fail | Paul Appleby, Virtana
Inside the Enterprise Browser Rebuilding Security for the AI Era | Bradon Rogers, Island
What Industrial AI Actually Looks Like | Kriti Sharma, Nexus Black
Similar Episodes
Related episodes from other podcasts
NVIDIA AI Podcast
Apr 2
NVIDIA’s Jacob Liberman on the Power of Agentic AI in the Enterprise - Ep. 250
Cognitive Revolution
Jul 1
1000 Designs a Day: Neural Concept's Thomas von Tschammer on AI-Native Engineering
In Good Company with Nicolai Tangen
Jun 17
Snowflake CEO: Scaling Data, AI Agents and the New Software Era
No Priors: Artificial Intelligence | Technology | Startups
May 11
Amex Global Business Travel: The World’s First AI Take Private with Long Lake CEO Alexander Taubman
No Priors: Artificial Intelligence | Technology | Startups
Apr 17
Scaling Global Organizations in the Age of AI with ServiceNow CEO Bill McDermott
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime