Skip to main content
NVIDIA AI Podcast

Powering the AI Inference Wave with EPRI's Ben Sooter - Ep. 292

32 min episode · 2 min read
·
Epri's Ben Sooter

Episode

32 min

Read time

2 min

Topics

Remote Work, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Inference vs. Training Power Split: Over the lifetime of an AI model, roughly 80–90% of total compute consumption occurs during inference, not training. Organizations planning energy infrastructure should weight capacity planning heavily toward inference workloads, not just the headline-grabbing training buildouts that currently dominate industry conversation and capital allocation.
  • Substation Co-location Strategy: Thousands of existing electrical substations across the US carry underutilized capacity, typically 3–10 megawatts of headroom. Siting micro data centers directly adjacent to these substations bypasses transmission interconnection queues, reduces permitting timelines, and avoids new steel-in-ground costs — accelerating time-to-power for inference deployments by a meaningful margin.
  • Distributed Aggregation Model: A single substation's 5-megawatt surplus may be insufficient for viable data center economics. EPRI's approach aggregates five nearby sites within a metro region into a 25-megawatt distributed cluster, treating it as one logical project — satisfying both utility grid constraints and the minimum scale thresholds that data center operators require to justify investment.
  • Demand Flexibility as Grid Asset: Substations often hold significantly more capacity than their rated surplus, except during annual peak demand days. Pairing micro data centers with battery storage and backup generation, then engineering load-shedding protocols that reroute compute tasks to other nodes during grid peaks, unlocks that larger envelope without requiring additional grid upgrades.
  • Agentic AI Reshapes Load Forecasting: Early inference load models assumed human-driven usage patterns — daytime peaks, overnight lows. The rapid emergence of autonomous AI agents that run continuous background tasks around the clock invalidates that assumption. Grid planners and data center operators should model inference loads as potentially flat or inverted curves, not standard residential demand profiles.

What It Covers

EPRI's Ben Sooter explains how micro data centers — small, distributed inference facilities of 3–20 megawatts — can be co-located at underutilized electrical substations across the US to meet the coming wave of AI inference demand without overloading transmission grids or requiring new infrastructure investment.

Key Questions Answered

  • Inference vs. Training Power Split: Over the lifetime of an AI model, roughly 80–90% of total compute consumption occurs during inference, not training. Organizations planning energy infrastructure should weight capacity planning heavily toward inference workloads, not just the headline-grabbing training buildouts that currently dominate industry conversation and capital allocation.
  • Substation Co-location Strategy: Thousands of existing electrical substations across the US carry underutilized capacity, typically 3–10 megawatts of headroom. Siting micro data centers directly adjacent to these substations bypasses transmission interconnection queues, reduces permitting timelines, and avoids new steel-in-ground costs — accelerating time-to-power for inference deployments by a meaningful margin.
  • Distributed Aggregation Model: A single substation's 5-megawatt surplus may be insufficient for viable data center economics. EPRI's approach aggregates five nearby sites within a metro region into a 25-megawatt distributed cluster, treating it as one logical project — satisfying both utility grid constraints and the minimum scale thresholds that data center operators require to justify investment.
  • Demand Flexibility as Grid Asset: Substations often hold significantly more capacity than their rated surplus, except during annual peak demand days. Pairing micro data centers with battery storage and backup generation, then engineering load-shedding protocols that reroute compute tasks to other nodes during grid peaks, unlocks that larger envelope without requiring additional grid upgrades.
  • Agentic AI Reshapes Load Forecasting: Early inference load models assumed human-driven usage patterns — daytime peaks, overnight lows. The rapid emergence of autonomous AI agents that run continuous background tasks around the clock invalidates that assumption. Grid planners and data center operators should model inference loads as potentially flat or inverted curves, not standard residential demand profiles.

Notable Moment

Sooter reveals that his original assumption — inference loads would mirror human daily activity patterns and smooth out naturally — collapsed almost immediately when he recognized that agentic AI systems operate continuously overnight, forcing a complete revision of EPRI's load modeling approach before field measurements even began.

Know someone who'd find this useful?

Episode Transcript

Welcome to the NVIDIA AI podcast. I'm Noah Kravitz. Today, we're talking microdata centers with Ben Suter, director of r and d at EPRI, the Electric Power Research Institute. The relationship between AI data centers and energy grids is an increasingly important one, to say the least. In a moment, we'll talk about how micro data centers can help strengthen that relationship. But first, a quick note about GTC San Jose. Join us at the world's premier AI conference. GTC San Jose is online and in person, March 16 through the nineteenth. From physical AI and AI factories to agentic AI and inference, GTC twenty twenty six will showcase the breakthrough shaping every industry. Learn more and register at nvidia.com/gtc. Ben Suter, welcome. Thank you so much for taking the time to join the NVIDIA AI podcast. Really glad to have you here. Yeah. Great to be here, Noah. Super excited. So, Ben, to to kinda set the table before we dive in, for listeners who don't know EPRI, can you briefly explain, well, first, who you are and what you do, and as part of that, what EPRI is and what EPRI does? Yeah. Absolutely. So EPRI is a is a sort of a unique organization. We're a a five zero one c three not for profit. It's an independent institute, focusing on r and d, collaborates with more than 400 companies across more than 40 countries, and really drives innovation to ensure sort of the public has reliable and affordable energy. So really awesome mission statement and been a really exciting place to work. You've been at EPRI for a couple of decades now. Yeah. I've been here a while. I just crossed over the 20 mark, which I I feel like is is ancient times in the way the corporate world works now. Right. Well, congratulations. And, you kinda this is exactly why I asked because I'm I'm thinking about data centers and and AI, but hearing you talk about things like nuclear and thinking, like, man, you, like, must have seen some things and worked on some projects. And thinking back, you know, over twenty years and how technology and, energy reliance and consumption must have evolved. I don't know if this is a fair question to ask, but can you kind of place our current moment in context to, you know, sort of what you've seen, with how the world uses energy and stuff you've worked on over the years? Yeah. Yeah. That's a that's a a great question, a great way to frame it. Because it and it it gets to why I probably have stayed here twenty years, which is that there's just so much, you know, change and and a lot of different exciting things that have evolved across the sector in the industry, that have have sort of landed us here today. Yeah. So so lot lots of stuff going on. It's been interesting as you come in, and I'm I'm …

Get the full transcript (5,862 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all NVIDIA AI Podcast transcripts →

You just read a 3-minute summary of a 29-minute episode.

Get NVIDIA AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from NVIDIA AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into NVIDIA AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from NVIDIA AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime