#308 Christopher Bergey: How Arm Enables AI to Run Directly on Devices
Episode
51 min
Read time
2 min
Topics
Productivity, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Heterogeneous Computing Architecture: Arm devices combine CPUs, GPUs, and NPUs in single SoCs, dynamically moving AI workloads between processors based on latency, performance, and power requirements, with jobs typically starting on CPU before routing to specialized accelerators.
- ✓Big-Little Power Management: Arm's architecture switches workloads between high-performance and low-power CPU cores, firing up computing elements only when triggered by events like motion detection, enabling devices like Meta's wristband to run AI for weeks on tiny batteries.
- ✓Memory Bandwidth Bottleneck: AI performance at the edge depends more on memory bandwidth and size than raw computing power. Integrated SoCs with unified memory systems up to 128GB outperform discrete solutions that split memory, making integration critical for edge AI.
- ✓Developer Ecosystem Scale: Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux without requiring specialized accelerator programming languages like CUDA.
What It Covers
Christopher Bergey explains how Arm's v9 architecture with scalable matrix extensions enables AI inference directly on edge devices like smartphones, wearables, and IoT products, balancing performance, power efficiency, and memory constraints.
Key Questions Answered
- •Heterogeneous Computing Architecture: Arm devices combine CPUs, GPUs, and NPUs in single SoCs, dynamically moving AI workloads between processors based on latency, performance, and power requirements, with jobs typically starting on CPU before routing to specialized accelerators.
- •Big-Little Power Management: Arm's architecture switches workloads between high-performance and low-power CPU cores, firing up computing elements only when triggered by events like motion detection, enabling devices like Meta's wristband to run AI for weeks on tiny batteries.
- •Memory Bandwidth Bottleneck: AI performance at the edge depends more on memory bandwidth and size than raw computing power. Integrated SoCs with unified memory systems up to 128GB outperform discrete solutions that split memory, making integration critical for edge AI.
- •Developer Ecosystem Scale: Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux without requiring specialized accelerator programming languages like CUDA.
Notable Moment
Bergey predicts AI will become as fundamental as touchscreens within a decade. Children who expect every screen to respond to touch will soon expect every device to understand natural language and anticipate their needs without manual configuration.
Episode Transcript
The ARM architecture has existed now for almost thirty years, and that started from early investments from companies like Apple, early adopters of the ARM architecture that made it the stalwart that it is, where companies like Nintendo, companies like Nokia back in the eighties and nineties as these smartphone revolution, all that kind of stuff really started around ARM, and then that's what's driven us to be where we are today. We have big CPUs and little CPUs, and we're actually moving the workloads back and forth because certain times you need the performance, certain times you don't. And so that's really the way these devices work where you to your doorbell example, you're looking for maybe some motion. You're looking for something. And then once you trigger that event, okay. Now let's fire up some more computing elements. In business, they say you can have better, cheaper, or faster, but you only get to pick two. What if you could have all three at the same time? That's exactly what Coher, Thompson Reuters, and Specialized Bikes have. Since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 percent less for compute, 70% less for storage, and 80% less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com/ionai.i on AI all run together, eyeonai. That's oracle.com/ionai. So it's great to be here, Craig. Thanks for inviting me. So my name is Chris Bergey. I'm the senior vice president general manager of the client line of business at Arm. And that means I very much focus on all of the rich rich edge devices that Arm, is so prominent in. So, things like smartphones, but, also, we are obviously making quite a bit of inroads into things like PCs. We participate all parts of your house, whether that's your TVs, your smart speakers, all kind of rich endpoints that that you have powered and and are quickly becoming AI enabled. And I think that's what we're gonna talk about today, Craig. Just a little bit about my background. I've spent almost thirty years in semiconductors. Various big companies started out of school at AMD and and and moved to Broadcom for almost a decade and and also did some startups in between. So, along it I can't believe it's gone by this this quickly, but, I guess, you know, semiconductors have never been this cool as, it seems like governments and, everyone really cares about semiconductors. So …
Get the full transcript (8,245 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 48-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Sep 8 · 54 min
Odd Lots
Why Laser Beams Are the Hottest New Tech in Defense
Sep 4
More from Eye on AI
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
Sep 3 · 38 min
The Rich Roll Podcast
Hunter Biden Is Out Of Secrets: A Story of Recovery, Exposure & Amends
Aug 31
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
“Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux”
Gear
by Arm
“Christopher Bergey explains how Arm's v9 architecture with scalable matrix extensions enables AI inference directly on edge devices like smartphones, wearables, and IoT products”
More from Eye on AI
We summarize every new episode. Want them in your inbox?
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
In 5 to 10 Years, Using Weapons Without AI Will Be Considered Unethical | Yaroslav Azhnyuk, The Fourth Law
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Similar Episodes
Related episodes from other podcasts
Odd Lots
Sep 4
Why Laser Beams Are the Hottest New Tech in Defense
The Rich Roll Podcast
Aug 31
Hunter Biden Is Out Of Secrets: A Story of Recovery, Exposure & Amends
Software Engineering Daily
Aug 18
How LLMs Are Reshaping Recommendation Systems
This Week in Startups
Aug 17
Bittensor creator Const on Affine, dTAO, "mining reasoning," and more | E2326
Latent Space
Jul 28
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime