Skip to main content
Eye on AI

#308 Christopher Bergey: How Arm Enables AI to Run Directly on Devices

51 min episode · 2 min read
·
Christopher Bergey

Episode

51 min

Read time

2 min

Topics

Productivity, Startups, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Heterogeneous Computing Architecture: Arm devices combine CPUs, GPUs, and NPUs in single SoCs, dynamically moving AI workloads between processors based on latency, performance, and power requirements, with jobs typically starting on CPU before routing to specialized accelerators.
  • Big-Little Power Management: Arm's architecture switches workloads between high-performance and low-power CPU cores, firing up computing elements only when triggered by events like motion detection, enabling devices like Meta's wristband to run AI for weeks on tiny batteries.
  • Memory Bandwidth Bottleneck: AI performance at the edge depends more on memory bandwidth and size than raw computing power. Integrated SoCs with unified memory systems up to 128GB outperform discrete solutions that split memory, making integration critical for edge AI.
  • Developer Ecosystem Scale: Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux without requiring specialized accelerator programming languages like CUDA.

What It Covers

Christopher Bergey explains how Arm's v9 architecture with scalable matrix extensions enables AI inference directly on edge devices like smartphones, wearables, and IoT products, balancing performance, power efficiency, and memory constraints.

Key Questions Answered

  • Heterogeneous Computing Architecture: Arm devices combine CPUs, GPUs, and NPUs in single SoCs, dynamically moving AI workloads between processors based on latency, performance, and power requirements, with jobs typically starting on CPU before routing to specialized accelerators.
  • Big-Little Power Management: Arm's architecture switches workloads between high-performance and low-power CPU cores, firing up computing elements only when triggered by events like motion detection, enabling devices like Meta's wristband to run AI for weeks on tiny batteries.
  • Memory Bandwidth Bottleneck: AI performance at the edge depends more on memory bandwidth and size than raw computing power. Integrated SoCs with unified memory systems up to 128GB outperform discrete solutions that split memory, making integration critical for edge AI.
  • Developer Ecosystem Scale: Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux without requiring specialized accelerator programming languages like CUDA.

Notable Moment

Bergey predicts AI will become as fundamental as touchscreens within a decade. Children who expect every screen to respond to touch will soon expect every device to understand natural language and anticipate their needs without manual configuration.

Know someone who'd find this useful?

Episode Transcript

The ARM architecture has existed now for almost thirty years, and that started from early investments from companies like Apple, early adopters of the ARM architecture that made it the stalwart that it is, where companies like Nintendo, companies like Nokia back in the eighties and nineties as these smartphone revolution, all that kind of stuff really started around ARM, and then that's what's driven us to be where we are today. We have big CPUs and little CPUs, and we're actually moving the workloads back and forth because certain times you need the performance, certain times you don't. And so that's really the way these devices work where you to your doorbell example, you're looking for maybe some motion. You're looking for something. And then once you trigger that event, okay. Now let's fire up some more computing elements. In business, they say you can have better, cheaper, or faster, but you only get to pick two. What if you could have all three at the same time? That's exactly what Coher, Thompson Reuters, and Specialized Bikes have. Since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 percent less for compute, 70% less for storage, and 80% less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com/ionai.i on AI all run together, eyeonai. That's oracle.com/ionai. So it's great to be here, Craig. Thanks for inviting me. So my name is Chris Bergey. I'm the senior vice president general manager of the client line of business at Arm. And that means I very much focus on all of the rich rich edge devices that Arm, is so prominent in. So, things like smartphones, but, also, we are obviously making quite a bit of inroads into things like PCs. We participate all parts of your house, whether that's your TVs, your smart speakers, all kind of rich endpoints that that you have powered and and are quickly becoming AI enabled. And I think that's what we're gonna talk about today, Craig. Just a little bit about my background. I've spent almost thirty years in semiconductors. Various big companies started out of school at AMD and and and moved to Broadcom for almost a decade and and also did some startups in between. So, along it I can't believe it's gone by this this quickly, but, I guess, you know, semiconductors have never been this cool as, it seems like governments and, everyone really cares about semiconductors. So …

Get the full transcript (8,245 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Eye on AI transcripts →

You just read a 3-minute summary of a 48-minute episode.

Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Tools

  • Arm supports 22 million software developers through frameworks like Clidy that abstract hardware complexity, enabling AI applications to run seamlessly across iOS, Android, Windows, and Linux

Gear

  • by Arm

    Christopher Bergey explains how Arm's v9 architecture with scalable matrix extensions enables AI inference directly on edge devices like smartphones, wearables, and IoT products

More from Eye on AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Eye on AI.

Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime