Skip to main content
a16z Podcast

Building the Cloud for an Agentic World | AWS CEO Matt Garman

56 min episode · 2 min read
·
Matt Garman

Episode

56 min

Read time

2 min

Topics

Investing, Startups, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • ✓GPU Allocation Strategy: AWS deliberately reserves GPU capacity for startups despite frontier labs like Anthropic and OpenAI consuming the majority. Currently, AWS fulfills roughly 60% of GPU requests, sometimes in alternate regions or configurations. Startups that become enterprises represent an estimated 30–40% of total AWS revenue, making early capacity access a calculated long-term investment.
  • ✓Agentic Infrastructure Requirements: Agents demand fundamentally different cloud primitives than human developers. Key gaps include lightweight compute sandboxes via Firecracker micro-VMs, time-boxed fine-grained permissions distinct from user IAM roles, and sub-three-second database spin-up times. Tail latencies at p99.9 matter significantly to agents, whereas human users rarely notice them, making performance tuning a priority.
  • ✓Trainium for Inference: Despite its name, Trainium 3 currently delivers AWS's strongest cost-performance ratio for inference workloads, not just training. The majority of Bedrock inference traffic runs on Trainium. Startups evaluating GPU costs should assess Trainium as a primary inference option, with capacity sold out through approximately end of 2026.
  • ✓Enterprise Agent Adoption Barrier: Most enterprise agents remain non-autonomous with humans in the loop. The primary blocker is not capability but trust—specifically the absence of reliable eval systems, production drift detection, and guardrails preventing agents from deleting production databases. AWS runs 45-day FTE engagements to teach enterprises to build evals and labeled datasets independently.
  • ✓Agentic Development Velocity: AWS internal "frontier teams" now operate with agents writing all code while engineers manage agent teams rather than writing code directly. This model has measurably accelerated new product deployment rates. Team sizes for product ownership have shrunk from roughly 10 engineers to 3–4, with faster project rotation replacing long-term feature ownership structures.

What It Covers

AWS CEO Matt Garman speaks with a16z about the infrastructure demands of an AI-driven world, covering how AWS allocates scarce GPU capacity across frontier labs and startups, the shift to agentic cloud architecture, Trainium chip performance, $220 billion in 2026 capital expenditure, and how agents are reshaping internal software development velocity.

Key Questions Answered

  • •GPU Allocation Strategy: AWS deliberately reserves GPU capacity for startups despite frontier labs like Anthropic and OpenAI consuming the majority. Currently, AWS fulfills roughly 60% of GPU requests, sometimes in alternate regions or configurations. Startups that become enterprises represent an estimated 30–40% of total AWS revenue, making early capacity access a calculated long-term investment.
  • •Agentic Infrastructure Requirements: Agents demand fundamentally different cloud primitives than human developers. Key gaps include lightweight compute sandboxes via Firecracker micro-VMs, time-boxed fine-grained permissions distinct from user IAM roles, and sub-three-second database spin-up times. Tail latencies at p99.9 matter significantly to agents, whereas human users rarely notice them, making performance tuning a priority.
  • •Trainium for Inference: Despite its name, Trainium 3 currently delivers AWS's strongest cost-performance ratio for inference workloads, not just training. The majority of Bedrock inference traffic runs on Trainium. Startups evaluating GPU costs should assess Trainium as a primary inference option, with capacity sold out through approximately end of 2026.
  • •Enterprise Agent Adoption Barrier: Most enterprise agents remain non-autonomous with humans in the loop. The primary blocker is not capability but trust—specifically the absence of reliable eval systems, production drift detection, and guardrails preventing agents from deleting production databases. AWS runs 45-day FTE engagements to teach enterprises to build evals and labeled datasets independently.
  • •Agentic Development Velocity: AWS internal "frontier teams" now operate with agents writing all code while engineers manage agent teams rather than writing code directly. This model has measurably accelerated new product deployment rates. Team sizes for product ownership have shrunk from roughly 10 engineers to 3–4, with faster project rotation replacing long-term feature ownership structures.

Notable Moment

Garman revealed that one county where AWS operates sees residents paying $5,000 less annually in taxes due to AWS's local tax contributions—yet the community remains unaware. He argued the cloud industry broadly fails to communicate these concrete economic benefits, allowing a few bad actors to define public perception.

Know someone who'd find this useful?

Episode Transcript

Agentic workflows tend to perform better on AWS than anywhere else. Compute sandboxes, gateways, agent permissions versus people. A lot of those are things that we have built and are building and thinking actively about. The top frontier labs gobble up every available GPUs. At the same time, you want to promote newer companies that are gonna become the enterprises of tomorrow. How are you thinking about balancing? From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do. These are the innovators that are at the edge of technology, understanding what's possible. We're very intentional about keeping capacity available for the startups. We recently announced we're gonna be buying 2,000,000 NVIDIA GPUs over the next couple of years. Your CapEx is same as what? 200 or something? $220,000,000,000 for '26. And we don't anticipate slowing down anytime soon because the demand is just massive. With all the debate around AI extension risk and the hugging face attack, what are CEOs asking you about all these things? Amazon plans to spend $220,000,000,000 in capital this year, and AWS CEO Matt Garman says they don't anticipate slowing down anytime soon. In this episode, a sixteen z's Raghuraguram sits down with Matt to unpack what's driving one of the largest infrastructure build outs in history and how AI is changing the cloud itself. They discuss how AWS decides who gets scarce GPU capacity. They're shifting bottlenecks from power and chips to memory and construction, and the company's bet on custom silicon like Tranium. They also get into what changes when agents increasingly write code and manage infrastructure. And that shares what AWS is seeing inside its own teams, where agents now write code and engineers increasingly manage teams of agents, accelerating how quickly new products can be built. Welcome to the pod, Matt. What a time we are living in, so I have a lot of topics to talk to you about. Awesome. Thanks for having me. I'm excited. Yeah. Absolutely. So let's start, actually, right from the beginning. So you were the first GM for EC two, and that was 2006, right? And today you guys are what, 160, 170,000,000,000 in revenue? Yeah, about $169,170,000,000,000, yep. Yeah, growing thirty Thirty seven. 7%. Yep. That's 37% at 169. Yeah. Crazy. There's a ton of... It's interesting to think about it from day one when we got the first dollar of revenue. But Yeah. Yeah. The interesting thing is we're still at the early stages of what the business can be, and what the opportunity is for customers. Most workloads, there's a huge amount of workloads still that live on prem today, and the amount of compute that people are doing every single day is more than it was the day before. And so, you see the tailwind from AI, you see the tailwind from migration into the cloud, and the business has grown really fast, and it's been a super fun thing …

Get the full transcript (11,803 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 53-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime