Skip to main content
Latent Space

The Professor of Outputmaxxing — Anjney Midha, AMP

59 min episode · 2 min read
·
Anjney Midha

Episode

59 min

Read time

2 min

Topics

Productivity, Relationships, Startups

AI-Generated Summary

Key Takeaways

  • GPU Cluster Utilization Benchmarks: Node utilization below 96% in AI clusters is indefensible — Google treated anything under 95% as an outage. MFU (model flop utilization) best-in-class sits at 60–70%. Most singleton clusters fall short of both metrics due to misalignment between capital providers and the engineers actually managing infrastructure at scale.
  • Community-Aligned Data Center Pricing: Up to 20% of US data centers face community backlash risk in the current year. A concrete mitigation: charge $4.50 per compute-hour instead of $4.00, routing the $0.50 margin directly to local communities as cash or electricity bill reductions. This converts opposition into partnership without requiring regulatory intervention.
  • Independent System Operator Model for Compute: Rather than owning assets, AMP pools demand from frontier labs and supply from trusted 20-plus-year data center operators at 1.3 gigawatt base load, targeting 6 gigawatts over four years. Labs receive guaranteed base load with flexible spike capacity — mirroring how PJM Interconnect coordinates uncorrelated industrial demand across the Northeast US grid.
  • Culture as a Fragile Daily Practice, Not a Moat: Anthropic's coding dominance traces directly to four years of resource scarcity forcing precise prioritization — coding as the singular P-zero toward AGI. Teams flush with capital skip this definition process. Culture requires daily action-based reinforcement; without hardship forcing trade-off clarity, lab cultures become brittle before reaching capability takeoff.
  • Trust Boundary as the Primary Chip Company Risk: Hardware startups like Matx adopt NVIDIA's open reference architecture to eliminate data center integration battles, focusing innovation on systems co-design at the logic tile level. The real bottleneck is trust access — chip tape-out cycles run two years, so without early visibility into next-generation model architectures, chips arrive mismatched to production workloads.

What It Covers

Anjney Midha, CEO of AMP, explains how his company operates as an independent system operator for compute — modeled on electric grid utilities like PJM Interconnect — pooling 1.3 gigawatts of supply across clouds and silicon to eliminate stranded capacity and serve frontier AI labs.

Key Questions Answered

  • GPU Cluster Utilization Benchmarks: Node utilization below 96% in AI clusters is indefensible — Google treated anything under 95% as an outage. MFU (model flop utilization) best-in-class sits at 60–70%. Most singleton clusters fall short of both metrics due to misalignment between capital providers and the engineers actually managing infrastructure at scale.
  • Community-Aligned Data Center Pricing: Up to 20% of US data centers face community backlash risk in the current year. A concrete mitigation: charge $4.50 per compute-hour instead of $4.00, routing the $0.50 margin directly to local communities as cash or electricity bill reductions. This converts opposition into partnership without requiring regulatory intervention.
  • Independent System Operator Model for Compute: Rather than owning assets, AMP pools demand from frontier labs and supply from trusted 20-plus-year data center operators at 1.3 gigawatt base load, targeting 6 gigawatts over four years. Labs receive guaranteed base load with flexible spike capacity — mirroring how PJM Interconnect coordinates uncorrelated industrial demand across the Northeast US grid.
  • Culture as a Fragile Daily Practice, Not a Moat: Anthropic's coding dominance traces directly to four years of resource scarcity forcing precise prioritization — coding as the singular P-zero toward AGI. Teams flush with capital skip this definition process. Culture requires daily action-based reinforcement; without hardship forcing trade-off clarity, lab cultures become brittle before reaching capability takeoff.
  • Trust Boundary as the Primary Chip Company Risk: Hardware startups like Matx adopt NVIDIA's open reference architecture to eliminate data center integration battles, focusing innovation on systems co-design at the logic tile level. The real bottleneck is trust access — chip tape-out cycles run two years, so without early visibility into next-generation model architectures, chips arrive mismatched to production workloads.

Notable Moment

Midha describes how a researcher who declined to join Periodic Labs for a higher-paying role later requested to return after a technical breakthrough. He refused — framing the rejection not as punitive but as a culture-preservation decision, arguing that mission alignment must be demonstrated before breakthroughs, not after.

Know someone who'd find this useful?

Episode Transcript

We're in Periodic Labs with Anj Middha, CEO, founder of AMP welcome. Thanks for having me. At Google, if utilize so there's two types of utilization usually, right, that you're measuring in these clusters. One is node allocation, and then the other is m f u. So node at utilization is usually, like, what percentage of cards in the data center are just, like, used. And that, if it's not at, like, 95% There's no excuse. There's no excuse. Right? Like, I think 95% of Google, which is where my cofounder, Seb, came from, he built the Borg export GQM scheduler at Google. And there, I think, 95% was considered an outage. So 96% node utilization is should be standard, and and most singleton clusters are not running at that. So that's one. And then MFU utilization should be, I would say, the best in class today is somewhere between 6070%. I think it's a leadership question. Right? Is is there an and and, fundamentally, it's an alignment question, which is are the people who are funding the cluster and then deploying the cluster actually aligned? And sometimes, theoretically, they are. But in practice, the number of people in the chain, the supply chain between, like, the capital and all the way to whoever's managing the cluster and then whoever's measuring what the output is are just so many, you know, degrees of separation away that, like, the you know have you heard that sort of, you know, radian metaphor, which is at the beginning of of of an arc, if you have two arcs that are two lines that are just off by a few degrees, that It spreads out. It spreads out, right, at scale. And I think what's happening is a lot of cluster implementations and infrastructure, a lot of Frontier Labs and other teams, that's what's happening is they're they they initialize the plan, which is kinda like North North Star with a team that wants to do good, but then they're required to scale so fast instead of iteratively that the waste, it just compounds really fast at scale. And so I I think we know the answer, which is just do iterative bring ups. You know? If you spend time with people who've been in the semiconductor industry or the data center industry for a long time, this is not new. And I don't think AI should be an excuse. Like, sure, something what is new? Okay. We have a lot of new capabilities, but that doesn't mean just abandon common sense. Common sense should always be in fashion. You know, like, AI scaling doesn't change the in fact, if anything, AI scaling should be putting a premium on the value of common sense and infrastructure because the the margin of error now is so much lower and the cost of wastage are so much higher. And the cost of wastage, by the way, is not just economic. I mean, obviously, I'm I'm an …

Get the full transcript (11,786 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Latent Space transcripts →

You just read a 3-minute summary of a 56-minute episode.

Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

company

  • Midha describes how a researcher who declined to join Periodic Labs for a higher-paying role later requested to return
  • modeled on electric grid utilities like PJM Interconnect — pooling 1.3 gigawatts of supply across clouds and silicon
  • Anjney Midha, CEO of AMP, explains how his company operates as an independent system operator for compute
  • Google treated anything under 95% as an outage
  • Hardware startups like Matx adopt NVIDIA's open reference architecture to eliminate data center integration battles
  • Anthropic's coding dominance traces directly to four years of resource scarcity forcing precise prioritization
  • Hardware startups like Matx adopt NVIDIA's open reference architecture to eliminate data center integration battles

More from Latent Space

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Latent Space.

Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime