Skip to main content
Odd Lots

Anjney Midha's Plan to Radically Lower the Price of Compute

50 min episode · 2 min read
·
Anjney Midha

Episode

50 min

Read time

2 min

Topics

Productivity, Investing, Startups

AI-Generated Summary

Key Takeaways

  • Compute Utilization Gap: Most independent data centers run below 70% node utilization, and model flop utilization (actual chip usage during workloads) can fall below 11%. Elon Musk's Colossus 2 cluster in Memphis ran at under 60% node utilization. Researchers should measure output efficiency, not chip headcount, when evaluating AI infrastructure investments.
  • True Cost of Leased Compute: Long-term GPU leases appear priced at $2.50–3.00 per hour, but because research demand is spiky and teams over-provision for peak loads, the effective cost balloons to $25–28 per hour. AMP's grid reallocates idle capacity to other users, returning the actual price paid closer to the marketed rate.
  • Verifiable Feedback Drives Model Progress: AI models improve fastest where task outcomes can be objectively verified — software passing unit tests and pull request reviews, or materials science predictions confirmed by X-ray diffraction. Subjective feedback like "that answer was wrong" produces minimal improvement; structured verification loops are what separate fast-progressing domains from stagnant ones.
  • Multiple Frontiers, Not One Winner: The AI landscape contains at least 17 distinct frontiers — software engineering, consumer chat, video generation, scientific discovery — each with different leaders. Anthropic leads coding with under 5,000 employees while Google's 60,000-person team remains close but behind. Corporate AI buyers will increasingly route queries to whichever model is cheapest for a given task, abstracting away brand entirely.
  • Model-Harness Co-Design: Breakthroughs like Claude Code result from simultaneous development of model capabilities and the surrounding tooling harness, not harness innovation alone. Teams build the harness to anticipate specific model improvements three months out, then remove third-party tool dependencies once the model internalizes those capabilities — collapsing task completion time by one to two minutes per operation.

What It Covers

Anjney Midha, founder of AMP PBC and early Anthropic backer, explains how software-based compute orchestration can reduce effective GPU costs from $25–28 per hour to the marketed rate of $2.50, by standardizing fragmented chip infrastructure into a unified grid modeled on electricity distribution.

Key Questions Answered

  • Compute Utilization Gap: Most independent data centers run below 70% node utilization, and model flop utilization (actual chip usage during workloads) can fall below 11%. Elon Musk's Colossus 2 cluster in Memphis ran at under 60% node utilization. Researchers should measure output efficiency, not chip headcount, when evaluating AI infrastructure investments.
  • True Cost of Leased Compute: Long-term GPU leases appear priced at $2.50–3.00 per hour, but because research demand is spiky and teams over-provision for peak loads, the effective cost balloons to $25–28 per hour. AMP's grid reallocates idle capacity to other users, returning the actual price paid closer to the marketed rate.
  • Verifiable Feedback Drives Model Progress: AI models improve fastest where task outcomes can be objectively verified — software passing unit tests and pull request reviews, or materials science predictions confirmed by X-ray diffraction. Subjective feedback like "that answer was wrong" produces minimal improvement; structured verification loops are what separate fast-progressing domains from stagnant ones.
  • Multiple Frontiers, Not One Winner: The AI landscape contains at least 17 distinct frontiers — software engineering, consumer chat, video generation, scientific discovery — each with different leaders. Anthropic leads coding with under 5,000 employees while Google's 60,000-person team remains close but behind. Corporate AI buyers will increasingly route queries to whichever model is cheapest for a given task, abstracting away brand entirely.
  • Model-Harness Co-Design: Breakthroughs like Claude Code result from simultaneous development of model capabilities and the surrounding tooling harness, not harness innovation alone. Teams build the harness to anticipate specific model improvements three months out, then remove third-party tool dependencies once the model internalizes those capabilities — collapsing task completion time by one to two minutes per operation.

Notable Moment

Midha reveals that Google's internal compute orchestration system, called Borg, achieved 99% chip utilization — up from 62% when his co-founder Sebastian Lobo joined. AMP is rebuilding that same software layer for the broader research ecosystem, where the industry average remains below 70%.

Know someone who'd find this useful?

Episode Transcript

Odd lots is brought to you by VanEck. For years, investors basically forgot about real assets, energy, gold, and infrastructure. But look what's driving markets now. Central banks loading up on gold, massive CapEx cycles, currencies doing weird things. These assets are at the center of it. Racks, the VanEck real asset ETF, is an actively managed one stop shop for real assets spanning gold, commodities, natural resource equities, and more. Go to vaneck.com/raaxpod to learn more. Fun disclosures later in this episode. So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now a global workforce of 300,000 can use AI to fill their HR questions, resolving 94% of common questions. Not noise, proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM. Small businesses are the pulse of every community. They bring people together, create opportunities, and drive growth. Chase for business helps business owners like you with personalized guidance and convenient digital tools all in one place. With that guidance and your determination, you can take your business farther and help build a brighter future for your community. Learn more at chase.com/business. Chase for business. Make more of what's yours. The Chase mobile app is available for select mobile devices. Message and data rates may apply. JPMorgan Chase Bank, NA. Member, FDIC. Copyright 2026. JPMorgan Chase and Company. Bloomberg Audio Studios. Podcasts, radio, news. Hello, and welcome to another episode of the Odd Thoughts podcast. I'm Tracy Alloway. And I'm Joe Wiesenthal. Joe, we like to talk a lot about, physical constraints Yes. On this show. Right? And this is one reason why AI is a really fascinating area for us right now because there are a lot of physical constraints on what is ultimately the sort of ephemeral technology. And I think that the tension between those two things is really interesting. Right? Right. Like, you type a prompt into chat GPT or Claude or whatever, and it's the sort of, like, disembodied digital platform. Disembodied. And you don't necessarily think about the power usage, the real resources, the transformers that have to go into data centers to get compute. The thing that I've been on my mind lately and I've written about it and I plan to write more is this idea that the canonical AI thought experiment is what happens if you tell an AI to make a lot of paper clips and then it destroys the world because in the pursuit of marshaling all of the world's resources, it just turns everything into paper clips because it doesn't know. And I have to ask. Is this canonical example, is this based on your traumatic fear of Clippy for me for so far? No. But that is you …

Get the full transcript (11,120 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Odd Lots transcripts →

You just read a 3-minute summary of a 47-minute episode.

Get Odd Lots summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

company

  • Anjney Midha, founder of AMP PBC and early Anthropic backer, explains how software-based compute orchestration can reduce effective GPU costs from $25–28 per hour to the marketed rate of $2.50
  • Anjney Midha, founder of AMP PBC and early Anthropic backer
  • Anthropic leads coding with under 5,000 employees while Google's 60,000-person team remains close but behind.

More from Odd Lots

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Finance Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Odd Lots.

Every Monday, we deliver AI summaries of the latest episodes from Odd Lots and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime