SF Compute: Commoditizing Compute to solve the GPU Bubble forever
Episode
72 min
Read time
3 min
Topics
Investing, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓GPU Economics vs CPU Economics: GPU customers exhibit extreme price sensitivity because scaling laws mean every additional GPU generates revenue, unlike CPU customers who hit flat capacity needs. A 10% price difference on $1 billion hardware equals $100 million in value, making customers willing to spend $50 million replicating software rather than paying margin premiums. This destroys traditional cloud provider business models that depend on high-margin software services layered on commodity hardware.
- ✓CoreWeave's Real Estate Model: CoreWeave succeeded by signing only long-term contracts with creditworthy customers like Microsoft and OpenAI (77% of revenue), then using these contracts to secure prime interest rates from lenders. This approach works because it treats GPUs as a banking business rather than a software company. Hyperscalers likely lose money reselling NVIDIA GPUs because they cannot capture the high margins their CPU businesses require while competing on price with specialized providers.
- ✓The Four Quadrants of GPU Risk: GPU providers face a risk matrix based on contract length and payment timing. Optimal position: long-term contracts paid upfront to low-risk customers (lowest interest rates). Worst position: short-term contracts at low prices paid over time (highest interest rates). Most GPU clouds died trying to charge high prices for short-term contracts, getting competed down by price-sensitive customers until margins disappeared completely.
- ✓SF Compute's Market Mechanism: The platform enables hourly GPU reservations that function like perishable inventory—prices drop continuously until compute clears the market, similar to day-old milk. Users can set limit prices (e.g., $4/hour maximum) to buy at spot rates (often $0.80-$0.90) while avoiding price spikes. This creates liquidity pockets where customers needing one month of compute can buy from others selling back unused portions of year-long contracts.
- ✓Automated Auditing Infrastructure: SF Compute runs LINPACK burn-in tests for 48 hours to seven days to kill dead GPUs before deployment, then implements active and passive performance checks during operation. The company controls clusters via BMC access (remote reimaging capability) and provides automated refunds when hardware fails. This standardization layer enables fungible GPU contracts necessary for a functioning marketplace, unlike bespoke cluster arrangements that lock in buyers.
What It Covers
Evan Conrad explains how SF Compute built a GPU marketplace by recognizing that GPU economics fundamentally differ from CPU clouds. CoreWeave succeeded by locking in long-term contracts and treating the business like real estate rather than software. SF Compute creates liquidity through hourly reservations, automated auditing, and plans for cash-settled futures to derisk the entire AI infrastructure market.
Key Questions Answered
- •GPU Economics vs CPU Economics: GPU customers exhibit extreme price sensitivity because scaling laws mean every additional GPU generates revenue, unlike CPU customers who hit flat capacity needs. A 10% price difference on $1 billion hardware equals $100 million in value, making customers willing to spend $50 million replicating software rather than paying margin premiums. This destroys traditional cloud provider business models that depend on high-margin software services layered on commodity hardware.
- •CoreWeave's Real Estate Model: CoreWeave succeeded by signing only long-term contracts with creditworthy customers like Microsoft and OpenAI (77% of revenue), then using these contracts to secure prime interest rates from lenders. This approach works because it treats GPUs as a banking business rather than a software company. Hyperscalers likely lose money reselling NVIDIA GPUs because they cannot capture the high margins their CPU businesses require while competing on price with specialized providers.
- •The Four Quadrants of GPU Risk: GPU providers face a risk matrix based on contract length and payment timing. Optimal position: long-term contracts paid upfront to low-risk customers (lowest interest rates). Worst position: short-term contracts at low prices paid over time (highest interest rates). Most GPU clouds died trying to charge high prices for short-term contracts, getting competed down by price-sensitive customers until margins disappeared completely.
- •SF Compute's Market Mechanism: The platform enables hourly GPU reservations that function like perishable inventory—prices drop continuously until compute clears the market, similar to day-old milk. Users can set limit prices (e.g., $4/hour maximum) to buy at spot rates (often $0.80-$0.90) while avoiding price spikes. This creates liquidity pockets where customers needing one month of compute can buy from others selling back unused portions of year-long contracts.
- •Automated Auditing Infrastructure: SF Compute runs LINPACK burn-in tests for 48 hours to seven days to kill dead GPUs before deployment, then implements active and passive performance checks during operation. The company controls clusters via BMC access (remote reimaging capability) and provides automated refunds when hardware fails. This standardization layer enables fungible GPU contracts necessary for a functioning marketplace, unlike bespoke cluster arrangements that lock in buyers.
- •Path to GPU Futures Market: SF Compute builds toward cash-settled GPU futures to derisk the entire AI supply chain. Currently, startups absorb hardware risk, pushing it to VCs who write inflated pre-revenue valuations to fund $50 million GPU purchases. A futures market would let data centers lock in prices without forcing customers into year-long commitments, reducing systemic risk similar to how agricultural futures stabilized farming economics by separating price risk from physical delivery.
Notable Moment
Conrad reveals SF Compute started accidentally when they signed a year-long GPU contract to train music models, planned to use one month and sublease eleven months, then owed $500,000 monthly with only $500,000 in the bank. For twelve months, failing to sell out the cluster meant bankruptcy. This forced survival mode transformed into understanding GPU brokerage economics better than anyone, eventually building the market infrastructure they originally needed as desperate customers.
Episode Transcript
Hey, everyone. Welcome to the Living Space podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my colleague Swyx, founder of Small AI. Hey. And today, we're so excited to be finally in the studio with Evan Konrad from SF Compute. Welcome. Hello, how good is it? How are you? I've been fortunate enough to be your friend before, you're famous, and also we've hung out at, like, various social things. So it's really cool to see that SF Compute, like, is is coming into its own thing and it's, you know, it's it's a significant presence, at least in the San Francisco community, which, of course, it's in the name. So you couldn't help but be. Indeed. Indeed. I think we have a long way to go, but yeah. Thanks. Of course. Yeah. One way I was thinking about kicking off this conversation is we will likely release this right after Core Reef IPO. Uh-huh. And and, I was watching I was looking doing some research on you. You did a talk at the curve. Yeah. I think I may have been viewer number 70. It was a great talk. More people should go see it, Evan Conrad at The Curve, but we're we have, like, three orders of magnitude more people. And I just wanted to to highlight, like, what is your analysis of what Core Eve did that that went so right for them? So locked in long term contracts and don't really do much short term at all. I think, like, a lot of people had this assumption that GPUs would work a lot like CPUs. And the, like, standard business model of any sort of CPU cloud is you buy commodity hardware, then you lay on services, that are mostly software, and that gives you high margins. And pretty much all your value comes from those services, not really the underlying compute in any capacity. And because it's commodity hardware and it's not actually that expensive, most of that can be, sort of on demand compute. And while you do want locked in contracts for folks, it's mostly just to sort of de risk your situation. It helps you plan revenue because you don't know if people are gonna scale up or down. But, fundamentally, people are, like, buying hourly, and that's how your business is structured. And you're gonna make 50% margins or higher. This, like, doesn't really work, in GPUs. And the reason why it doesn't work is because you end up with, like, super price sensitive customers. And that isn't because necessarily, it's just way more expensive, though that's totally the case. So in a CPU cloud, you might have, like, you know, let's say, if you had a million dollars of hardware, in GPUs, you have a billion dollars of hardware. And so your customers are buying, at much higher volumes than you'd otherwise expect. And it's also smaller customers who are buying at higher months of …
Get the full transcript (15,719 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 69-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
Venture Stories
Worldbuilders: The Largest Infrastructure Project in History with Evan Conrad (SF Compute)
Mar 13
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Odd Lots
Why Cerebras CEO Andrew Feldman Built The World's Largest Computer Chip
May 21
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
Venture Stories
Mar 13
Worldbuilders: The Largest Infrastructure Project in History with Evan Conrad (SF Compute)
Odd Lots
May 21
Why Cerebras CEO Andrew Feldman Built The World's Largest Computer Chip
Invest Like the Best with Patrick O'Shaughnessy
May 13
Krishna Rao - Anthropic's CFO on Compute, Scaling to $30B ARR, and the Returns to Frontier Intelligence - [Invest Like the Best, EP.471]
This Week in Startups
May 1
Can an AI Agent Legally Own a Company? Christian van der Henst's Wild Experiment| E2283
Modern Wisdom
Feb 21
#1062 - Dave Evans - It’s time to rethink your entire life plan
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime