Skip to main content
20VC (20 Minute VC)

20VC: "Anti-Data Centres is a Chinese Psyop" | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI's Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron

71 min episode · 3 min read
·
Thomas Sohmers

Episode

71 min

Read time

3 min

Topics

Productivity, Investing, Startups

AI-Generated Summary

Key Takeaways

  • Inference vs. Training Economics: Inference is memory-bound, not compute-bound — unlike training, each token requires reading all model weights sequentially, making parallelization impossible. This distinction matters for hardware investment: training scales with FLOP increases, but inference scales with memory bandwidth. Between 2014–2024, GPU FLOPs improved 120x while memory bandwidth improved only 17x, creating a structural bottleneck that defines the entire inference hardware market.
  • Token Margin Reality: Frontier AI companies are far more profitable than public perception suggests. Anthropic reportedly operates at 80% gross margin on API revenue. Cached tokens cost roughly one-thousandth of uncached tokens to process, yet providers charge near-standard rates for cache reads — generating outsized margin. If OpenAI and Anthropic stopped training today, both would become massively profitable overnight, undermining the "burning cash" narrative.
  • KV Cache Scale Problem: At long context lengths, individual user session data can exceed total model weight storage. A 10-trillion-parameter model carries roughly 5TB of weights, but a single user session at 100GB means just 50 concurrent users generate more context data than the model itself. Operators must tier storage across accelerator memory, host RAM, and NVMe flash — each tier adding latency and architectural complexity that defines inference system design.
  • Data Center Anti-Sentiment as Strategic Vulnerability: Bipartisan opposition to data centers in the US is built on factually incorrect premises — a single In-N-Out burger location uses more water than major US data centers, and golf courses consume orders of magnitude more. Meanwhile, China faces no equivalent political resistance and is adding gigawatts of new capacity. Republican governors who previously supported data center growth are reversing positions based on constituent misinformation, creating a concrete competitive disadvantage.
  • Pacing the Frontier Carries Concentration Risk: Regulatory frameworks limiting AI compute — such as restrictions on matrix multiplications or mandatory audits — risk concentrating capability among a small number of state-aligned actors. Sohmers argues this mirrors historical patterns where those advocating for regulation assume they will become the regulators. The more likely outcome is bureaucratic capture that halts progress while leaving existing capability concentrated in governments and a handful of corporations.

What It Covers

Thomas Sohmers, co-founder of Positron AI (recently valued at $5B after an $875M Series C), breaks down the technical and geopolitical realities shaping AI infrastructure — covering inference economics, KV caching mechanics, data center politics, frontier model scaling, and why anti-data-center sentiment may serve Chinese strategic interests.

Key Questions Answered

  • Inference vs. Training Economics: Inference is memory-bound, not compute-bound — unlike training, each token requires reading all model weights sequentially, making parallelization impossible. This distinction matters for hardware investment: training scales with FLOP increases, but inference scales with memory bandwidth. Between 2014–2024, GPU FLOPs improved 120x while memory bandwidth improved only 17x, creating a structural bottleneck that defines the entire inference hardware market.
  • Token Margin Reality: Frontier AI companies are far more profitable than public perception suggests. Anthropic reportedly operates at 80% gross margin on API revenue. Cached tokens cost roughly one-thousandth of uncached tokens to process, yet providers charge near-standard rates for cache reads — generating outsized margin. If OpenAI and Anthropic stopped training today, both would become massively profitable overnight, undermining the "burning cash" narrative.
  • KV Cache Scale Problem: At long context lengths, individual user session data can exceed total model weight storage. A 10-trillion-parameter model carries roughly 5TB of weights, but a single user session at 100GB means just 50 concurrent users generate more context data than the model itself. Operators must tier storage across accelerator memory, host RAM, and NVMe flash — each tier adding latency and architectural complexity that defines inference system design.
  • Data Center Anti-Sentiment as Strategic Vulnerability: Bipartisan opposition to data centers in the US is built on factually incorrect premises — a single In-N-Out burger location uses more water than major US data centers, and golf courses consume orders of magnitude more. Meanwhile, China faces no equivalent political resistance and is adding gigawatts of new capacity. Republican governors who previously supported data center growth are reversing positions based on constituent misinformation, creating a concrete competitive disadvantage.
  • Pacing the Frontier Carries Concentration Risk: Regulatory frameworks limiting AI compute — such as restrictions on matrix multiplications or mandatory audits — risk concentrating capability among a small number of state-aligned actors. Sohmers argues this mirrors historical patterns where those advocating for regulation assume they will become the regulators. The more likely outcome is bureaucratic capture that halts progress while leaving existing capability concentrated in governments and a handful of corporations.
  • Local Models Increase, Not Decrease, Cloud Token Volume: The assumption that on-device or enterprise-hosted smaller models reduce demand for frontier API calls is likely inverted. A local model continuously processing email, calendar, and messages will autonomously generate far more upstream requests to cloud-hosted frontier models than a human manually prompting ChatGPT. The bottleneck today is human initiation of requests — removing that bottleneck through local agents multiplies total token consumption across the entire stack.

Notable Moment

Sohmers describes spending several hours testing GPT-6 Astra on a chip design task — taking an encryption block from specification through full RTL-to-GDS flow using TSMC N3 process design kits. A task requiring two to three weeks for a skilled engineer was completed in roughly 50 hours, meeting timing at over one gigahertz.

Know someone who'd find this useful?

Episode Transcript

I would say I'm overall opposed to, you know, the pace of the frontier direction it's going in. The scariest thing to me on on the political spectrum and the way all of this being treated is that it's now become a almost unifying issue on left and right about being anti data centers. I think that is almost entirely a Chinese SIOP. A single in and out uses, you know, more more water than, you know, the largest data centers in The United States. This is 20 VC with me, Harry Stebbings. Now we have probably one of the most important times in history for technology. Have the biggest model providers saying we need to pace the frontier. But what does that actually mean in reality? How possible is it? What does it mean for the threat from China? What does it mean for the infrastructure layer moving forwards? We have a true expert of the space on the show today in the form of Thomas Somers. He's the cofounder and chairman of Positron AI. They just raised an $875,000,000 series c at a $5,000,000,000 valuation. They've got some of the best investors in the business, including the one and only Gavin Baker at Atredi's and many more great names. Thomas did not hold back in this episode, and it's this beautiful combination of incredible education on the infrastructure that powers this economy for AI. And then also, I don't know how to say it, but analytical gossip would be a more intellectual way of saying incredible discussion about what we can expect in the next few months from the biggest players in this space. But before we dive into the show today, the best model for your application might not exist yet. The most ambitious AI teams are training open models to beat the frontier in their domain. Now on November 3 in San Francisco, Forge by Fireworks brings those teams together. People like Jensen Huang, CEO of NVIDIA, Michelle Cataster, president and head of AI at Replit, Lin Kwao, CEO of Fireworks, and more to hear how AI leaders are taking control of their differentiation, margins, and road map by owning their intelligence. Learn how they're building it with inference, training, and intelligent routing. Forge is free to attend, but space is limited. Apply today at fireworks.ai/forge. While Fireworks AI powers product intelligence, Asana keeps the work moving. Most companies have tried AI. Most aren't seeing results. Not because AI doesn't work, it's because AI hasn't reached the workflows yet. That's the gap Asana is built to close. Asana is the operating system for human agent teams. Your easy button for AI productivity across every team. Ready to go AI teammates, prebuilt for marketing, ops, and IT. No prompt engineering, no setup. They show up where the work is happening, already onboarded in your workflows, ready to deliver. With Asana, your whole company can work on the same plan towards the same goal, whether you're a team …

Get the full transcript (13,644 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all 20VC (20 Minute VC) transcripts →

You just read a 3-minute summary of a 68-minute episode.

Get 20VC (20 Minute VC) summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by OpenAI

    A local model continuously processing email, calendar, and messages will autonomously generate far more upstream requests to cloud-hosted frontier models than a human manually prompting ChatGPT.
  • by OpenAI

    Sohmers describes spending several hours testing GPT-6 Astra on a chip design task — taking an encryption block from specification through full RTL-to-GDS flow using TSMC N3 process design kits.
  • SPONSORS: Fireworks AI
  • SPONSORS: Asana
  • SPONSORS: Superhuman

company

  • Positron AIBy guest
    Thomas Sohmers, co-founder of Positron AI (recently valued at $5B after an $875M Series C)
  • Frontier AI companies are far more profitable than public perception suggests. Anthropic reportedly operates at 80% gross margin on API revenue.
  • If OpenAI and Anthropic stopped training today, both would become massively profitable overnight, undermining the 'burning cash' narrative.

More from 20VC (20 Minute VC)

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into 20VC (20 Minute VC).

Every Monday, we deliver AI summaries of the latest episodes from 20VC (20 Minute VC) and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime