Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Episode
78 min
Read time
3 min
Topics
Productivity, Remote Work, Startups
AI-Generated Summary
Key Takeaways
- ✓Latency vs. Throughput Trade-off: GPU hardware forces a fundamental choice between low latency and high throughput. Batching more users' work together maximizes GPU utilization but increases per-request wait time. Sail deliberately optimizes for throughput, accepting slower individual responses in exchange for dramatically lower cost per token — a viable strategy only when users aren't actively waiting on results.
- ✓Scavenger Chip Strategy: Rather than competing with Anthropic or OpenAI for Blackwell allocations, Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix. Because other providers lack kernel expertise for non-NVIDIA architectures, Sail exploits the perception gap — treating "no bad chips, only bad pricing" as a core acquisition principle to build aggregate compute capacity.
- ✓Distributed 1-Megawatt Data Centers: Inference workloads, unlike training, don't require co-located clusters. Sail targets small, distributed one-megawatt facilities — roughly eight refrigerator-sized racks — that larger operators ignore. These facilities accept 95% uptime, skip redundant fiber and diesel generators, and can integrate intermittent solar and wind power, cutting overhead costs that traditional data centers treat as non-negotiable requirements.
- ✓KV Cache as the Primary Inefficiency: The KV cache — which stores every token in an active conversation — frequently exceeds model weight size and remains largely uncompressed. DeepSeek's published research shows roughly one order-of-magnitude compression progress annually, suggesting two or more orders of magnitude of remaining inefficiency. Targeting KV cache compression, not compute scaling, represents the highest-leverage optimization available in current inference stacks.
- ✓Background Agents as the Dominant Workload: Movva projects the inference mix shifting from roughly 50/50 real-time versus background today to 90% background within a few years. Background tasks — deep research across 10,000+ sources, autonomous cybersecurity pen-testing, proactive personal agents — consume tokens without human attention as the bottleneck, making token volume theoretically unbounded compared to human-in-the-loop workflows.
What It Covers
Neil Movva, founder of Sail Research, explains how his "token factory" targets background AI agents running hours or days rather than real-time chatbots. By scavenging underpriced chips, tolerating 95% uptime data centers, and optimizing software kernels for throughput over latency, Sail aims to reduce inference costs by 10x or more.
Key Questions Answered
- •Latency vs. Throughput Trade-off: GPU hardware forces a fundamental choice between low latency and high throughput. Batching more users' work together maximizes GPU utilization but increases per-request wait time. Sail deliberately optimizes for throughput, accepting slower individual responses in exchange for dramatically lower cost per token — a viable strategy only when users aren't actively waiting on results.
- •Scavenger Chip Strategy: Rather than competing with Anthropic or OpenAI for Blackwell allocations, Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix. Because other providers lack kernel expertise for non-NVIDIA architectures, Sail exploits the perception gap — treating "no bad chips, only bad pricing" as a core acquisition principle to build aggregate compute capacity.
- •Distributed 1-Megawatt Data Centers: Inference workloads, unlike training, don't require co-located clusters. Sail targets small, distributed one-megawatt facilities — roughly eight refrigerator-sized racks — that larger operators ignore. These facilities accept 95% uptime, skip redundant fiber and diesel generators, and can integrate intermittent solar and wind power, cutting overhead costs that traditional data centers treat as non-negotiable requirements.
- •KV Cache as the Primary Inefficiency: The KV cache — which stores every token in an active conversation — frequently exceeds model weight size and remains largely uncompressed. DeepSeek's published research shows roughly one order-of-magnitude compression progress annually, suggesting two or more orders of magnitude of remaining inefficiency. Targeting KV cache compression, not compute scaling, represents the highest-leverage optimization available in current inference stacks.
- •Background Agents as the Dominant Workload: Movva projects the inference mix shifting from roughly 50/50 real-time versus background today to 90% background within a few years. Background tasks — deep research across 10,000+ sources, autonomous cybersecurity pen-testing, proactive personal agents — consume tokens without human attention as the bottleneck, making token volume theoretically unbounded compared to human-in-the-loop workflows.
- •Open Source Distillation is Structurally Unavoidable: A growing percentage of GitHub repositories are generated by tools like Claude Code, meaning frontier model capabilities diffuse into open-source training data passively and continuously. Movva argues that training a frontier-class model solely on high-quality open-source code outputs is plausible today, making it structurally impossible for closed labs to permanently maintain capability advantages through access restrictions alone.
Notable Moment
Movva argues that losing access to TSMC would be far less catastrophic than geopolitical discourse suggests. He contends that Intel's best processes trail TSMC by at most two times on performance-per-watt — a gap far smaller than commonly assumed — and that successive TSMC nodes themselves deliver only modest incremental efficiency gains per generation.
Episode Transcript
Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth. Ramp customers grow revenue 3.2 times faster than the average American business. Visa, Vercel, Cursor, Stripe, Notion, ElevenLab, Shopify, and 70,000 other businesses all now run on Ramp. Mine does too and so should yours. Learn more at ramp.com/invest. OpenAI, Cursor, Anthropic, Perplexity, and Vercel all have something in common. They all use Work OS. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, RBAC, and audit logs. Instead of spending months building these mission critical capabilities yourself, you can just use Work OS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit workos.com to get started. Felix by Rogo is a personal finance agent that turns a single prompt into finished client ready work using your firm's own templates, context, and standards. Send Felix an email like, take these comments and turn them for me, or update my tracker with the context of these emails, And Felix sends back finished PowerPoint decks, Excel models, and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo.ai/felix. Hello and welcome everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and wanna go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at colossus.com. Patrick O'Shaughnessy is is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Clients of positive sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. My guest today is Neil Mova, the founder of Sail Research. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run-in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters much more, and Neil has built the entire company around driving the cost of a token as low as it can possibly go. What makes this conversation special is it's one of the most detailed tours I've ever done through the full stack of …
Get the full transcript (17,173 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
Browse all Invest Like the Best with Patrick O'Shaughnessy transcripts →
You just read a 3-minute summary of a 75-minute episode.
Get Invest Like the Best with Patrick O'Shaughnessy summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Invest Like the Best with Patrick O'Shaughnessy
Ben Thompson on Big Tech, China, and the AI Boom Running Out of Money - [Invest Like the Best, EP.487]
Aug 18 · 76 min
David Senra
Micky Malka, Founder of Ribbit Capital
Aug 2
More from Invest Like the Best with Patrick O'Shaughnessy
Eric Vishria - A Decade of Lessons Investing in Software & Hardware - [Invest Like the Best, EP.486]
Aug 11 · 65 min
The TWIML AI Podcast
How AI Learns to Smell with Alex Wiltschko - #771
Jul 8
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
by Anthropic
“A growing percentage of GitHub repositories are generated by tools like Claude Code, meaning frontier model capabilities diffuse into open-source training data passively and continuously.”
“SPONSORS: Ramp”
“SPONSORS: WorkOS”
“SPONSORS: Rogo (Felix)”
“SPONSORS: Vanta”
“SPONSORS: Ridgeline”
Gear
by NVIDIA
“Rather than competing with Anthropic or OpenAI for Blackwell allocations, Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
by Google
“Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
by AWS
“Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
company
- Sail ResearchBy guest
“Neil Movva, founder of Sail Research, explains how his "token factory" targets background AI agents running hours or days rather than real-time chatbots.”
“Rather than competing with Anthropic or OpenAI for Blackwell allocations, Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
“Rather than competing with Anthropic or OpenAI for Blackwell allocations, Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
“Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
“Sail buys underpriced chips from AMD, TPUs, Trainium, and emerging vendors like Etch or D-Matrix.”
“DeepSeek's published research shows roughly one order-of-magnitude compression progress annually, suggesting two or more orders of magnitude of remaining inefficiency.”
“A growing percentage of GitHub repositories are generated by tools like Claude Code, meaning frontier model capabilities diffuse into open-source training data passively and continuously.”
“Movva argues that losing access to TSMC would be far less catastrophic than geopolitical discourse suggests.”
More from Invest Like the Best with Patrick O'Shaughnessy
We summarize every new episode. Want them in your inbox?
Ben Thompson on Big Tech, China, and the AI Boom Running Out of Money - [Invest Like the Best, EP.487]
Eric Vishria - A Decade of Lessons Investing in Software & Hardware - [Invest Like the Best, EP.486]
Gavin Baker - AI Market Jitters - [Invest Like the Best, EP.485]
Sam Altman - How to Make an Abundant Future - [Invest Like the Best, EP.484]
Matthew Smith — Natural Gas: The Next Bottleneck - [Invest Like the Best, EP.483]
Similar Episodes
Related episodes from other podcasts
David Senra
Aug 2
Micky Malka, Founder of Ribbit Capital
The TWIML AI Podcast
Jul 8
How AI Learns to Smell with Alex Wiltschko - #771
No Priors: Artificial Intelligence | Technology | Startups
Aug 6
Chasing Trillion-Dollar Companies, Founder Ambition, Token Budgets, and Regulatory Capture with Sarah & Elad
20VC (20 Minute VC)
Aug 1
20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile
The Mel Robbins Podcast
Jul 13
Find Your Purpose & Live a Meaningful Life Today with the #1 Happiness Expert
Explore Related Topics
This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Invest Like the Best with Patrick O'Shaughnessy.
Every Monday, we deliver AI summaries of the latest episodes from Invest Like the Best with Patrick O'Shaughnessy and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime