Skip to main content
20VC (20 Minute VC)

20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks

77 min episode · 3 min read
·
Lin Qiao

Episode

77 min

Read time

3 min

Topics

Productivity, Remote Work, Startups

AI-Generated Summary

Key Takeaways

  • Specialized vs. General Intelligence: The majority of the world's data is private enterprise data locked inside applications — it never reaches public internet training sets. Companies that activate this proprietary data through fine-tuned, specialized models will outperform general-purpose frontier models on their specific workflows. Fireworks reports that the majority of its 40 trillion daily tokens already come from customized models, not off-the-shelf deployments.
  • Token Economics Forecast: Token costs will fall approximately 10x over the next three years, driven by easing supply chain constraints, GPU manufacturing scale, and model efficiency improvements. This cost reduction will trigger roughly 100x growth in token usage — the same dynamic seen historically when utility costs drop and consumption explodes. Enterprises currently blocked by cost forecasting will unlock AI deployment at scale once affordability crosses a threshold.
  • Open Source as Strategic Infrastructure: Open-weight models give enterprises full control over model weights, enabling customization without dependency on any single vendor. Fireworks built its entire business on the bet that open models would reach quality thresholds sufficient for production use — a bet that paid off. Enterprises should evaluate open models not just on cost but on the control they provide for tuning, deployment, and avoiding vendor lock-in or geopolitical supply chain risk.
  • Product-Market Fit vs. Durable Business: In the SaaS era, finding product-market fit and building a durable business were nearly equivalent. In AI, they are separate problems. Companies can find strong customer demand but scale into bankruptcy because inference costs make unit economics unworkable at volume. Founders must model cost-per-token against revenue per user before scaling, not after, to avoid the trap of growth that destroys margins.
  • Distributed Reinforcement Learning Training: Fireworks and Cursor co-developed a decoupled RL training architecture that separates weight-updating from rollout inference, running across five to six global data center regions using scattered GPUs rather than a single expensive interconnected cluster. This approach enables capital-efficient post-training at startup scale, reducing dependency on 100,000-chip hyperscaler infrastructure while maintaining numerically sound reward freshness through rapid weight synchronization.

What It Covers

Lin Qiao, founder and CEO of Fireworks AI, argues that the future of intelligence is not one dominant AGI model but millions of specialized models built on private enterprise data. Fireworks processes 40 trillion tokens daily, targets 20-100x token growth, and predicts a 10x cost reduction driving 100x usage expansion over three years.

Key Questions Answered

  • Specialized vs. General Intelligence: The majority of the world's data is private enterprise data locked inside applications — it never reaches public internet training sets. Companies that activate this proprietary data through fine-tuned, specialized models will outperform general-purpose frontier models on their specific workflows. Fireworks reports that the majority of its 40 trillion daily tokens already come from customized models, not off-the-shelf deployments.
  • Token Economics Forecast: Token costs will fall approximately 10x over the next three years, driven by easing supply chain constraints, GPU manufacturing scale, and model efficiency improvements. This cost reduction will trigger roughly 100x growth in token usage — the same dynamic seen historically when utility costs drop and consumption explodes. Enterprises currently blocked by cost forecasting will unlock AI deployment at scale once affordability crosses a threshold.
  • Open Source as Strategic Infrastructure: Open-weight models give enterprises full control over model weights, enabling customization without dependency on any single vendor. Fireworks built its entire business on the bet that open models would reach quality thresholds sufficient for production use — a bet that paid off. Enterprises should evaluate open models not just on cost but on the control they provide for tuning, deployment, and avoiding vendor lock-in or geopolitical supply chain risk.
  • Product-Market Fit vs. Durable Business: In the SaaS era, finding product-market fit and building a durable business were nearly equivalent. In AI, they are separate problems. Companies can find strong customer demand but scale into bankruptcy because inference costs make unit economics unworkable at volume. Founders must model cost-per-token against revenue per user before scaling, not after, to avoid the trap of growth that destroys margins.
  • Distributed Reinforcement Learning Training: Fireworks and Cursor co-developed a decoupled RL training architecture that separates weight-updating from rollout inference, running across five to six global data center regions using scattered GPUs rather than a single expensive interconnected cluster. This approach enables capital-efficient post-training at startup scale, reducing dependency on 100,000-chip hyperscaler infrastructure while maintaining numerically sound reward freshness through rapid weight synchronization.
  • Sovereign and Enterprise Intelligence Ownership: Every company will eventually own its own intelligence stack — not as an option but as a competitive and operational necessity. The analogy is software: no company runs entirely on off-the-shelf SaaS for its core differentiation. Enterprises should begin identifying which parts of their AI stack require proprietary control versus commodity infrastructure, prioritizing model tuning on internal data as the first step toward intelligence independence.

Notable Moment

Lin Qiao recounts a conversation with Jensen Huang where Jensen stated that no company is truly generalized — every company exists because it solves a unique problem. Qiao found this observation profound because it reframes the entire AGI debate: if every business is inherently specialized, then one universal intelligence model cannot serve all of them optimally.

Know someone who'd find this useful?

You just read a 3-minute summary of a 74-minute episode.

Get 20VC (20 Minute VC) summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from 20VC (20 Minute VC)

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Investing Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into 20VC (20 Minute VC).

Every Monday, we deliver AI summaries of the latest episodes from 20VC (20 Minute VC) and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime