Skip to main content
Odd Lots

Understanding the Most Viral Chart in Artificial Intelligence

56 min episode · 2 min read
·
Joel Becker,Chris Painter

Episode

56 min

Read time

2 min

Topics

Productivity, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Time Horizon Methodology: METR measures AI capability by timing skilled human engineers completing identical tasks, then testing AI on the same tasks. The "time horizon" is the task length at which AI succeeds 50% of the time. Claude Opus 4.6 reaches 11 hours 59 minutes, nearly doubling GPT Codex's previous 5-hour 50-minute benchmark.
  • 50% Threshold vs. 80%: METR defaults to the 50% success threshold rather than 80% for statistical reasons: measuring at 50% requires fewer samples and is least sensitive to scoring noise. The 80% chart shows the same doubling pace but at roughly one-fifth the task length, meaning current 80% performance matches today's 50% performance within approximately eight months.
  • Doubling Rate Revision: METR initially published a seven-month capability doubling time but revised it to four months after newer models consistently matched the faster trend. Compute investment has grown at essentially the same exponential rate as capability gains, and already-committed data center buildouts through 2027-2028 make a near-term slowdown unlikely regardless of other variables.
  • Benchmark vs. Real-World Gap: AI time horizon scores overestimate real-world productivity gains for several reasons: holistic code quality standards differ from automated scoring, real tasks involve larger codebases and collaboration, and verification of AI-generated work requires extra time without the engineer's original context. These frictions are real but not considered fundamental barriers to eventual productivity gains.
  • Chinese Model Gap: Chinese models including Qwen do not appear on METR's main time horizon charts because they trail US frontier models by an estimated nine to twelve months on task capability. METR also notes Chinese benchmark scores may overstate actual held-out task performance relative to US models, making the capability gap potentially larger than raw benchmark comparisons suggest.

What It Covers

METR, a 30-person San Francisco nonprofit, created the most viral chart in AI: a "time horizon" graph measuring how AI models perform on engineering tasks scaled by human completion time. Claude Opus 4.6 now completes tasks requiring nearly 12 human hours at 50% success rate, doubling roughly every four months.

Key Questions Answered

  • Time Horizon Methodology: METR measures AI capability by timing skilled human engineers completing identical tasks, then testing AI on the same tasks. The "time horizon" is the task length at which AI succeeds 50% of the time. Claude Opus 4.6 reaches 11 hours 59 minutes, nearly doubling GPT Codex's previous 5-hour 50-minute benchmark.
  • 50% Threshold vs. 80%: METR defaults to the 50% success threshold rather than 80% for statistical reasons: measuring at 50% requires fewer samples and is least sensitive to scoring noise. The 80% chart shows the same doubling pace but at roughly one-fifth the task length, meaning current 80% performance matches today's 50% performance within approximately eight months.
  • Doubling Rate Revision: METR initially published a seven-month capability doubling time but revised it to four months after newer models consistently matched the faster trend. Compute investment has grown at essentially the same exponential rate as capability gains, and already-committed data center buildouts through 2027-2028 make a near-term slowdown unlikely regardless of other variables.
  • Benchmark vs. Real-World Gap: AI time horizon scores overestimate real-world productivity gains for several reasons: holistic code quality standards differ from automated scoring, real tasks involve larger codebases and collaboration, and verification of AI-generated work requires extra time without the engineer's original context. These frictions are real but not considered fundamental barriers to eventual productivity gains.
  • Chinese Model Gap: Chinese models including Qwen do not appear on METR's main time horizon charts because they trail US frontier models by an estimated nine to twelve months on task capability. METR also notes Chinese benchmark scores may overstate actual held-out task performance relative to US models, making the capability gap potentially larger than raw benchmark comparisons suggest.

Notable Moment

When asked about fully autonomous AI-to-AI collaboration today, METR's Joel Becker described current systems as eventually "falling on their faces" without human idea generation — the human still provides the concept while AI handles execution, meaning true autonomous research loops remain beyond present capability.

Know someone who'd find this useful?

Episode Transcript

Hey, Fidelity. What's it cost to invest with the Fidelity app? Start with as little as $1 with no account fees or trade commissions on US stocks and ETFs. That's music to my ears. I can only talk. Investing involves risk including risk of loss. Zero account fees apply to retail brokerage accounts only. Seller assessment fee not included. A limited number of ETFs are subject to a transaction based service fee of $100. See full list at fidelity.com/commissions. Fidelity Broker Services LLC, member NYSE SIPC. If you follow markets, you know the value of long term thinking. You plan, you diversify, you prepare for volatility. But in life, even the best strategies can't prevent every bad day. A fire, a loss, a disruption that demands immediate attention. When that happens, what matters isn't just what you planned, it's who shows up. That's where Cincinnati Insurance comes in. For more than seventy five years, they've helped individuals and businesses navigate life's toughest moments with care, expertise, and personal attention. Together with independent agents, Cincinnati Insurance focuses on relationships, not transactions. Their approach is grounded in experience, follow through, and trust built over time. Bad days happen. And when they do, you deserve an insurance partner who understands risk, respects what you've built, and is ready to help you move forward. The Cincinnati Insurance Companies. Let them make your bad day better. Find an independent agent at cinfin.com. So there's a lot of noise about AI, but time's too tight for more promises. So let's talk about results. At IBM, we work with our employees to integrate technology right into the systems they need. Now a global workforce of 300,000 can use AI to fill their HR questions, resolving 94% of common questions. Not noise. Proof of how we can help companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business. IBM. Bloomberg Audio Studios. Podcasts, radio, news. Hello, and welcome to another episode of the Odd Lots podcast. I'm Joe Weisenthal. And I'm Tracy Alloway. Tracy, one thing about AI is that, lots of lines that go up. Yes. Famously, there is perhaps one line that has captured the attention more than others when it comes to lines going up. Yes. But we're recording this April 7. Did you see the, anthropic revenue chart, by the way? Oh. It's just, like, straight. Yeah. Okay. It's just on the number of lines going up. I mean, there are some some really Let me caveat that. Yeah. Okay. Up until recently, there was one chart Yeah. Of a line going up exponentially that became, I think it's fair to say, the most viral chart in AI. Right? Yes. I would absolutely agree with that. So one of the many lines that go up, or there are various lines that sort of capture this, is, essentially just measures of AI progress and what they could do, what the models are …

Get the full transcript (13,304 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Odd Lots transcripts →

You just read a 3-minute summary of a 53-minute episode.

Get Odd Lots summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from Odd Lots

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Finance Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Odd Lots.

Every Monday, we deliver AI summaries of the latest episodes from Odd Lots and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime