Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Episode
78 min
Read time
2 min
Topics
Productivity, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Independent Benchmarking Economics: Artificial Analysis runs evaluations costing hundreds to thousands of dollars monthly, using mystery shopper policies with unidentified accounts to prevent labs from optimizing specific endpoints. They maintain independence by never accepting payment for better rankings while monetizing through enterprise subscriptions and private benchmarking services.
- ✓Intelligence Cost Deflation: GPT-4 level intelligence now costs 100-1000x less than at launch, yet total AI spending increases simultaneously. This paradox occurs because frontier models use 10x more tokens through reasoning chains and agentic workflows, creating a smile curve where both cheap commodity intelligence and expensive frontier capabilities grow.
- ✓Hallucination Measurement Innovation: The Omniscience Index scores models from negative 100 to positive 100, deducting points for incorrect answers rather than rewarding guesses. Claude models show lowest hallucination rates at 15-20%, while intelligence level shows no correlation with hallucination tendency, revealing post-training recipe differences between labs.
- ✓Agentic Benchmark Methodology: GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management. Models in custom harnesses outperform their official chatbot versions, with Gemini 3 Pro using 95% confidence intervals requiring multiple evaluation runs.
- ✓Hardware Efficiency Reality: Blackwell generation GPUs deliver 2-3x throughput gains over Hopper for most workloads, not the marketed 4x, with actual improvements varying by model sparsity. Total parameter count correlates more strongly with knowledge retention than active parameters, suggesting sparse models like Kimi K2 at 3% activation still benefit from larger total sizes.
What It Covers
George Cameron and Micah-Hill Smith explain how Artificial Analysis became the independent benchmarking standard for AI models, covering their methodology for measuring intelligence, speed, cost, hallucination rates, and openness across hundreds of models and providers.
Key Questions Answered
- •Independent Benchmarking Economics: Artificial Analysis runs evaluations costing hundreds to thousands of dollars monthly, using mystery shopper policies with unidentified accounts to prevent labs from optimizing specific endpoints. They maintain independence by never accepting payment for better rankings while monetizing through enterprise subscriptions and private benchmarking services.
- •Intelligence Cost Deflation: GPT-4 level intelligence now costs 100-1000x less than at launch, yet total AI spending increases simultaneously. This paradox occurs because frontier models use 10x more tokens through reasoning chains and agentic workflows, creating a smile curve where both cheap commodity intelligence and expensive frontier capabilities grow.
- •Hallucination Measurement Innovation: The Omniscience Index scores models from negative 100 to positive 100, deducting points for incorrect answers rather than rewarding guesses. Claude models show lowest hallucination rates at 15-20%, while intelligence level shows no correlation with hallucination tendency, revealing post-training recipe differences between labs.
- •Agentic Benchmark Methodology: GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management. Models in custom harnesses outperform their official chatbot versions, with Gemini 3 Pro using 95% confidence intervals requiring multiple evaluation runs.
- •Hardware Efficiency Reality: Blackwell generation GPUs deliver 2-3x throughput gains over Hopper for most workloads, not the marketed 4x, with actual improvements varying by model sparsity. Total parameter count correlates more strongly with knowledge retention than active parameters, suggesting sparse models like Kimi K2 at 3% activation still benefit from larger total sizes.
Notable Moment
The team revealed they ran DeepSeek V3 evaluations on Boxing Day 2024 in New Zealand, immediately recognizing it as a breakthrough moment before the world noticed weeks later with R1. Their early detection came from systematic tracking of global players beyond mainstream attention.
You just read a 3-minute summary of a 75-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
Jun 4 · 75 min
The Peter Attia Drive
#367 - Tylenol, pregnancy, and autism: What recent studies show and how to interpret the data
Oct 6
More from Latent Space
🔬Scaling Past Informal AI - Carina Hong, Axiom Math
Jun 3 · 93 min
Software Engineering Daily
The Hardware Bottleneck AI Can’t Fix
Jun 2
More from Latent Space
We summarize every new episode. Want them in your inbox?
Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs
🔬Scaling Past Informal AI - Carina Hong, Axiom Math
⚡️Satya Nadella: No Priors x Latent Space Crossover Special at Microsoft Build
GitHub's plan for Agents — Kyle Daigle, GitHub
Why Video Agent models are next — Ethan He, xAI Grok Imagine
Similar Episodes
Related episodes from other podcasts
The Peter Attia Drive
Oct 6
#367 - Tylenol, pregnancy, and autism: What recent studies show and how to interpret the data
Software Engineering Daily
Jun 2
The Hardware Bottleneck AI Can’t Fix
The Daily (NYT)
May 20
Trump’s Taxpayer-Funded Revenge Plan
Investing for Beginners
May 18
How AI Is Changing Investing— with David Trainer
Stuff You Should Know
May 16
Did Mallory Make it to the Top of Everest First?
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for up to 3 shows.
Start My Monday DigestNo credit card · Unsubscribe anytime