Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah-Hill Smith
Episode
78 min
Read time
2 min
Topics
Productivity, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Independent Benchmarking Economics: Artificial Analysis runs evaluations costing hundreds to thousands of dollars monthly, using mystery shopper policies with unidentified accounts to prevent labs from optimizing specific endpoints. They maintain independence by never accepting payment for better rankings while monetizing through enterprise subscriptions and private benchmarking services.
- ✓Intelligence Cost Deflation: GPT-4 level intelligence now costs 100-1000x less than at launch, yet total AI spending increases simultaneously. This paradox occurs because frontier models use 10x more tokens through reasoning chains and agentic workflows, creating a smile curve where both cheap commodity intelligence and expensive frontier capabilities grow.
- ✓Hallucination Measurement Innovation: The Omniscience Index scores models from negative 100 to positive 100, deducting points for incorrect answers rather than rewarding guesses. Claude models show lowest hallucination rates at 15-20%, while intelligence level shows no correlation with hallucination tendency, revealing post-training recipe differences between labs.
- ✓Agentic Benchmark Methodology: GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management. Models in custom harnesses outperform their official chatbot versions, with Gemini 3 Pro using 95% confidence intervals requiring multiple evaluation runs.
- ✓Hardware Efficiency Reality: Blackwell generation GPUs deliver 2-3x throughput gains over Hopper for most workloads, not the marketed 4x, with actual improvements varying by model sparsity. Total parameter count correlates more strongly with knowledge retention than active parameters, suggesting sparse models like Kimi K2 at 3% activation still benefit from larger total sizes.
What It Covers
George Cameron and Micah-Hill Smith explain how Artificial Analysis became the independent benchmarking standard for AI models, covering their methodology for measuring intelligence, speed, cost, hallucination rates, and openness across hundreds of models and providers.
Key Questions Answered
- •Independent Benchmarking Economics: Artificial Analysis runs evaluations costing hundreds to thousands of dollars monthly, using mystery shopper policies with unidentified accounts to prevent labs from optimizing specific endpoints. They maintain independence by never accepting payment for better rankings while monetizing through enterprise subscriptions and private benchmarking services.
- •Intelligence Cost Deflation: GPT-4 level intelligence now costs 100-1000x less than at launch, yet total AI spending increases simultaneously. This paradox occurs because frontier models use 10x more tokens through reasoning chains and agentic workflows, creating a smile curve where both cheap commodity intelligence and expensive frontier capabilities grow.
- •Hallucination Measurement Innovation: The Omniscience Index scores models from negative 100 to positive 100, deducting points for incorrect answers rather than rewarding guesses. Claude models show lowest hallucination rates at 15-20%, while intelligence level shows no correlation with hallucination tendency, revealing post-training recipe differences between labs.
- •Agentic Benchmark Methodology: GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management. Models in custom harnesses outperform their official chatbot versions, with Gemini 3 Pro using 95% confidence intervals requiring multiple evaluation runs.
- •Hardware Efficiency Reality: Blackwell generation GPUs deliver 2-3x throughput gains over Hopper for most workloads, not the marketed 4x, with actual improvements varying by model sparsity. Total parameter count correlates more strongly with knowledge retention than active parameters, suggesting sparse models like Kimi K2 at 3% activation still benefit from larger total sizes.
Notable Moment
The team revealed they ran DeepSeek V3 evaluations on Boxing Day 2024 in New Zealand, immediately recognizing it as a breakthrough moment before the world noticed weeks later with R1. Their early detection came from systematic tracking of global players beyond mainstream attention.
Episode Transcript
This is kind of a full circle moment for us in a way. Because Yeah. The, like, first time artificial analysis got mentioned on a podcast was you and Alyssa all made a space. Amazing. Which was January 2024. I I don't even remember doing that, but yeah. It was it was very influential to me. Yeah. I'm looking at AI news for Jan seventeen or Jan sixteen twenty twenty four. I said, this gem of a models and host comparison site was just launched. And, and then I put in a few screenshots. And I said, it's an independent third party. It clearly outlines the quality versus throughput trade off. Mhmm. And it breaks out by model and hosting provider. I did give you shit for missing fireworks. And, how do you have a model benchmarking thing without fireworks? But you had together, you had perplexity. And, I think we just started chatting there. Welcome, George and Micah, to Linspace. You've I've been following your progress. Congrats on an amazing year. You guys have really come together to be the presumptive new gardener of AI. Right? Which is something that Yeah. But you can't, pay us for better results. Yes. Exactly. Very important. Start off. Straight into it. Let's go. Start off with a spicy take. Okay. How do I pay you? Let's get right into that. How do you make money? Well, very happy to talk about that. That. So it's been a, like, big journey the last couple of years. Artificial analysis is gonna be two years old in January 2026, which is pretty soon now. We first run, like, the website for free, obviously, and give away a ton ton of data to help developers and companies navigate AI and make decisions about models, providers, technologies across the AI stack for building stuff. We're very committed to doing that and intend to keep doing that. We have, along the way, built a business that is working out pretty sustainably. We've got just over 20 people now. And two main customer groups. So we wanna be who enterprise look to for data and insights on AI. So we want to help them with their decisions about models and technologies for building stuff. And then on the other side, we do private benchmarking for companies throughout the AI stack who build AI stuff. So no one pays to be on the website. We've been very clear about that from the very start, because there's no use doing what we do unless it's independent AI benchmarking. Yeah. But turns out a bunch of our stuff can be pretty useful to companies building AI stuff. And is it like, I'm a Fortune 500. I need advisers on objective analysis, and I call you guys and you pull up a custom report for me. You come into my office and give me a workshop. What what what kind of engagement is that? So we have a benchmarking insight subscription, which looks …
Get the full transcript (15,681 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 75-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
The Peter Attia Drive
#367 - Tylenol, pregnancy, and autism: What recent studies show and how to interpret the data
Oct 6
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
Stuff You Should Know
Selects: Blacksmiths? You got that right!
Sep 5
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
“GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management.”
“Agentic Benchmark Methodology: GDP-VAL AA uses 220 sub-tasks across 44 white-collar job scenarios, running models through their open-source STIRRUP harness with code execution, web search, and context management.”
“The Omniscience Index scores models from negative 100 to positive 100, deducting points for incorrect answers rather than rewarding guesses.”
“George Cameron and Micah-Hill Smith explain how Artificial Analysis became the independent benchmarking standard for AI models, covering their methodology for measuring intelligence, speed, cost, hallucination rates, and openness across hundreds of models and providers.”
Gear
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
The Peter Attia Drive
Oct 6
#367 - Tylenol, pregnancy, and autism: What recent studies show and how to interpret the data
Stuff You Should Know
Sep 5
Selects: Blacksmiths? You got that right!
This Week in Startups
Aug 17
Bittensor creator Const on Affine, dTAO, "mining reasoning," and more | E2326
a16z Podcast
Aug 17
Stripe’s AI Strategy: Build More, Not Less
This Week in Startups
Aug 7
How AI splits startups into winners and losers | E2322
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime