Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Episode
39 min
Read time
2 min
Topics
Productivity, Investing, Startups
AI-Generated Summary
Key Takeaways
- ✓Public Benchmark Distortion: When Meta released Llama 4, it scored well on all major public benchmarks but underperformed on Vals' private held-out benchmarks. Because public benchmark questions and rubrics are open-source, labs can optimize directly against them. Enterprises and policymakers should treat self-reported public benchmark scores with skepticism and prioritize private, held-out evaluation results instead.
- ✓Token Spend vs. Salary Spend: During a one-month unlimited-access experiment at Vals, engineers consumed up to 6 billion tokens per day, totaling roughly $1.5 million in token costs—10x the team's salary spend for that same month. Enterprises should audit actual token consumption by team and task before setting usage budgets, rather than applying arbitrary per-engineer caps like $100 or $300 daily.
- ✓Private Repo Benchmarking via ValSmith: Vals released ValSmith, a tool allowing companies to upload their private GitHub repositories and generate internal coding benchmarks to identify which AI coding agents deliver the highest ROI for their specific codebase. Counterintuitively, Claude Sonnet often costs more than Opus due to token inefficiency, making model selection non-obvious without running task-specific evals first.
- ✓Recursive Self-Improvement (RSI) Measurement: Vals released an RSI index that benchmarks how well frontier models can contribute to training the next model version, covering pretraining, post-training, and harness-level engineering proxies. As RSI accelerates, this benchmark provides the first standardized, cross-model language for tracking compounding capability gains—a metric Krishnan identifies as the most geopolitically consequential evaluation category.
- ✓Benchmark Deprecation as Standard Practice: Vals actively retires saturated benchmarks rather than maintaining high scores on obsolete tests, operating under the principle that benchmarks must reflect the current state of the world—similar to how lawyers retake bar exams and doctors recertify. Enterprises building internal evals should build deprecation schedules into their evaluation programs to prevent hill-climbing on stale criteria.
What It Covers
a16z's Ben Horowitz joins Vals.ai founder Rayan Krishnan to examine why public AI benchmarks fail to measure true model capabilities, how independent third-party evaluation firms like Vals fill that gap, and why enterprises face existential pressure to quantify ROI as token spend approaches—and sometimes exceeds—employee salary costs.
Key Questions Answered
- •Public Benchmark Distortion: When Meta released Llama 4, it scored well on all major public benchmarks but underperformed on Vals' private held-out benchmarks. Because public benchmark questions and rubrics are open-source, labs can optimize directly against them. Enterprises and policymakers should treat self-reported public benchmark scores with skepticism and prioritize private, held-out evaluation results instead.
- •Token Spend vs. Salary Spend: During a one-month unlimited-access experiment at Vals, engineers consumed up to 6 billion tokens per day, totaling roughly $1.5 million in token costs—10x the team's salary spend for that same month. Enterprises should audit actual token consumption by team and task before setting usage budgets, rather than applying arbitrary per-engineer caps like $100 or $300 daily.
- •Private Repo Benchmarking via ValSmith: Vals released ValSmith, a tool allowing companies to upload their private GitHub repositories and generate internal coding benchmarks to identify which AI coding agents deliver the highest ROI for their specific codebase. Counterintuitively, Claude Sonnet often costs more than Opus due to token inefficiency, making model selection non-obvious without running task-specific evals first.
- •Recursive Self-Improvement (RSI) Measurement: Vals released an RSI index that benchmarks how well frontier models can contribute to training the next model version, covering pretraining, post-training, and harness-level engineering proxies. As RSI accelerates, this benchmark provides the first standardized, cross-model language for tracking compounding capability gains—a metric Krishnan identifies as the most geopolitically consequential evaluation category.
- •Benchmark Deprecation as Standard Practice: Vals actively retires saturated benchmarks rather than maintaining high scores on obsolete tests, operating under the principle that benchmarks must reflect the current state of the world—similar to how lawyers retake bar exams and doctors recertify. Enterprises building internal evals should build deprecation schedules into their evaluation programs to prevent hill-climbing on stale criteria.
Notable Moment
Krishnan revealed that during Vals' internal token-maximizing experiment, one engineer consumed 6 billion tokens in a single day. When Krishnan calculated the monthly total, the team had spent roughly ten times more on tokens than on employee salaries—prompting Vals to build ValSmith specifically to solve their own runaway AI spending problem.
Episode Transcript
Every time a new trillion dollar industry emerges, there's a need for this independent testing group. When Meta released Llama four on our held out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities. What's the limit of what you can achieve? And then within that, how are you going about it? In an ideal world, take a Frontier model and have it train the next version of itself. But, obviously, that's very expensive and slow. And so what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the models. Evaluations, as they become more complex, have a fewer sample size, but a larger set of criteria or expectations of them. Where do you see the gap that's happening today? The government kind of has an inclination of what it's afraid of, be it biohacking or cyberhacking. But then there becomes the question of, can the model do it? And then can you get the model to do it? What do you think the landscape will look like? AI models keep getting better, but the tests we use to measure them can become obsolete almost as quickly. In this episode, I'm joined by a16z's Van Horowitz and Jennifer Lee for a conversation with Valve's founder and CEO, Rayam Krishnan, about the increasingly difficult problem of measuring AI. We get into why public benchmarks can give a distorted picture of model capabilities, why independent evaluation matters, and what it takes to test a new model in a few hours before it launches. But this is becoming about much more than model leaderboards. As companies spend more on AI, they need to know which models and agents actually perform best for their own work and whether that intelligence is worth what they're paying for it. Ryan also explains why benchmarks need to evolve alongside the models, how VALS is measuring recursive self improvement, and why EVALs could eventually become a shared language for AI capabilities, risk, and policy. So I'll start a question from when WILL start starting 2024 after your team discovered that all the public benchmarks are just not sufficient enough to measure model progress, and there needs to be a new methodology and approach coming to keep us on the frontier and help model labs continue to hill climb, take us back to the the inception of laws and what you see was missing in the market then. Yeah. Yeah. Mean, so I had a background doing research in particular building benchmarks and evaluations. And so what was very clear to me was the very tight relationship between what it takes to build new systems for generation and actually new mechanisms for evaluation. In fact, in order to get one, often need to get better at the other. And actually, of the biggest drivers for model capability is having a new legible way to evaluate …
Get the full transcript (7,749 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 36-minute episode.
Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from a16z Podcast
OpenAI Researchers on the Future of Mathematical Reasoning
Sep 8 · 65 min
Hard Fork
The Ezra Klein Show: How Fast Will A.I. Agents Rip Through the Economy?
Mar 27
More from a16z Podcast
Can Open Source Keep AI Power From Concentrating?
Sep 7 · 8 min
20VC (20 Minute VC)
20VC: The AI Boom Will Create Enormous Roadkill: Who Wins & Loses | Why Founders Should Never Take Multi-Stage Money at Seed | Why Triple, Triple, Double, Double is Good Enough
Aug 8
More from a16z Podcast
We summarize every new episode. Want them in your inbox?
OpenAI Researchers on the Future of Mathematical Reasoning
Can Open Source Keep AI Power From Concentrating?
Your AI Doctor Is Coming | Julie Yoo
Aaron Levie on Why Open AI Wins
Fei Fei Li: The Race to Build World Models For AI
Similar Episodes
Related episodes from other podcasts
Hard Fork
Mar 27
The Ezra Klein Show: How Fast Will A.I. Agents Rip Through the Economy?
20VC (20 Minute VC)
Aug 8
20VC: The AI Boom Will Create Enormous Roadkill: Who Wins & Loses | Why Founders Should Never Take Multi-Stage Money at Seed | Why Triple, Triple, Double, Double is Good Enough
Masters of Scale
Jul 16
Build better relationships at work
Masters of Scale
Apr 25
Possible: Netflix co-founder Reed Hastings: stories, schools, superpowers
The Prof G Pod
Apr 22
Raging Moderates: How Trump’s Iran War Could Break the GOP (ft. Ben Shapiro)
Explore Related Topics
This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into a16z Podcast.
Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime