Skip to main content
a16z Podcast

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

39 min episode · 2 min read
·
Rayan Krishnan,Jennifer Lee

Episode

39 min

Read time

2 min

Topics

Productivity, Investing, Startups

AI-Generated Summary

Key Takeaways

  • Public Benchmark Distortion: When Meta released Llama 4, it scored well on all major public benchmarks but underperformed on Vals' private held-out benchmarks. Because public benchmark questions and rubrics are open-source, labs can optimize directly against them. Enterprises and policymakers should treat self-reported public benchmark scores with skepticism and prioritize private, held-out evaluation results instead.
  • Token Spend vs. Salary Spend: During a one-month unlimited-access experiment at Vals, engineers consumed up to 6 billion tokens per day, totaling roughly $1.5 million in token costs—10x the team's salary spend for that same month. Enterprises should audit actual token consumption by team and task before setting usage budgets, rather than applying arbitrary per-engineer caps like $100 or $300 daily.
  • Private Repo Benchmarking via ValSmith: Vals released ValSmith, a tool allowing companies to upload their private GitHub repositories and generate internal coding benchmarks to identify which AI coding agents deliver the highest ROI for their specific codebase. Counterintuitively, Claude Sonnet often costs more than Opus due to token inefficiency, making model selection non-obvious without running task-specific evals first.
  • Recursive Self-Improvement (RSI) Measurement: Vals released an RSI index that benchmarks how well frontier models can contribute to training the next model version, covering pretraining, post-training, and harness-level engineering proxies. As RSI accelerates, this benchmark provides the first standardized, cross-model language for tracking compounding capability gains—a metric Krishnan identifies as the most geopolitically consequential evaluation category.
  • Benchmark Deprecation as Standard Practice: Vals actively retires saturated benchmarks rather than maintaining high scores on obsolete tests, operating under the principle that benchmarks must reflect the current state of the world—similar to how lawyers retake bar exams and doctors recertify. Enterprises building internal evals should build deprecation schedules into their evaluation programs to prevent hill-climbing on stale criteria.

What It Covers

a16z's Ben Horowitz joins Vals.ai founder Rayan Krishnan to examine why public AI benchmarks fail to measure true model capabilities, how independent third-party evaluation firms like Vals fill that gap, and why enterprises face existential pressure to quantify ROI as token spend approaches—and sometimes exceeds—employee salary costs.

Key Questions Answered

  • Public Benchmark Distortion: When Meta released Llama 4, it scored well on all major public benchmarks but underperformed on Vals' private held-out benchmarks. Because public benchmark questions and rubrics are open-source, labs can optimize directly against them. Enterprises and policymakers should treat self-reported public benchmark scores with skepticism and prioritize private, held-out evaluation results instead.
  • Token Spend vs. Salary Spend: During a one-month unlimited-access experiment at Vals, engineers consumed up to 6 billion tokens per day, totaling roughly $1.5 million in token costs—10x the team's salary spend for that same month. Enterprises should audit actual token consumption by team and task before setting usage budgets, rather than applying arbitrary per-engineer caps like $100 or $300 daily.
  • Private Repo Benchmarking via ValSmith: Vals released ValSmith, a tool allowing companies to upload their private GitHub repositories and generate internal coding benchmarks to identify which AI coding agents deliver the highest ROI for their specific codebase. Counterintuitively, Claude Sonnet often costs more than Opus due to token inefficiency, making model selection non-obvious without running task-specific evals first.
  • Recursive Self-Improvement (RSI) Measurement: Vals released an RSI index that benchmarks how well frontier models can contribute to training the next model version, covering pretraining, post-training, and harness-level engineering proxies. As RSI accelerates, this benchmark provides the first standardized, cross-model language for tracking compounding capability gains—a metric Krishnan identifies as the most geopolitically consequential evaluation category.
  • Benchmark Deprecation as Standard Practice: Vals actively retires saturated benchmarks rather than maintaining high scores on obsolete tests, operating under the principle that benchmarks must reflect the current state of the world—similar to how lawyers retake bar exams and doctors recertify. Enterprises building internal evals should build deprecation schedules into their evaluation programs to prevent hill-climbing on stale criteria.

Notable Moment

Krishnan revealed that during Vals' internal token-maximizing experiment, one engineer consumed 6 billion tokens in a single day. When Krishnan calculated the monthly total, the team had spent roughly ten times more on tokens than on employee salaries—prompting Vals to build ValSmith specifically to solve their own runaway AI spending problem.

Know someone who'd find this useful?

Episode Transcript

Every time a new trillion dollar industry emerges, there's a need for this independent testing group. When Meta released Llama four on our held out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks, it was showing incredible capabilities. What's the limit of what you can achieve? And then within that, how are you going about it? In an ideal world, take a Frontier model and have it train the next version of itself. But, obviously, that's very expensive and slow. And so what we're doing is forming a set of proxies for every part of the process it takes to build the next version of the models. Evaluations, as they become more complex, have a fewer sample size, but a larger set of criteria or expectations of them. Where do you see the gap that's happening today? The government kind of has an inclination of what it's afraid of, be it biohacking or cyberhacking. But then there becomes the question of, can the model do it? And then can you get the model to do it? What do you think the landscape will look like? AI models keep getting better, but the tests we use to measure them can become obsolete almost as quickly. In this episode, I'm joined by a16z's Van Horowitz and Jennifer Lee for a conversation with Valve's founder and CEO, Rayam Krishnan, about the increasingly difficult problem of measuring AI. We get into why public benchmarks can give a distorted picture of model capabilities, why independent evaluation matters, and what it takes to test a new model in a few hours before it launches. But this is becoming about much more than model leaderboards. As companies spend more on AI, they need to know which models and agents actually perform best for their own work and whether that intelligence is worth what they're paying for it. Ryan also explains why benchmarks need to evolve alongside the models, how VALS is measuring recursive self improvement, and why EVALs could eventually become a shared language for AI capabilities, risk, and policy. So I'll start a question from when WILL start starting 2024 after your team discovered that all the public benchmarks are just not sufficient enough to measure model progress, and there needs to be a new methodology and approach coming to keep us on the frontier and help model labs continue to hill climb, take us back to the the inception of laws and what you see was missing in the market then. Yeah. Yeah. Mean, so I had a background doing research in particular building benchmarks and evaluations. And so what was very clear to me was the very tight relationship between what it takes to build new systems for generation and actually new mechanisms for evaluation. In fact, in order to get one, often need to get better at the other. And actually, of the biggest drivers for model capability is having a new legible way to evaluate …

Get the full transcript (7,749 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all a16z Podcast transcripts →

You just read a 3-minute summary of a 36-minute episode.

Get a16z Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from a16z Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Business Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into a16z Podcast.

Every Monday, we deliver AI summaries of the latest episodes from a16z Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime