
AI Summary
→ WHAT IT COVERS a16z's Ben Horowitz joins Vals.ai founder Rayan Krishnan to examine why public AI benchmarks fail to measure true model capabilities, how independent third-party evaluation firms like Vals fill that gap, and why enterprises face existential pressure to quantify ROI as token spend approaches—and sometimes exceeds—employee salary costs. → KEY INSIGHTS - **Public Benchmark Distortion:** When Meta released Llama 4, it scored well on all major public benchmarks but underperformed on Vals' private held-out benchmarks. Because public benchmark questions and rubrics are open-source, labs can optimize directly against them. Enterprises and policymakers should treat self-reported public benchmark scores with skepticism and prioritize private, held-out evaluation results instead. - **Token Spend vs. Salary Spend:** During a one-month unlimited-access experiment at Vals, engineers consumed up to 6 billion tokens per day, totaling roughly $1.5 million in token costs—10x the team's salary spend for that same month. Enterprises should audit actual token consumption by team and task before setting usage budgets, rather than applying arbitrary per-engineer caps like $100 or $300 daily. - **Private Repo Benchmarking via ValSmith:** Vals released ValSmith, a tool allowing companies to upload their private GitHub repositories and generate internal coding benchmarks to identify which AI coding agents deliver the highest ROI for their specific codebase. Counterintuitively, Claude Sonnet often costs more than Opus due to token inefficiency, making model selection non-obvious without running task-specific evals first. - **Recursive Self-Improvement (RSI) Measurement:** Vals released an RSI index that benchmarks how well frontier models can contribute to training the next model version, covering pretraining, post-training, and harness-level engineering proxies. As RSI accelerates, this benchmark provides the first standardized, cross-model language for tracking compounding capability gains—a metric Krishnan identifies as the most geopolitically consequential evaluation category. - **Benchmark Deprecation as Standard Practice:** Vals actively retires saturated benchmarks rather than maintaining high scores on obsolete tests, operating under the principle that benchmarks must reflect the current state of the world—similar to how lawyers retake bar exams and doctors recertify. Enterprises building internal evals should build deprecation schedules into their evaluation programs to prevent hill-climbing on stale criteria. → NOTABLE MOMENT Krishnan revealed that during Vals' internal token-maximizing experiment, one engineer consumed 6 billion tokens in a single day. When Krishnan calculated the monthly total, the team had spent roughly ten times more on tokens than on employee salaries—prompting Vals to build ValSmith specifically to solve their own runaway AI spending problem. 💼 SPONSORS None detected 🏷️ AI Benchmarking, Model Evaluation, Enterprise AI ROI, AI Policy & Regulation, Recursive Self-Improvement