The PhD students who became the judges of the AI industry
Episode
26 min
Read time
2 min
Topics
Investing, Fundraising & VC, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Dynamic vs. Static Benchmarks: Static benchmarks like Humanity's Last Exam become obsolete once models train on their questions — a problem called overfitting. Arena counters this by generating hundreds of thousands of fresh, never-repeated user conversations daily, making it structurally impossible for model providers to "teach to the test" and forcing genuine capability improvements instead.
- ✓Leaderboard Neutrality Structure: Arena's neutrality is methodological, not just policy-based. Scores are calculated via an open-source pipeline from real user votes — Arena staff cannot manually alter rankings. No model provider can pay to appear, improve, or be removed from the public leaderboard, and all public models are evaluated at no cost to maintain independence from investors.
- ✓Style Control Methodology: Arena developed a technique called style control that statistically factors out superficial response traits — length, markdown formatting, sycophancy — from leaderboard scores, the same way social science studies control for confounding variables. This prevents models from gaming rankings by sounding polished rather than being genuinely useful or accurate.
- ✓Occupational Segmentation for Enterprise: Arena segments its 60M monthly conversations by occupation and use case — 28% coding, 6% legal, 6% medical — and offers enterprises an analytical tool to identify which model performs best for their specific domain. Enterprises can privately test models during development without public score release, enabling faster, data-driven model upgrade decisions.
- ✓Agentic Evaluation Expansion: Arena launched WebDev Arena (Corena) to evaluate AI agents on end-to-end tasks like building web applications, tool calling, and navigating codebases. The roadmap extends to Python and C++ coding agents, multimodal editing, deep research, and multi-step planning tasks — tracking AI capability shifts from single-turn chat toward long-horizon autonomous workflows.
What It Covers
Arena (formerly LM Arena and Chatbot Arena), cofounded by Berkeley PhD students Anastasios Angelopoulos and Wei-Lin Chiang, operates the de facto public leaderboard for frontier AI models. Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation, Arena uses 5M+ monthly users across 150 countries to rank AI models in real time.
Key Questions Answered
- •Dynamic vs. Static Benchmarks: Static benchmarks like Humanity's Last Exam become obsolete once models train on their questions — a problem called overfitting. Arena counters this by generating hundreds of thousands of fresh, never-repeated user conversations daily, making it structurally impossible for model providers to "teach to the test" and forcing genuine capability improvements instead.
- •Leaderboard Neutrality Structure: Arena's neutrality is methodological, not just policy-based. Scores are calculated via an open-source pipeline from real user votes — Arena staff cannot manually alter rankings. No model provider can pay to appear, improve, or be removed from the public leaderboard, and all public models are evaluated at no cost to maintain independence from investors.
- •Style Control Methodology: Arena developed a technique called style control that statistically factors out superficial response traits — length, markdown formatting, sycophancy — from leaderboard scores, the same way social science studies control for confounding variables. This prevents models from gaming rankings by sounding polished rather than being genuinely useful or accurate.
- •Occupational Segmentation for Enterprise: Arena segments its 60M monthly conversations by occupation and use case — 28% coding, 6% legal, 6% medical — and offers enterprises an analytical tool to identify which model performs best for their specific domain. Enterprises can privately test models during development without public score release, enabling faster, data-driven model upgrade decisions.
- •Agentic Evaluation Expansion: Arena launched WebDev Arena (Corena) to evaluate AI agents on end-to-end tasks like building web applications, tool calling, and navigating codebases. The roadmap extends to Python and C++ coding agents, multimodal editing, deep research, and multi-step planning tasks — tracking AI capability shifts from single-turn chat toward long-horizon autonomous workflows.
Notable Moment
When asked whether investor relationships with OpenAI, Google, and Anthropic compromise neutrality, the cofounders argued the opposite: those companies actively want truthful rankings because accurate evaluations serve their own scientific and product development needs, making them structurally motivated to support honest results.
Episode Transcript
Presented by dot tech domains, where tech founders find sharp, memorable names for their tech startups. Hello, and welcome back to Equity, TechCrunch's flagship podcast about the business of startups. I'm Rebecca Balan, and this is the episode where we bring on industry experts to help us explore a trend in the tech world and dive deep. AI models are multiplying fast. Competition is stiff, and the question is, which one will be the best and who gets to decide that? Well, Arena, formerly LM Arena, has emerged as the de facto public leaderboard for Frontier LLMs. They're influencing funding, launches, and PR cycles. So today, we're joined by arena cofounders Anastasios Angelopoulos and Wei Lin Chang. Anastasios Whelan, welcome to the show. Thanks so much for having us. Yeah. Thanks. Tell me a little give me, like, a quick background about both of you so our listeners can know who they're who they're talking to. Yeah. So Waelin and I I'm Anastasis. I'm the CEO of Arena. Waelin is our CTO. Waelin and I met in graduate school, you know, three years ago or so when we, were at that time just PhD students at University of California, Berkeley, working on trying to deal with the consequences of CHaD GPT and understand how do we evaluate LLMs. At the time, CHaD GPT had, you know, just recently been released, and there were some new models coming out. We didn't know how to compare them, in particular, on the distribution of real world users. So it's not just about a static benchmark and understanding how well a model takes a test, but can we measure its intelligence in the real world? We've been working on this since early twenty twenty three, like Anas SaaS. At that time, we were, like, PhD student at Berkeley, you know, trying to understand how can we develop new methodology, new system to evaluate chatbots. That was initially started as a research project. A bunch of, you know, PhD students just, like, come together as a team. And, basically, we build a research prototype that allows anyone on the Internet to come and visit a website. We call it shopper arena, and you can ask any questions. There will be two anonymized model, response, and you get to choose which one you prefer. And we use this kinda, like, AOWISE preference feedback data to construct the evaluations and then release published leaderboards, covering, you know, all the best frontier AIs, you know, as we have been tracking this over time since the day of ChartGPT just came out, Cloud and Gemini, and these days, many, many others. Okay. So you've recently changed. So it was Chatbot Arena, then it's LM Arena, and you've recently dropped the LM. Now it's just Arena. What We dropped it. It was cleaner. It's cleaner. That's it. It was cleaner. You know? People What if people just keep calling you l m arena? What if it doesn't stick like Twitter, …
Get the full transcript (4,378 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 23-minute episode.
Get Equity summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Equity
An ex-Anthropic researcher’s doomsday warning comes at a very interesting time
Sep 11 · 39 min
20VC (20 Minute VC)
20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena
Aug 3
More from Equity
ControlAI's Connor Leahy on why superintelligence is ‘not a weapon, it's an adversary’
Sep 9 · 43 min
This Week in Startups
How the 1% Will Own Compute (and What It Means for You)
May 13
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“Arena launched WebDev Arena (Corena) to evaluate AI agents on end-to-end tasks like building web applications, tool calling, and navigating codebases.”
“Arena (formerly LM Arena and Chatbot Arena), cofounded by Berkeley PhD students Anastasios Angelopoulos and Wei-Lin Chiang, operates the de facto public leaderboard for frontier AI models.”
“Arena launched WebDev Arena (Corena) to evaluate AI agents on end-to-end tasks like building web applications, tool calling, and navigating codebases.”
company
“Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation”
“Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation”
“Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation”
“Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation”
“Backed by a16z, Kleiner Perkins, OpenAI, Google, and Anthropic at a $1.7B valuation”
other
“Static benchmarks like Humanity's Last Exam become obsolete once models train on their questions — a problem called overfitting.”
More from Equity
We summarize every new episode. Want them in your inbox?
An ex-Anthropic researcher’s doomsday warning comes at a very interesting time
ControlAI's Connor Leahy on why superintelligence is ‘not a weapon, it's an adversary’
Apple's Ternus era begins as Nvidia bets on the whole AI stack
We're ‘dangerously close’ to dead internet theory according to Pangram's CEO
Will TikTok and YouTube follow Meta’s new rules for teens?
Similar Episodes
Related episodes from other podcasts
20VC (20 Minute VC)
Aug 3
20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena
This Week in Startups
May 13
How the 1% Will Own Compute (and What It Means for You)
Latent Space
Jan 6
[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Latent Space
Dec 31
[State of Evals] LMArena's $100M Vision — Anastasios Angelopoulos, LMArena
The Daily (NYT)
Jun 17
The Battle Over A.I. in the Classroom
Explore Related Topics
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Equity.
Every Monday, we deliver AI summaries of the latest episodes from Equity and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime