Skip to main content
No Priors: Artificial Intelligence | Technology | Startups

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

36 min episode · 2 min read
·
Openai Research Scientist Noam

Episode

36 min

Read time

2 min

Topics

Investing, Startups, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • Benchmark evaluation methodology: Standard benchmark grids comparing models on single scores are misleading because they ignore test-time compute allocation. When OpenAI released a recent model, initial skepticism faded once users discovered it was more compute-efficient than its predecessor—not weaker. Evaluators should plot performance against a token, cost, or time budget rather than reporting a single number.
  • Safety framework gap: Responsible scaling policies and preparedness frameworks were designed before test-time compute scaling existed. A model's dangerous capability ceiling is now a direct function of inference budget—$10 versus $10,000 versus $10,000,000 produces meaningfully different outputs. No current policy explicitly defines which budget level triggers safety thresholds, leaving a structural blind spot.
  • Performance plateau timelines: Modern frontier models, when scaffolded properly, can continue improving on benchmarks for weeks without plateauing—unlike GPT-3-era models that saturated quickly. This makes "run until plateau" an impractical evaluation standard. A viable alternative is extrapolating performance curves from lower budgets (e.g., $10–$100) to project behavior at $10,000 scale.
  • Unexplored capability overhang: Frontier models already contain capabilities that researchers have not fully mapped because the model release cycle (every two to three months) is shorter than the time required to push models to their limits. The Erdős unit distance conjecture was disproved using an internal OpenAI model at a relatively low inference budget before anyone had systematically tested what $100,000 of compute on a public model could produce.
  • Research taste as the bottleneck: Models accelerate coding, optimization, and algorithm implementation—Brown estimates a 5–10x speed gain on his poker solver work—but consistently fail at generating novel research directions without human steering. The current constraint is not raw reasoning capacity but the absence of genuine research taste, which remains the non-automatable input researchers should protect and develop.

What It Covers

OpenAI research scientist Noam Brown joins Sarah Guo on No Priors to explain why standard benchmark grids misrepresent modern AI model capabilities, how test-time compute scaling breaks existing safety evaluation frameworks, and what the current ceiling of frontier models actually looks like in practice.

Key Questions Answered

  • Benchmark evaluation methodology: Standard benchmark grids comparing models on single scores are misleading because they ignore test-time compute allocation. When OpenAI released a recent model, initial skepticism faded once users discovered it was more compute-efficient than its predecessor—not weaker. Evaluators should plot performance against a token, cost, or time budget rather than reporting a single number.
  • Safety framework gap: Responsible scaling policies and preparedness frameworks were designed before test-time compute scaling existed. A model's dangerous capability ceiling is now a direct function of inference budget—$10 versus $10,000 versus $10,000,000 produces meaningfully different outputs. No current policy explicitly defines which budget level triggers safety thresholds, leaving a structural blind spot.
  • Performance plateau timelines: Modern frontier models, when scaffolded properly, can continue improving on benchmarks for weeks without plateauing—unlike GPT-3-era models that saturated quickly. This makes "run until plateau" an impractical evaluation standard. A viable alternative is extrapolating performance curves from lower budgets (e.g., $10–$100) to project behavior at $10,000 scale.
  • Unexplored capability overhang: Frontier models already contain capabilities that researchers have not fully mapped because the model release cycle (every two to three months) is shorter than the time required to push models to their limits. The Erdős unit distance conjecture was disproved using an internal OpenAI model at a relatively low inference budget before anyone had systematically tested what $100,000 of compute on a public model could produce.
  • Research taste as the bottleneck: Models accelerate coding, optimization, and algorithm implementation—Brown estimates a 5–10x speed gain on his poker solver work—but consistently fail at generating novel research directions without human steering. The current constraint is not raw reasoning capacity but the absence of genuine research taste, which remains the non-automatable input researchers should protect and develop.

Notable Moment

Brown describes asking a model to verify a basic poker calculation—$100 in the pot, player folds—and receiving the answer $92. When challenged, the model defended the wrong figure as close enough. This specific failure mode, models confidently rationalizing errors, drove his emphasis on systematic verification over trust.

Know someone who'd find this useful?

Episode Transcript

With GPT three, you couldn't scale test time compute. Like, if you gave it a budget of $10,000,000 and said, okay. Well, let's see what GPT three can do. It really can't do that much. The precarrenous frameworks and responsible scaling policies, they don't really account for the amount of test time compute. They just say, okay. Well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically. If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10,000,000, it could do even more. More. At what budget should you evaluate these models? The policies that exist today don't really address that question. Hi, listeners. I'm Sarah Gore, and welcome back to No Priors. Today, I'm here with Noam Brown, one of our godfathers of AI reasoning. We talk about the broken state of evaluations, very large scale test time compute, how he thinks about recursive self improvement, and what's next on the horizon for competition at the frontier. Welcome. Noam, I'm so excited to have you back. That's great to be back. Yeah. You are our first guest. I'm very proud of my taste in friends and researchers for, for the pod, given, you know, how important, you know, inference time scaling has become to the industry. You should be proud too having actually pioneered it. Played apart. Yeah. My among many others. You just wrote this essay that really resonated about large scale test time compute and why, the industry is not evaluating these models as robustly as it should be. What was the motivation for it? Yeah. The motivation was we released 5.5, and the initial reaction was kinda skepticism that it was a substantially better model. This to be fair, that only lasted for a few hours before people had some time to play around with it and and try it out themselves, and they saw that it was actually substantially better. But I think a lot of the skepticism came from the benchmark grid that was published. Basically, whenever a new model is released, there is this benchmark grid where they show all these different benchmarks on on the x axis and then the performance of different models on the y axis, and you can just, like, compare different models. It's like a single number for a model on a single benchmark. And if you look on paper at the difference between, like, five point five and five point four or or other models, it wasn't it was an improvement, but it wasn't a huge improvement. It was only a few percentage points in some benchmarks. So people looked at that, and they were skeptical that it was actually a better model. Once they played around with it, the story changed. I think the …

Get the full transcript (7,615 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all No Priors: Artificial Intelligence | Technology | Startups transcripts →

You just read a 3-minute summary of a 33-minute episode.

Get No Priors: Artificial Intelligence | Technology | Startups summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from No Priors: Artificial Intelligence | Technology | Startups

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into No Priors: Artificial Intelligence | Technology | Startups.

Every Monday, we deliver AI summaries of the latest episodes from No Priors: Artificial Intelligence | Technology | Startups and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime