Skip to main content
The AI Breakdown

The Best Way to Test New AI Models

49 min episode · 2 min read
·

Episode

49 min

Read time

2 min

Topics

Relationships, Investing, Fundraising & VC

AI-Generated Summary

Key Takeaways

  • ✓Task Selection Framework: Choose four to six benchmark tasks representing your actual work distribution, prioritizing high-frequency tasks where errors are identifiable. Critically, include at least one "wish list" task — something AI previously handled poorly — so new model releases can be evaluated against previously unachievable use cases rather than only confirming existing capabilities.
  • ✓Blind Testing Protocol: Anonymize model outputs before scoring by labeling them A through G, using a script, spreadsheet randomization, or a trusted colleague. Nufar's live test revealed her own predictions about which model produced which output were frequently wrong, confirming that named-model bias significantly distorts evaluation when identities are visible during scoring.
  • ✓AI Judge Limitations: Using a separate model family as evaluator reduces nepotism bias — models tend to favor outputs from their own family. However, Nufar's results showed her scores and the Gemini judge disagreed on four of five tasks, suggesting human judgment should remain primary, with AI scoring treated as supplementary data rather than a deciding factor.
  • ✓Three-Outcome Decision Structure: After benchmarking, choose one of three outcomes — switch entirely, split usage across models by task type, or stay with the current setup. Cost, latency, company policy, subscription tier, and habit-formation costs all factor into the decision independently of raw performance scores, meaning the highest-scoring model is not automatically the correct choice.
  • ✓OpenRouter for Multi-Model Testing: Running benchmarks through OpenRouter with a single API key allows simultaneous evaluation of seven or more models while automatically tracking per-task cost and response time. Nufar's results showed Fable 5.1 scored highest with the AI judge but carried costs an order of magnitude above GPT-4.1 Sol, which matched Fable on her personal scoring at a fraction of the price.

What It Covers

Nufar Gaspar and NLW present a five-step framework for building personal AI benchmarks, demonstrating how to evaluate models like Claude Opus 5.5, GPT-4.1 Sol, and Grok against individual work tasks rather than relying on published benchmark scores that carry limited real-world relevance.

Key Questions Answered

  • •Task Selection Framework: Choose four to six benchmark tasks representing your actual work distribution, prioritizing high-frequency tasks where errors are identifiable. Critically, include at least one "wish list" task — something AI previously handled poorly — so new model releases can be evaluated against previously unachievable use cases rather than only confirming existing capabilities.
  • •Blind Testing Protocol: Anonymize model outputs before scoring by labeling them A through G, using a script, spreadsheet randomization, or a trusted colleague. Nufar's live test revealed her own predictions about which model produced which output were frequently wrong, confirming that named-model bias significantly distorts evaluation when identities are visible during scoring.
  • •AI Judge Limitations: Using a separate model family as evaluator reduces nepotism bias — models tend to favor outputs from their own family. However, Nufar's results showed her scores and the Gemini judge disagreed on four of five tasks, suggesting human judgment should remain primary, with AI scoring treated as supplementary data rather than a deciding factor.
  • •Three-Outcome Decision Structure: After benchmarking, choose one of three outcomes — switch entirely, split usage across models by task type, or stay with the current setup. Cost, latency, company policy, subscription tier, and habit-formation costs all factor into the decision independently of raw performance scores, meaning the highest-scoring model is not automatically the correct choice.
  • •OpenRouter for Multi-Model Testing: Running benchmarks through OpenRouter with a single API key allows simultaneous evaluation of seven or more models while automatically tracking per-task cost and response time. Nufar's results showed Fable 5.1 scored highest with the AI judge but carried costs an order of magnitude above GPT-4.1 Sol, which matched Fable on her personal scoring at a fraction of the price.

Notable Moment

After completing her blind evaluation across six tasks and seven models, Nufar discovered she was ready to shift her primary tool away from Claude — her long-standing default — toward GPT, a conclusion she described as genuinely unexpected and one she would not have reached without running the structured blind test.

Know someone who'd find this useful?

Episode Transcript

We are currently in the fall glut of new models. Over the past several weeks, we've gotten a slew of new closed models at the Frontier, OPUS five five, Astra six, Soul six one, with more on the way. And from those same labs, we've also gotten more cost effective and faster models like Sonnet 5.5. And then as if all that wasn't enough, we've also gotten a number of new open weight models, which are increasingly relevant for lots and lots of different types of businesses. Of And, course, every time one of these models gets released, it's released with benchmarks. But benchmarks don't really tell us much about how it's going to be relevant for us personally. For that, you need to build your own personal AI benchmark, a system to tell if and how a model matters for your own personal AI stack. And that is the goal of today's operator focused episode. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. Alright, friends. Quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, robots and pencils, Harbor, and Blitzy. To get an ad free version of the show, go to patreon.com/aidailybrief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors@aidailybrief.ai. One more quick note before we get into this. This is a recording of the webinar that was hosted by me and Nufar Gaspar last week. Given that already we've gotten another couple new models this week, I think the system that is discussed here continues to be extremely pertinent. Hopefully you find this useful, and tomorrow we will be back with a normal episode. But for now, let's talk about building a personal AI benchmark. We prepared some very fun materials to you, so we will share everything at the end. You'll see throughout what are the materials. But basically, everything that you see me doing and talking about, you will get everything in order to try and do a similar process for yourself and do leverage the Q and A box because we have a large number, so we're unable to open microphones. But Dan is here and Nathaniel is here and everybody will try to cater to your questions as best we can. Without further ado, Nathaniel, some motivation on why we are here. Yeah. Thanks everyone for being here. I think for those of you who are regular listeners of AIDB, you'll know that this is a soapbox y issue for me when it comes to benchmarks. Every time we get a new model, which by the way is now, I think, eleven days on average for a Frontier model, one of my standard cautions is basically, first, to be skeptical of the published benchmarks. Not so much because we think that any of these companies are lying. That's not the nature of what they …

Get the full transcript (9,575 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The AI Breakdown transcripts →

You just read a 3-minute summary of a 46-minute episode.

Get The AI Breakdown summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from The AI Breakdown

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The AI Breakdown.

Every Monday, we deliver AI summaries of the latest episodes from The AI Breakdown and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime