Sonnet 5 review: I ran 64 generations to find out if it's worth it
Episode
25 min
Read time
2 min
Topics
Investing, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Benchmark design: Build repeatable AI evals using frozen inputs, blind scoring, and a structured rubric rather than one-off vibe checks. Claude Code can scan past session history stored on your desktop to suggest relevant benchmark tasks tailored to your actual workflows, making setup faster and more personalized than starting from scratch.
- ✓Sonnet 5 pricing window: Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens, with Anthropic confirming prices rise after summer 2025. For teams running high-volume agentic workloads, testing and locking in usage now captures near-Opus performance at a significant cost discount before the pricing structure changes.
- ✓Model-by-task routing: No single frontier model wins across all tasks. GPT-5.5 produces the most comprehensive PRDs, Sonnet 4.6 performs best for UI prototyping and conversational agents, and Opus 4.8 handles dense, complex UI generation. Routing prompts to the right model by task type outperforms defaulting to one model for everything.
- ✓Human vs. LLM judgment gap: When the host's 70% human-weighted scores were combined with 30% automated LLM scores, rankings flipped significantly from pure LLM evaluation. LLM judges cluster scores near the middle of the scale and miss visual taste signals, making human review essential for design and writing quality assessments.
- ✓Agentic benchmark saturation: Standard multi-step coding tasks no longer differentiate frontier models because GPT-5.5, Gemini 2.5 Pro, Opus 4.8, and Sonnet 5 all score similarly. Effective agentic evals need harder, more specialized tasks. Retiring saturated benchmarks and replacing them with higher-difficulty challenges is necessary to surface meaningful capability differences between models.
What It Covers
Host introduces the "How I AI Bench," a repeatable evaluation framework testing Claude Sonnet 5 against GPT-5.5, Gemini 2.5 Pro, Opus 4.8, and Sonnet 4.6 across 64 generations spanning PRD writing, UI prototyping, agentic coding, and voice personality tasks, revealing surprising model rankings.
Key Questions Answered
- •Benchmark design: Build repeatable AI evals using frozen inputs, blind scoring, and a structured rubric rather than one-off vibe checks. Claude Code can scan past session history stored on your desktop to suggest relevant benchmark tasks tailored to your actual workflows, making setup faster and more personalized than starting from scratch.
- •Sonnet 5 pricing window: Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens, with Anthropic confirming prices rise after summer 2025. For teams running high-volume agentic workloads, testing and locking in usage now captures near-Opus performance at a significant cost discount before the pricing structure changes.
- •Model-by-task routing: No single frontier model wins across all tasks. GPT-5.5 produces the most comprehensive PRDs, Sonnet 4.6 performs best for UI prototyping and conversational agents, and Opus 4.8 handles dense, complex UI generation. Routing prompts to the right model by task type outperforms defaulting to one model for everything.
- •Human vs. LLM judgment gap: When the host's 70% human-weighted scores were combined with 30% automated LLM scores, rankings flipped significantly from pure LLM evaluation. LLM judges cluster scores near the middle of the scale and miss visual taste signals, making human review essential for design and writing quality assessments.
- •Agentic benchmark saturation: Standard multi-step coding tasks no longer differentiate frontier models because GPT-5.5, Gemini 2.5 Pro, Opus 4.8, and Sonnet 5 all score similarly. Effective agentic evals need harder, more specialized tasks. Retiring saturated benchmarks and replacing them with higher-difficulty challenges is necessary to surface meaningful capability differences between models.
Notable Moment
The live leaderboard reveal produced an unexpected result: Gemini 2.5 Pro, a model the host had nearly forgotten was included in the test, tied for first place on the automated scoring, while the newly released Sonnet 5 landed at the bottom of the host's personal preference ranking.
Episode Transcript
We've got a new model, people, and it's from Anthropic. Now is it Mythos? No. Is it Fable? No. But it is Claude SONNET five. Anthropic is claiming it's the most agentic SONNET model yet, and we will get opus level tasks at Sonnet level prices. Now I've been testing a lot of models, and I'm starting to get bored of doing the vibe check. What I wanna start developing is a set of benchmarks we can regularly test these new models against that you'll care about. So today, I'm going to be introducing the How I AI Bench, a set of AI and Claro graded benchmarks that are gonna tell us if this model and any model is good at writing PRDs, solving bugs, and one shotting designs. I'm gonna show you exactly how I built this benchmark using Claude Code, and we're gonna see on a blind test what comes out on top. Let's get to it. This episode is brought to you by Runway, a new kind of creative platform that has everything you need to generate any image, video, or piece of content you want, all in one place. With Runway, it's now possible to go from initial idea to a finished deliverable in a matter of minutes. From turning low fidelity product shots into campaign ready imagery all the way through putting together big brand films, Runway can help your team scale your creative ambitions while keeping your budgets and timelines from doing the same. Runway brings together the world's most advanced AI models, which is why enterprises like Microsoft, Robinhood, Amazon, and Adobe, along with studios like Lionsgate and Legendary, all use Runway to ship real work every day. Try it yourself at runwayml.com/howiai, promo code how iai. Quickly, before we get to our evals, let's just talk about the headlines of Sonnet five, this new model. Anthropic is pitching it as close to the performance of Opus four eight, but much less expensive. So as you can see here, it's not quite at this 69% on Agenca coating, SuiteBench Pro, or the 82% on terminal bench 2.1, but it's not that far behind. And I suspect that most of us are not going to notice the difference. It's also supposed to be really good at computer work and knowledge work, and so this should be an everyday model that people reach for. In my episode with Felix from Anthropic, he says that we're all abusing opus, and we should definitely be using the sonnet models more, and we are going to put SONNET five to the test against that proposition. Now what do they say that SONNET five is really good at? Well, it's really good at agentic tool use. So you're gonna get slightly longer running tool runs, longer running sessions than you would with SONNET four six at a lower cost than doing the same comparable task with Opus. So you're gonna see here, you know, SONNET four six, a lower …
Get the full transcript (4,213 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 22-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Aug 10 · 43 min
No Priors: Artificial Intelligence | Technology | Startups
Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
Jun 26
More from How I AI
Build an AI code review bot in 30 minutes with Vercel Eve
Aug 5 · 24 min
The AI Breakdown
Fable 5 Raises the Bar for AI Ambition
Jun 10
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
- Claude CodeRecommended
by Anthropic
“Claude Code can scan past session history stored on your desktop to suggest relevant benchmark tasks tailored to your actual workflows, making setup faster and more personalized than starting from scratch.”
More from How I AI
We summarize every new episode. Want them in your inbox?
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Build an AI code review bot in 30 minutes with Vercel Eve
ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)
From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun
Claude Opus 5 review: this model is brilliant (but annoying)
Similar Episodes
Related episodes from other podcasts
No Priors: Artificial Intelligence | Technology | Startups
Jun 26
Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
The AI Breakdown
Jun 10
Fable 5 Raises the Bar for AI Ambition
Huberman Lab
Apr 13
How Women Can Improve Their Fertility & Hormone Health | Dr. Natalie Crawford
The AI Breakdown
Apr 1
Introducing Maturity Maps — A New Way to Measure AI Adoption
Latent Space
Jan 8
Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime