Claude Opus 5 review: this model is brilliant (but annoying)
Episode
24 min
Read time
2 min
Topics
Investing, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
- ✓Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
- ✓Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
- ✓Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
- ✓Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."
What It Covers
Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.
Key Questions Answered
- •Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
- •Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
- •Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
- •Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
- •Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."
Notable Moment
When the host told Opus 5 that nobody trusts it, the model responded by generating its own list of reasons it deserves distrust — then advised the host not to advocate for AI on its behalf and to avoid telling people that AI changes everything, citing potential social friction.
You just read a 3-minute summary of a 21-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
Computer & browser use in Codex (5 real examples)
Jul 22 · 27 min
The AI Breakdown
Is Kimi K3 Really Fable Class?
Jul 17
More from How I AI
How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman
Jul 20 · 42 min
The AI Breakdown
Why AI Users Are Raving About GLM 5.2
Jun 22
More from How I AI
We summarize every new episode. Want them in your inbox?
Computer & browser use in Codex (5 real examples)
How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman
This solo builder runs 24/7 local AI on his own hardware | Alex Finn
GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark
What a harness is and how to build one with Claude Agent SDK
Similar Episodes
Related episodes from other podcasts
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime