Claude Opus 5 review: this model is brilliant (but annoying)
Episode
24 min
Read time
2 min
Topics
Investing, Fundraising & VC, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
- ✓Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
- ✓Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
- ✓Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
- ✓Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."
What It Covers
Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.
Key Questions Answered
- •Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
- •Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
- •Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
- •Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
- •Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."
Notable Moment
When the host told Opus 5 that nobody trusts it, the model responded by generating its own list of reasons it deserves distrust — then advised the host not to advocate for AI on its behalf and to avoid telling people that AI changes everything, citing potential social friction.
Episode Transcript
You guys, I'm tired. What I'm tired of is models coming out every week. New models, new benchmarks, new frontier intelligence, new things to test. It's been a little bit of a run the past month. We've seen Fable come and go and come again. We've seen GPT five six. We've seen SONNET five. Lots of so many fives recently and just so many models. And I've been lucky I've been able to test these models, been able to play with them for, you know, sometimes days, sometimes weeks. It just depends on who I'm working with. And it's been really interesting and exciting to have access to all this frontier intelligence. But I think we have an intelligence overhang. I really think that we're running out of and by we, I mean, the average coder, average software engineer, average creator, average builder, average consumer, average business person. I think we're running out of ways to truly leverage this incremental intelligence. So this is my hypothesis. In the next year, we're always talking a lot more about speed, talking more about cost. We're talking more about open source, and we're gonna be talking a little less about intelligence. Although, I think we might be talking about specific types of intelligence other than software engineering. But despite being tired, today, we are going to talk about Opus five, baby. Opus five is here. So we got point two additional Opus points, Opus Opals, whatever. However, we're tracking the increments here on Opus. Opus five is here. I've been able to test it a little bit. I have some opinions. Now some of the stuff that I cover this episode is gonna be a little different than what I've done in the past. Yes. We're gonna do the how IAI benchmark live. And, yes, we are gonna look at the prototypes. We're gonna look at PRDs, and we're gonna look at agent personality. But I'm also going to put on my large language model psychologist hat, and we're gonna talk about Opus's personality. And we're gonna talk about Opus's personality relative to GPT's personality because I think this is super interesting. If you're thinking about what is the difference really between these models and you don't wanna look at the difference in terms of benchmark capability. You really wanna understand what these labs are going for, why these models are being built, and how they're being tuned. Looking at their personality at this moment where intelligence is very high is super fun. So we're gonna do a little of that. We're gonna do the Howat AI benchmark. We might do some live coding. We're not gonna cover too much of the specs in the model because read the blog post. Read the blog post. We'll link to it in the show notes. What we're really gonna talk about is, is Opus five good? Am I gonna swap it in? And how is it different than the other Frontier models on the market? …
Get the full transcript (4,168 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 21-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Anthropic
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
by OpenAI
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
by Anthropic
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
by Google
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
by Anthropic
“Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.”
by Anthropic
“Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.”
More from How I AI
We summarize every new episode. Want them in your inbox?
GPT-6 Astra is a banger - here’s everything I’ve built
Grok Bot vs. OpenClaw: How I replaced my entire agent stack
How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take
Similar Episodes
Related episodes from other podcasts
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime