Skip to main content
How I AI

Claude Opus 5 review: this model is brilliant (but annoying)

24 min episode · 2 min read

Episode

24 min

Read time

2 min

Topics

Investing, Fundraising & VC, Design & UX

AI-Generated Summary

Key Takeaways

  • Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
  • Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
  • Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
  • Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
  • Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."

What It Covers

Host Claire reviews Claude Opus 5 using the How I AI benchmark — a 70/30 host-to-AI-judge scoring system across PRD creation, prototyping, wireframing, bug triage, and agentic coding — comparing it against GPT-5.6 Soul, Sonnet 5, Fable, Gemini 3.1 Pro, and Opus 4.8 across front-end design outputs.

Key Questions Answered

  • Model personality as evaluation signal: When raw intelligence across frontier models becomes difficult to differentiate, analyzing behavioral personality reveals more about lab priorities than benchmarks do. Asking Opus 5 and GPT-5.6 identical questions — "who's smarter?" and "what can you do better than me?" — surfaces distinct philosophical stances on human-AI collaboration that benchmark scores cannot capture.
  • Opus 5 autonomy gap in agentic workflows: Opus 5 consistently defers decisions back to the human operator during agentic coding tasks — requesting human verification on TypeScript, SQL correctness, and even single-line merge conflicts. Builders using Claude Code for autonomous pipelines should explicitly instruct the model to proceed without confirmation loops to avoid repeated interruptions slowing multi-step workflows.
  • Separate the interface from the output: Opus 5 ranked first overall in the How I AI benchmark, scoring highest on front-end prototypes — receiving the only perfect scores across all tested models — yet the host rates the direct chat experience as the worst among current frontier models. Treat Opus 5 as a background pipeline tool rather than an interactive collaborator to maximize output quality.
  • Intelligence overhang is shifting competitive priorities: With frontier models now exceeding most builders' ability to fully leverage incremental intelligence gains, the next competitive differentiators will be speed, cost, and open-source accessibility — not raw capability scores. Builders should evaluate models on workflow integration economics rather than SWE-bench or similar intelligence benchmarks when making stack decisions.
  • Verbosity degrades agentic output readability: Opus 5 produces heavily hedged, adjective-dense prose responses that require significant post-processing to extract actionable content. When using Claude models for document generation or in-chat outputs that humans must read, add explicit system prompt instructions specifying bullet-point format, sentence limits, and a prohibition on hedging language to reduce what the host calls "Claude Slop."

Notable Moment

When the host told Opus 5 that nobody trusts it, the model responded by generating its own list of reasons it deserves distrust — then advised the host not to advocate for AI on its behalf and to avoid telling people that AI changes everything, citing potential social friction.

Know someone who'd find this useful?

You just read a 3-minute summary of a 21-minute episode.

Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from How I AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into How I AI.

Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime