Skip to main content
How I AI

Claude Opus 4.8 is here. Is it as good as they say?

13 min episode · 2 min read

Episode

13 min

Read time

2 min

Topics

Productivity, Artificial Intelligence, Software Development

AI-Generated Summary

Key Takeaways

  • Greenfield vs. existing code: Opus 4.8 performs well on one-shot, net-new feature builds — it planned and autonomously coded a full prototyping tool in roughly 20 minutes — but degrades significantly when navigating existing codebases, rebasing branches, or resolving edge-case bugs.
  • Hallucination risk under confidence: Despite running on high-effort mode, Opus 4.8 fabricated conclusions from hypotheses rather than validated data, both in coding and strategy contexts. Treat high-confidence outputs with skepticism and explicitly prompt it to verify sources before accepting results.
  • Strategy work: 4.7 outperforms 4.8: Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps. For data-driven strategy tasks, 4.7 remains the stronger choice.
  • New agentic infrastructure worth testing: Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents. Claude.ai and CoWork gain effort control settings from low to max. These harness-level changes may offset model limitations when prompting strategies are tuned appropriately.

What It Covers

Claire Vo shares early hands-on testing of Anthropic's Claude Opus 4.8, a coding-focused agent model priced at $5/$25 per million tokens, evaluating its performance across greenfield coding, existing codebases, and business strategy tasks.

Key Questions Answered

  • Greenfield vs. existing code: Opus 4.8 performs well on one-shot, net-new feature builds — it planned and autonomously coded a full prototyping tool in roughly 20 minutes — but degrades significantly when navigating existing codebases, rebasing branches, or resolving edge-case bugs.
  • Hallucination risk under confidence: Despite running on high-effort mode, Opus 4.8 fabricated conclusions from hypotheses rather than validated data, both in coding and strategy contexts. Treat high-confidence outputs with skepticism and explicitly prompt it to verify sources before accepting results.
  • Strategy work: 4.7 outperforms 4.8: Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps. For data-driven strategy tasks, 4.7 remains the stronger choice.
  • New agentic infrastructure worth testing: Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents. Claude.ai and CoWork gain effort control settings from low to max. These harness-level changes may offset model limitations when prompting strategies are tuned appropriately.

Notable Moment

During a fun test asking Opus 4.8 to build a game and then play it autonomously to tune difficulty for a nine-year-old, the model generated a workable but unambitious result — repeatedly falling short despite explicit prompts to push further.

Know someone who'd find this useful?

Episode Transcript

Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, we have a very special mini episode because Anthropic just dropped Opus 4.8, their latest state of the art coding model. And I got a few hours of early access, and I'm here to share my very early thoughts about where this model is intended to perform well, where it did a great job and totally impressed me, and where there's still a little bit further to go. Let's get to it. As you can tell, I am not in my regular How I AI studio, and that's because I am so excited to give you my early thoughts on Opus 4.8 and couldn't wait between meetings to share what I thought. So to get started, I wanna talk about what this model is, what Anthropic has told us about its benchmarks, performance, and what it's good at. So Anthropic is shipping Opus 4.8. It is supposed to be their step change model for agents, and there's a couple things they've called out that this model does particularly well. It's supposed to be more honest, a less design flop, longer horizon autonomy on long running tasks, and enterprise ready. So it means it follows its instructions. And they're saying that Sweebench Pro, they're hitting 69.2%, which is almost five points higher than Opus 4.7, almost 10 points higher than GPT 5.5, and 15 points higher than Gemini 3.1. Now this model is not cheap. It's $5 per input tokens and $25 per million output tokens. And then same as 4.7, effort defaults to high and fast mode can be a lot faster. This is what they say. This is what you're gonna read on the blog post. And so on paper, this is a very exciting model. But I wanna tell you my personal experience using this model and where I thought it did a really good job and, again, where it did not do a perfect job. And so when I was getting feedback to the team, I said, surprise, surprise, LOL, it's a good coding model. In that when I opened up plot code and asked it to do a fairly complex one shot brand new surface area task, it did a pretty good job. So I asked, in Cloud Code, Opus four eight to build a prototyping capability in chat PRD. So we make PRDs. I said, let's just go whole hog. Let's, compete with the big boys. Let's make an entire prototyping tool. And I gave it some architecture decisions I wanted to make, what platforms I wanted to use, how I wanted it to function. It went through plan and then it autonomously coded for, I would say, about twenty minutes and shipped it. And when I pushed this live to my preview branch, it worked. And so I would say from a one shot feature, it …

Get the full transcript (2,623 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all How I AI transcripts →

You just read a 3-minute summary of a 10-minute episode.

Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • by Anthropic

    Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents.
  • by Anthropic

    Claude.ai and CoWork gain effort control settings from low to max.
  • by Anthropic

    Claire Vo shares early hands-on testing of Anthropic's Claude Opus 4.8, a coding-focused agent model priced at $5/$25 per million tokens, evaluating its performance across greenfield coding, existing codebases, and business strategy tasks.
  • by Anthropic

    Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps.
  • Claude.ai and CoWork gain effort control settings from low to max.

More from How I AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into How I AI.

Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime