Claude Opus 4.8 is here. Is it as good as they say?
Episode
13 min
Read time
2 min
Topics
Productivity, Artificial Intelligence, Software Development
AI-Generated Summary
Key Takeaways
- ✓Greenfield vs. existing code: Opus 4.8 performs well on one-shot, net-new feature builds — it planned and autonomously coded a full prototyping tool in roughly 20 minutes — but degrades significantly when navigating existing codebases, rebasing branches, or resolving edge-case bugs.
- ✓Hallucination risk under confidence: Despite running on high-effort mode, Opus 4.8 fabricated conclusions from hypotheses rather than validated data, both in coding and strategy contexts. Treat high-confidence outputs with skepticism and explicitly prompt it to verify sources before accepting results.
- ✓Strategy work: 4.7 outperforms 4.8: Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps. For data-driven strategy tasks, 4.7 remains the stronger choice.
- ✓New agentic infrastructure worth testing: Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents. Claude.ai and CoWork gain effort control settings from low to max. These harness-level changes may offset model limitations when prompting strategies are tuned appropriately.
What It Covers
Claire Vo shares early hands-on testing of Anthropic's Claude Opus 4.8, a coding-focused agent model priced at $5/$25 per million tokens, evaluating its performance across greenfield coding, existing codebases, and business strategy tasks.
Key Questions Answered
- •Greenfield vs. existing code: Opus 4.8 performs well on one-shot, net-new feature builds — it planned and autonomously coded a full prototyping tool in roughly 20 minutes — but degrades significantly when navigating existing codebases, rebasing branches, or resolving edge-case bugs.
- •Hallucination risk under confidence: Despite running on high-effort mode, Opus 4.8 fabricated conclusions from hypotheses rather than validated data, both in coding and strategy contexts. Treat high-confidence outputs with skepticism and explicitly prompt it to verify sources before accepting results.
- •Strategy work: 4.7 outperforms 4.8: Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps. For data-driven strategy tasks, 4.7 remains the stronger choice.
- •New agentic infrastructure worth testing: Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents. Claude.ai and CoWork gain effort control settings from low to max. These harness-level changes may offset model limitations when prompting strategies are tuned appropriately.
Notable Moment
During a fun test asking Opus 4.8 to build a game and then play it autonomously to tune difficulty for a nine-year-old, the model generated a workable but unambitious result — repeatedly falling short despite explicit prompts to push further.
Episode Transcript
Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, we have a very special mini episode because Anthropic just dropped Opus 4.8, their latest state of the art coding model. And I got a few hours of early access, and I'm here to share my very early thoughts about where this model is intended to perform well, where it did a great job and totally impressed me, and where there's still a little bit further to go. Let's get to it. As you can tell, I am not in my regular How I AI studio, and that's because I am so excited to give you my early thoughts on Opus 4.8 and couldn't wait between meetings to share what I thought. So to get started, I wanna talk about what this model is, what Anthropic has told us about its benchmarks, performance, and what it's good at. So Anthropic is shipping Opus 4.8. It is supposed to be their step change model for agents, and there's a couple things they've called out that this model does particularly well. It's supposed to be more honest, a less design flop, longer horizon autonomy on long running tasks, and enterprise ready. So it means it follows its instructions. And they're saying that Sweebench Pro, they're hitting 69.2%, which is almost five points higher than Opus 4.7, almost 10 points higher than GPT 5.5, and 15 points higher than Gemini 3.1. Now this model is not cheap. It's $5 per input tokens and $25 per million output tokens. And then same as 4.7, effort defaults to high and fast mode can be a lot faster. This is what they say. This is what you're gonna read on the blog post. And so on paper, this is a very exciting model. But I wanna tell you my personal experience using this model and where I thought it did a really good job and, again, where it did not do a perfect job. And so when I was getting feedback to the team, I said, surprise, surprise, LOL, it's a good coding model. In that when I opened up plot code and asked it to do a fairly complex one shot brand new surface area task, it did a pretty good job. So I asked, in Cloud Code, Opus four eight to build a prototyping capability in chat PRD. So we make PRDs. I said, let's just go whole hog. Let's, compete with the big boys. Let's make an entire prototyping tool. And I gave it some architecture decisions I wanted to make, what platforms I wanted to use, how I wanted it to function. It went through plan and then it autonomously coded for, I would say, about twenty minutes and shipped it. And when I pushed this live to my preview branch, it worked. And so I would say from a one shot feature, it …
Get the full transcript (2,623 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 10-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
Aug 24 · 44 min
The AI Breakdown
What To Build First With Claude Design
Apr 20
More from How I AI
I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take
Aug 18 · 27 min
Cognitive Revolution
AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
Jan 9
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Anthropic
“Claude Code now supports dynamic workflows enabling hundreds of parallel sub-agents.”
by Anthropic
“Claire Vo shares early hands-on testing of Anthropic's Claude Opus 4.8, a coding-focused agent model priced at $5/$25 per million tokens, evaluating its performance across greenfield coding, existing codebases, and business strategy tasks.”
by Anthropic
“Side-by-side testing on a business strategy prompt showed Opus 4.7 anchored responses in specific numbers and structured data, while 4.8 produced vague, hand-wavy roadmaps.”
“Claude.ai and CoWork gain effort control settings from low to max.”
More from How I AI
We summarize every new episode. Want them in your inbox?
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
I tested Grok Bot, Grok 4.6, and Cursor Origin - here’s my honest take
How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder
Claude Code for normal people: skills, voice mode, and how to collaborate with AI
Build an AI code review bot in 30 minutes with Vercel Eve
Similar Episodes
Related episodes from other podcasts
The AI Breakdown
Apr 20
What To Build First With Claude Design
Cognitive Revolution
Jan 9
AMA Part 1: Is Claude Code AGI? Are we in a bubble? Plus Live Player Analysis
Software Engineering Daily
Aug 4
AI-Powered Threats to the Software Supply Chain
The AI Breakdown
Jul 27
Where Claude Opus 5 Fits in Your Model Rotation
Hard Fork
Jul 24
OpenAI Models Go Rogue + Kimi K3 Freakout + A.I. Superforecasting
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime