Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?
Episode
25 min
Read time
2 min
Topics
Leadership, Design & UX, Marketing
AI-Generated Summary
Key Takeaways
- ✓Model-specific workflows: Opus 4.5 creates structured to-do lists before coding (redesign listing page, improve layout, enhance post display, add SEO), while Gemini 3 executes immediately without planning steps. This planning capability produces 20-30% better design details and functional improvements across components.
- ✓Design quality hierarchy: Opus 4.5 delivers superior results with asset selection from existing repositories, placeholder images for missing content, hover animations with call-to-action arrows, and reading time estimates. Gemini 3 produces serviceable designs with glass morphism cards and hero sections but lacks refinement in spacing and asset handling.
- ✓SEO implementation differences: Gemini 3 adds JSON-LD schema, breadcrumbs, semantic HTML, and related articles to individual posts. Opus 4.5 implements metadata, Open Graph tags, and structured data but skips JSON-LD. Codex 5.1 provides minimal SEO with basic metadata and schema.org embedding only.
- ✓Model specialization strategy: Use different models for different workflow stages rather than one model for everything. Opus 4.5 excels at front-end design, Codex 5.1 performs better on back-end engineering tasks, and Gemini 3 handles mid-tier design work requiring less detailed planning and implementation steps.
What It Covers
Claire Vo tests three leading AI coding models—Gemini 3 Pro, Claude Opus 4.5, and GPT-5.1 Codex—by having each redesign an existing blog page to determine which performs best at front-end design work.
Key Questions Answered
- •Model-specific workflows: Opus 4.5 creates structured to-do lists before coding (redesign listing page, improve layout, enhance post display, add SEO), while Gemini 3 executes immediately without planning steps. This planning capability produces 20-30% better design details and functional improvements across components.
- •Design quality hierarchy: Opus 4.5 delivers superior results with asset selection from existing repositories, placeholder images for missing content, hover animations with call-to-action arrows, and reading time estimates. Gemini 3 produces serviceable designs with glass morphism cards and hero sections but lacks refinement in spacing and asset handling.
- •SEO implementation differences: Gemini 3 adds JSON-LD schema, breadcrumbs, semantic HTML, and related articles to individual posts. Opus 4.5 implements metadata, Open Graph tags, and structured data but skips JSON-LD. Codex 5.1 provides minimal SEO with basic metadata and schema.org embedding only.
- •Model specialization strategy: Use different models for different workflow stages rather than one model for everything. Opus 4.5 excels at front-end design, Codex 5.1 performs better on back-end engineering tasks, and Gemini 3 handles mid-tier design work requiring less detailed planning and implementation steps.
Notable Moment
Codex 5.1 generated a purple-to-blue gradient background—the stereotypical AI design aesthetic—and selected a white logo that was illegible against the colored background, demonstrating poor visual design judgment despite being OpenAI's leading coding model.
Episode Transcript
Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, I have a really fun mini episode where I'm gonna answer the question on everyone's mind. Which of these new models is actually the best designer? I'm gonna take a page on my site that I don't think is particularly well designed and have Gemini three, Opus four five, and Codex five one duke it out and see which one can redesign my page better one shot. Let's get to it. This episode is brought to you by Lovable. If you've ever had an idea for an app but didn't know where to start, Lovable is for you. Lovable lets you build working apps and websites by simply chatting with AI. Then you can customize it, add automations, and deploy it to a live domain. It's perfect for marketers spinning up tools, product managers prototyping new ideas, or founders launching their next business. Unlike no code tools, Lovable isn't about static pages. It builds full apps with real functionality, and it's fast. What used to take weeks, months, or even years, you can now do over the weekend. So if you've been sitting on an idea, now's the time to bring it to life. Get started for free at lovable.dev. That's lovable.dev. If you've been paying attention in the last couple weeks, it seems like every single model provider has released a brand new coding model. And what I heard the most from people is, sure, they're fast and sure, they're great and sure, they're beating benchmarks, but they are all really good at design. If you've been on x or social media, you've probably seen these beautifully designed landing pages, apps, and user experience components generated using Gemini three or Opus four five or even codex five one. And I thought, let's put these side by side and actually see which one's better at redesigning an existing page. I think it's easy to one shot something and make it look beautiful, especially if you're a great prompter and know exactly what to say as a designer. But if you have an existing site and you wanna make it better, who's your trusted design engineer? Which of these models is really gonna do the trick? And I'm gonna show you what I think today in a couple minutes on which of these models is the better designer or redesigner of a page that I don't think is really great. So this is the chat PRD blog. It is not very good. I don't think this is a very beautiful site. It's not my favorite. I think it could be a lot better. And it could be a lot better from a functional perspective, but it can also be a lot better from a design perspective. And, you know, if I had a team, which I have a little small one, but if …
Get the full transcript (4,376 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 22-minute episode.
Get How I AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from How I AI
Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy
Sep 7 · 50 min
The Vergecast
Pencils down: We share our vibe-coded websites
Aug 10
More from How I AI
GPT-6 Astra is a banger - here’s everything I’ve built
Sep 3 · 32 min
The AI Breakdown
Why AI Users Are Raving About GLM 5.2
Jun 22
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
by Google
“Claire Vo tests three leading AI coding models—Gemini 3 Pro, Claude Opus 4.5, and GPT-5.1 Codex—by having each redesign an existing blog page to determine which performs best at front-end design work.”
by Anthropic
“Claire Vo tests three leading AI coding models—Gemini 3 Pro, Claude Opus 4.5, and GPT-5.1 Codex—by having each redesign an existing blog page to determine which performs best at front-end design work.”
“SPONSORS: Lovable (https://lovable.dev)”
by OpenAI
“Claire Vo tests three leading AI coding models—Gemini 3 Pro, Claude Opus 4.5, and GPT-5.1 Codex—by having each redesign an existing blog page to determine which performs best at front-end design work.”
More from How I AI
We summarize every new episode. Want them in your inbox?
Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy
GPT-6 Astra is a banger - here’s everything I’ve built
Grok Bot vs. OpenClaw: How I replaced my entire agent stack
How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)
I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)
Similar Episodes
Related episodes from other podcasts
The Vergecast
Aug 10
Pencils down: We share our vibe-coded websites
The AI Breakdown
Jun 22
Why AI Users Are Raving About GLM 5.2
The AI Breakdown
Mar 26
Why AI Needs Better Benchmarks
The Startup Ideas Podcast
Nov 26
Reviewing Claude Opus 4.5
Hard Fork
Jul 3
Fable Ban Reversed + Dr. Dana Suskind on Parenting With A.I. + Prediction Market Drama
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into How I AI.
Every Monday, we deliver AI summaries of the latest episodes from How I AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime