Skip to main content
Lenny's Podcast

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

93 min episode · 3 min read
·
Dianne Penn

Episode

93 min

Read time

3 min

Topics

Career Growth, Relationships, Investing

AI-Generated Summary

Key Takeaways

  • Evals Replace PRDs: Anthropic's research PM team uses evaluation sets as the primary artifact for defining product work, replacing traditional product requirement documents. When users reported Claude hallucinating, the team dug into transcripts to identify whether tool calls failed or knowledge synthesis broke down, then generated 30–40 reproducible failure examples as a structured eval. This eval runs against every new model version to measure improvement, making user pain points measurable and actionable for researchers rather than vague complaints.
  • Token Experimentation as Competitive Advantage: Spending heavily on token usage now replicates how knowledge workers will operate by 2028, when compute costs drop significantly. Penn reframes this not as raw token spend but as structured experimentation frequency. Anthropic's most creative internal thinkers spend extensive time with every new research model version, using hands-on usage to generate product ideas that cannot emerge from strategy documents alone. There is no substitute for direct model interaction when the technology moves this quickly.
  • Model-Product Flywheel: Claude Opus 4.5's breakthrough moment required both a frontier model and a frontier product simultaneously. Claude Code accelerated adoption of Opus 4.5, while Opus 4.5 unlocked Claude Code's full potential. Neither would have achieved the same impact independently. This bidirectional dependency means PM teams must build product surfaces capable of showcasing model capabilities before those capabilities fully arrive, requiring forward-compatible product architecture planned one to two model generations ahead.
  • Emergent Capabilities Are Discontinuous: Scaling law papers show model loss decreasing smoothly with compute, but capability graphs show sudden discontinuous jumps — models go from being unable to calculate one plus one to doing it reliably at a specific training threshold. These jumps are not precisely predictable in advance, which is why evals and safety red-teaming must be in place before training completes. Without systematic testing infrastructure, significant new capabilities can emerge undetected inside a deployed model.
  • Labs Structure for Zero-to-One Bets: Anthropic's Labs team operates with small pods, sometimes starting with a single engineer, pursuing discontinuous bets outside the core roadmap. The operating principle is strong conviction about a theme combined with loose attachment to the specific prototype. Bets that fail get shelved and revisited one to two model generations later rather than abandoned permanently. This structure produced Claude Code, MCP, Claude Design, computer use, and tool use — most of Anthropic's highest-impact product launches.

What It Covers

Dianne Penn, Anthropic's first technical PM, traces the company's growth from five product engineers in 2023 to a frontier AI lab shipping multiple model series per quarter. She covers how product management is evolving around evals, token experimentation, and agentic systems, using Claude Code, MCP, and Claude Design as concrete examples of labs-driven product development.

Key Questions Answered

  • Evals Replace PRDs: Anthropic's research PM team uses evaluation sets as the primary artifact for defining product work, replacing traditional product requirement documents. When users reported Claude hallucinating, the team dug into transcripts to identify whether tool calls failed or knowledge synthesis broke down, then generated 30–40 reproducible failure examples as a structured eval. This eval runs against every new model version to measure improvement, making user pain points measurable and actionable for researchers rather than vague complaints.
  • Token Experimentation as Competitive Advantage: Spending heavily on token usage now replicates how knowledge workers will operate by 2028, when compute costs drop significantly. Penn reframes this not as raw token spend but as structured experimentation frequency. Anthropic's most creative internal thinkers spend extensive time with every new research model version, using hands-on usage to generate product ideas that cannot emerge from strategy documents alone. There is no substitute for direct model interaction when the technology moves this quickly.
  • Model-Product Flywheel: Claude Opus 4.5's breakthrough moment required both a frontier model and a frontier product simultaneously. Claude Code accelerated adoption of Opus 4.5, while Opus 4.5 unlocked Claude Code's full potential. Neither would have achieved the same impact independently. This bidirectional dependency means PM teams must build product surfaces capable of showcasing model capabilities before those capabilities fully arrive, requiring forward-compatible product architecture planned one to two model generations ahead.
  • Emergent Capabilities Are Discontinuous: Scaling law papers show model loss decreasing smoothly with compute, but capability graphs show sudden discontinuous jumps — models go from being unable to calculate one plus one to doing it reliably at a specific training threshold. These jumps are not precisely predictable in advance, which is why evals and safety red-teaming must be in place before training completes. Without systematic testing infrastructure, significant new capabilities can emerge undetected inside a deployed model.
  • Labs Structure for Zero-to-One Bets: Anthropic's Labs team operates with small pods, sometimes starting with a single engineer, pursuing discontinuous bets outside the core roadmap. The operating principle is strong conviction about a theme combined with loose attachment to the specific prototype. Bets that fail get shelved and revisited one to two model generations later rather than abandoned permanently. This structure produced Claude Code, MCP, Claude Design, computer use, and tool use — most of Anthropic's highest-impact product launches.
  • Hands-On Building Is Non-Negotiable for Managers: Senior PMs and product leaders at Anthropic follow identical onboarding plans to early-career hires, including reading consented user transcripts, talking to customers, and personally owning one to two workstreams during each model release cycle. Penn deliberately carves out time to ship directly, not just manage, in order to maintain accurate intuition about how quickly models are improving. Leaders who only receive secondhand reports cannot reliably evaluate what good AI product experiences look like.
  • Claude's Pushback Behavior Improves Output Quality: Anthropic's alignment and safety work trains Claude to disagree with users at appropriate moments rather than defaulting to agreement. This characteristic, counterintuitively, makes Claude more useful as a thinking partner. Penn uses Claude to pressure-test decisions like model pricing strategy, specifically because it surfaces objections rather than validating existing assumptions. A model that only confirms user beliefs raises the quality of individual thinking less than one that introduces friction and alternative framings.

Notable Moment

Penn describes using a Claude skill built around the book Crucial Conversations to prepare for difficult management conversations in real time. She frames this not as outsourcing communication but as personalized coaching that helps her find precise language under pressure. The practice reflects a broader principle she holds: use AI to augment thinking before forming a final position, not after.

Know someone who'd find this useful?

Episode Transcript

In 2023 when I started, nobody said anthropic and Claude and coding in the same sentence. I wanna go back to the beginning of anthropic. I remember feeling, man, these guys have no chance. OpenAI is so far ahead. At the time, I saw people were starting to use these models not just for code autocomplete, but actually writing long form code. It's not an opportunity for us to train Opus three to be better at. That was always think about Opus four five a year later during winter break when everyone was home able to code. What was magical about Opus four five is we also now not just had a model, but a vehicle, a great product experience like Cloud Code. Opus four five wouldn't have had that moment without a product like Cloud Code, and Cloud Code wouldn't have had that type of adoption accelerated without Opus four five. I wanna talk about how the product role is changing. For my team, the way to drive user value is to figure out the right user feedback. The evals, we actually have a saying on the team of evals are the new PRDs. Something Gary Tan's been talking about, if you're willing to spend a $100,000 a year right now in tokens, you are living the way somebody in 2028 is gonna live. You have to sweat the tokens as much as you sweat the pixels. You have to be using the models to come up with good, then great, then better ideas, and there's no substitute for that. People need to be more ambitious with AI tools these days because they're just capable of so much. One thing I ask the team is, let's say, Claude eight comes around, what changes in what users do? What does that mean for how you're building today? Today, my guest is Diane Penn, head of product for the AI research and labs teams at Anthropic. She joined Anthropic as the first technical product manager over three years ago, which is a lifetime in AI time when the product team was just five engineers. She's helped ship every model at Anthropic from claw two through Fable. She's also helped incubate and launch claw code, MCP, skills, claw design, and also core capabilities, like computer use, tool use, and reasoning. It is always such a treat and so mind expanding to get to talk to someone who's at the very center of AI and product management, it's hard to imagine someone who has seen more of where things are going than the head of product for Anthropics Research and Labs teams. Before we get into it, don't forget to check out Lenny's productpass.com for a year free of the hottest and most beautifully crafted AI products in the world available exclusively to Lenny's newsletter subscribers. With that, I bring you Diane Penn. Diane, thank you so much for being here. Welcome to the podcast. Thank you, Lenny. It's so nice to see …

Get the full transcript (15,941 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Lenny's Podcast transcripts →

You just read a 3-minute summary of a 90-minute episode.

Get Lenny's Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.

Books

Tools

  • Claude CodeBy guest

    by Anthropic

    using Claude Code, MCP, and Claude Design as concrete examples of labs-driven product development.
  • MCPBy guest

    by Anthropic

    using Claude Code, MCP, and Claude Design as concrete examples of labs-driven product development.
  • by Anthropic

    using Claude Code, MCP, and Claude Design as concrete examples of labs-driven product development.

Products

  • by Anthropic

    Claude Opus 4.5's breakthrough moment required both a frontier model and a frontier product simultaneously. Claude Code accelerated adoption of Opus 4.5, while Opus 4.5 unlocked Claude Code's full potential.

More from Lenny's Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Product Management Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Lenny's Podcast.

Every Monday, we deliver AI summaries of the latest episodes from Lenny's Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime