How to measure AI developer productivity in 2025 | Nicole Forsgren
Episode
67 min
Read time
2 min
Topics
Productivity, Investing, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Productivity metrics fail with AI: Lines of code becomes meaningless as a metric because AI generates verbose code easily. Companies must track which code comes from humans versus AI to measure survivability rates, quality, and avoid training bias loops in fine-tuned systems.
- ✓DORA and Space frameworks need adaptation: DORA's four metrics (deployment frequency, lead time, MTTR, change fail rate) still assess pipeline performance, but miss AI-era feedback loops. Space framework remains relevant because it measures satisfaction, performance, activity, communication, and efficiency without prescribing specific metrics.
- ✓Trust and code review dominate workflow: Developers now spend significantly more time reviewing AI-generated code than writing it. Teams must evaluate for hallucinations, reliability, and style consistency. Non-deterministic LLM outputs require validation that deterministic compilers never needed, fundamentally changing daily work structure.
- ✓Flow state requires new approaches: Senior engineers create effective workflows by architecting systems upfront, assigning parallel tasks to multiple AI agents with clear API conventions, then reviewing integrated results. This planning-heavy approach produces near-production code faster than traditional iterative coding methods.
- ✓Quick wins start with listening: Before implementing tools or automation, conduct listening tours asking developers about yesterday's friction points. Companies often discover process changes (like replacing physical approval walks with emails) that eliminate delays without engineering investment, delivering immediate productivity gains.
What It Covers
Nicole Forsgren explains how AI tools are accelerating code generation but not overall developer productivity proportionally, due to persistent bottlenecks in processes, broken builds, and unreliable systems that create friction throughout development workflows.
Key Questions Answered
- •Productivity metrics fail with AI: Lines of code becomes meaningless as a metric because AI generates verbose code easily. Companies must track which code comes from humans versus AI to measure survivability rates, quality, and avoid training bias loops in fine-tuned systems.
- •DORA and Space frameworks need adaptation: DORA's four metrics (deployment frequency, lead time, MTTR, change fail rate) still assess pipeline performance, but miss AI-era feedback loops. Space framework remains relevant because it measures satisfaction, performance, activity, communication, and efficiency without prescribing specific metrics.
- •Trust and code review dominate workflow: Developers now spend significantly more time reviewing AI-generated code than writing it. Teams must evaluate for hallucinations, reliability, and style consistency. Non-deterministic LLM outputs require validation that deterministic compilers never needed, fundamentally changing daily work structure.
- •Flow state requires new approaches: Senior engineers create effective workflows by architecting systems upfront, assigning parallel tasks to multiple AI agents with clear API conventions, then reviewing integrated results. This planning-heavy approach produces near-production code faster than traditional iterative coding methods.
- •Quick wins start with listening: Before implementing tools or automation, conduct listening tours asking developers about yesterday's friction points. Companies often discover process changes (like replacing physical approval walks with emails) that eliminate delays without engineering investment, delivering immediate productivity gains.
Notable Moment
Forsgren reveals that in companies tracking AI coding tool usage, developers using AI assistants regularly not only received more generated code, but their own manual coding output doubled compared to the AI contribution, suggesting AI primarily unblocks developers rather than replacing their work.
Episode Transcript
Lot of companies are trying to measure productivity for their teams. Most productivity metrics are a lie. If the goal is more lines of code, I can prompt something to write the longest piece of code ever. It's just too easy to gain that system. How do I know if my edge team is moving fast enough, if they can move faster, if they're just not performing as well as they can? Most teams can move faster, but faster for what? We can ship trash faster every single day. We need strategy and really smart decisions to know what to ship. One of the biggest issues we're gonna probably have with AI is learning how much to trust code that it generates. We can't just put in a command and get something back and accept it. We really need to evaluate it. You know, are we seeing hallucinations? What's the reliability? Does it meet the style that we would typically write? So much of the time is not gonna be spent reviewing code versus writing code. There's some real opportunity there to not just rethink workflows, but rethink how we structure our days and how we structure our work. Now we can also make a forty five minute work block useful because getting into the flow is actually kind of handed off at least in part to the machine, or the machine can help us get back into the flow by reminding us of context and generating diagrams of the system. What's just, like, one thing that you think an eng team, a product team can do this week, next week to get more done? Honestly, I think the best thing you can do Today, my guest is Nicole Forsgren. With so much talk about how AI is increasing developer productivity, more and more people are asking, how do we measure this productivity gain? And are these AI tools actually helping us or hurting how our developers work? Nicole has been at the forefront of this space longer than anyone. She created the most used frameworks for measuring developer experience called Dora and Space. She wrote the most important book in the space called Accelerate and is about to publish her newest book called Frictionless, which gives you a guide to helping your team move faster and do more in this emerging AI world. Her core thesis is that AI indeed accelerates coding, but developers aren't speeding up as much as you think because they still have to deal with broken builds and unreliable tools and processes and a bunch of new bottlenecks that are emerging. In our conversation, we chat about her current best and very specific advice for how to measure productivity gains from AI, signs that your team could be moving faster, what companies get wrong when trying to measure engineering productivity, how AI tools are both helping and hurting engineers including getting into flow states, her seven step process for setting up a developer experience team …
Get the full transcript (13,274 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 64-minute episode.
Get Lenny's Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Lenny's Podcast
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
Sep 8 · 82 min
NVIDIA AI Podcast
Harrison Chase of LangChain on Deep Agents, LangSmith, and Earning Trust | NVIDIA AI Podcast Ep. 297
May 6
More from Lenny's Podcast
Why companies are becoming a series of loops | Anish Acharya (a16z)
Sep 6 · 79 min
Software Engineering Daily
SmartBear and Multi-Agent QA
May 5
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
other
“DORA's four metrics (deployment frequency, lead time, MTTR, change fail rate) still assess pipeline performance, but miss AI-era feedback loops.”
“Space framework remains relevant because it measures satisfaction, performance, activity, communication, and efficiency without prescribing specific metrics.”
More from Lenny's Podcast
We summarize every new episode. Want them in your inbox?
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
Why companies are becoming a series of loops | Anish Acharya (a16z)
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
How to close $100K+ enterprise deals, step by step | Jen Abel
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Similar Episodes
Related episodes from other podcasts
NVIDIA AI Podcast
May 6
Harrison Chase of LangChain on Deep Agents, LangSmith, and Earning Trust | NVIDIA AI Podcast Ep. 297
Software Engineering Daily
May 5
SmartBear and Multi-Agent QA
The Prof G Pod
May 5
China Decode: China Is Beating the U.S. in Space?!
Eye on AI
Apr 14
#332 Dan Faulkner: The Code Is Clean. The App Is Broken. Why AI Development Has an Integrity Problem
Software Engineering Daily
Mar 5
Organizational Context for AI Coding Agents with Dennis Pilarinos
Explore Related Topics
This podcast is featured in Best Product Management Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Lenny's Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Lenny's Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime