Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Episode
86 min
Read time
2 min
Topics
Investing, Fundraising & VC, Leadership
AI-Generated Summary
Key Takeaways
- ✓Non-deterministic challenge: AI products differ fundamentally from traditional software because both input (user behavior via natural language) and output (LLM responses) are unpredictable. This means product builders cannot map workflows deterministically like booking.com, requiring new approaches to handle uncertainty on both ends of the interaction.
- ✓Agency-control tradeoff: Start AI products with high human control and low AI autonomy, then gradually increase agency as trust builds. For customer support, begin with routing suggestions humans review, progress to draft responses, then eventually autonomous ticket resolution—taking four to six months minimum for meaningful ROI.
- ✓CCCD framework: Continuous Calibration Continuous Development replaces traditional CI/CD for AI. Scope capabilities with curated data, deploy with evaluation metrics, then analyze emerging behavior patterns users exhibit that weren't predicted. Iterate when new data distribution patterns stop appearing, signaling readiness for increased autonomy.
- ✓Leadership requirement: Successful AI adoption requires CEO-level engagement. The Rackspace CEO blocks 4-6 AM daily for "catching up with AI," rebuilding intuitions from scratch. Leaders must accept their decade of experience may not apply and become the "dumbest person in the room" willing to learn from everyone.
- ✓Evaluation balance: False dichotomy exists between pre-deployment evals and production monitoring—both are essential. Evals catch known failure modes during development. Production monitoring reveals emerging patterns through implicit signals like answer regeneration, which indicates customer dissatisfaction even without explicit thumbs-down feedback. High-transaction products need both approaches simultaneously.
What It Covers
Aishwarya Reganti and Kiriti Badham share lessons from 50+ AI deployments at OpenAI, Google, and Amazon, explaining why most AI products fail due to non-determinism and agency-control tradeoffs, plus their framework for building successful AI systems.
Key Questions Answered
- •Non-deterministic challenge: AI products differ fundamentally from traditional software because both input (user behavior via natural language) and output (LLM responses) are unpredictable. This means product builders cannot map workflows deterministically like booking.com, requiring new approaches to handle uncertainty on both ends of the interaction.
- •Agency-control tradeoff: Start AI products with high human control and low AI autonomy, then gradually increase agency as trust builds. For customer support, begin with routing suggestions humans review, progress to draft responses, then eventually autonomous ticket resolution—taking four to six months minimum for meaningful ROI.
- •CCCD framework: Continuous Calibration Continuous Development replaces traditional CI/CD for AI. Scope capabilities with curated data, deploy with evaluation metrics, then analyze emerging behavior patterns users exhibit that weren't predicted. Iterate when new data distribution patterns stop appearing, signaling readiness for increased autonomy.
- •Leadership requirement: Successful AI adoption requires CEO-level engagement. The Rackspace CEO blocks 4-6 AM daily for "catching up with AI," rebuilding intuitions from scratch. Leaders must accept their decade of experience may not apply and become the "dumbest person in the room" willing to learn from everyone.
- •Evaluation balance: False dichotomy exists between pre-deployment evals and production monitoring—both are essential. Evals catch known failure modes during development. Production monitoring reveals emerging patterns through implicit signals like answer regeneration, which indicates customer dissatisfaction even without explicit thumbs-down feedback. High-transaction products need both approaches simultaneously.
Notable Moment
Air Canada's chatbot hallucinated a refund policy that didn't exist, and the company had to honor it legally. This incident illustrates why constraining AI autonomy matters—74% of enterprises cite reliability concerns as their biggest barrier to deploying customer-facing AI products, preferring productivity tools with lower risk.
Episode Transcript
Worked on a guest post together. They had this really key insight that building AI products is very different from building non AI products. Most people tend to ignore the non determinism. You don't know how the user might behave with your product, and you also don't know how the LLM might respond to that. The second difference is the agency control trade off. Every time you hand over decision making capabilities to agentic systems, you're kind of relinquishing some amount of control on your end. This significantly changes the way you should be building product. So we recommend building step by step. When you start small, it forces you to think about what is the problem that I'm gonna solve. In all these advancements of the AI, one easy slippery slope is to keep thinking about complexities of the solution and forget the problem that you're trying to solve. It's not about being the first company to have an agent among your competitors. It's about have you built the right flywheels in place so that you can improve over time. What kind of ways of working do you see in companies that build AI products successfully? I used to work with the CEO of now Rackspace. He would have this block every day in the morning, which would say catching up with AI four to 6AM. Leaders have to get back to being hands on. You must be comfortable with the fact that your intuitions might not be right, and you probably are the dumbest person in the room and you wanna learn from everyone. What do you think the next year of AI is gonna look like? Persistence is extremely valuable. Successful companies right now building in any new area. They are going through the pain of learning this, implementing this, and understanding what works and what doesn't work. Pain is the new mode. Today, my guests are Aishwarya Reganti and Kiriti Badham. Kiriti works on codecs at OpenAI and has spent the last decade building AI and ML infrastructure at Google and at Kumo. Ash was an early AI researcher at Alexa and Microsoft and and has published over 35 research papers. Together, they've led and supported over 50 AI product deployments across companies like Amazon, Databricks, OpenAI, Google, and both startups and large enterprises. Together, they also teach the number one rated AI course on Maven, where they teach product leaders all of the key lessons they've learned about building successful AI products. The goal of this episode is to save you and your team a lot of pain and suffering and wasted time trying to build your AI product. Whether you are already struggling to make your product work or want to avoid that struggle, this episode is for you. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. It helps tremendously. And if you become an annual subscriber of my newsletter, you get a …
Get the full transcript (16,682 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 83-minute episode.
Get Lenny's Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Lenny's Podcast
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
Sep 8 · 82 min
Dwarkesh Podcast
The rise and fall of agent civilizations
Aug 31
More from Lenny's Podcast
Why companies are becoming a series of loops | Anish Acharya (a16z)
Sep 6 · 79 min
The Prof G Pod
What SpaceX's IPO Means for Tech Stocks, and Coping With Panic Attacks
Jul 13
More from Lenny's Podcast
We summarize every new episode. Want them in your inbox?
How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
Why companies are becoming a series of loops | Anish Acharya (a16z)
AI’s third era: the rise of persistent AI coworkers | Tara Seshan (OpenAI’s product lead)
How to close $100K+ enterprise deals, step by step | Jen Abel
OpenAI’s Head of Design: This is the best time in history to be a designer | Ian Silber
Similar Episodes
Related episodes from other podcasts
Dwarkesh Podcast
Aug 31
The rise and fall of agent civilizations
The Prof G Pod
Jul 13
What SpaceX's IPO Means for Tech Stocks, and Coping With Panic Attacks
The Prof G Pod
Jul 10
The Week: Who Does the Market Actually Work For?
All-In with Chamath, Jason, Sacks & Friedberg
Jun 9
Bill Maris: How Google Could Crush AI Competitors, Why Small Funds Win, and AI's Atari Stage
The School of Greatness
May 22
Why You Keep Choosing the Wrong Person (And How to Finally Stop) | Faith Jenkins
Explore Related Topics
This podcast is featured in Best Product Management Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Investing & Markets Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Lenny's Podcast.
Every Monday, we deliver AI summaries of the latest episodes from Lenny's Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime