Building Snipd: The AI Podcast App for Learning
Episode
77 min
Read time
3 min
Topics
Productivity, Startups, Design & UX
AI-Generated Summary
Key Takeaways
- ✓Triple-tap snipping feature: Users capture podcast insights by triple-tapping headphones, triggering AI summarization that saves the moment with transcript and audio to a knowledge library. This addresses the core problem that 90-99% of podcast content gets forgotten within minutes of listening, transforming passive consumption into active learning without requiring users to take manual notes or screenshots.
- ✓Cost optimization through model tiering: Snipd uses cheaper models for initial processing of over one million transcribed podcasts, then pipes outputs to premium models as judges for quality control. For quote selection, they generate five candidates with a budget model, then use a superior model to choose the best one, balancing scale economics with output quality across their massive audio processing pipeline.
- ✓Streaming LLM corrections with regex: When streaming chat responses to users, Snipd cannot pipe streams through additional LLM layers for formatting corrections due to API limitations. Instead, they maintain countless regex patterns that correct formatting errors in real-time as tokens stream, ensuring clean user experience. This represents practical AI engineering where traditional programming complements LLM capabilities rather than replacing it entirely.
- ✓Speaker diarization through heuristics: Snipd improves open-source diarization models by applying podcast-specific rules: speakers appearing for under thirty seconds in hour-long episodes are likely ad voices, not guests or hosts. They combine embedding-based clustering with LLM orchestration to assign speaker names, achieving better accuracy than tools like Descript by leveraging domain knowledge about podcast structure and high-quality studio audio.
- ✓Discovery through voice interfaces: Voice AI enables hooking into existing podcast listening habits rather than creating new notification-based triggers like Duolingo. When episodes end, an AI companion can initiate two-to-three minute conversations that force active processing of key takeaways, dramatically improving retention and application of knowledge. This backgroundable interaction stays within users' existing flow rather than requiring separate app opens or text-based engagement.
What It Covers
Kevin Ben Smith, founder of Snipd, explains how his four-person team built an AI-powered podcast app focused on learning and knowledge retention. The conversation covers their technical architecture using Python, Google Cloud, and Flutter, their transition from self-hosted models to API-based LLMs, and their product philosophy of making AI invisible while solving real user problems around podcast discovery and retention.
Key Questions Answered
- •Triple-tap snipping feature: Users capture podcast insights by triple-tapping headphones, triggering AI summarization that saves the moment with transcript and audio to a knowledge library. This addresses the core problem that 90-99% of podcast content gets forgotten within minutes of listening, transforming passive consumption into active learning without requiring users to take manual notes or screenshots.
- •Cost optimization through model tiering: Snipd uses cheaper models for initial processing of over one million transcribed podcasts, then pipes outputs to premium models as judges for quality control. For quote selection, they generate five candidates with a budget model, then use a superior model to choose the best one, balancing scale economics with output quality across their massive audio processing pipeline.
- •Streaming LLM corrections with regex: When streaming chat responses to users, Snipd cannot pipe streams through additional LLM layers for formatting corrections due to API limitations. Instead, they maintain countless regex patterns that correct formatting errors in real-time as tokens stream, ensuring clean user experience. This represents practical AI engineering where traditional programming complements LLM capabilities rather than replacing it entirely.
- •Speaker diarization through heuristics: Snipd improves open-source diarization models by applying podcast-specific rules: speakers appearing for under thirty seconds in hour-long episodes are likely ad voices, not guests or hosts. They combine embedding-based clustering with LLM orchestration to assign speaker names, achieving better accuracy than tools like Descript by leveraging domain knowledge about podcast structure and high-quality studio audio.
- •Discovery through voice interfaces: Voice AI enables hooking into existing podcast listening habits rather than creating new notification-based triggers like Duolingo. When episodes end, an AI companion can initiate two-to-three minute conversations that force active processing of key takeaways, dramatically improving retention and application of knowledge. This backgroundable interaction stays within users' existing flow rather than requiring separate app opens or text-based engagement.
- •Multimodal future with raw audio: Current pipelines using Whisper transcription plus separate diarization will be replaced by feeding raw audio files directly into multimodal LLMs like Gemini 2.0 Flash. These models will output transcripts, speaker labels, timestamps, and metadata in one pass. The transition awaits cost parity with current self-hosted pipelines, but represents the inevitable direction as transformers continue consuming specialized audio processing tasks.
Notable Moment
Smith reveals that Substack-hosted podcasts cannot play on Apple Watch due to platform restrictions, affecting the Latent Space podcast itself. Despite reaching out to Substack and major podcasters, the company shows no interest in fixing this limitation. This restriction exists across all podcast apps, not just Snipd, demonstrating how Substack treats podcasting as an afterthought despite being a creator platform.
Episode Transcript
Hey. I'm here in New York with Kevin Ben Smith of Snits. Welcome. Hi. Hi. Amazing to be here. Yeah. This is our first ever, I think, outdoors, podcast recording. Well, it's quite a location for the first time, I have to say. I was actually unsure because, you know, it's cold. It's like I checked the temperature. It's like, kind of one one degree Celsius, but it's not that bad with the sun. No. It's quite nice. Yeah. Yeah. Especially with our beautiful tea. With the tea. Yeah. Perfect. We're gonna talk about Snips. I'm a Snips user. I had to basically you know, apart from Twitter, it's like the the number one use app on my phone. Nice. When I when I wake up in the morning, I I open Snips and I, you know, see what's what's new. And I I I think in terms of time spent or usage on my phone, like, it's it's it's I think it's number one or number two. Nice. Nice. So so I I really have to talk about it. Also, because I think, like, people interested in AI wanna think about, like, how can they and we're an AI podcast. We have to talk about the AI podcast app. But before we get there, we just finished the AI Engineer Summit, and you came for the the two days. How was it? It was quite incredible. I mean, for me, the most valuable was just being in the same room with like minded people who are building the future and who are seeing the future. You know, especially when it comes to AI agents, it's so often I have conversations with friends who are not in the AI world, and it's, like, so quickly it happens that you it sounds like you're you're talking in science fiction. And, it's just crazy talk. It was, you know, it's so refreshing to to talk with so many other people who already see these things and, yeah, be inspired then by them and not always feel like like, okay, I think I'm just crazy and, like, this would never happen. It really is happening. And, for me, it was very valuable. So day two more relevant more relevant for you than day one? Yeah. Day two. So day two was the engineering track. Yep. That was definitely the most valuable for me, like, also as a practitioner myself. Especially, there were one or two talks that had to do with voice AI and AI agents with voice. Okay. So that was, quite fascinating. Also, spoke with the speakers afterwards. Yeah. And, yeah, they were also very open and and, you know, this this sharing attitude that's, I think in general quite prevalent in the AI community. I also learned a lot like really practical things that I can now take away with me. Yeah. I mean, on on my side, I I think I watched only like half of the talks because …
Get the full transcript (15,271 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 74-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
The Mel Robbins Podcast
Find Your Purpose & Live a Meaningful Life Today with the #1 Happiness Expert
Jul 13
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The Tim Ferriss Show
#872: Graham Duncan — Talent Is the Best Asset Class (Repost)
Jul 1
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
by Duolingo
“Voice AI enables hooking into existing podcast listening habits rather than creating new notification-based triggers like Duolingo.”
by OpenAI
“Current pipelines using Whisper transcription plus separate diarization will be replaced by feeding raw audio files directly into multimodal LLMs”
“The conversation covers their technical architecture using Python, Google Cloud, and Flutter”
by Google
“The conversation covers their technical architecture using Python, Google Cloud, and Flutter”
by Google
“The conversation covers their technical architecture using Python, Google Cloud, and Flutter”
by Google
“Current pipelines using Whisper transcription plus separate diarization will be replaced by feeding raw audio files directly into multimodal LLMs like Gemini 2.0 Flash.”
by Substack
“Smith reveals that Substack-hosted podcasts cannot play on Apple Watch due to platform restrictions, affecting the Latent Space podcast itself.”
by Descript
“They combine embedding-based clustering with LLM orchestration to assign speaker names, achieving better accuracy than tools like Descript”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
The Mel Robbins Podcast
Jul 13
Find Your Purpose & Live a Meaningful Life Today with the #1 Happiness Expert
The Tim Ferriss Show
Jul 1
#872: Graham Duncan — Talent Is the Best Asset Class (Repost)
Eye on AI
Mar 27
#328 Kevin Tian: Exploring Doppel's AI-Native Social Engineering Defense Platform
Foundr
Dec 11
613: Why Most Beauty Brands Fail - and How to Beat The Rest | DIBS Beauty Founder
a16z Podcast
Sep 11
What It Takes to Build a Startup | Andrew Chen & Matt Perault
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime