[State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify
Episode
28 min
Read time
2 min
Topics
Career Growth, Productivity, Relationships
AI-Generated Summary
Key Takeaways
- ✓IPO Market Requirements: Companies now need $600M+ revenue to go public, up from previous thresholds. The DBT-Fivetran merger combined two profitable companies approaching $600M to accelerate their path to liquidity, not because the modern data stack failed. Both companies exceeded revenue targets and remain essential infrastructure for frontier AI labs managing training datasets.
- ✓Seed Funding Dysfunction: Founders raise $100M+ seed rounds at billion-dollar valuations with seven-day decision windows but no concrete twelve to twenty-four month roadmap. They pitch long-term visions without near-term milestones, making it impossible for investors to assess execution capability. This creates hiring advantages through inflated equity values but sets teams up for failure if exits fall below funding amounts.
- ✓Memory and Personalization Gap: AI applications suffer from high churn because they lack effective memory management and continual learning. Cursor rules represent primitive memory implementation. True personalization requires models that update weights based on user interactions, creating stateful inference systems. This applies equally to consumer apps and enterprise tools where models must learn company-specific terminology and workflows continuously.
- ✓Research-Application Integration: The most successful AI companies like Harvey and Sierra hire researchers to solve hard technical problems that directly unlock product capabilities. Harvey advanced RAG implementations for legal search, while Sierra focused on rule-following for customer support. This tight coupling between research breakthroughs and application value creates defensible competitive advantages that pure application layers cannot replicate.
- ✓Data Infrastructure Scaling: Modern data tools like DBT and Fivetran scale effectively to AI workloads despite concerns. Frontier labs use these tools within weeks of formation. Training dataset management requires more ad hoc, less predictable workloads than traditional analytics, but existing infrastructure handles the scale. GPU data loading efficiency matters more than database architecture for preventing idle compute time.
What It Covers
Sarah Catanzaro from Amplify Partners discusses the DBT-Fivetran merger as IPO preparation rather than industry decline, critiques the $100M+ seed funding trend with unclear roadmaps, and identifies memory management, continual learning, and personalization as critical infrastructure opportunities while dismissing RL environments as potentially overvalued.
Key Questions Answered
- •IPO Market Requirements: Companies now need $600M+ revenue to go public, up from previous thresholds. The DBT-Fivetran merger combined two profitable companies approaching $600M to accelerate their path to liquidity, not because the modern data stack failed. Both companies exceeded revenue targets and remain essential infrastructure for frontier AI labs managing training datasets.
- •Seed Funding Dysfunction: Founders raise $100M+ seed rounds at billion-dollar valuations with seven-day decision windows but no concrete twelve to twenty-four month roadmap. They pitch long-term visions without near-term milestones, making it impossible for investors to assess execution capability. This creates hiring advantages through inflated equity values but sets teams up for failure if exits fall below funding amounts.
- •Memory and Personalization Gap: AI applications suffer from high churn because they lack effective memory management and continual learning. Cursor rules represent primitive memory implementation. True personalization requires models that update weights based on user interactions, creating stateful inference systems. This applies equally to consumer apps and enterprise tools where models must learn company-specific terminology and workflows continuously.
- •Research-Application Integration: The most successful AI companies like Harvey and Sierra hire researchers to solve hard technical problems that directly unlock product capabilities. Harvey advanced RAG implementations for legal search, while Sierra focused on rule-following for customer support. This tight coupling between research breakthroughs and application value creates defensible competitive advantages that pure application layers cannot replicate.
- •Data Infrastructure Scaling: Modern data tools like DBT and Fivetran scale effectively to AI workloads despite concerns. Frontier labs use these tools within weeks of formation. Training dataset management requires more ad hoc, less predictable workloads than traditional analytics, but existing infrastructure handles the scale. GPU data loading efficiency matters more than database architecture for preventing idle compute time.
Notable Moment
Catanzaro reveals her failed prediction on data catalogs, which she believed would become essential infrastructure. Instead, companies like Snowflake and DBT built cataloging as features that proved sufficient for humans. She suggests the real opportunity may have been building metadata services for machines and microservices rather than human discoverability, or focusing on governance over discovery.
Episode Transcript
Okay. We're here with Sarah and Zarrow from Amplify. Welcome. Thank you. First time on the pod. To be here. I Too long. I know. I know. We've we've known each other for so long. Yeah. Yeah. Never made an appearance. And also made the transition from data to AI, I guess. I I don't know if if, I did. I don't know if you were always, like, as deep on on on AI, but I'll be there's a lot of simpatico. Yeah. I've always actually kind of oscillated between data and AI. Sure. Like, arguably, I started my career in quote, unquote, AI. It was just more, like, symbolic systems back then. But as you said, I think, like, they're they're so symbiotic. Like, it it's almost hard to divorce them. That's actually what brought me into data. I was like, I want to better understand what happens when I write a SQL query. So Yeah. Let's briefly touch on data because I I think obviously that's that's a lot of where you and I first met. D B T 5 Tran. That was so cool. I mean, or Yeah. How do how do you how do you think about the end of the modern data stack? Okay. So so, like, a lot of people look at the, like, DBT five Tran, merger and, like, talk about the end of the modern data stack, and I think that is, like, a fundamentally wrong take. Both of these companies were growing, you know, very healthily. Both of these companies And you've do you fund a DBT? We funded DBT. So so, like, both of the companies were actually, like, beating their revenue targets. I think what you're more seeing is, you know, IPO environment wherein companies are expected to have far more than, you know, like, a 100,000,000 revenue. And so What would you say the bar is now? 300? No. Like, above 600. 600. Yeah. Yeah. And the combined company is 400? I believe that they'll actually be close to 600. I don't have the exact number. But they clearly just getting ready Yes. For IPO. So so so, you know, basically, like, the merger was a way to accelerate that pass to liquidity. You know, as you might remember And they were the presumptive winners in their categories anyway. So Exactly. You know, I think one of the things that has actually, pleasantly surprised me, and this speaks to, again, the symbiotic relationship between node data and AI. Many of the big frontier labs are actually using both DBT and Fivetran. I recall talking to folks at, Thinking Machines, like, within weeks of the company's formation, and DBT was already an important part of their stack. Certainly, like, training datasets need to be managed. We need insight into what users are doing on these platforms. And in fact, like, the way in which you would analyze interactions with an agent or analyze interactions with an LLM is even …
Get the full transcript (5,253 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 25-minute episode.
Get Latent Space summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Aug 26 · 83 min
a16z Podcast
Inside Cursor: The Anatomy of a Generational Startup
Aug 27
More from Latent Space
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
Aug 21 · 69 min
The Diary of a CEO
Most Replayed Moment: Fear Is A Skill You Can Train! Lessons From The World's Greatest Climber
Aug 14
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links.
Tools
“Cursor rules represent primitive memory implementation”
company
“The most successful AI companies like Harvey and Sierra hire researchers to solve hard technical problems that directly unlock product capabilities. Sierra focused on rule-following for customer support”
“Instead, companies like Snowflake and DBT built cataloging as features that proved sufficient for humans”
“The DBT-Fivetran merger combined two profitable companies approaching $600M to accelerate their path to liquidity”
“The most successful AI companies like Harvey and Sierra hire researchers to solve hard technical problems that directly unlock product capabilities. Harvey advanced RAG implementations for legal search”
“The DBT-Fivetran merger combined two profitable companies approaching $600M to accelerate their path to liquidity”
“Sarah Catanzaro from Amplify Partners discusses the DBT-Fivetran merger as IPO preparation rather than industry decline”
More from Latent Space
We summarize every new episode. Want them in your inbox?
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Simulation: the new Scaling Law — Joon Sung Park, Simile AI
🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
Similar Episodes
Related episodes from other podcasts
a16z Podcast
Aug 27
Inside Cursor: The Anatomy of a Generational Startup
The Diary of a CEO
Aug 14
Most Replayed Moment: Fear Is A Skill You Can Train! Lessons From The World's Greatest Climber
How I Built This
Aug 6
Advice Line: "Strategy Sessions"
Biotech Hangout
Aug 3
Episode 191 - July 31, 2026
Pivot
Jul 28
Paramount Merger Pause, Meta AI Optimism, and the Peptide Boom
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
You're clearly into Latent Space.
Every Monday, we deliver AI summaries of the latest episodes from Latent Space and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime