Skip to main content
Software Engineering Daily

Turbopuffer with Simon Hørup Eskildsen

50 min episode · 2 min read
·
Simon H

Episode

50 min

Read time

2 min

Topics

Fundraising & VC, Artificial Intelligence, Software Development

AI-Generated Summary

Key Takeaways

  • Storage architecture economics: TurboPuffer uses S3 object storage at 2¢ per gigabyte versus traditional in-memory vector databases at $2-5 per gigabyte, achieving 100x cost reduction while maintaining sub-second query performance through strategic caching layers.
  • Cluster-based indexing for disk: Graph-based vector indexes require hundreds of milliseconds per jump on S3, making them impractable. Cluster-based indexes fetch centroids and clusters in just three round trips, enabling cold queries under one second on object storage.
  • Production recall monitoring: TurboPuffer samples 1% of production queries to measure recall accuracy against exact results, maintaining 90-95% recall across real-world datasets. This catches edge cases that academic benchmarks miss, ensuring consistent search quality at scale.
  • Namespace sharding primitive: TurboPuffer maps each namespace to one shard with separate S3 prefixes, supporting over 100 million namespaces. Each namespace can use customer-managed encryption keys, providing isolation equivalent to separate buckets without coordination overhead.

What It Covers

Simon Eskildsen explains how TurboPuffer reduces vector database costs by 95% using object storage instead of memory, enabling companies like Cursor and Notion to scale AI search economically at 2¢ per gigabyte.

Key Questions Answered

  • Storage architecture economics: TurboPuffer uses S3 object storage at 2¢ per gigabyte versus traditional in-memory vector databases at $2-5 per gigabyte, achieving 100x cost reduction while maintaining sub-second query performance through strategic caching layers.
  • Cluster-based indexing for disk: Graph-based vector indexes require hundreds of milliseconds per jump on S3, making them impractable. Cluster-based indexes fetch centroids and clusters in just three round trips, enabling cold queries under one second on object storage.
  • Production recall monitoring: TurboPuffer samples 1% of production queries to measure recall accuracy against exact results, maintaining 90-95% recall across real-world datasets. This catches edge cases that academic benchmarks miss, ensuring consistent search quality at scale.
  • Namespace sharding primitive: TurboPuffer maps each namespace to one shard with separate S3 prefixes, supporting over 100 million namespaces. Each namespace can use customer-managed encryption keys, providing isolation equivalent to separate buckets without coordination overhead.

Notable Moment

Eskildsen discovered the vector database cost problem when calculating that storing Readwise article embeddings would cost $30,000 monthly versus $3,000 for their entire Postgres database, revealing a 10x cost amplification blocking AI feature adoption.

Know someone who'd find this useful?

Episode Transcript

Vector search has become a foundational technology for AI applications, enabling everything from semantic code search to contextual retrieval for large language models. However, a major challenge with vector databases has been the cost as data storage scales. TurboPuffer is a vector database that focuses on speed, cost, and scalability. It was created by Simon Harroub Eskelson and Justin Lee in 2023 and has seen adoption from high profile companies such as Cursor and Notion. Siman joins the podcast with Gregor Vann to discuss the origin of turbopuffer, its unique technical design, the economics of vector storage, and more. Gregor Vann is a CTO and founder currently working at the intersection of communication, security, and AI and is based in Singapore. His latest venture, wintik.ai, reimagines what email can be in the AI era. For more on Gregor, find him at van.hk or on LinkedIn. Hello, welcome to Software Engineering Daily. My guest today is Simon Eskelson. Thank you for having me and Gregor. Yeah, great to have you here Simon. We're here today to talk all about TurboPuffer, which is a company some of our audience may have heard of equally, maybe quite a few haven't yet, and it's gonna be why we're here today and talking all about it. However, you probably are using this product behind the scenes without you realizing. So we're gonna get into that as well. But I think to begin with, Simon, you've got a really interesting kind of backstory, and we always dive into those with our guests. So just maybe talk us through very briefly how you kind of got to turbopuffer, and I'll just call out you did spend quite a chunk of time at Shopify, and I think that would be interesting to understand how that's helped lead into what TurboPuffer is today as well. Yeah. That's right. I started my career at Shopify and moved to Canada as a result where Shopify was built. I moved from Denmark where I grew up and worked on infrastructure at Shopify for almost a decade. When I joined in 2013, it was in the hundreds of requests per second. And when I left, we regularly saw peaks in north of 1,000,000 requests per second. And as part of that, I worked on more or less every single aspect of the infrastructure. Generally, the bottlenecks for that kind of scale tend to appear in the data layer. Those are the most persistent ones. So I spent, yeah, almost the entire time working on that layer between the Rails app and the databases and sometimes inside the databases themselves, but mostly on top of them. So, yeah, almost ten years working on every single aspect of database scalability at Shopify. And then after that, I spent about two years working in small increments at my friends' companies on whatever infrastructure problems they had. It turns out in 2022, 2023, it's mostly tuning Postgres auto vacuum was the persistent infrastructure problem. And it was …

Get the full transcript (10,768 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Software Engineering Daily transcripts →

You just read a 3-minute summary of a 47-minute episode.

Get Software Engineering Daily summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • SPONSORS: Stream, url: getstream.io/podcast
  • TurbopufferBy guest
    Simon Eskildsen explains how TurboPuffer reduces vector database costs by 95% using object storage instead of memory, enabling companies like Cursor and Notion to scale AI search economically at 2¢ per gigabyte.
  • SPONSORS: Redis, url: redis.io/genai

company

  • enabling companies like Cursor and Notion to scale AI search economically at 2¢ per gigabyte
  • Eskildsen discovered the vector database cost problem when calculating that storing Readwise article embeddings would cost $30,000 monthly versus $3,000 for their entire Postgres database
  • SPONSORS: Capital One
  • enabling companies like Cursor and Notion to scale AI search economically at 2¢ per gigabyte

More from Software Engineering Daily

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's AI & Machine Learning Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into Software Engineering Daily.

Every Monday, we deliver AI summaries of the latest episodes from Software Engineering Daily and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime