Skip to main content
The Changelog

Kaizen! Let it crash (Friends)

101 min episode · 2 min read
·
Gerhard Lazu

Episode

101 min

Read time

2 min

Topics

Startups, Software Development, Philosophy & Wisdom

AI-Generated Summary

Key Takeaways

  • Memory fragmentation solution: Varnish crashed 43 times in three months from MP3 files (30-100MB each) causing memory fragmentation. Moving large files from malloc memory storage to file-based cache with pre-allocated disk space eliminated crashes while maintaining 93% cache hit ratio across 15 global regions.
  • Concurrency misconfiguration impact: Setting Fly.io proxy concurrency to connections instead of requests caused 2,700 long-running connections to block new traffic in Newark region. HTTP/2 clients experienced response body timeouts while headers returned successfully, resolved by switching concurrency mode and explicitly setting 60-second idle timeouts.
  • Thread pool architecture benefits: Varnish runs as daemon with multiple threads, so out-of-memory kills only restart individual threads within two seconds rather than entire VM. This let-it-crash philosophy from Erlang ecosystem enables system stability despite component failures, with zero thread failures recorded after five days uptime.
  • Bandwidth abuse detection: Episode 456 generated 30 terabytes from San Jose alone in 60 days, with 10,000+ distinct IPs downloading repeatedly. Honeycomb observability reveals patterns like 170,000 favicon requests in two hours and weekly Python/Go clients scraping all MP3s, requiring vmod-throttle implementation for rate limiting.
  • Regional traffic optimization: San Jose and Tokyo handle highest CDN load at 2.29 gigabits per second peak. Automated hourly checks using hurl test all 15 regions, downloading full MP3s to validate response times under 100 seconds. Fly.io allows per-region instance sizing but requires manual scaling after initial deployment.

What It Covers

Gerhard Lazu debugs Changelog's CDN infrastructure after 43 out-of-memory crashes since October, implementing file-based caching for MP3s, fixing Fly.io proxy misconfigurations, and discovering massive bandwidth abuse from 10,000+ IPs downloading episode 456 repeatedly.

Key Questions Answered

  • Memory fragmentation solution: Varnish crashed 43 times in three months from MP3 files (30-100MB each) causing memory fragmentation. Moving large files from malloc memory storage to file-based cache with pre-allocated disk space eliminated crashes while maintaining 93% cache hit ratio across 15 global regions.
  • Concurrency misconfiguration impact: Setting Fly.io proxy concurrency to connections instead of requests caused 2,700 long-running connections to block new traffic in Newark region. HTTP/2 clients experienced response body timeouts while headers returned successfully, resolved by switching concurrency mode and explicitly setting 60-second idle timeouts.
  • Thread pool architecture benefits: Varnish runs as daemon with multiple threads, so out-of-memory kills only restart individual threads within two seconds rather than entire VM. This let-it-crash philosophy from Erlang ecosystem enables system stability despite component failures, with zero thread failures recorded after five days uptime.
  • Bandwidth abuse detection: Episode 456 generated 30 terabytes from San Jose alone in 60 days, with 10,000+ distinct IPs downloading repeatedly. Honeycomb observability reveals patterns like 170,000 favicon requests in two hours and weekly Python/Go clients scraping all MP3s, requiring vmod-throttle implementation for rate limiting.
  • Regional traffic optimization: San Jose and Tokyo handle highest CDN load at 2.29 gigabits per second peak. Automated hourly checks using hurl test all 15 regions, downloading full MP3s to validate response times under 100 seconds. Fly.io allows per-region instance sizing but requires manual scaling after initial deployment.

Notable Moment

One episode from August 2021 about OAuth complexity has been downloaded over one million times, generating 400 gigabytes every four hours from thousands of Asian IP addresses. The team suspects speed testing or archiving bots rather than genuine listeners, forcing implementation of throttling mechanisms to control bandwidth costs.

Know someone who'd find this useful?

Episode Transcript

Welcome to Change Log and Friends, a weekly talk show about how good systems become bad systems. Thanks as always to our partners at fly2.io, the platform for devs who just wanna ship, build fast, run any code fearlessly at fly to IO. Okay. Let's Kaizen. Well, friends, I I don't know about you, but something bothers me about GitHub Actions. I love the fact that it's there. I love the fact that it's so ubiquitous. I love the fact that agents that do my coding for me believe that my CICD workflow begins with drafting TOML files for GitHub Actions. That's great. It's all great until, yes, until your builds start moving like molasses. GitHub Actions is slow. It's just the way it is. That's how it works. I'm sorry. But But I'm not sorry because our friends at Namespace, they fix that. Yes. We use namespace. So to do all of our builds so much faster. Namespace is like GitHub actions, but faster. I mean, like, way faster. It caches everything smartly. It caches your dependencies, your Docker layers, your build artifacts, so your CI can run super fast. You get shorter feedback loops, happier developers because we love our time, and you get fewer, I'll be back after this coffee and my build finishes. So that's that's not cool. The best part is it's drop in. It works right alongside your existing GitHub actions with almost zero config. It's a one line change. So you get speed up your builds, you get to let your team, and you can finally stop pretending that build time is focus time. It's not. Learn more, go to namespace.so. That's namespace.so. Just like it sounds, like it said, Go there. Check them out. We use them. We love them, and you should too. Namespace.so. How else would you learn? Let it crash. Exactly. The best things happen when things fail. Seriously. If it's in a controlled way. Right? I think that's like something which isn't isn't said. It's implied. It has to be a controlled failure where you have the boundary and things will not blow up. I mean, they'll blow up, but, like, you know, like the fireworks sort of blowing up where it's a controlled explosion. Yeah. Right. Tiny little crashes to learn from. Welcome everyone to Kaizen 22 with the incomparable Gerhard Lazu. He's here to let us know how he lets it crash. It's like that song. Let it snow. Let it snow. Let it snow. Only you know how to replace. Hey, Gerhard. How are you? Hey, Jared. I'm good. Thank you. Thank you. Had a great holiday. It was a great couple of weeks where I've managed to finally disconnect. It's been, I don't know, like twenty years since I had two weeks completely off. Nice. Even my holidays are only a week. So this was very different, very enjoyable, and I feel so refreshed. So I'm firing on all cylinders. You unplugged, and now …

Get the full transcript (18,191 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The Changelog transcripts →

You just read a 3-minute summary of a 98-minute episode.

Get The Changelog summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • Honeycomb observability reveals patterns like 170,000 favicon requests in two hours and weekly Python/Go clients scraping all MP3s, requiring vmod-throttle implementation for rate limiting.
  • Varnish crashed 43 times in three months from MP3 files (30-100MB each) causing memory fragmentation. Moving large files from malloc memory storage to file-based cache with pre-allocated disk space eliminated crashes while maintaining 93% cache hit ratio across 15 global regions.
  • Automated hourly checks using hurl test all 15 regions, downloading full MP3s to validate response times under 100 seconds.

company

  • Setting Fly.io proxy concurrency to connections instead of requests caused 2,700 long-running connections to block new traffic in Newark region. Fly.io allows per-region instance sizing but requires manual scaling after initial deployment.

More from The Changelog

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The Changelog.

Every Monday, we deliver AI summaries of the latest episodes from The Changelog and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime