Skip to main content
The Changelog

Voices of Oxide (Interview)

76 min episode · 2 min read
·
Cliff Biffle,Kyle Galbraith

Episode

76 min

Read time

2 min

Topics

Career Growth, Leadership, Design & UX

AI-Generated Summary

Key Takeaways

  • Firmware Architecture: Oxide runs 64-70 instances of Hubris operating system across every rack component, from sub-50-cent microcontrollers to service processors. Each compute sled contains two copies minimum—one for service processing and one for root of trust security—because no single chip currently provides both required feature sets.
  • Update System Complexity: Self-service updates replace hundreds of software components across 32 sleds, 2 switches, and power controllers while maintaining system availability. The process takes approximately two hours and requires careful orchestration to avoid intermediate states where incompatible software versions communicate, using a plan-execute pattern for safety validation.
  • Rust Type Safety Benefits: The team uses Dropshot to generate OpenAPI specs from code and Progenitor to generate clients, ensuring API changes that break backwards compatibility fail at compile time rather than runtime. This approach catches incompatible enum variants and schema changes before deployment, eliminating entire classes of upgrade failures.
  • Uniform Compensation Model: Oxide pays all employees identical salaries regardless of role, with equity varying only by join date. This eliminates negotiation stress and prevents the $100,000 salary gaps common at companies like Google, where managers discover significant pay disparities among same-level team members after promotion cycles.
  • Design System Integration: The company uses a single UI design system across web console, marketing website, and physical hardware, maintaining consistent colors and elements. Industrial design decisions prioritize manufacturability at scale over prototype aesthetics, avoiding the common trap where mass production compromises initial design quality through cost-cutting measures.

What It Covers

Oxide Computer Company engineers discuss their custom server rack architecture, including Hubris operating system development, self-service update system challenges, and design philosophy. The team covers technical decisions around Rust, firmware development, and building hardware from first principles.

Key Questions Answered

  • Firmware Architecture: Oxide runs 64-70 instances of Hubris operating system across every rack component, from sub-50-cent microcontrollers to service processors. Each compute sled contains two copies minimum—one for service processing and one for root of trust security—because no single chip currently provides both required feature sets.
  • Update System Complexity: Self-service updates replace hundreds of software components across 32 sleds, 2 switches, and power controllers while maintaining system availability. The process takes approximately two hours and requires careful orchestration to avoid intermediate states where incompatible software versions communicate, using a plan-execute pattern for safety validation.
  • Rust Type Safety Benefits: The team uses Dropshot to generate OpenAPI specs from code and Progenitor to generate clients, ensuring API changes that break backwards compatibility fail at compile time rather than runtime. This approach catches incompatible enum variants and schema changes before deployment, eliminating entire classes of upgrade failures.
  • Uniform Compensation Model: Oxide pays all employees identical salaries regardless of role, with equity varying only by join date. This eliminates negotiation stress and prevents the $100,000 salary gaps common at companies like Google, where managers discover significant pay disparities among same-level team members after promotion cycles.
  • Design System Integration: The company uses a single UI design system across web console, marketing website, and physical hardware, maintaining consistent colors and elements. Industrial design decisions prioritize manufacturability at scale over prototype aesthetics, avoiding the common trap where mass production compromises initial design quality through cost-cutting measures.

Notable Moment

One engineer revealed they joined Oxide specifically because the company was not fully remote, accepting the position in February 2020. Within weeks, the pandemic forced complete remote work, creating an ironic situation where their primary reason for joining immediately disappeared yet they stayed for four years.

Know someone who'd find this useful?

Episode Transcript

Well, friends, it is your favorite podcast, the change log. Yes. Jared and I have a special episode for you. Jared and I and team went to Emeryville, California at the invitation of Oxide. We went to Oxide's HQ. They have an annual conference every year. It's an internal conference called Oxcon, and they invited us out to celebrate and to peel back the layers to have a good look at the inside of Oxide. So they recently raised a series b round of a $100,000,000, and they also just signed a purchase order, a massive purchase order that is truly helping them cross the chasm. Now we also have three awesome conversations conversations for you. First up is Cliff Biffle. He's in charge of all things Hubris. Hubris is Oxide's operating system, and Cliff is also in charge of pretty much everything that happens before the CPU empowers on. Next up is Dave Pacheco. He's in charge of all things update, and update is the system for which they update the system. It's really important. And last up is Ben Leonard. Ben is in charge of all things design and brain for oxide. And if you like how oxide looks, I do. Well, that's Ben. A massive thank you to our friends and our partners over at Fly. Check them out at fly.io. Alright. Let's do this. What's up, friends? I'm here with Kyle Galbraith, cofounder and CEO of Deepo. Deepo is the only build platform looking to make your builds as fast as possible. But, Kyle, this is an issue because GitHub Actions is the number one CI CI provider out there, but not everyone's a fan. Explain that. I think when you're thinking about GitHub Actions, it's really quite jarring how you can have such a wildly popular CI provider, and yet it's lacking some of the basic functionality or tools that you need to actually be able to debug your builds or deployments. And so back in June, we essentially took a stab at that problem in particular with Depo's GitHub Action runners. What we've observed over time is effectively GitHub Actions, when it comes to, like, actually debugging a build, is pretty much useless. The job logs in GitHub Actions UI is pretty much where your dreams go to die. Like, they're collapsed by default. They have no resource metrics. When jobs fail, you're essentially left playing detective, like, clicking each little drop down on each step in your job to figure out, like, okay. Where did this actually go wrong? And so what we set out to do with our own GitHub Actions observability is essentially we built a real observability solution around GitHub Actions. Okay. So how does it work? All of the logs by default for a job that runs on a depot GitHub Action runner, they're uncollapsed. You can search them. You can detect if there's been out of memory errors. You can see all of the resource contention that was …

Get the full transcript (16,355 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The Changelog transcripts →

You just read a 3-minute summary of a 73-minute episode.

Get The Changelog summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • The team uses Dropshot to generate OpenAPI specs from code and Progenitor to generate clients, ensuring API changes that break backwards compatibility fail at compile time rather than runtime.
  • The team uses Dropshot to generate OpenAPI specs from code and Progenitor to generate clients, ensuring API changes that break backwards compatibility fail at compile time rather than runtime.
  • Oxide runs 64-70 instances of Hubris operating system across every rack component, from sub-50-cent microcontrollers to service processors.

company

  • Oxide Computer Company engineers discuss their custom server rack architecture, including Hubris operating system development, self-service update system challenges, and design philosophy.

More from The Changelog

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best Cybersecurity Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into The Changelog.

Every Monday, we deliver AI summaries of the latest episodes from The Changelog and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime