Skip to main content
The TWIML AI Podcast

Scaling Agentic Inference Across Heterogeneous Compute with Zain Asgar - #757

48 min episode · 2 min read
·
Zain Asgar

Episode

48 min

Read time

2 min

Topics

Startups, Fundraising & VC, Artificial Intelligence

AI-Generated Summary

Key Takeaways

  • Workload Disaggregation Strategy: Gimlet splits agent workflows into granular components, assigns performance-critical pieces to premium hardware like B200s, and offloads less critical tasks to lower-cost accelerators, optimizing cost per token while maintaining SLA requirements through dynamic resource allocation.
  • Kernel Optimization Performance: LLM-based automatic kernel synthesis delivers single-digit improvements on mature H100 hardware but achieves 20-40% gains on newer B200/RTX 6000 systems and over 2x speedups on AMD/Intel/Apple hardware where optimization frameworks remain underdeveloped.
  • Hardware Utilization Economics: Most GPU deployments show only 30% utilization, wasting two-thirds of capacity. Heterogeneous orchestration captures the majority of cost savings by efficiently packing workloads across different hardware types based on compute cost, memory bandwidth, and capacity requirements.
  • Multi-Agent Kernel Generation: The system uses hardware-in-the-loop testing where supervisor agents generate candidate kernels, execute them on target hardware with profiling and correctness checks, then iteratively optimize based on performance data until convergence, caching verified kernels offline.

What It Covers

Zain Asgar explains how Gimlet Labs optimizes AI inference costs through heterogeneous compute orchestration, using workload disaggregation, MLIR compilation, and LLM-generated kernel optimization across NVIDIA, AMD, and Intel hardware platforms.

Key Questions Answered

  • Workload Disaggregation Strategy: Gimlet splits agent workflows into granular components, assigns performance-critical pieces to premium hardware like B200s, and offloads less critical tasks to lower-cost accelerators, optimizing cost per token while maintaining SLA requirements through dynamic resource allocation.
  • Kernel Optimization Performance: LLM-based automatic kernel synthesis delivers single-digit improvements on mature H100 hardware but achieves 20-40% gains on newer B200/RTX 6000 systems and over 2x speedups on AMD/Intel/Apple hardware where optimization frameworks remain underdeveloped.
  • Hardware Utilization Economics: Most GPU deployments show only 30% utilization, wasting two-thirds of capacity. Heterogeneous orchestration captures the majority of cost savings by efficiently packing workloads across different hardware types based on compute cost, memory bandwidth, and capacity requirements.
  • Multi-Agent Kernel Generation: The system uses hardware-in-the-loop testing where supervisor agents generate candidate kernels, execute them on target hardware with profiling and correctness checks, then iteratively optimize based on performance data until convergence, caching verified kernels offline.

Notable Moment

Asgar reveals that AI training infrastructure has regressed to the supercomputer era with fully vertically integrated rack-scale systems reaching 600 kilowatts, while inference workloads benefit from disaggregated commodity hardware approaches that enable sustainable scaling.

Know someone who'd find this useful?

Episode Transcript

If you take a look at training hardware, it's kind of gone the way of, like, building supercomputers. Right? Like, you know, people don't talk about building machines anymore. They're like, here's my entire rack. Right? This starts looking like, you know, what Cray was doing. So in some ways, you know, you could be like, oh, we have kind of regressed back to the supercomputer era. And I I don't know if I use that word positively right. We're building this, like, fully vertically integrated systems. I'm not sure that's the route for for inference. I think inference is much better served as a large scale workload where you can utilize a bunch of relatively commodity hardware and be able to, like, scale out efficiently. Alright, everyone. Welcome to another episode of the TwinWell AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Zane Asghar. Zane is cofounder and CEO at Genet Labs and an adjunct professor of computer science at Stanford University. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Zane, welcome to the podcast. Hi, Sam. Thanks for having me here. Super excited to be here. Excited to have you on the show and looking forward to digging into our conversation. We'll be talking about the work you're doing around heterogeneous inference for agentic systems. To get us going, I'd love to have you share a little bit about your background. As I mentioned, I'm cofounder and CEO of Gimlet and also adjunct faculty of computer science at Stanford. Prior to this, I was a general manager at New Relic through an acquisition of my previous startup, Pixi. And, actually, you know, a bunch of people from Pixi are are now at Gimlet as well. I was an EIR benchmark, capital where, the idea for Pixi came out of. I was in Google Research and spent a lot of time at NVIDIA. So I've kind of focused on, like, efficient compute and being able to orchestrate and run compute efficiently on on large scale clusters. And where did the idea for Gimlet come from? What's your what are you going after there? So when we started Gimlet a couple years ago, we had a we had a focus on, like, how do we actually make AI workloads, you know, at least 10 times more efficient. Right? And part of the part of the challenge over here has been that, you know, there's you've seen us, like, huge explosion in AI, AI workloads, especially around Adjunctic AI where you're, you know, consuming, like, 10 x more more tokens. And, really, if you wanna be able to keep this somewhat sustainable, you need to have these big leaps on improvements. So that was our original focus at Gimlet. And when we started off, we were really thinking about how do we get models to run on things like your laptop and, you know, Raspberry Pis or or whatever. …

Get the full transcript (9,308 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all The TWIML AI Podcast transcripts →

You just read a 3-minute summary of a 45-minute episode.

Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

More from The TWIML AI Podcast

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.

You're clearly into The TWIML AI Podcast.

Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime