Scaling Agentic Inference Across Heterogeneous Compute with Zain Asgar - #757
Episode
48 min
Read time
2 min
Topics
Startups, Fundraising & VC, Artificial Intelligence
AI-Generated Summary
Key Takeaways
- ✓Workload Disaggregation Strategy: Gimlet splits agent workflows into granular components, assigns performance-critical pieces to premium hardware like B200s, and offloads less critical tasks to lower-cost accelerators, optimizing cost per token while maintaining SLA requirements through dynamic resource allocation.
- ✓Kernel Optimization Performance: LLM-based automatic kernel synthesis delivers single-digit improvements on mature H100 hardware but achieves 20-40% gains on newer B200/RTX 6000 systems and over 2x speedups on AMD/Intel/Apple hardware where optimization frameworks remain underdeveloped.
- ✓Hardware Utilization Economics: Most GPU deployments show only 30% utilization, wasting two-thirds of capacity. Heterogeneous orchestration captures the majority of cost savings by efficiently packing workloads across different hardware types based on compute cost, memory bandwidth, and capacity requirements.
- ✓Multi-Agent Kernel Generation: The system uses hardware-in-the-loop testing where supervisor agents generate candidate kernels, execute them on target hardware with profiling and correctness checks, then iteratively optimize based on performance data until convergence, caching verified kernels offline.
What It Covers
Zain Asgar explains how Gimlet Labs optimizes AI inference costs through heterogeneous compute orchestration, using workload disaggregation, MLIR compilation, and LLM-generated kernel optimization across NVIDIA, AMD, and Intel hardware platforms.
Key Questions Answered
- •Workload Disaggregation Strategy: Gimlet splits agent workflows into granular components, assigns performance-critical pieces to premium hardware like B200s, and offloads less critical tasks to lower-cost accelerators, optimizing cost per token while maintaining SLA requirements through dynamic resource allocation.
- •Kernel Optimization Performance: LLM-based automatic kernel synthesis delivers single-digit improvements on mature H100 hardware but achieves 20-40% gains on newer B200/RTX 6000 systems and over 2x speedups on AMD/Intel/Apple hardware where optimization frameworks remain underdeveloped.
- •Hardware Utilization Economics: Most GPU deployments show only 30% utilization, wasting two-thirds of capacity. Heterogeneous orchestration captures the majority of cost savings by efficiently packing workloads across different hardware types based on compute cost, memory bandwidth, and capacity requirements.
- •Multi-Agent Kernel Generation: The system uses hardware-in-the-loop testing where supervisor agents generate candidate kernels, execute them on target hardware with profiling and correctness checks, then iteratively optimize based on performance data until convergence, caching verified kernels offline.
Notable Moment
Asgar reveals that AI training infrastructure has regressed to the supercomputer era with fully vertically integrated rack-scale systems reaching 600 kilowatts, while inference workloads benefit from disaggregated commodity hardware approaches that enable sustainable scaling.
Episode Transcript
If you take a look at training hardware, it's kind of gone the way of, like, building supercomputers. Right? Like, you know, people don't talk about building machines anymore. They're like, here's my entire rack. Right? This starts looking like, you know, what Cray was doing. So in some ways, you know, you could be like, oh, we have kind of regressed back to the supercomputer era. And I I don't know if I use that word positively right. We're building this, like, fully vertically integrated systems. I'm not sure that's the route for for inference. I think inference is much better served as a large scale workload where you can utilize a bunch of relatively commodity hardware and be able to, like, scale out efficiently. Alright, everyone. Welcome to another episode of the TwinWell AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Zane Asghar. Zane is cofounder and CEO at Genet Labs and an adjunct professor of computer science at Stanford University. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Zane, welcome to the podcast. Hi, Sam. Thanks for having me here. Super excited to be here. Excited to have you on the show and looking forward to digging into our conversation. We'll be talking about the work you're doing around heterogeneous inference for agentic systems. To get us going, I'd love to have you share a little bit about your background. As I mentioned, I'm cofounder and CEO of Gimlet and also adjunct faculty of computer science at Stanford. Prior to this, I was a general manager at New Relic through an acquisition of my previous startup, Pixi. And, actually, you know, a bunch of people from Pixi are are now at Gimlet as well. I was an EIR benchmark, capital where, the idea for Pixi came out of. I was in Google Research and spent a lot of time at NVIDIA. So I've kind of focused on, like, efficient compute and being able to orchestrate and run compute efficiently on on large scale clusters. And where did the idea for Gimlet come from? What's your what are you going after there? So when we started Gimlet a couple years ago, we had a we had a focus on, like, how do we actually make AI workloads, you know, at least 10 times more efficient. Right? And part of the part of the challenge over here has been that, you know, there's you've seen us, like, huge explosion in AI, AI workloads, especially around Adjunctic AI where you're, you know, consuming, like, 10 x more more tokens. And, really, if you wanna be able to keep this somewhat sustainable, you need to have these big leaps on improvements. So that was our original focus at Gimlet. And when we started off, we were really thinking about how do we get models to run on things like your laptop and, you know, Raspberry Pis or or whatever. …
Get the full transcript (9,308 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 45-minute episode.
Get The TWIML AI Podcast summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from The TWIML AI Podcast
World Models and the Future of Spatial AI with Justin Johnson - #775
Sep 1 · 66 min
Invest Like the Best with Patrick O'Shaughnessy
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Aug 25
More from The TWIML AI Podcast
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Aug 25 · 55 min
This Week in Startups
Cerebras's IPO goes vertical, and the death of OpenClaw? | E2287
May 11
More from The TWIML AI Podcast
We summarize every new episode. Want them in your inbox?
World Models and the Future of Spatial AI with Justin Johnson - #775
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773
Why Models Are AI’s Next Training Dataset with Damian Borth - #772
How AI Learns to Smell with Alex Wiltschko - #771
Similar Episodes
Related episodes from other podcasts
Invest Like the Best with Patrick O'Shaughnessy
Aug 25
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
This Week in Startups
May 11
Cerebras's IPO goes vertical, and the death of OpenClaw? | E2287
Eye on AI
Jan 4
#311 Stefano Ermon: Why Diffusion Language Models Will Define the Next Generation of LLMs
David Senra
Sep 9
Mati Staniszewski on ElevenLabs, Voice AI & Building the Communication Layer for AI
Eye on AI
Sep 8
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into The TWIML AI Podcast.
Every Monday, we deliver AI summaries of the latest episodes from The TWIML AI Podcast and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime