#313 Nick Pandher: How Inference-First Infrastructure Is Powering the Next Wave of AI
Episode
56 min
Read time
2 min
Topics
Productivity, Startups, Fundraising & VC
AI-Generated Summary
Key Takeaways
- ✓Inference Cost Optimization: NeoCloud providers reduce inference costs through power-efficient accelerators like Qualcomm AI 100 Ultra that consume less energy per rack while maintaining throughput, ideal for workloads running continuously versus hyperscaler GPU solutions designed primarily for training.
- ✓Enterprise Proof of Value Framework: Organizations should score 100-plus AI use cases before POC deployment, selecting highest-value automation opportunities first rather than tackling hardest problems initially, then progress through POC to pilot to production with validated assumptions at each stage.
- ✓Private Model Deployment: Enterprises deploy open-weight models like OpenAI's OSS 120 in private NeoCloud environments to maintain data sovereignty and regulatory compliance, avoiding concerns about proprietary information sharing while achieving near-ChatGPT-5 capability in controlled infrastructure.
- ✓Serverless Inference Platform: Qualcomm's inference stack on Cirascale enables developers to deploy foundational models with API endpoints in minutes without configuring GPU infrastructure, supporting fine-tuning capabilities and eliminating CUDA dependency for inference workloads unlike training requirements.
What It Covers
Nick Pandher from Cirascale explains how NeoCloud providers deliver inference-first infrastructure optimized for enterprise AI workloads, focusing on Qualcomm's AI accelerators, serverless deployment models, and cost-effective alternatives to hyperscalers for production inference.
Key Questions Answered
- •Inference Cost Optimization: NeoCloud providers reduce inference costs through power-efficient accelerators like Qualcomm AI 100 Ultra that consume less energy per rack while maintaining throughput, ideal for workloads running continuously versus hyperscaler GPU solutions designed primarily for training.
- •Enterprise Proof of Value Framework: Organizations should score 100-plus AI use cases before POC deployment, selecting highest-value automation opportunities first rather than tackling hardest problems initially, then progress through POC to pilot to production with validated assumptions at each stage.
- •Private Model Deployment: Enterprises deploy open-weight models like OpenAI's OSS 120 in private NeoCloud environments to maintain data sovereignty and regulatory compliance, avoiding concerns about proprietary information sharing while achieving near-ChatGPT-5 capability in controlled infrastructure.
- •Serverless Inference Platform: Qualcomm's inference stack on Cirascale enables developers to deploy foundational models with API endpoints in minutes without configuring GPU infrastructure, supporting fine-tuning capabilities and eliminating CUDA dependency for inference workloads unlike training requirements.
Notable Moment
Pandher reveals mortgage application processing can shrink from 21 days to three days using multimodal AI models with OCR to automatically parse documents and flag missing information for underwriters, demonstrating concrete automation value in regulated financial services.
Episode Transcript
Ciriscale is a NeoCloud specializing in lots of different areas around AI, but most specifically today, we're gonna focus on inference. I'm VP of product. And as part of my role as VP of product is, working with our development teams to build out unique capabilities for AI services around inference and, other areas as well. My background comes out of the GPU space. I was at NVIDIA very early on at AMD and several startups that have touched, AI models and other type of models in their use cases as well. Although I don't think of that as the cloud provider's job. Is that unusual that you're getting involved in, you know, optimizing inference, as opposed to just providing the compute? Okay. Yeah. So, Nick, why don't you start by introducing yourself to listeners, and then we'll talk about inference. Absolutely. So hello, everybody. I'm Nick Pander with Cirascale. Cirascale is a NeoCloud specializing in lots of different areas around AI. But most specifically today, we're gonna focus on inference. I'm VP of product. And as part of my role as VP of product is working with our development teams to build out unique capabilities for AI services around inference and other areas as well. My background comes out of the GPU space. I was at NVIDIA very early on at AMD and several startups that have touched AI models and other type of models in their use cases as well. Yeah. Tell us first what Cirascale does. Absolutely. So we're a more white glove type of cloud services provider. What we're focusing on is trying to understand customer specific issues and when they're looking for outside cloud services. So, typically, we're taking customers who are coming from an on premise environment or a hyperscaler and are just looking for services and capabilities that may not be part of what they're getting today. So if we look at it from an AI point of view, it's being able to deliver the correct AI fleets for given use cases. It's covering any sort of compute needs that might be required for pre or post processing data. It's delivering services that, you know, help bring high performance, reliable type of solutions, specifically, you know, if we look at inference being a differentiated area. We really wanted to, you know, increase those services and expand from our AI accelerated fleet where customers typically bring their own stacks and bring their own solutions into more ready to go solutions that a customer can take and utilize for, you know, training, fine tuning, and also for for inference capabilities as well. Yeah. And is that primarily your differentiator primarily in the software, or is it that you have, you know, specialized GPUs tuned for inference? Yes. So that's a great question. So we a few different facets there. One is we're all about operating a cloud service that's taking the highest performance with the capabilities a customer is looking for with the highest reliability …
Get the full transcript (10,107 words) + summary by email — free
One-time email with the complete transcript and AI summary of this episode. No account needed.
One email, no spam. We’ll also show you what SignalCast does.
You just read a 3-minute summary of a 53-minute episode.
Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.
Pick Your Podcasts — FreeKeep Reading
More from Eye on AI
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
Sep 8 · 54 min
This Week in Startups
Dr. Mark Hyman on Function Health & GLP-1 microdosing | E2334
Sep 4
More from Eye on AI
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
Sep 3 · 38 min
Beyond Biotech
How Evox Therapeutics is targeting CNS diseases with exosomes
Sep 4
Books, tools, and gear mentioned in this episode
SignalCast may earn commission on purchases via these links. As an Amazon Associate, SignalCast earns from qualifying purchases.
Tools
by Qualcomm
“Qualcomm's inference stack on Cirascale enables developers to deploy foundational models with API endpoints in minutes without configuring GPU infrastructure”
Gear
by Qualcomm
“NeoCloud providers reduce inference costs through power-efficient accelerators like Qualcomm AI 100 Ultra that consume less energy per rack while maintaining throughput”
Products
by OpenAI
“Enterprises deploy open-weight models like OpenAI's OSS 120 in private NeoCloud environments to maintain data sovereignty and regulatory compliance”
company
“Nick Pandher from Cirascale explains how NeoCloud providers deliver inference-first infrastructure optimized for enterprise AI workloads”
More from Eye on AI
We summarize every new episode. Want them in your inbox?
86% of What Coding Agents Do Is Just Reading — Not Solving | Alexander Whedon of Subquadratic
From 10 Drones a Month to Nearly 100,000 — Inside Ukraine's Largest Drone Manufacturer | Marko Kushnir, General Cherry
In 5 to 10 Years, Using Weapons Without AI Will Be Considered Unethical | Yaroslav Azhnyuk, The Fourth Law
Inside Ukraine's Azov Drone R&D: The Engineer Building AI Weapons 18 km From the Front Line | Alexander Palamarchuk
95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Similar Episodes
Related episodes from other podcasts
This Week in Startups
Sep 4
Dr. Mark Hyman on Function Health & GLP-1 microdosing | E2334
Beyond Biotech
Sep 4
How Evox Therapeutics is targeting CNS diseases with exosomes
Invest Like the Best with Patrick O'Shaughnessy
Aug 25
Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]
Beyond Biotech
Aug 21
Beyond biology: Nanobiotix's physics-first approach to cancer
This Week in Startups
Aug 19
Neurosymbolic AI outperforms chatbots and product search | E2327
Explore Related Topics
This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.
Read this week's Startups & Product Podcast Insights — cross-podcast analysis updated weekly.
You're clearly into Eye on AI.
Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.
Start My Monday DigestNo credit card · Unsubscribe anytime