Skip to main content
Eye on AI

Why People Are Paying 10x More for AI | Sid Sheth, d-Matrix

50 min episode · 2 min read
·
Sid Sheth

Episode

50 min

Read time

2 min

Topics

Fundraising & VC, Leadership, Marketing

AI-Generated Summary

Key Takeaways

  • Premium Token Pricing: Two distinct inference tiers now exist in the market. Standard throughput-based inference runs approximately $2 per million tokens, while low-latency interactive inference commands $20 per million tokens. Enterprises and developers willingly pay this 10x premium for real-time responsiveness in coding tools like Claude Code and agent-to-agent communication workflows.
  • Memory Bandwidth as the Differentiator: Low-latency inference requires compute architectures that place memory and processing together, delivering an order-of-magnitude more memory bandwidth than GPU-based HBM solutions. d-Matrix, Groq, and Cerebras each use this approach. Companies evaluating inference hardware should benchmark memory bandwidth directly, not just raw compute FLOPS, when targeting interactive applications.
  • Rack-Scale Expertise as a Competitive Requirement: Hyperscalers, NeoCloud providers, and frontier labs now purchase and plan compute capacity in rack units, typically 50–150 kilowatts per rack. Chip vendors without rack-scale engineering expertise cannot support deployment or troubleshooting at customer sites. d-Matrix acquired Giga IO specifically to build this capability and accelerate deployment timelines.
  • Mandatory AI Adoption as Organizational Strategy: d-Matrix mandates Claude Code and Claude for Work enterprise tools across every department — hardware, software, marketing, operations, and corporate communications. A dedicated internal AI team governs rollout. The zero-data-retention policy on Anthropic's enterprise tier addresses proprietary data concerns, making company-wide deployment viable without security compromises.
  • Agentic Task Abstraction Trajectory: Current AI agents handle discrete tasks like database queries or presentation generation. The next evolution involves teams of agents collectively executing high-level strategic directives — such as identifying decision-makers, reshaping product positioning, and building full sales strategies autonomously. Companies building agent-first structures with minimal headcount, like the $1.6B revenue two-person operation cited, represent an accelerating segment.

What It Covers

Sid Sheth, CEO of d-Matrix, explains how a premium token economy is emerging around low-latency AI inference, where users pay 10x more ($20 vs $2 per million tokens) for real-time interactivity. He covers d-Matrix's chiplet architecture, the Giga IO acquisition, agentic enterprise adoption, and sovereign AI infrastructure demand.

Key Questions Answered

  • Premium Token Pricing: Two distinct inference tiers now exist in the market. Standard throughput-based inference runs approximately $2 per million tokens, while low-latency interactive inference commands $20 per million tokens. Enterprises and developers willingly pay this 10x premium for real-time responsiveness in coding tools like Claude Code and agent-to-agent communication workflows.
  • Memory Bandwidth as the Differentiator: Low-latency inference requires compute architectures that place memory and processing together, delivering an order-of-magnitude more memory bandwidth than GPU-based HBM solutions. d-Matrix, Groq, and Cerebras each use this approach. Companies evaluating inference hardware should benchmark memory bandwidth directly, not just raw compute FLOPS, when targeting interactive applications.
  • Rack-Scale Expertise as a Competitive Requirement: Hyperscalers, NeoCloud providers, and frontier labs now purchase and plan compute capacity in rack units, typically 50–150 kilowatts per rack. Chip vendors without rack-scale engineering expertise cannot support deployment or troubleshooting at customer sites. d-Matrix acquired Giga IO specifically to build this capability and accelerate deployment timelines.
  • Mandatory AI Adoption as Organizational Strategy: d-Matrix mandates Claude Code and Claude for Work enterprise tools across every department — hardware, software, marketing, operations, and corporate communications. A dedicated internal AI team governs rollout. The zero-data-retention policy on Anthropic's enterprise tier addresses proprietary data concerns, making company-wide deployment viable without security compromises.
  • Agentic Task Abstraction Trajectory: Current AI agents handle discrete tasks like database queries or presentation generation. The next evolution involves teams of agents collectively executing high-level strategic directives — such as identifying decision-makers, reshaping product positioning, and building full sales strategies autonomously. Companies building agent-first structures with minimal headcount, like the $1.6B revenue two-person operation cited, represent an accelerating segment.

Notable Moment

Sheth describes using Claude to run d-Matrix's merger and acquisition strategy, including integration planning for the Giga IO deal. He notes that analysis previously requiring entire banking advisory teams now produces detailed integration reports in roughly fifteen minutes, fundamentally changing how small companies can execute M&A.

Know someone who'd find this useful?

Episode Transcript

We talk about organizations being entirely run by agents, right, how far away from that, eventually, you have to run large portions of most companies. Do you see emerging companies, startups that are agent first? The New York Times had a piece recently about a guy and his brother running $1,600,000,000 sales company that basically a agent first company. Do you see that segment of the economy growing? It's really an opportunity for everyone to go into something new. I think at the d matrix, we are being very proactive about the use of AI across the whole organization for various different use cases. We have an AI team in the company that is entirely focused on building AI across the board. Right? So it is not an option. It is mandated. Tell us what you're talking about at Humanex. I mean, I I read that you've had some acquisitions since the last time we spoke. We're moving into a a larger system, not just providing chips. And I'm interested in that. You know, my interest is generally in algorithmic research. Mhmm. But I do try and follow hardware, and you guys are are a player in that space. Mhmm. I think I told you last time. I talked periodically to Andrew Feldman at Cerebras, to Rodrigo Liang at Semba Nova. I haven't spoken to Brock. If I'm not mistaken, those you and those three are the primary players in this new, internship space or maybe not. I mean, if you could give me, an overview on what you guys are doing and how you're differentiating. The inferencing space specifically, right, is, kind of bifurcating. Right? Because, I think everyone looked at inference as one, you know, kind of, you know, singular entity. Right? But I think there is lots of nuance. And when it comes to inference, there's not a one size fits all. Yep. Something we've been saying for for a long time. So, you know, depending on where you are doing the inferencing, what your target markets are, what applications you're going after, you can't really build a single chip that serves all of influencing's needs. Right? And, so you gotta really kinda break it down. And, not every company is playing in every pocket of the market. You can't. Right? GPUs, you know, just because they are so general, the ecosystem is so broad. Clearly, GPUs can go into many different applications, for inference. Right? But, they're not very, very efficient because of that. Right? I mean so, yes, they're very broad in their applicability, their general purpose. But but when it comes to certain breakaway applications where applications need, a very specific metric, fully optimized, then GPUs will not do well in that market. And I think so I think the way to look at it is, okay. Yeah. The inferencing market is gonna be the largest part of AI compute. It's expected to be over a trillion dollar market, you know, call it in …

Get the full transcript (8,924 words) + summary by email — free

One-time email with the complete transcript and AI summary of this episode. No account needed.

One email, no spam. We’ll also show you what SignalCast does.

Browse all Eye on AI transcripts →

You just read a 3-minute summary of a 47-minute episode.

Get Eye on AI summarized like this every Monday — plus up to 2 more podcasts, free.

Pick Your Podcasts — Free

Keep Reading

Books, tools, and gear mentioned in this episode

SignalCast may earn commission on purchases via these links.

Tools

  • Claude CodeRecommended

    by Anthropic

    Enterprises and developers willingly pay this 10x premium for real-time responsiveness in coding tools like Claude Code and agent-to-agent communication workflows.
  • Claude for WorkRecommended

    by Anthropic

    d-Matrix mandates Claude Code and Claude for Work enterprise tools across every department — hardware, software, marketing, operations, and corporate communications.

company

  • d-Matrix acquired Giga IO specifically to build this capability and accelerate deployment timelines.
  • d-Matrix, Groq, and Cerebras each use this approach.
  • d-Matrix, Groq, and Cerebras each use this approach.
  • The zero-data-retention policy on Anthropic's enterprise tier addresses proprietary data concerns, making company-wide deployment viable without security compromises.
  • d-MatrixBy guest
    Sid Sheth, CEO of d-Matrix, explains how a premium token economy is emerging around low-latency AI inference

More from Eye on AI

We summarize every new episode. Want them in your inbox?

Similar Episodes

Related episodes from other podcasts

Explore Related Topics

This podcast is featured in Best AI Podcasts (2026) — ranked and reviewed with AI summaries.

You're clearly into Eye on AI.

Every Monday, we deliver AI summaries of the latest episodes from Eye on AI and 192+ other podcasts. Free for one show.

Start My Monday Digest

No credit card · Unsubscribe anytime