Why People Are Paying 10x More for AI | Sid Sheth, d-Matrix
The AI chip market looks monolithic from the outside - NVIDIA dominates, and everyone else is fighting for scraps. But d-Matrix's CEO Sid Sheth argues that the market is quietly splitting into two distinct tiers, and the one that's exploding right now is the one NVIDIA's architecture isn't built for. In this episode, Sid joins Craig Smith to explain the "premium token economy": a new class of AI inference where interactivity is the product, users pay ten times more per million tokens for instant responses, and the memory bandwidth limits of GPU-based systems create a structural ceiling that purpose-built architectures don't have.
The conversation is unusually candid about what AI actually looks like at the executive level: Sid describes using Claude as a sounding board for M&A strategy, producing full integration plans in 15 minutes that used to require entire banking advisory teams, and watching AI shift from a tool that echoed his ideas back at him to one that genuinely disagrees, flags what he missed, and pushes back with enough confidence to be useful. He also makes the case that we're at the beginning of a shift from individual agents to what he calls "organizational AI" - teams of agents running entire company functions at a high level of abstraction - and that the infrastructure bet d-Matrix is making positions them directly in the path of that wave.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.
Confronting AI’s Next Big Challenge: Inference Compute
While AI training garners most of the spotlight — and investment — the demands ofAI inferenceare shaping up to be an even bigger challenge. In this episode ofThe New Stack Makers, Sid Sheth, founder and CEO of d-Matrix, argues that inference is anything but one-size-fits-all. Different use cases — from low-cost to high-interactivity or throughput-optimized — require tailored hardware, and existing GPU architectures aren’t built to address all these needs simultaneously.
“The world of inference is going to be truly heterogeneous,” Sheth said, meaning specialized hardware will be required to meet diverse performance profiles. A major bottleneck? The distance between memory and compute. Inference, especially in generative AI and agentic workflows, requires constant memory access, so minimizing the distance data must travel is key to improving performance and reducing cost.
To address this, d-Matrix developed Corsair, a modular platform where memory and compute are vertically stacked — “like pancakes” — enabling faster, more efficient inference. The result is scalable, flexible AI infrastructure purpose-built for inference at scale.
Learn more from The New Stack about inference compute and AI
Scaling AI Inference at the Edge with Distributed PostgreSQL
Deep Infra Is Building an AI Inference Cloud for Developers
Join our community of newsletter subscribers to stay on top of the news and at the top of your game
#251 Sid Sheth: How d-Matrix is Disrupting AI Inference in 2025
This episode is sponsored by the DFINITY Foundation.
DFINITY Foundation's mission is to develop and contribute technology that enables the Internet Computer (ICP) blockchain and its ecosystem, aiming to shift cloud computing into a fully decentralized state.
Find out more at https://internetcomputer.org/
In this episode of Eye on AI, we sit down with Sid Sheth, CEO and Co-Founder of d-Matrix, to explore how his company is revolutionizing AI inference hardware and taking on industry giants like NVIDIA.
Sid shares his journey from building multi-billion-dollar businesses in semiconductors to founding d-Matrix—a startup focused on generative AI inference, chiplet-based architecture, and ultra-low latency AI acceleration.
We break down:
Why the future of AI lies in inference, not training
How d-Matrix's Corsair PCIe accelerator outperforms NVIDIA's H200
The role of in-memory compute and high bandwidth memory in next-gen AI chips
How d-Matrix integrates seamlessly into hyperscaler and enterprise cloud environments
Why AI infrastructure is becoming heterogeneous and what that means for developers
The global outlook on inference chips—from the US to APAC and beyond
How Sid plans to build the next NVIDIA-level company from the ground up.
Whether you're building in AI infrastructure, investing in semiconductors, or just curious about the future of generative AI at scale, this episode is packed with value.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Intro
(02:46) Introducing Sid Sheth
(05:27) Why He Started d-Matrix
(07:28) Lessons from Building a $2.5B Chip Business
(11:52) How d-Matrix Prototypes New Chips
(15:06) Working with Hyperscalers Like Google & Amazon
(17:27) What's Inside the Corsair AI Accelerator
(21:12) How d-Matrix Beats NVIDIA on Chip Efficiency
(24:10) The Memory Bandwidth Advantage Explained
(26:27) Running Massive AI Models at High Speed
(30:20) Why Inference Isn't One-Size-Fits-All
(32:40) The Future of AI Hardware
(36:28) Supporting Llama 3 and Other Open Models
(40:16) Is the Inference Market Big Enough?
(43:21) Why the US Is Still the Key Market
(46:39) Can India Compete in the AI Chip Race?
(49:09) Will China Catch Up on AI Hardware?