cd /news/ai-chips/d-matrix-raptor-xpus-join-nvidia-mgx… · home topics ai-chips article
[ARTICLE · art-126202] src=storagereview.com ↗ pub= topic=ai-chips verified=true sentiment=↑ positive

d-Matrix Raptor XPUs Join NVIDIA MGX Racks Through NVLink Fusion, With First Systems Due Q4 2027

D-Matrix will integrate its next-generation Raptor inference XPUs into NVIDIA's MGX rack architecture via NVLink Fusion under a multi-year collaboration, with initial availability expected in the fourth quarter of 2027. The first product is a d-Matrix rack built on the MGX reference design with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet, targeting AI labs, hyperscalers, and neoclouds. d-Matrix CEO Sid Sheth said the deal gives customers "a faster, lower-risk path to deploy and scale ultralow-latency inference," with Raptor handling the decode phase alongside GPUs running prefill.

by read3 min views1 publishedSep 10, 2026
d-Matrix Raptor XPUs Join NVIDIA MGX Racks Through NVLink Fusion, With First Systems Due Q4 2027
Image: Storagereview (auto-discovered)

d-Matrix will put its next-generation Raptor inference XPUs into NVIDIA’s MGX rack architecture using NVLink Fusion, under a collaboration with NVIDIA that the company describes as a multi-year product roadmap. The first product is a d-Matrix rack built on the MGX reference design with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet, aimed at AI labs, hyperscalers, and neoclouds selling what d-Matrix calls premium, ultra-low-latency token services. Initial availability of Raptor XPUs in the MGX rack is expected in the fourth quarter of 2027.

The argument for the deal is centered around time and risk rather than silicon. Getting a custom accelerator into production at AI factory scale means sourcing and validating a scale-up interconnect, a rack design, power delivery, liquid cooling, scale-out networking, and a supply chain. NVLink Fusion is NVIDIA’s program for letting third-party XPU and CPU designers connect into that stack instead of building it. For d-Matrix, that means Raptor gets the same NVLink scale-up domain, MGX rack, cooling, and supply chain that NVIDIA’s own systems use, and data center operators can stand up one-rack architecture that carries GPUs, CPUs, and XPUs.

“Demand for inference is soaring, but capital, time and energy remain finite,” said Sid Sheth, cofounder and CEO of d-Matrix. “With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference.” NVIDIA CEO Jensen Huang framed it from the other side: “With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms, expanding accelerator choice for customers building the next generation of AI factories.”

The interconnect itself is sixth-generation NVLink, which NVIDIA rates at 3.6TB/s per GPU or XPU and 260TB/s of aggregate bandwidth across a 72-accelerator all-to-all domain, more than 14 times the bandwidth of PCIe Gen6 by NVIDIA’s comparison. d-Matrix plans to use it to connect Raptor XPUs into a single high-bandwidth scale-up domain, with the racks built from modular, cable-free MGX trays. Astera Labs is also part of the design, supplying connectivity for the system as an existing member of the NVLink Fusion ecosystem, which also includes Arm, Intel, Fujitsu, SiFive, Marvell, MediaTek, Samsung, Alchip, GUC, Cadence, Synopsys, Ayar Labs, and Lightmatter. We covered MediaTek’s NVLink Fusion XPU work and AWS bringing NVLink Fusion to Trainium in the last two weeks; d-Matrix is the latest to take the same route.

Disaggregated Inference: GPUs for Prefill, Raptor for Decode #

The deployment model d-Matrix is pitching is heterogeneous disaggregation. Rather than replacing GPUs, a Raptor rack sits alongside a Vera Rubin NVL72 and takes the phase of inference it is built for. For AI coding assistants, the example d-Matrix uses, the GPUs handle the compute-heavy prefill phase while the Raptor XPUs run the latency-sensitive decode phase, where interactivity is what the customer is paying for. The same split applies to real-time chatbots and voice agents, and it is the reason d-Matrix keeps describing the target as a premium token economy: workloads where buyers pay more for speed.

Raptor is the follow-on to d-Matrix’s Corsair XPU, which is in production today. It extends the company’s memory-centric design with what d-Matrix calls a first-of-its-kind 3D DRAM stacking approach, pairing a DRAM chip with an SRAM compute chip in a single two-story package. Co-founder and CTO Sudeep Bhoja previewed the technology at Hot Chips 2026, and the technical details have been published through IEEE. d-Matrix says Raptor was designed from the start for NVLink Fusion and MGX integration, is expected to tape out before the end of this year, is under evaluation at hyperscalers and frontier labs, and is backed by more than 100 patents. The company is demonstrating the design at the AI Infra Summit next week.

── more in #ai-chips 4 stories · sorted by recency
── more on @d-matrix 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/d-matrix-raptor-xpus…] indexed:0 read:3min 2026-09-10 ·