d-Matrix Partners With NVIDIA to Bring Inference XPUs Into AI Factory Racks
d-Matrix will integrate Raptor with NVIDIA’s latest rack architecture, which includes Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking.
AI inference chipmaker d-Matrix has announced a multi-year collaboration with NVIDIA that will integrate its next-generation inference XPUs into NVIDIA’s widely deployed AI factory infrastructure.
The collaboration will initially centre on d-Matrix’s Raptor XPU being integrated into an NVIDIA MGX rack-scale system enabled by NVLink Fusion. The system is for AI labs, hyperscalers and neocloud providers running latency-sensitive AI services where faster responses can command a premium.

As an NVIDIA NVLink Fusion partner, d-Matrix will integrate Raptor with NVIDIA’s latest rack architecture, which includes Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs and Spectrum-X Ethernet networking.
“Being integrated into NVIDIA’s latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform. That’s the future d-Matrix has been building toward—ultra-low latency, energy-efficient inference XPUs and GPUs working together, at rack scale, to deliver premium AI experiences,” Sid Sheth, d-Matrix Founder and CEO, said.
“With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms — expanding accelerator choice for customers building the next generation of AI factories,” Jensen Huang, NVIDIA Founder and CEO, added.
The partnership comes as AI providers increasingly explore heterogeneous infrastructure, using different processors for different stages of inference. d-Matrix said its rack architecture can pair its Raptor XPUs with NVIDIA GPUs, allowing operators to optimise workloads based on performance and latency requirements.
For AI coding applications, for example, GPUs can handle the compute-intensive prefill stage, while Raptor XPUs can accelerate the decode stage, where response latency is critical.

Raptor is the successor to d-Matrix’s Corsair XPU and uses the company’s memory-centric architecture. Its key technology combines a DRAM memory chip with an SRAM compute chip in a single package using a 3D stacking approach.
Raptor is expected to tape out before the end of 2026, with initial availability of Raptor XPUs integrated into NVIDIA MGX racks expected in Q4 2027. The company said the platform is already being evaluated by AI hyperscalers and frontier labs and is backed by more than 100 patents.
The collaboration gives d-Matrix access to NVIDIA’s established rack architecture and supply chain while giving customers another option for scaling inference workloads alongside NVIDIA GPUs.
Earlier this year, d-Matrix acquired GigaIO’s data centre business, a systems engineering organisation with deep expertise in rack-scale infrastructure and high-performance interconnects.





