2,048 DIMC cores · 150 TB/s performance memory · up to 9,600 TFLOPS (4-bit) · PCIe Gen5 · Raptor adds 3DIMC 3D DRAM
Single full-height, full-length PCIe Gen5 x16 inference accelerator built on TSMC 6nm with two ASIC packages of four chiplets each.
Two cards bridged package-to-package via the DMX Bridge for all-to-all chiplet connectivity and faster token generation.
Next-generation accelerator and successor to Corsair, introducing d-Matrix's 3D In-Memory Compute (3DIMC) technology.
d-Matrix (Santa Clara, CA)
Corsair / Raptor
AI inference accelerator
(in-memory compute)
DIMC digital in-memory compute
TSMC 6nm · 2 ASICs
4 chiplets each
2,048 (card)
4,096 (dual-card)
9,600 TFLOPS (card)
19,200 TFLOPS (dual)
2 GB · 150 TB/s
Up to 256 GB
400 GB/s bandwidth
PCIe Gen5 x16
128 GB/s bi-directional
~38 TOPS/W
600 W (card)
1,200 W (dual-card)
Block floating point ·
MX / Microscaling (OCP)
Redfish · PLDM · SPDM
secure boot
Aviator stack · JetStream networking
Corsair full production
June 2026
Available
Quote
d-Matrix is a Santa Clara-based startup pioneering Digital In-Memory Compute (DIMC) for datacenter AI inference. Rather than repurposing a GPU, its Corsair accelerator integrates compute directly into memory to break the "memory wall" — the dominant bottleneck in autoregressive decode, where every generated token re-reads the full weight matrix and KV cache.
Each Corsair PCIe card packs 2,048 DIMC compute cores with 2 GB of on-chip performance memory delivering 150 TB/s of bandwidth — an order of magnitude beyond today's HBM. A two-tier memory system adds up to 256 GB of off-chip capacity memory (LPDDR5X, 400 GB/s) for large models and long context windows. Peak compute is 2,400 TFLOPS at 8-bit and 9,600 TFLOPS at 4-bit; pairing two cards over the DMX Bridge doubles everything to 4,096 cores, 19,200 TFLOPS, and 512 GB capacity memory at a 1,200 W envelope. Built on TSMC 6nm, each card mounts two ASIC packages of four chiplets in an all-to-all topology.
d-Matrix positions Corsair as a GPU complement for the decode phase of disaggregated inference — GPUs handle prefill and attention, while Corsair serves the latency-sensitive token-generation workload at ~38 TOPS/W. The platform claims up to 10× better interactive performance, 3× energy efficiency, and 3× cost-performance versus GPU alternatives, and entered full production in June 2026, winning the 2026 AI Breakthrough "AI Processor Innovation Award."
The roadmap successor, Raptor, introduces 3DIMC (3D In-Memory Compute) — billed as the world's first 3D DRAM solution for AI inference. Co-developed with Alchip on advanced 2.5D/3D CoWoS packaging and orchestrated by the AndesCore AX46MPV RISC-V vector CPU, Raptor stacks compute and memory in 3D to cut data-movement energy and latency ten-fold, targeting up to 10× faster inference than HBM4-based solutions for generative and agentic AI. The technology has been validated on d-Matrix's Pavehawk test silicon.
Need d-Matrix Corsair / Raptor inference accelerators?
Contact QS Compute for availability, configuration, and volume pricing.
Request Quote