Specifications
Model
DX-H1 Quattro
Vendor
DEEPX
AI Performance
100 TOPS (INT8)
AI Processors
4x DX-H1 NPU in a Quattro parallel-processing layout
Host Interface
PCIe Gen3 x16 (BIOS bifurcation x4x4x4x4 required)
Memory
16 GB LPDDR5 at 6000 MT/s
Power Consumption
20 W TDP
Efficiency
Approximately 5 TOPS per watt
Thermal Solution
Heatsink
Form Factor
Low-profile PCIe card
Dimensions
68.1 x 165.9 x 16 mm (with heatsink)
OS Support
Windows 11 / 10
Linux Support
Debian-based Linux - Ubuntu 24.04 / 22.04 / 20.04 LTS
Additional Platforms
Yocto Project and Docker
AI Frameworks
Ultralytics, PyTorch, TensorFlow, ONNX, Keras
System Support
x86 and ARM-based architectures
Installation
Standard PCIe slot - drop-in upgrade for existing servers
Best Fit
Data-center and edge AI inference workloads
Overview
The DEEPX DX-H1 Quattro is a low-profile PCIe inference accelerator that packages four DX-H1 NPUs on a single card, delivering 100 TOPS of INT8 performance inside a 20 W thermal design power. That ratio is the point of the product: it lets an existing x86 or ARM server add substantial AI inference capacity inside the constraints of a standard PCIe slot and a low-profile chassis.
Because the card is a standard PCIe Gen3 x16 device - hosting is done through x4x4x4x4 bifurcation - it installs into existing servers rather than requiring a specialised AI platform. DEEPX pairs the card with the DXNN SDK and framework support for Ultralytics, PyTorch, TensorFlow, ONNX and Keras, so ONNX and supported PyTorch models can be compiled to the NPUs without redesigning the application stack.
Key Benefits
Extreme energy efficiency. 100 TOPS inside 20 W makes the DX-H1 Quattro a leader for AI inference per watt, cutting both power draw and rack thermal load. Four NPUs, one slot. The Quattro layout gives parallel processing across four DX-H1 devices for multi-stream workloads. Drop-in integration. A standard low-profile PCIe card with x86 and ARM host support and Windows/Linux/Yocto/Docker coverage. Framework friendly. Ultralytics, PyTorch, TensorFlow, ONNX and Keras are all supported through the DXNN toolchain.
Applications
Data-center AI inference, edge server vision analytics, smart factory inspection, video analytics offload, private inferencing nodes, and retrofitting existing x86 or ARM servers with dedicated NPU capacity.
Request a Quote — DEEPX DX-H1 QUATTRO — 100 TOPS INT8 LOW-PROFILE PCIE AI INFERENCE CARD
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →