Specifications

Model

DX-H1 Quattro

Vendor

DEEPX

AI Performance

100 TOPS (INT8)

AI Processors

4x DX-H1 NPU in a Quattro parallel-processing layout

Host Interface

PCIe Gen3 x16 (BIOS bifurcation x4x4x4x4 required)

Memory

16 GB LPDDR5 at 6000 MT/s

Power Consumption

20 W TDP

Efficiency

Approximately 5 TOPS per watt

Thermal Solution

Heatsink

Form Factor

Low-profile PCIe card

Dimensions

68.1 x 165.9 x 16 mm (with heatsink)

OS Support

Windows 11 / 10

Linux Support

Debian-based Linux - Ubuntu 24.04 / 22.04 / 20.04 LTS

Additional Platforms

Yocto Project and Docker

AI Frameworks

Ultralytics, PyTorch, TensorFlow, ONNX, Keras

System Support

x86 and ARM-based architectures

Installation

Standard PCIe slot - drop-in upgrade for existing servers

Best Fit

Data-center and edge AI inference workloads

Overview

The DEEPX DX-H1 Quattro is a low-profile PCIe inference accelerator that packages four DX-H1 NPUs on a single card, delivering 100 TOPS of INT8 performance inside a 20 W thermal design power. That ratio is the point of the product: it lets an existing x86 or ARM server add substantial AI inference capacity inside the constraints of a standard PCIe slot and a low-profile chassis.

Because the card is a standard PCIe Gen3 x16 device - hosting is done through x4x4x4x4 bifurcation - it installs into existing servers rather than requiring a specialised AI platform. DEEPX pairs the card with the DXNN SDK and framework support for Ultralytics, PyTorch, TensorFlow, ONNX and Keras, so ONNX and supported PyTorch models can be compiled to the NPUs without redesigning the application stack.

Key Benefits

Extreme energy efficiency. 100 TOPS inside 20 W makes the DX-H1 Quattro a leader for AI inference per watt, cutting both power draw and rack thermal load. Four NPUs, one slot. The Quattro layout gives parallel processing across four DX-H1 devices for multi-stream workloads. Drop-in integration. A standard low-profile PCIe card with x86 and ARM host support and Windows/Linux/Yocto/Docker coverage. Framework friendly. Ultralytics, PyTorch, TensorFlow, ONNX and Keras are all supported through the DXNN toolchain.

Applications

Data-center AI inference, edge server vision analytics, smart factory inspection, video analytics offload, private inferencing nodes, and retrofitting existing x86 or ARM servers with dedicated NPU capacity.

Request a Quote — DEEPX DX-H1 QUATTRO — 100 TOPS INT8 LOW-PROFILE PCIE AI INFERENCE CARD

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

DEEPX DX-M2 Edge AI NPU DEEPX DX-M1 Edge AI NPU Hailo-8 Century PCIe Card Axelera Edge 130p