Specifications

Product

DEEPX DX-M2 (XM2 silicon)

Type

Edge generative-AI NPU / AI accelerator

AI Performance

80 TOPS (INT8)

Power Budget

5 W

LLM Support

20 B to 100 B parameter models using mixture-of-experts (MoE)

Token Rate

20–30 tokens/s on-device generation

Predecessor

DX-M1 / DX-M1M — 25 TOPS at 1–5 W

MoE Compiler

DXNN SDK / DX-COM compiler with MoE support

Framework Support

PyTorch, ONNX, TensorFlow, Keras

Target Workloads

On-device LLM and multimodal generative inference

Efficiency

≈16 TOPS/W — data-centre class inference without a data centre

Silicon Status

XM2 tape-out scheduled around March 2027

Ecosystem

DX-M1 / DX-M1+ modules, V2H DEEPX SBC, ROCK 5B+ and workstation integrations

Storage Integration

Apacer 4 TB PCIe SSD + 2 × DX-M1 on a single M.2 board

Development Suite

DX-AS (All Suite) fully integrated version-aligned package

Deployment Targets

Robotics, physical AI, smart factory, smart city, mobility

Form Factor

Module / embedded integration (M.2 family shares the DXNN stack)

Compliance

RoHS

Lead Time

Engineering samples on request

MOQ

1 unit (samples)

Overview

The DEEPX DX-M2 targets the problem that defines edge AI in 2026: running a real large language model locally, at a power budget an embedded system can actually supply. It delivers 80 TOPS inside a 5 W envelope, and — unlike vision-only NPUs — it is specified for generative workloads, supporting mixture-of-experts models from 20 B up to 100 B parameters.

The practical outcome is 20–30 tokens/s of on-device generation. That is the difference between an edge box that forwards prompts to a cloud API and one that answers without a network round trip, which is what robotics, industrial safety and privacy-sensitive deployments need.

The DX-M2 continues DEEPX's DXNN software line, so models compiled for the 25 TOPS DX-M1 / DX-M1M modules port forward through the same SDK and DX-COM compiler path, and the DX-AS All Suite package keeps toolchain versions aligned.

Key Benefits

80 TOPS in 5 W — roughly 16 TOPS/W, so no active cooling or 12 V rail is required; generative-AI class capability with 20–100 B parameter MoE support; 20–30 tokens/s on-device LLM generation; DXNN SDK continuity from the DX-M1 family for forward porting.

Applications

On-device LLM assistants in robotics and physical AI; real-time industrial safety and visual inspection with language-model reasoning; smart-city and mobility inference without cloud egress; privacy-preserving deployments where data cannot leave the device.

Request a Quote — DEEPX DX-M2 — 80 TOPS GENERATIVE AI NPU AT 5 W

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

DEEPX DX-M1 — 25 TOPS M.2 Edge AI Accelerator MemryX MX3 — Edge AI Accelerator Hailo-8 M.2 — 26 TOPS Edge AI Accelerator