Specifications
Product
DEEPX DX-M2 (XM2 silicon)
Type
Edge generative-AI NPU / AI accelerator
AI Performance
80 TOPS (INT8)
Power Budget
5 W
LLM Support
20 B to 100 B parameter models using mixture-of-experts (MoE)
Token Rate
20–30 tokens/s on-device generation
Predecessor
DX-M1 / DX-M1M — 25 TOPS at 1–5 W
MoE Compiler
DXNN SDK / DX-COM compiler with MoE support
Framework Support
PyTorch, ONNX, TensorFlow, Keras
Target Workloads
On-device LLM and multimodal generative inference
Efficiency
≈16 TOPS/W — data-centre class inference without a data centre
Silicon Status
XM2 tape-out scheduled around March 2027
Ecosystem
DX-M1 / DX-M1+ modules, V2H DEEPX SBC, ROCK 5B+ and workstation integrations
Storage Integration
Apacer 4 TB PCIe SSD + 2 × DX-M1 on a single M.2 board
Development Suite
DX-AS (All Suite) fully integrated version-aligned package
Deployment Targets
Robotics, physical AI, smart factory, smart city, mobility
Form Factor
Module / embedded integration (M.2 family shares the DXNN stack)
Compliance
RoHS
Lead Time
Engineering samples on request
MOQ
1 unit (samples)
Overview
The DEEPX DX-M2 targets the problem that defines edge AI in 2026: running a real large language model locally, at a power budget an embedded system can actually supply. It delivers 80 TOPS inside a 5 W envelope, and — unlike vision-only NPUs — it is specified for generative workloads, supporting mixture-of-experts models from 20 B up to 100 B parameters.
The practical outcome is 20–30 tokens/s of on-device generation. That is the difference between an edge box that forwards prompts to a cloud API and one that answers without a network round trip, which is what robotics, industrial safety and privacy-sensitive deployments need.
The DX-M2 continues DEEPX's DXNN software line, so models compiled for the 25 TOPS DX-M1 / DX-M1M modules port forward through the same SDK and DX-COM compiler path, and the DX-AS All Suite package keeps toolchain versions aligned.
Key Benefits
80 TOPS in 5 W — roughly 16 TOPS/W, so no active cooling or 12 V rail is required; generative-AI class capability with 20–100 B parameter MoE support; 20–30 tokens/s on-device LLM generation; DXNN SDK continuity from the DX-M1 family for forward porting.
Applications
On-device LLM assistants in robotics and physical AI; real-time industrial safety and visual inspection with language-model reasoning; smart-city and mobility inference without cloud egress; privacy-preserving deployments where data cannot leave the device.
Request a Quote — DEEPX DX-M2 — 80 TOPS GENERATIVE AI NPU AT 5 W
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →