Models & Pricing
Cambricon MLU370-X8
- Dual SiYuan 370 chiplet package
- 256 TOPS (INT8)
- 128 TOPS (INT16)
- 96 TFLOPS (FP16 / BF16 / FP32)
- MLU-Link interconnect
- 250W TDP, 7nm MLUarch03
Cambricon MLU370-S4
- Single-chip inference card
- MLUarch03 architecture
- High-density cloud inference
- FP32 / FP16 / BF16 / INT16 / INT8 / INT4
- AIDC deployment
- Fits standard server PCIe slots
Cambricon MLU290-M5
- 512 INT8 TOPS — flagship training card
- 256 TOPS (INT16)
- 64 CINT32 TOPS
- 32GB HBM2 memory
- 1,228 GB/s memory bandwidth
- Multi-card MLU-Link training clusters
Cambricon MLU220-M.2
- MLUv02 edge architecture
- ~8 TOPS theoretical peak
- Finger-sized standard M.2 card
- Low-power edge intelligence
- Video analytics at the edge
- Companion to the cloud MLU line
Specifications
Architecture
Cambricon MLUarch03 (MLU370 series)
Process Node
7nm
Card Construction
Dual SiYuan 370 chips in a multi-chiplet package (MLU370-X8)
Peak Performance
256 TOPS (INT8) / 128 TOPS (INT16) / 96 TFLOPS (FP16, BF16, FP32)
Supported Precisions
FP32, FP16, BF16, INT16, INT8, INT4
Interconnect
MLU-Link for multi-card scaling
TDP
250W (MLU370-X8); 350W class GPU comparison in vendor benchmarks
Flagship Specs
MLU290-M5 — 512 INT8 TOPS, 256 INT16 TOPS, 64 CINT32 TOPS
Flagship Memory
32GB HBM2 at 1,228 GB/s bandwidth
Edge Module
MLU220-M.2 — MLUv02, approximately 8 TOPS in an M.2 form factor
Edge Sum Module
MLU220-SOM smart module for embedded designs
Reference Platform
NF5468M5 server with Intel Xeon Gold 5218, MLU370 SDK 1.2.0
Software Stack
Cambricon NeuWare / MLU370 SDK
Intelligent Accelerator Systems
Fanqiang 1000 intelligent accelerator (MLU220 series modules)
Target Workloads
Cloud AI training and inference, AIDC deployments, edge video analytics
Product Strategy
Cloud-edge-device unified architecture across the MLU family
Market Position
Leading domestic Chinese AI accelerator supplier and NVIDIA alternative
Business Traction
Posted first profit in 2025 as domestic AI processor demand surged
Overview
Cambricon is China’s most prominent pure-play AI accelerator company and the MLU370-X8 is its training-focused workhorse card. The design is notable for its construction: two SiYuan 370 dies are packaged together as chiplets and linked by Cambricon’s MLU-Link interconnect, yielding a 7nm MLUarch03 accelerator that delivers 256 TOPS at INT8, 128 TOPS at INT16 and 96 TFLOPS across FP32, FP16 and BF16 — with native INT4 support as well. TDP is 250W.
Cambricon’s own benchmark data compares the MLU370-X8 against 350W-class GPUs, a positioning that reflects the intended market: organisations seeking datacenter AI throughput without dependence on export-controlled NVIDIA or AMD silicon. The company validated the card in a reference NF5468M5 server running Intel Xeon Gold 5218 processors with MLU370 SDK 1.2.0. The MLU370 series fills out with the MLU370-S4 single-chip inference card for high-density AIDC deployment, and the architecture family extends upward to the MLU290-M5 flagship, which carries 512 INT8 TOPS, 256 INT16 TOPS and 64 CINT32 TOPS alongside 32GB of HBM2 delivering 1,228 GB/s.
What distinguishes Cambricon from smaller AI chip startups is the cloud-edge-device span. The same product philosophy reaches down to the MLU220-M.2, a finger-sized standard M.2 accelerator card based on the MLUv02 architecture with roughly 8 TOPS of peak performance, and the MLU220-SOM smart module for embedded designs. Intelligent accelerator systems such as the Fanqiang 1000 wrap those modules for turnkey edge deployment. For a systems integrator, the appeal is a coherent software stack — Cambricon NeuWare — across training cards, inference cards and edge modules, rather than three unrelated toolchains.
The commercial signal matters too: Cambricon posted its first profit as demand for domestic Chinese AI processors surged, which is a meaningful indicator of production maturity and supply availability rather than a roadmap-only story. For buyers outside China, the practical consideration is software ecosystem compatibility and support, but for organisations building sovereign or regionally-sourced AI infrastructure, the MLU370-X8 and MLU290-M5 are among the few genuinely shipping alternatives at this performance tier.
Target Applications
Cloud AI training and inference clusters, AIDC (AI datacenter) deployments, sovereign and regionally-sourced AI infrastructure, high-density inference with the MLU370-S4, large-model training with MLU290-M5 HBM2 cards, and edge video analytics with MLU220 modules.
Related Products
Request a Quote — Cambricon MLU370-X8
QS Compute — sourcing advisory and B2B supply for Cambricon MLU370-X8 / S4 accelerator cards, MLU290-M5 HBM2 cards and MLU220 edge modules. Lead times and allocation on request.
Get Your Quote →