China’s leading general-purpose GPU line · 3rd-generation MUSA architecture · 450 W datacenter card and the 1,000-card KUAE training cluster
Moore Threads’ datacenter AI accelerator card, the compute engine behind the KUAE training cluster.
Entry-level, slot-powered GPU for display and light inference at the edge.
Consumer and workstation GPU with the first PCIe Gen5 interface in its class.
Moore Threads’ domestically-produced AI training cluster, built from MTT S4000 cards.
Moore Threads (Beijing, China)
MTT S4000 / S30 / S80 / X300
3rd-generation MUSA GPU
48 GB
768 GB/s
PCIe 5.0 ×16
25 TFLOPS (S4000)
14.7 TFLOPS (S80)
50 TFLOPS
200 TFLOPS
200 TOPS
450 W (S4000)
255 W (S80) / 40 W (S30)
MTLink
240 GB/s I/O bandwidth
H.265, H.264, AV1
48 × 1080p30
H.265, H.264, AV1, AVS2, VP9
96 × 1080p30
4 × DisplayPort 1.4a
MUSA Security Engine 2.0
TEE, encryption/decryption
Hardware virtualization
GPU elastic partitioning, SR-IOV
MUSA, DirectX, Vulkan, OpenGL, OpenGL ES
KUAE Intelligent Computing Center
1,000 cards
Quote
Moore Threads is China’s leading general-purpose GPU vendor, built on its in-house MUSA architecture and now on its third GPU generation. For AI buyers outside the mainstream NVIDIA stack, the relevant parts are the MTT S4000 datacenter accelerator and the KUAE training cluster it powers — a fully domestic, 1,000-card AI training system.
The MTT S4000 pairs 48 GB of high-bandwidth memory at 768 GB/s with a PCIe 5.0 ×16 host interface and rates 25 TFLOPS FP32, 50 TFLOPS TF32, 200 TFLOPS FP16/BF16 and 200 TOPS INT8 at a 450 W TGP. It supports multi-modal inference and AIGC serving alongside graphics workloads, carries hardware virtualization with SR-IOV isolation and GPU elastic partitioning, and exposes a MUSA Security Engine 2.0 with TEE support. Chip-to-chip communication runs over MTLink, and the card provides 240 GB/s of I/O bandwidth with 48-channel 1080p30 encode and 96-channel decode.
At the lower end, the MTT S30 is a 40 W, slot-powered single-slot half-height card with 1,024 shaders and 4 GB of GDDR6, suited to display and light edge inference, while the MTT S80 brings 4,096 MUSA cores, 14.7 TFLOPS FP32 and 16 GB of GDDR6 over PCIe Gen5. All parts are programmed through the MUSA software stack, with tooling to ease migration from CUDA. Availability and export terms are quoted on request.
Need Moore Threads MTT S4000?
Contact QS Compute for availability, configuration, and volume pricing.
Request Quote