MPN: MTT S4000
Moore Threads MTT S4000
48 GB · 200 TFLOPS FP16 · PCIe 5.0

China’s leading general-purpose GPU line · 3rd-generation MUSA architecture · 450 W datacenter card and the 1,000-card KUAE training cluster

MTT S4000

Moore Threads’ datacenter AI accelerator card, the compute engine behind the KUAE training cluster.

MTT S30

Entry-level, slot-powered GPU for display and light inference at the edge.

MTT S80

Consumer and workstation GPU with the first PCIe Gen5 interface in its class.

KUAE Cluster

Moore Threads’ domestically-produced AI training cluster, built from MTT S4000 cards.

Specifications

Vendor

Moore Threads (Beijing, China)

Product Family

MTT S4000 / S30 / S80 / X300

Architecture

3rd-generation MUSA GPU

Memory (S4000)

48 GB
768 GB/s

Bus Interface

PCIe 5.0 ×16

FP32

25 TFLOPS (S4000)
14.7 TFLOPS (S80)

TF32 Tensor

50 TFLOPS

FP16 / BF16

200 TFLOPS

INT8

200 TOPS

Max TGP

450 W (S4000)
255 W (S80) / 40 W (S30)

Chip Interconnect

MTLink
240 GB/s I/O bandwidth

Video Encode

H.265, H.264, AV1
48 × 1080p30

Video Decode

H.265, H.264, AV1, AVS2, VP9
96 × 1080p30

Display

4 × DisplayPort 1.4a

Security

MUSA Security Engine 2.0
TEE, encryption/decryption

Virtualization

Hardware virtualization
GPU elastic partitioning, SR-IOV

APIs

MUSA, DirectX, Vulkan, OpenGL, OpenGL ES

Cluster

KUAE Intelligent Computing Center
1,000 cards

Price

Quote

Overview

Moore Threads is China’s leading general-purpose GPU vendor, built on its in-house MUSA architecture and now on its third GPU generation. For AI buyers outside the mainstream NVIDIA stack, the relevant parts are the MTT S4000 datacenter accelerator and the KUAE training cluster it powers — a fully domestic, 1,000-card AI training system.

The MTT S4000 pairs 48 GB of high-bandwidth memory at 768 GB/s with a PCIe 5.0 ×16 host interface and rates 25 TFLOPS FP32, 50 TFLOPS TF32, 200 TFLOPS FP16/BF16 and 200 TOPS INT8 at a 450 W TGP. It supports multi-modal inference and AIGC serving alongside graphics workloads, carries hardware virtualization with SR-IOV isolation and GPU elastic partitioning, and exposes a MUSA Security Engine 2.0 with TEE support. Chip-to-chip communication runs over MTLink, and the card provides 240 GB/s of I/O bandwidth with 48-channel 1080p30 encode and 96-channel decode.

At the lower end, the MTT S30 is a 40 W, slot-powered single-slot half-height card with 1,024 shaders and 4 GB of GDDR6, suited to display and light edge inference, while the MTT S80 brings 4,096 MUSA cores, 14.7 TFLOPS FP32 and 16 GB of GDDR6 over PCIe Gen5. All parts are programmed through the MUSA software stack, with tooling to ease migration from CUDA. Availability and export terms are quoted on request.

Need Moore Threads MTT S4000?

Contact QS Compute for availability, configuration, and volume pricing.

Request Quote