Specifications
Product Line
Enflame CloudBlazer accelerator cards: S60, T20, T21 and L600
S60 Generation
Third-generation AI inference acceleration card for large-scale data-centre deployment
S60 Silicon
SCORPIO-AO compute chiplet, manufactured by TSMC on the N6NTO-HPC node
S60 Workloads
Large language models, search and recommendation, image and text generation, speech recognition
T20
Second-generation CloudBlazer training card (PCIe)
T21
Second-generation CloudBlazer training card (OAM form factor)
i20
Second-generation CloudBlazer inference card
L600
New-generation training and inference integrated AI chip, launched at WAIC 2025
L600 Positioning
Company states performance surpasses the NVIDIA H20 for its target workloads
DTU Architecture
32 AI compute cores in 4 scalable intelligent clusters, 40 data transfer engines, 4 high-speed interconnects
On-chip Memory
2x HBM2 providing 512 GB/s bandwidth (DTU generation)
GCU-CARE Core
20 TFLOPS FP32, 80 TFLOPS FP16 / BF16, 1024-bit bus, fully programmable VLIW
Per-core Throughput
FP32 256 MAC (20 TFLOPS); FP16/BF16/INT16/INT8 1024 MAC (80 TOPS per core)
Card Interconnect
GCU-LARE — 200 GB/s bi-directional per card, 800 GB/s within a single server, sub-1 microsecond latency
Cluster Topology
8 or 16 cards per node, 64 cards per rack, 2D torus topology, no RDMA required inside the rack
Host Interface
PCIe Gen 4.0 with CXL support on DTU generation parts
Board Power
225W (PCIe T10-class) and 300W (OAM T11-class) training cards
Precision Support
FP32, FP16, BF16, INT8, INT16, INT32 and full mixed-precision
Software Frameworks
PyTorch, PaddlePaddle, TensorFlow and ONNX with the TopsPlatform runtime
Verified Models
PaddleNLP llama2-13B adapted and optimised on S60 with a GCU inference interface matching GPUs
Migration Path
Source-level device change only — no model rewrite required for supported frameworks
Overview
Enflame Technology's CloudBlazer line is China's most commercially established domestic AI accelerator family, now in its third generation. The S60 is the current inference workhorse: a card built on the SCORPIO-AO compute chiplet fabricated by TSMC on its N6NTO-HPC process, targeted squarely at large-scale deployment of large language models, search and recommendation systems, and generative image, text and speech workloads.
The architecture descends from Enflame's DTU design, which packs 32 scalable intelligent processors into four clusters with dedicated data transfer engines and 2.5D-packaged HBM2. Each GCU-CARE core delivers 20 FP32 TFLOPS and 80 INT8 TOPS, and the GCU-LARE interconnect moves 200 GB/s bi-directionally per card with sub-microsecond latency, allowing 8 to 16 cards per node and 64 cards per rack over a 2D torus without RDMA.
Training coverage spans the second-generation CloudBlazer T20 (PCIe) and T21 (OAM) cards at 225W and 300W respectively, while the L600 — launched at WAIC 2025 — is a new-generation part that integrates both training and inference, with Enflame claiming performance beyond the NVIDIA H20 for its target segment.
The practical barrier for accelerator adoption is software, and Enflame's answer is a GCU interface that mirrors the GPU programming model: PaddleNLP has adapted llama2-13B to run on S60 with only device changes, and PyTorch, PaddlePaddle, TensorFlow and ONNX are supported through the TopsPlatform runtime. QS Compute can source CloudBlazer cards and Enflame-based 8-card servers and racks.
Key Benefits
Third-generation inference silicon on a leading 6nm-class node; strong LLM and recommendation throughput per card; a CUDA-like programming model with documented single-change migration from GPU code; 64-card rack scaling over a lossless 2D torus without RDMA; a domestic-supply-chain option for regulated deployments.
Applications
LLM inference serving and search/recsys acceleration at scale; generative image, text and speech inference; AI training clusters; sovereign and regulated AI infrastructure; mixed training and inference nodes on L600.
Request a Quote — ENFLAME CLOUDBLAZER SERIES — S60 / T20 / T21 / L600 AI ACCELERATOR CARDS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →