Specifications
Products
Biren BR100 (OAM module) and Biren BR104 (PCIe card), both general-purpose GPU accelerators
Transistor Count
77 billion (BR100 dual-chip module); 38.5 billion per BR104 monolithic die
Process Node
TSMC N7
FP32 Performance
256 TFLOPS peak (BR100); 128 TFLOPS peak (BR104)
TF32+ Performance
512 TFLOPS peak (BR100); 256 TFLOPS peak (BR104)
BF16 Performance
1024 TFLOPS peak (BR100); 512 TFLOPS peak (BR104)
INT8 Performance
2048 TOPS peak (BR100) — 2 INT8 PFLOPS; 1024 TOPS peak (BR104)
Memory — BR100
64GB HBM2E, 4096-bit interface, 1.64 TB/s bandwidth
Memory — BR104
32GB HBM2E, 2048-bit interface, 819 GB/s bandwidth
Form Factor
BR100: OAM module; BR104: FHFL dual-wide PCIe card
Board Power
550W (BR100); 300W (BR104)
Host Interface
PCIe 5.0 x16 with CXL support for accelerators
Interconnect
BLink proprietary link — 192 GB/s per BR104 supporting 3x8 ports; up to 8-way per system on BR100
Precision Support
INT8, INT16, INT32, FP16, BF16, TF32+ and FP32 (no FP64)
Software Stack
BIRENSUPA compiler, runtime and CUDA-compatibility layer with PyTorch and TensorFlow integration
Secure Instances
Up to 4 secure virtual instances per BR104
Reference Platforms
Biren offers reference cluster designs to system builders
Company Milestone
First mainland Chinese GPU company to complete a Hong Kong Stock Exchange listing
Roadmap
Successor BR-series generations in development for 2026-2027, targeting H100-class performance
Overview
The Biren BR100 and BR104 are China's most ambitious general-purpose GPU accelerators, designed for AI training, inference and HPC workloads in data centres that cannot or will not standardise on NVIDIA. Biren claims 2 INT8 PFLOPS and 1024 BF16 TFLOPS for the flagship BR100, a dual-chip module in OAM form factor that is positioned directly against NVIDIA's H100 class.
The BR104 is the same architecture in a single monolithic die, delivered as a 300W FHFL dual-wide PCIe card with 32GB of HBM2E at 819 GB/s — roughly half the compute and memory bandwidth of the BR100 but deployable in standard servers without OAM carriers. Both parts use PCIe 5.0 x16 with CXL support, and both rely on Biren's proprietary BLink interconnect for multi-GPU scaling: up to 3-way on BR104 and 8-way per system with the BR100.
The software story is the BIRENSUPA stack, which layers a CUDA-compatibility path on top of Biren's own compiler and runtime, with PyTorch and TensorFlow framework integration. The goal is to let teams move NVIDIA-targeted code across with minimal rewrite. Biren supplies reference cluster designs to system builders alongside the accelerators.
Biren became the first mainland Chinese GPU company to list on the Hong Kong Stock Exchange, and has stated that successor BR-series products targeting H100-class performance are in development for 2026-2027 production. QS Compute can quote BR100 modules and BR104 cards as individual parts or as part of configured GPU servers.
Key Benefits
A credible non-NVIDIA accelerator path with 2 INT8 PFLOPS per OAM module; 64GB HBM2E at 1.64 TB/s in a single module; CUDA-compatibility layer that reduces porting cost from NVIDIA codebases; both OAM and standard PCIe form factors from one architecture; secure multi-tenant virtualisation with up to four instances per card.
Applications
Large-scale AI model training clusters; LLM inference serving; HPC and scientific simulation excluding FP64-dependent workloads; sovereign and regulated data-centre AI deployment; secure multi-tenant inference.
Request a Quote — BIREN BR100 / BR104 — 2 PFLOPS GPGPU AI ACCELERATORS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →