2.5D CoWoS Chiplet · 102 B Transistors · 1,040 PCUs + 1,040 PMUs · Three-Tier Memory (SRAM / HBM / DDR) · Dataflow Architecture
SN40L
SambaNova Systems
Reconfigurable Dataflow Unit (RDU), Cerulean
TSMC 5 nm
2.5D CoWoS, chiplet-based
102 Billion
2 × SN40L RDD per socket
1,040 PCUs + 1,040 PMUs
638 TFLOPS per socket
520 MB
64 GB HBM
1.5 TB (3rd tier)
Three-tier: SRAM / HBM / DDR
8-RDU node (SambaRack)
Trillion-parameter training + inference
The SambaNova SN40L is a Reconfigurable Dataflow Unit — not a GPU. Fabricated on TSMC 5 nm and assembled as a 2.5D chip-on-wafer-on-substrate package containing two SN40L Reconfigurable Dataflow Dies plus HBM, each RDU socket delivers 638 BF16 TFLOPS from 1,040 Pattern Compute Units paired with 1,040 Pattern Memory Units, across 102 billion transistors.
The architectural bet is memory, not math. The SN40L carries a three-tier memory system: 520 MB of on-chip SRAM for weights and activations that fit, 64 GB of HBM for the working set, and up to 1.5 TB of high-capacity DDR as a third tier that the compiler addresses transparently. For models that are memory-bandwidth-bound — which is most large language model inference — a dataflow engine with deep local memory avoids the cost of repeatedly moving tensors between a small cache and off-package DRAM.
SambaNova deploys the SN40L in 8-RDU nodes inside the SambaRack, and the platform has been benchmarked publicly against NVIDIA H200 and B200 on LLM inference throughput per dollar. QS Compute can source SN40L nodes and full SambaRack configurations for enterprise and research buyers in APAC — contact us for configuration and pricing.
Need SambaNova SN40L RDU?
Enterprise hardware — configured, tested, deployed. Contact us for pricing and availability.
Request Quote