Tensor Contraction Processor · TSMC 5 nm · 256 MB SRAM · 1.5 TB/s HBM3 · PCIe Gen5 x16 · 12VHPWR · SR-IOV
RNGD (Renegade)
FuriosaAI
Tensor Contraction Processor (TCP)
TSMC 5 nm
1.0 GHz
256 TFLOPS
512 TFLOPS
512 TOPS
1,024 TOPS
48 GB HBM3
1.5 TB/s
16
256 MB
PCIe Gen5 x16
180 W
12VHPWR
PCIe dual-slot full-height 3/4 length
8 instances
SR-IOV
Passive
The FuriosaAI RNGD (pronounced "renegade") is a data center inference accelerator built on Furiosa's Tensor Contraction Processor architecture — a coarse-grained reconfigurable dataflow design in which the compiler maps each tensor operation onto a network of eight processing elements rather than relying on a fixed-pipeline tensor core. The result is 512 TFLOPS FP8 / 512 TOPS INT8 (and 1,024 TOPS INT4) inside a 180 W envelope.
The memory configuration is the headline: 48 GB of HBM3 with 16 channels delivering 1.5 TB/s, backed by 256 MB of on-chip SRAM. That combination lets a single card hold and serve large language models that would otherwise require two or more conventional GPUs, without splitting the model across devices. RNGD supports up to eight multi-instances and SR-IOV, so one physical card can be partitioned between tenants or workloads.
The card is a passive-cooled PCIe dual-slot full-height 3/4-length part on a 12VHPWR connector, targeting standard server airflow. FuriosaAI ships the RNGD software stack (RBLN SDK, compiler and runtime) with ONNX and PyTorch model import paths. Its 180 W TDP is roughly an eighth of a flagship training GPU, which is the core of the pitch: inference tokens per watt rather than peak training throughput. QS Compute can source RNGD cards and RNGD-based servers for APAC buyers — contact us for pricing and lead time.
Need FuriosaAI RNGD?
Enterprise hardware — configured, tested, deployed. Contact us for pricing and availability.
Request Quote