Specifications
Configuration
8x NVIDIA Rubin SXM GPUs on a single NVLink baseboard
Optional Variant
HGX Vera Rubin NVL8 — 8x Rubin SXM with a single-socket Vera CPU
GPU Memory
288 GB HBM4 per GPU
Memory Bandwidth
22 TB/s per GPU
NVLink
Sixth-generation NVIDIA NVLink, 3.6 TB/s per GPU
NVFP4 Inference
50 PFLOPS per GPU; up to 400 PFLOPS per 8-GPU system
NVFP4 Training
35 PFLOPS per GPU; up to 270 PFLOPS per system
FP8 / FP6 Training
17.5 PFLOPS per GPU; 130 PFLOPS per system
INT8
250 TOPS per GPU
FP16 / BF16
4 PFLOPS per GPU; 32 PFLOPS per system
TF32
2 PFLOPS per GPU; 16 POPS per system
FP32
130 TFLOPS per GPU; 1,040 TFLOPS per system
FP64
33 TFLOPS per GPU; 265 TFLOPS per system
System CPU (Vera variant)
88 custom NVIDIA Olympus cores, up to 1.5 TB LPDDR5X
Memory Subsystem
HBM4 built on the 2,048-bit interface generation
Target Workload
Mixture-of-experts pretraining and long-context agentic AI
Platform
OEM 8-GPU servers from NVIDIA HGX partners
Generation
Successor to HGX B200 / HGX B300 Blackwell platforms
Overview
NVIDIA HGX Rubin NVL8 is the eighth-generation Rubin baseboard for OEM 8-GPU servers. Each of the eight Rubin SXM GPUs carries 288 GB of HBM4 with 22 TB/s of memory bandwidth, connected by sixth-generation NVLink running at 3.6 TB/s per GPU.
The platform is explicitly aimed at mixture-of-experts pretraining and long-context agentic inference: NVIDIA positions Rubin NVL8 as training next-generation agentic models with roughly four times fewer GPUs than the Blackwell generation, driven by about 4x more NVFP4 training FLOPS, 1.4x more high-speed HBM capacity and 1.7x more NVLink bandwidth versus HGX B200.
A Vera CPU variant pairs the same eight-GPU baseboard with a single-socket NVIDIA Vera CPU using 88 custom Olympus cores and up to 1.5 TB of LPDDR5X, giving OEMs an all-NVIDIA compute node for rack-scale deployments. It supersedes the HGX B200/B300 8-GPU boards used across current AI clusters.
Key Benefits
288 GB HBM4 per GPU at 22 TB/s. 3.6 TB/s NVLink per GPU with full all-to-all connectivity. Up to 400 PFLOPS NVFP4 inference per node. Vera CPU variant for an all-NVIDIA node design.
Applications
Frontier-scale LLM and MoE pretraining, long-context agentic AI inference, HPC simulation, sovereign AI clusters, and rack-scale AI factory deployments from OEM HGX partners.
Request a Quote — NVIDIA HGX RUBIN NVL8 — 8-GPU RUBIN PLATFORM, 288 GB HBM4 PER GPU
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →