Specifications
GPUs
72× NVIDIA Rubin GPUs
36× NVIDIA Vera CPUs
Inference (NVFP4)
3.6 EFLOPS
Training (NVFP4)
2.5 EFLOPS
GPU Memory
20.7 TB HBM4
Memory Bandwidth
1.6 PB/s aggregate
Scale-Up
260 TB/s NVLink 6
Scale-Out
ConnectX-9 SuperNIC
1.6 Tb/s per GPU
DPU / Switch
BlueField-4 DPUs
Spectrum-6 Ethernet
Power
~190 kW (Max Q)
~230 kW (Max P)
Cooling
100% liquid · 45°C inlet water
Rack
~4,000 lbs · ~1,300 chips
1.3M components
Assembly
Full rack in ~2 hours
Confidential Computing
3rd-gen rack-scale
Variants
HGX Rubin NVL8
DGX SuperPOD
Availability
Full production · partners H2 2026
Overview
The NVIDIA Vera Rubin NVL72 is the flagship rack-scale system of the Rubin generation, packing 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled enclosure that exceeds 200 kW per rack. It delivers 3.6 exaFLOPS of NVFP4 inference and 2.5 exaFLOPS of NVFP4 training — a generational leap over the Blackwell-era GB200 NVL72.
At its core is sixth-generation NVLink, providing 260 TB/s of scale-up bandwidth per rack — more than the entire internet's bandwidth, per NVIDIA — with built-in in-network compute to accelerate collective operations. Scale-out is handled by ConnectX-9 SuperNICs (1.6 Tb/s per GPU) and BlueField-4 DPUs, with Spectrum-6 Ethernet switching for AI-optimized fabrics.
The system's 100% liquid cooling with 45°C inlet water enables free cooling without chillers in most climates, and its modular cable-free tray design allows a full rack to be assembled in roughly two hours. Rubin systems are available from Cisco, Dell, HPE, Lenovo, and Supermicro, with first cloud deployments on AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda, Nebius, and Nscale in the second half of 2026.
Key Benefits
<strong>Gigascale AI Performance:</strong> 3.6 exaFLOPS of FP4 inference in a single rack, purpose-built for trillion-parameter training and frontier reasoning models. <strong>NVLink 6 Scale-Up:</strong> 260 TB/s of GPU-to-GPU bandwidth with in-network compute for collective operations — the fastest GPU interconnect ever shipped. <strong>Efficient Liquid Cooling:</strong> 45°C inlet water supports free cooling in most climates, cutting total cost of ownership. <strong>Rack-Scale Confidential Computing:</strong> First platform to secure data across CPU, GPU, and NVLink domains simultaneously. <strong>Rapid Deployment:</strong> Modular cable-free trays enable ~2-hour rack assembly and ~5-minute compute tray swaps.
Request a Quote — NVIDIA Vera Rubin NVL72
QS Compute — global B2B supply for NVIDIA edge AI and data center hardware. Enterprise procurement, direct manufacturer channels.
Get Your Quote →