Specifications

GPUs

72× NVIDIA Rubin GPUs
36× NVIDIA Vera CPUs

Inference (NVFP4)

3.6 EFLOPS

Training (NVFP4)

2.5 EFLOPS

GPU Memory

20.7 TB HBM4

Memory Bandwidth

1.6 PB/s aggregate

Scale-Up

260 TB/s NVLink 6

Scale-Out

ConnectX-9 SuperNIC
1.6 Tb/s per GPU

DPU / Switch

BlueField-4 DPUs
Spectrum-6 Ethernet

Power

~190 kW (Max Q)
~230 kW (Max P)

Cooling

100% liquid · 45°C inlet water

Rack

~4,000 lbs · ~1,300 chips
1.3M components

Assembly

Full rack in ~2 hours

Confidential Computing

3rd-gen rack-scale

Variants

HGX Rubin NVL8
DGX SuperPOD

Availability

Full production · partners H2 2026

Overview

The NVIDIA Vera Rubin NVL72 is the flagship rack-scale system of the Rubin generation, packing 72 Rubin GPUs and 36 Vera CPUs into a single liquid-cooled enclosure that exceeds 200 kW per rack. It delivers 3.6 exaFLOPS of NVFP4 inference and 2.5 exaFLOPS of NVFP4 training — a generational leap over the Blackwell-era GB200 NVL72.

At its core is sixth-generation NVLink, providing 260 TB/s of scale-up bandwidth per rack — more than the entire internet's bandwidth, per NVIDIA — with built-in in-network compute to accelerate collective operations. Scale-out is handled by ConnectX-9 SuperNICs (1.6 Tb/s per GPU) and BlueField-4 DPUs, with Spectrum-6 Ethernet switching for AI-optimized fabrics.

The system's 100% liquid cooling with 45°C inlet water enables free cooling without chillers in most climates, and its modular cable-free tray design allows a full rack to be assembled in roughly two hours. Rubin systems are available from Cisco, Dell, HPE, Lenovo, and Supermicro, with first cloud deployments on AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda, Nebius, and Nscale in the second half of 2026.

Key Benefits

<strong>Gigascale AI Performance:</strong> 3.6 exaFLOPS of FP4 inference in a single rack, purpose-built for trillion-parameter training and frontier reasoning models. <strong>NVLink 6 Scale-Up:</strong> 260 TB/s of GPU-to-GPU bandwidth with in-network compute for collective operations — the fastest GPU interconnect ever shipped. <strong>Efficient Liquid Cooling:</strong> 45°C inlet water supports free cooling in most climates, cutting total cost of ownership. <strong>Rack-Scale Confidential Computing:</strong> First platform to secure data across CPU, GPU, and NVLink domains simultaneously. <strong>Rapid Deployment:</strong> Modular cable-free trays enable ~2-hour rack assembly and ~5-minute compute tray swaps.

Request a Quote — NVIDIA Vera Rubin NVL72

QS Compute — global B2B supply for NVIDIA edge AI and data center hardware. Enterprise procurement, direct manufacturer channels.

Get Your Quote →