Specifications

Configuration

8x NVIDIA Rubin SXM GPUs on a single NVLink baseboard

Optional Variant

HGX Vera Rubin NVL8 — 8x Rubin SXM with a single-socket Vera CPU

GPU Memory

288 GB HBM4 per GPU

Memory Bandwidth

22 TB/s per GPU

NVLink

Sixth-generation NVIDIA NVLink, 3.6 TB/s per GPU

NVFP4 Inference

50 PFLOPS per GPU; up to 400 PFLOPS per 8-GPU system

NVFP4 Training

35 PFLOPS per GPU; up to 270 PFLOPS per system

FP8 / FP6 Training

17.5 PFLOPS per GPU; 130 PFLOPS per system

INT8

250 TOPS per GPU

FP16 / BF16

4 PFLOPS per GPU; 32 PFLOPS per system

TF32

2 PFLOPS per GPU; 16 POPS per system

FP32

130 TFLOPS per GPU; 1,040 TFLOPS per system

FP64

33 TFLOPS per GPU; 265 TFLOPS per system

System CPU (Vera variant)

88 custom NVIDIA Olympus cores, up to 1.5 TB LPDDR5X

Memory Subsystem

HBM4 built on the 2,048-bit interface generation

Target Workload

Mixture-of-experts pretraining and long-context agentic AI

Platform

OEM 8-GPU servers from NVIDIA HGX partners

Generation

Successor to HGX B200 / HGX B300 Blackwell platforms

Overview

NVIDIA HGX Rubin NVL8 is the eighth-generation Rubin baseboard for OEM 8-GPU servers. Each of the eight Rubin SXM GPUs carries 288 GB of HBM4 with 22 TB/s of memory bandwidth, connected by sixth-generation NVLink running at 3.6 TB/s per GPU.

The platform is explicitly aimed at mixture-of-experts pretraining and long-context agentic inference: NVIDIA positions Rubin NVL8 as training next-generation agentic models with roughly four times fewer GPUs than the Blackwell generation, driven by about 4x more NVFP4 training FLOPS, 1.4x more high-speed HBM capacity and 1.7x more NVLink bandwidth versus HGX B200.

A Vera CPU variant pairs the same eight-GPU baseboard with a single-socket NVIDIA Vera CPU using 88 custom Olympus cores and up to 1.5 TB of LPDDR5X, giving OEMs an all-NVIDIA compute node for rack-scale deployments. It supersedes the HGX B200/B300 8-GPU boards used across current AI clusters.

Key Benefits

288 GB HBM4 per GPU at 22 TB/s. 3.6 TB/s NVLink per GPU with full all-to-all connectivity. Up to 400 PFLOPS NVFP4 inference per node. Vera CPU variant for an all-NVIDIA node design.

Applications

Frontier-scale LLM and MoE pretraining, long-context agentic AI inference, HPC simulation, sovereign AI clusters, and rack-scale AI factory deployments from OEM HGX partners.

Request a Quote — NVIDIA HGX RUBIN NVL8 — 8-GPU RUBIN PLATFORM, 288 GB HBM4 PER GPU

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Vera Rubin NVL72 Rack NVIDIA B300 NVL8 SuperPOD NVIDIA DGX B300