Specifications

Cores

72 Arm Neoverse V2 cores, each with 4x 128-bit SVE2 vector units

Core clock

Up to 3.44 GHz

L3 cache

114 MB shared L3 across the core complex

Coherency fabric

NVIDIA Scalable Coherency Fabric (SCF) delivering 3.2 TB/s of bisection bandwidth, double traditional CPU designs

Vector/SIMD

Arm SVE2 at 128 bits, 4 units per core for wide vector workloads

Memory type

High-bandwidth LPDDR5X, a server-first use of low-power memory

Memory capacity

Up to 480 GB per Grace CPU; 960 GB in the two-CPU Grace CPU Superchip configuration

Memory bandwidth

Roughly 500 GB/s per Grace CPU

GPU interconnect

900 GB/s NVLink-C2C coherent link to Hopper or Blackwell GPUs

Superchip option

Grace CPU Superchip pairs two Grace CPUs for 144 Neoverse V2 cores and 960 GB LPDDR5X

Platform integration

GH200 Grace Hopper and GB200/GB10 Grace Blackwell combinations use Grace as host CPU

Performance claim

Up to 2x the performance of leading x86 CPUs in the same power envelope

Process

TSMC 4N custom silicon

Memory efficiency

LPDDR5X delivers server-class bandwidth at markedly lower power than DDR5 RDIMM

Software

Full Arm64 server software support with NVIDIA CUDA, DOCA and Grace performance tuning guides

Deployment

Grace-based nodes for AI training, inference, HPC and large-scale data analytics

Overview

The NVIDIA Grace CPU is a purpose-built Arm server processor designed to host GPUs rather than simply feed them over PCIe. It carries 72 Arm Neoverse V2 cores running up to 3.44 GHz, each with four 128-bit SVE2 vector units, sharing 114 MB of L3 cache across NVIDIA's Scalable Coherency Fabric. That fabric supplies 3.2 TB/s of bisection bandwidth — double what conventional CPU designs achieve — which is what keeps a 72-core mesh fed.

The memory subsystem is the more radical choice. Grace uses LPDDR5X rather than DDR5 RDIMMs, yielding roughly 500 GB/s of bandwidth per CPU at far lower power, with up to 480 GB per processor and 960 GB in the dual-CPU Grace CPU Superchip configuration. NVIDIA rates the design at up to twice the performance of leading x86 CPUs in the same power envelope.

The part that matters most in AI systems is the interconnect. Grace attaches to Hopper or Blackwell GPUs over a 900 GB/s NVLink-C2C coherent link, so the CPU and GPU share a single memory space without PCIe round-trips. That combination appears as GH200 Grace Hopper and GB200/GB10 Grace Blackwell, and it is the reason Grace shows up as the host processor throughout NVIDIA's rack-scale AI platforms.

Key Benefits

72 Neoverse V2 cores with SVE2 and 3.2 TB/s coherency fabric give Grace x86-class host throughput at lower power, letting more of the rack budget go to accelerators. LPDDR5X memory delivers around 500 GB/s per CPU with markedly better energy efficiency than DDR5 RDIMM. 900 GB/s NVLink-C2C coherence removes the CPU-GPU copy bottleneck for data ingestion, KV cache management and memory-bound preprocessing, and the dual-CPU Superchip configuration doubles both cores and capacity.

Applications

AI training and inference head nodes in GH200 and GB200 systems; GPU-adjacent data preprocessing and tokenisation; KV cache and memory-tier management for inference serving; HPC clusters; large-scale data analytics; and Arm-native cloud infrastructure where performance per watt governs rack density.

Request a Quote — NVIDIA GRACE CPU — 72-CORE ARM NEOVERSE V2 DATA CENTER PROCESSOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Ampere Altra Max M128-30 AmpereOne A192-32X AMD EPYC 9006 Venice Intel Xeon 6952P Intel Xeon 6760P