Specifications
Cores
72 Arm Neoverse V2 cores, each with 4x 128-bit SVE2 vector units
Core clock
Up to 3.44 GHz
L3 cache
114 MB shared L3 across the core complex
Coherency fabric
NVIDIA Scalable Coherency Fabric (SCF) delivering 3.2 TB/s of bisection bandwidth, double traditional CPU designs
Vector/SIMD
Arm SVE2 at 128 bits, 4 units per core for wide vector workloads
Memory type
High-bandwidth LPDDR5X, a server-first use of low-power memory
Memory capacity
Up to 480 GB per Grace CPU; 960 GB in the two-CPU Grace CPU Superchip configuration
Memory bandwidth
Roughly 500 GB/s per Grace CPU
GPU interconnect
900 GB/s NVLink-C2C coherent link to Hopper or Blackwell GPUs
Superchip option
Grace CPU Superchip pairs two Grace CPUs for 144 Neoverse V2 cores and 960 GB LPDDR5X
Platform integration
GH200 Grace Hopper and GB200/GB10 Grace Blackwell combinations use Grace as host CPU
Performance claim
Up to 2x the performance of leading x86 CPUs in the same power envelope
Process
TSMC 4N custom silicon
Memory efficiency
LPDDR5X delivers server-class bandwidth at markedly lower power than DDR5 RDIMM
Software
Full Arm64 server software support with NVIDIA CUDA, DOCA and Grace performance tuning guides
Deployment
Grace-based nodes for AI training, inference, HPC and large-scale data analytics
Overview
The NVIDIA Grace CPU is a purpose-built Arm server processor designed to host GPUs rather than simply feed them over PCIe. It carries 72 Arm Neoverse V2 cores running up to 3.44 GHz, each with four 128-bit SVE2 vector units, sharing 114 MB of L3 cache across NVIDIA's Scalable Coherency Fabric. That fabric supplies 3.2 TB/s of bisection bandwidth — double what conventional CPU designs achieve — which is what keeps a 72-core mesh fed.
The memory subsystem is the more radical choice. Grace uses LPDDR5X rather than DDR5 RDIMMs, yielding roughly 500 GB/s of bandwidth per CPU at far lower power, with up to 480 GB per processor and 960 GB in the dual-CPU Grace CPU Superchip configuration. NVIDIA rates the design at up to twice the performance of leading x86 CPUs in the same power envelope.
The part that matters most in AI systems is the interconnect. Grace attaches to Hopper or Blackwell GPUs over a 900 GB/s NVLink-C2C coherent link, so the CPU and GPU share a single memory space without PCIe round-trips. That combination appears as GH200 Grace Hopper and GB200/GB10 Grace Blackwell, and it is the reason Grace shows up as the host processor throughout NVIDIA's rack-scale AI platforms.
Key Benefits
72 Neoverse V2 cores with SVE2 and 3.2 TB/s coherency fabric give Grace x86-class host throughput at lower power, letting more of the rack budget go to accelerators. LPDDR5X memory delivers around 500 GB/s per CPU with markedly better energy efficiency than DDR5 RDIMM. 900 GB/s NVLink-C2C coherence removes the CPU-GPU copy bottleneck for data ingestion, KV cache management and memory-bound preprocessing, and the dual-CPU Superchip configuration doubles both cores and capacity.
Applications
AI training and inference head nodes in GH200 and GB200 systems; GPU-adjacent data preprocessing and tokenisation; KV cache and memory-tier management for inference serving; HPC clusters; large-scale data analytics; and Arm-native cloud infrastructure where performance per watt governs rack density.
Request a Quote — NVIDIA GRACE CPU — 72-CORE ARM NEOVERSE V2 DATA CENTER PROCESSOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →