Specifications
Core Configurations
8 to 128 Arm Neoverse N4 cores per die
Maximum Core Count
128 Neoverse N4 cores on a single die
Maximum Frequency
Up to 3.8 GHz
Process Node
TSMC N3P
Private L2 Cache
Up to 2 MB per core
L1 Cache
64 KB L1 instruction + 64 KB L1 data per core
Shared System Cache
Up to 256 MB system-level cache per die
Memory Support
DDR5 or LPDDR6 — the first Neoverse CSS design with LPDDR6
PCIe I/O
Up to 128 lanes of PCIe Gen 7 / Gen 6
Coherent Interconnect
CXL 4.0 support
Compute vs CSS N3
Up to 2X CSS N3 performance
Efficiency
Up to 1.25X performance per watt vs CSS N3
Memory Bandwidth
Up to 1.75X CSS N3
Target Designs
Custom server CPUs, DPUs, networking silicon and specialised AI infrastructure chips
Prior CSS Deployments
Microsoft Cobalt 100, Google Axion, Alibaba Yitian and multiple DPU products on earlier Neoverse generations
Announced
September 8, 2026
Production Status
Pre-integrated subsystem offered for customer tape-out; no production N4-based chip named at launch
Overview
The Arm Neoverse CSS N4 is Arm's most configurable compute subsystem to date, giving chip designers a pre-integrated platform for custom server CPUs, data processing units (DPUs) and AI infrastructure silicon. Customers configure core count, cache, memory type and I/O, then tape out — removing much of the integration risk that would otherwise sit between an architecture licence and production silicon.
CSS N4 scales from 8 to 128 Neoverse N4 cores on a single die at up to 3.8 GHz, fabricated on TSMC's N3P node. Each core carries up to 2 MB of private L2 alongside 64 KB L1 instruction and data caches, backed by up to 256 MB of shared system-level cache per die. It is the first Neoverse CSS generation to support LPDDR6 memory alongside DDR5, and the first to offer PCIe Gen 7 connectivity with CXL 4.0.
Relative to the previous CSS N3 generation, Arm quotes up to 2X the performance, up to 1.25X performance per watt and up to 1.75X memory bandwidth. The platform is designed for hyperscale and sovereign AI infrastructure where the CPU has to orchestrate accelerators, feed long-context inference pipelines and host agentic control planes. QS Compute supplies Neoverse-based server platforms, DPU-equipped systems and rack-scale AI infrastructure — request a quote for your configuration.
Key Benefits
Configuration freedom: scale from 8 to 128 N4 cores per die without re-architecting. Memory headroom: first Neoverse CSS with LPDDR6, plus DDR5 options. Faster I/O: 128 lanes of PCIe Gen 7 with CXL 4.0. Rack density: up to 2X CSS N3 performance at 1.25X performance per watt, so more compute per rack watt.
Applications
Custom cloud and hyperscale server CPUs, DPUs and SmartNICs, AI infrastructure control plane processors, sovereign and national AI clusters, networking and storage offload silicon, accelerator orchestration for agentic AI inference.
Request a Quote — ARM NEOVERSE CSS N4 — 128-CORE COMPUTE SUBSYSTEM FOR CUSTOM AI SILICON
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →