MPN: CS-4 (Nexus Platform)
Cerebras CS-4
Rack-Scale AI System · 750 PFLOPS

3× WSE-3 Turbo · 132 GB SRAM · 160.5 PB/s fabric · 2µs I/O latency — Cerebras's first true rack-scale system

Product Lineup

CS-4 (Nexus)

First true rack-scale Cerebras system — three WSE-3 Turbo wafers in modular 'backpacks'.

CS-3 (2024)

Single-WSE system — smallest unit of the WSE ecosystem, 16U liquid-cooled.

Specifications

Vendor

Cerebras Systems (Sunnyvale, CA)

Product

CS-4 rack-scale system

Platform

Nexus (modular, multi-generation)

Wafer Engines

3× WSE-3 Turbo per rack

Compute

750 PFLOPS sparse FP16

On-Wafer SRAM

132 GB (3× 44 GB)

Memory Bandwidth

129.6 PB/s

Fabric Bandwidth

160.5 PB/s

I/O Bandwidth

900 GB/s

I/O Latency

2 µs

Networking

RoCEv2 scale-up/out; optional chained topology

Cooling

Liquid-cooled (backpack modules)

Predecessor

CS-3 (16U, single WSE-3, 125 PFLOPS)

Roadmap

CS-5 & CS-6 on Nexus platform

Announced

Aug 2026

Stock

Available

Price

Quote

Overview

The Cerebras CS-4 is the company's first true rack-scale system, replacing the CS-3 (a single-wafer, 16U liquid-cooled unit) with a modular architecture that houses three WSE-3 Turbo wafers in a single rack — delivering 750 PFLOPS of sparse FP16, 132 GB of on-wafer SRAM, and 160.5 PB/s of aggregate fabric bandwidth.

At the heart of the redesign is the Nexus platform, a highly modular, multi-generation rack that places power supplies, fans, and supporting hardware at the front and the wafers at the rear. Each WSE mounts in a self-contained, vertically-standing backpack that connects it to the rack's power and liquid-cooling loops plus the wafer I/O modules housing the networking — decoupling networking hardware from the wafers so it can be upgraded independently. Cerebras has already committed Nexus as the basis for the upcoming CS-5 and CS-6 systems.

For scale-up/scale-out networking, the initial CS-4 uses RDMA over Converged Ethernet v2 (RoCEv2), with UALink or Ultra Ethernet as likely future options. An optional chained topology links backpacks directly without switches to minimize wafer-to-wafer latency to 2 microseconds. Compared to the CS-3, the CS-4 delivers roughly six times the per-system performance — three times the wafers, each running twice as fast.

Request a Quote — Cerebras CS-4 Rack-Scale System

QS Compute — global B2B supply of AI computing hardware. Price on request, stock availability confirmed on inquiry.

Request Quote