Specifications

Accelerators per Rack

72x AMD Instinct MI455X

Total HBM4 Memory

~31 TB per rack

Aggregate Memory Bandwidth

~1.4 PB/s per rack

Inference Throughput

~2.9 FP4 exaFLOPS per rack

GPU Compute Units

18,000 per rack

Memory per Accelerator

432 GB HBM4 per MI455X

Bandwidth per Accelerator

19.6 TB/s per MI455X

Per-GPU 4-bit Peak

~40 PFLOPS theoretical 4-bit per MI455X

CPU

AMD EPYC 'Venice' (6th Gen EPYC) — 4,600 CPU cores per rack

Scale-up Interconnect

UALink at 260 TB/s — an open standard rather than a proprietary fabric

Scale-out Fabric

Ultra Ethernet

Networking Silicon

AMD Pensando DPUs

Rack Price Band

~$5.0M to $5.5M per rack as disclosed at AMD Advancing AI 2026

Production Status

In full production as of August 2026; first shipments late Q3 2026 ramping through Q4 2026 into H1 2027

Commercial Model

Reference design that OEM and ODM partners build and sell under their own brands

Named Adopters

Microsoft (Azure large-scale inference), OpenAI, Meta and Oracle

Versus NVIDIA GB200 NVL72

~31 TB vs ~13.4 TB memory, ~1.4 PB/s vs ~0.58 PB/s bandwidth, ~2.9 vs ~1.44 FP4 exaFLOPS

Announced

Detailed at AMD Advancing AI 2026 on July 23, 2026

Overview

AMD Helios is AMD's first complete rack-scale AI system and its direct answer to NVIDIA's NVL72 class. A single rack holds 72 Instinct MI455X accelerators with 432 GB of HBM4 each, giving roughly 31 TB of rack memory and ~1.4 PB/s of aggregate bandwidth against about 2.9 FP4 exaFLOPS of inference compute and 18,000 GPU compute units.

Where Helios differs architecturally is in its commitment to open interconnects. Scale-up runs over UALink at 260 TB/s instead of a proprietary GPU fabric, scale-out runs over Ultra Ethernet, and the accelerators are paired with EPYC Venice CPUs and Pensando DPUs. That open-fabric posture is aimed squarely at buyers who want rack-scale AI capacity without single-vendor lock-in on the interconnect.

AMD says Helios entered full production in August 2026, with first shipments in late Q3 2026 ramping through Q4 2026 into the first half of 2027. It is a reference design rather than a finished AMD-branded product: OEM and ODM partners build and sell the systems, with Microsoft, OpenAI, Meta and Oracle named as early adopters and a disclosed rack price band of roughly $5.0M-$5.5M. QS Compute supplies rack-scale AI systems, GPU servers and accelerator infrastructure — request a quote for your deployment.

Key Benefits

Density: 72 accelerators, ~31 TB HBM4 and ~1.4 PB/s per rack. Open fabric: UALink scale-up and Ultra Ethernet scale-out instead of proprietary interconnect. Memory headroom: 432 GB HBM4 per accelerator for trillion-parameter models. Multi-vendor supply: reference design available through OEM/ODM partners with a broad early-adopter roster.

Applications

Trillion-parameter model training and inference, large-scale enterprise and sovereign AI clusters, Azure-class cloud inference capacity, Mixture-of-Experts serving, and open-standard AI datacentre builds.

Request a Quote — AMD HELIOS RACK — 72X MI455X RACK-SCALE AI SYSTEM

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Vera Rubin NVL72 NVIDIA DGX B300 Cerebras CS-4