Specifications
Accelerators per Rack
72x AMD Instinct MI455X
Total HBM4 Memory
~31 TB per rack
Aggregate Memory Bandwidth
~1.4 PB/s per rack
Inference Throughput
~2.9 FP4 exaFLOPS per rack
GPU Compute Units
18,000 per rack
Memory per Accelerator
432 GB HBM4 per MI455X
Bandwidth per Accelerator
19.6 TB/s per MI455X
Per-GPU 4-bit Peak
~40 PFLOPS theoretical 4-bit per MI455X
CPU
AMD EPYC 'Venice' (6th Gen EPYC) — 4,600 CPU cores per rack
Scale-up Interconnect
UALink at 260 TB/s — an open standard rather than a proprietary fabric
Scale-out Fabric
Ultra Ethernet
Networking Silicon
AMD Pensando DPUs
Rack Price Band
~$5.0M to $5.5M per rack as disclosed at AMD Advancing AI 2026
Production Status
In full production as of August 2026; first shipments late Q3 2026 ramping through Q4 2026 into H1 2027
Commercial Model
Reference design that OEM and ODM partners build and sell under their own brands
Named Adopters
Microsoft (Azure large-scale inference), OpenAI, Meta and Oracle
Versus NVIDIA GB200 NVL72
~31 TB vs ~13.4 TB memory, ~1.4 PB/s vs ~0.58 PB/s bandwidth, ~2.9 vs ~1.44 FP4 exaFLOPS
Announced
Detailed at AMD Advancing AI 2026 on July 23, 2026
Overview
AMD Helios is AMD's first complete rack-scale AI system and its direct answer to NVIDIA's NVL72 class. A single rack holds 72 Instinct MI455X accelerators with 432 GB of HBM4 each, giving roughly 31 TB of rack memory and ~1.4 PB/s of aggregate bandwidth against about 2.9 FP4 exaFLOPS of inference compute and 18,000 GPU compute units.
Where Helios differs architecturally is in its commitment to open interconnects. Scale-up runs over UALink at 260 TB/s instead of a proprietary GPU fabric, scale-out runs over Ultra Ethernet, and the accelerators are paired with EPYC Venice CPUs and Pensando DPUs. That open-fabric posture is aimed squarely at buyers who want rack-scale AI capacity without single-vendor lock-in on the interconnect.
AMD says Helios entered full production in August 2026, with first shipments in late Q3 2026 ramping through Q4 2026 into the first half of 2027. It is a reference design rather than a finished AMD-branded product: OEM and ODM partners build and sell the systems, with Microsoft, OpenAI, Meta and Oracle named as early adopters and a disclosed rack price band of roughly $5.0M-$5.5M. QS Compute supplies rack-scale AI systems, GPU servers and accelerator infrastructure — request a quote for your deployment.
Key Benefits
Density: 72 accelerators, ~31 TB HBM4 and ~1.4 PB/s per rack. Open fabric: UALink scale-up and Ultra Ethernet scale-out instead of proprietary interconnect. Memory headroom: 432 GB HBM4 per accelerator for trillion-parameter models. Multi-vendor supply: reference design available through OEM/ODM partners with a broad early-adopter roster.
Applications
Trillion-parameter model training and inference, large-scale enterprise and sovereign AI clusters, Azure-class cloud inference capacity, Mixture-of-Experts serving, and open-standard AI datacentre builds.
Request a Quote — AMD HELIOS RACK — 72X MI455X RACK-SCALE AI SYSTEM
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →