Specifications

Compute

Up to 2.517 PFLOPs of MXFP8 compute per package

Memory

144 GB of HBM3e per chip

Memory Bandwidth

4.9 TB/s (1.7x Trainium2)

Memory Capacity Uplift

1.5x Trainium2

Supported Data Types

MXFP8 and MXFP4

Process Node

AWS's first 3 nm AI chip

Generation

Fourth-generation AWS AI accelerator platform

Server SKUs

Trn3 NL32x2 Switched and Trn3 NL72x2 Switched; the flagship configuration is the Trn3 Gen2 UltraServer

Perf vs Trn2 UltraServer

Up to 4.4x higher performance

Bandwidth vs Trn2

3.9x higher memory bandwidth

Efficiency vs Trn2

Up to 4x better performance per watt

Amazon Bedrock

AWS's fastest accelerator on Bedrock — up to 3x faster than Trainium2 with over 5x higher output tokens per megawatt

Availability

General availability December 2025; production silicon shipping to anchor customers from Q1 2026

Named Customers

Anthropic (Project Rainier, scaling toward a million Trainium chips) and Uber

Workload Fit

Dense and expert-parallel training, Mixture-of-Experts, reinforcement learning, reasoning models and long-context architectures

Successor

Trainium4 (Trn4) announced

Overview

Trainium3 is AWS's fourth-generation AI accelerator and its first built on a 3 nm process. Each package delivers up to 2.517 PFLOPs of MXFP8 compute paired with 144 GB of HBM3e running at 4.9 TB/s — 1.5x the memory capacity and 1.7x the bandwidth of Trainium2. Support for MXFP4 alongside MXFP8 lets it trade precision for throughput on the long-context and reasoning-heavy workloads that dominate current inference demand.

The chip ships into two rack topologies, Trn3 NL32x2 Switched and Trn3 NL72x2 Switched, with the Trn3 Gen2 UltraServer as the flagship configuration. Against a Trn2 UltraServer, AWS quotes up to 4.4x higher performance, 3.9x higher memory bandwidth and 4x better performance per watt. On Amazon Bedrock it is described as the fastest accelerator available, up to 3x faster than Trainium2 with more than 5x the output tokens per megawatt.

Trainium3 is purpose-built for training and serving frontier-scale models: dense and expert-parallel training, Mixture-of-Experts, reinforcement learning and long-context architectures. It reached general availability in December 2025 with production silicon to anchor customers from Q1 2026, and named deployments include Anthropic's Project Rainier and Uber. QS Compute supplies AI training accelerators, GPU servers and rack-scale AI infrastructure — get in touch for a quote.

Key Benefits

Training throughput: 2.517 PFLOPs MXFP8 per package. Large fast memory: 144 GB HBM3e at 4.9 TB/s. Rack-scale options: NL32x2 and NL72x2 Switched topologies plus Gen2 UltraServers. Token economics: 5x+ output tokens per megawatt versus Trainium2.

Applications

Frontier model pre-training and post-training, Mixture-of-Experts training at scale, reinforcement learning, long-context and reasoning model serving, and cost-optimised generative AI inference on Amazon Bedrock.

Request a Quote — AWS TRAINIUM3 — 2.5 PFLOP MXFP8 TRAINING ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Rubin GPU — Next-Gen AI Accelerator Cerebras WSE-3 Turbo NVIDIA B300 Blackwell Ultra