Specifications
Compute
Up to 2.517 PFLOPs of MXFP8 compute per package
Memory
144 GB of HBM3e per chip
Memory Bandwidth
4.9 TB/s (1.7x Trainium2)
Memory Capacity Uplift
1.5x Trainium2
Supported Data Types
MXFP8 and MXFP4
Process Node
AWS's first 3 nm AI chip
Generation
Fourth-generation AWS AI accelerator platform
Server SKUs
Trn3 NL32x2 Switched and Trn3 NL72x2 Switched; the flagship configuration is the Trn3 Gen2 UltraServer
Perf vs Trn2 UltraServer
Up to 4.4x higher performance
Bandwidth vs Trn2
3.9x higher memory bandwidth
Efficiency vs Trn2
Up to 4x better performance per watt
Amazon Bedrock
AWS's fastest accelerator on Bedrock — up to 3x faster than Trainium2 with over 5x higher output tokens per megawatt
Availability
General availability December 2025; production silicon shipping to anchor customers from Q1 2026
Named Customers
Anthropic (Project Rainier, scaling toward a million Trainium chips) and Uber
Workload Fit
Dense and expert-parallel training, Mixture-of-Experts, reinforcement learning, reasoning models and long-context architectures
Successor
Trainium4 (Trn4) announced
Overview
Trainium3 is AWS's fourth-generation AI accelerator and its first built on a 3 nm process. Each package delivers up to 2.517 PFLOPs of MXFP8 compute paired with 144 GB of HBM3e running at 4.9 TB/s — 1.5x the memory capacity and 1.7x the bandwidth of Trainium2. Support for MXFP4 alongside MXFP8 lets it trade precision for throughput on the long-context and reasoning-heavy workloads that dominate current inference demand.
The chip ships into two rack topologies, Trn3 NL32x2 Switched and Trn3 NL72x2 Switched, with the Trn3 Gen2 UltraServer as the flagship configuration. Against a Trn2 UltraServer, AWS quotes up to 4.4x higher performance, 3.9x higher memory bandwidth and 4x better performance per watt. On Amazon Bedrock it is described as the fastest accelerator available, up to 3x faster than Trainium2 with more than 5x the output tokens per megawatt.
Trainium3 is purpose-built for training and serving frontier-scale models: dense and expert-parallel training, Mixture-of-Experts, reinforcement learning and long-context architectures. It reached general availability in December 2025 with production silicon to anchor customers from Q1 2026, and named deployments include Anthropic's Project Rainier and Uber. QS Compute supplies AI training accelerators, GPU servers and rack-scale AI infrastructure — get in touch for a quote.
Key Benefits
Training throughput: 2.517 PFLOPs MXFP8 per package. Large fast memory: 144 GB HBM3e at 4.9 TB/s. Rack-scale options: NL32x2 and NL72x2 Switched topologies plus Gen2 UltraServers. Token economics: 5x+ output tokens per megawatt versus Trainium2.
Applications
Frontier model pre-training and post-training, Mixture-of-Experts training at scale, reinforcement learning, long-context and reasoning model serving, and cost-optimised generative AI inference on Amazon Bedrock.
Request a Quote — AWS TRAINIUM3 — 2.5 PFLOP MXFP8 TRAINING ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →