Specifications
Family Structure
Two chips: TPU 8t for training, TPU 8i for inference
TPU 8t HBM
216 GB of HBM per chip
TPU 8t Bandwidth
6,500 GB/s of HBM bandwidth
TPU 8t SRAM
128 MB of on-chip SRAM
TPU 8t Compute
12.6 PFLOPS peak FP4 per chip
TPU 8t vs Ironwood
3x the processing of TPU v7 Ironwood at 2x performance per watt
TPU 8t Superpod
9,600 chips with 2 PB of shared HBM memory
TPU 8t Superpod Compute
121 EFLOPS of FP4 compute
TPU 8i Memory
288 GB per chip
TPU 8i Compute
About 10,100 FP4 TFLOPS per chip
TPU 8i Interconnect
Boardfly interconnect — roughly 50% lower latency on communication-intensive workloads
TPU 8i Price/Performance
Up to 80% better performance per dollar than Ironwood at low-latency targets
Training Price/Performance
Up to 2.7x performance per dollar vs Ironwood for large-scale training
Energy Efficiency
Up to 2x better performance per watt on both chips
Scale Ceiling
Designed to scale beyond 1 million chips
Networking
Google Virgo interconnect with speedier Lustre storage tiers
Availability
Announced April 22, 2026 at Google Cloud Next; general availability stated for later in 2026
Overview
Google's eighth-generation TPU is the first to split training and inference into two distinct architectures. Pre-training, post-training and real-time serving have diverged enough in operational intensity that a single balanced chip no longer sits at the optimum for all three, so TPU 8t targets large-scale training while TPU 8i targets low-latency post-training and serving.
TPU 8t carries 216 GB of HBM at 6,500 GB/s with 128 MB of on-chip SRAM and 12.6 PFLOPS of peak FP4 compute — three times the processing of Ironwood at twice the performance per watt. A TPU 8t superpod spans 9,600 chips sharing 2 PB of HBM and delivering 121 EFLOPS of FP4 compute. Google quotes up to 2.7x better training performance per dollar than Ironwood.
TPU 8i attacks what Google calls the latency wall: autoregressive decoding and long chain-of-thought reasoning, where traditional throughput-optimised designs stall. It brings 288 GB per chip, roughly 10,100 FP4 TFLOPS and the new Boardfly interconnect, which cuts latency on communication-intensive workloads by around 50% and yields up to 80% better performance per dollar than Ironwood at low-latency targets. Both chips are designed to scale past a million units. QS Compute supplies AI accelerators and rack-scale training and inference infrastructure — request a quote.
Key Benefits
Workload-specialised: separate silicon for training and for low-latency serving. Massive memory pools: 216 GB per TPU 8t, 288 GB per TPU 8i. Lower serving latency: Boardfly interconnect cuts communication latency ~50%. Better economics: up to 2.7x training and 80% inference performance per dollar gains over Ironwood.
Applications
Frontier model pre-training, reinforcement learning and post-training, low-latency agentic inference, MoE serving, long-context reasoning workloads, and million-chip-scale AI supercomputers.
Request a Quote — GOOGLE TPU 8T & TPU 8I — EIGHTH-GENERATION TRAINING AND INFERENCE TPUS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →