Specifications

Family Structure

Two chips: TPU 8t for training, TPU 8i for inference

TPU 8t HBM

216 GB of HBM per chip

TPU 8t Bandwidth

6,500 GB/s of HBM bandwidth

TPU 8t SRAM

128 MB of on-chip SRAM

TPU 8t Compute

12.6 PFLOPS peak FP4 per chip

TPU 8t vs Ironwood

3x the processing of TPU v7 Ironwood at 2x performance per watt

TPU 8t Superpod

9,600 chips with 2 PB of shared HBM memory

TPU 8t Superpod Compute

121 EFLOPS of FP4 compute

TPU 8i Memory

288 GB per chip

TPU 8i Compute

About 10,100 FP4 TFLOPS per chip

TPU 8i Interconnect

Boardfly interconnect — roughly 50% lower latency on communication-intensive workloads

TPU 8i Price/Performance

Up to 80% better performance per dollar than Ironwood at low-latency targets

Training Price/Performance

Up to 2.7x performance per dollar vs Ironwood for large-scale training

Energy Efficiency

Up to 2x better performance per watt on both chips

Scale Ceiling

Designed to scale beyond 1 million chips

Networking

Google Virgo interconnect with speedier Lustre storage tiers

Availability

Announced April 22, 2026 at Google Cloud Next; general availability stated for later in 2026

Overview

Google's eighth-generation TPU is the first to split training and inference into two distinct architectures. Pre-training, post-training and real-time serving have diverged enough in operational intensity that a single balanced chip no longer sits at the optimum for all three, so TPU 8t targets large-scale training while TPU 8i targets low-latency post-training and serving.

TPU 8t carries 216 GB of HBM at 6,500 GB/s with 128 MB of on-chip SRAM and 12.6 PFLOPS of peak FP4 compute — three times the processing of Ironwood at twice the performance per watt. A TPU 8t superpod spans 9,600 chips sharing 2 PB of HBM and delivering 121 EFLOPS of FP4 compute. Google quotes up to 2.7x better training performance per dollar than Ironwood.

TPU 8i attacks what Google calls the latency wall: autoregressive decoding and long chain-of-thought reasoning, where traditional throughput-optimised designs stall. It brings 288 GB per chip, roughly 10,100 FP4 TFLOPS and the new Boardfly interconnect, which cuts latency on communication-intensive workloads by around 50% and yields up to 80% better performance per dollar than Ironwood at low-latency targets. Both chips are designed to scale past a million units. QS Compute supplies AI accelerators and rack-scale training and inference infrastructure — request a quote.

Key Benefits

Workload-specialised: separate silicon for training and for low-latency serving. Massive memory pools: 216 GB per TPU 8t, 288 GB per TPU 8i. Lower serving latency: Boardfly interconnect cuts communication latency ~50%. Better economics: up to 2.7x training and 80% inference performance per dollar gains over Ironwood.

Applications

Frontier model pre-training, reinforcement learning and post-training, low-latency agentic inference, MoE serving, long-context reasoning workloads, and million-chip-scale AI supercomputers.

Request a Quote — GOOGLE TPU 8T & TPU 8I — EIGHTH-GENERATION TRAINING AND INFERENCE TPUS

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Rubin GPU — Next-Gen AI Accelerator AMD Instinct MI455X — 432GB HBM4 Accelerator Microsoft Maia 200 AI Accelerator