Specifications

HBM Capacity

192 GB HBM3E per chip

HBM Bandwidth

7.37 - 7.38 TB/s per chip

FP8 Throughput

4,614 FP8 TFLOPS per chip

BF16 Throughput

2,307 BF16 TFLOPS per chip

Native Precision

First TPU generation with native FP8 hardware support

Interconnect

1.2 TB/s bidirectional ICI per chip

Superpod Size

9,216 chips

Superpod Compute

42.5 FP8 ExaFLOPS

Package TDP

600 W, liquid cooled

Board Density

One TPU board carries 4 TPU chip packages

Networking per Chip

4 OSFP cages for ICI plus 1 CDFP PCIe cage to the host CPU

Generation Uplift

10x the peak performance of TPU v5p; more than 4x the performance per chip of Trillium (v6e)

Bytes per FLOP

~3.2 bytes/FLOP — tuned for inference rather than training

Availability

Generally available April 22, 2026, announced at Google Cloud Next 2026

Anchor Customer

Anthropic — an initial phase of roughly 400,000 units with a further ~600,000 committed

Software

Google Cloud AI Hypercomputer stack with JAX/XLA tooling

Overview

Ironwood is Google's seventh-generation TPU and the first it explicitly positions as inference-first. Each chip carries 192 GB of HBM3E at up to 7.37 TB/s and delivers 4,614 FP8 TFLOPS — a memory-to-compute ratio tuned for serving large models rather than training them, which is why Ironwood also marks the first TPU generation with native FP8 support.

Scale is the other half of the story. A single Ironwood superpod packs 9,216 chips reaching 42.5 FP8 ExaFLOPS, with 1.2 TB/s of bidirectional ICI per chip, four OSFP cages per chip for the interconnect and one CDFP PCIe cage to the host CPU. Four TPU packages sit on each TPU board. Packaged at 600 W with liquid cooling, Ironwood roughly doubles performance per chip against Trillium and delivers about ten times the peak performance of TPU v5p.

Google made Ironwood generally available on April 22, 2026 at Google Cloud Next. Anthropic is the anchor customer with an initial phase of roughly 400,000 units plus a further ~600,000 committed, a deployment measured in gigawatts. QS Compute supplies AI accelerators, GPU servers and rack-scale compute for inference-heavy deployments — contact us for pricing and availability.

Key Benefits

Inference-optimised balance: 192 GB HBM3E at 7.4 TB/s against 4,614 FP8 TFLOPS. Native FP8: first TPU with hardware FP8 support. Pod-scale fabric: 1.2 TB/s ICI per chip across 9,216-chip pods. Proven at scale: in production for frontier model serving.

Applications

Serving frontier-scale language and multimodal models, low-latency inference at low batch sizes, long-context serving, and large-scale AI deployments where memory capacity per chip is the binding constraint.

Request a Quote — GOOGLE IRONWOOD TPU V7 — 192GB HBM3E INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Rubin GPU — Next-Gen AI Accelerator AMD Instinct MI455X — 432GB HBM4 Accelerator Cerebras WSE-3 Turbo