Specifications
HBM Capacity
192 GB HBM3E per chip
HBM Bandwidth
7.37 - 7.38 TB/s per chip
FP8 Throughput
4,614 FP8 TFLOPS per chip
BF16 Throughput
2,307 BF16 TFLOPS per chip
Native Precision
First TPU generation with native FP8 hardware support
Interconnect
1.2 TB/s bidirectional ICI per chip
Superpod Size
9,216 chips
Superpod Compute
42.5 FP8 ExaFLOPS
Package TDP
600 W, liquid cooled
Board Density
One TPU board carries 4 TPU chip packages
Networking per Chip
4 OSFP cages for ICI plus 1 CDFP PCIe cage to the host CPU
Generation Uplift
10x the peak performance of TPU v5p; more than 4x the performance per chip of Trillium (v6e)
Bytes per FLOP
~3.2 bytes/FLOP — tuned for inference rather than training
Availability
Generally available April 22, 2026, announced at Google Cloud Next 2026
Anchor Customer
Anthropic — an initial phase of roughly 400,000 units with a further ~600,000 committed
Software
Google Cloud AI Hypercomputer stack with JAX/XLA tooling
Overview
Ironwood is Google's seventh-generation TPU and the first it explicitly positions as inference-first. Each chip carries 192 GB of HBM3E at up to 7.37 TB/s and delivers 4,614 FP8 TFLOPS — a memory-to-compute ratio tuned for serving large models rather than training them, which is why Ironwood also marks the first TPU generation with native FP8 support.
Scale is the other half of the story. A single Ironwood superpod packs 9,216 chips reaching 42.5 FP8 ExaFLOPS, with 1.2 TB/s of bidirectional ICI per chip, four OSFP cages per chip for the interconnect and one CDFP PCIe cage to the host CPU. Four TPU packages sit on each TPU board. Packaged at 600 W with liquid cooling, Ironwood roughly doubles performance per chip against Trillium and delivers about ten times the peak performance of TPU v5p.
Google made Ironwood generally available on April 22, 2026 at Google Cloud Next. Anthropic is the anchor customer with an initial phase of roughly 400,000 units plus a further ~600,000 committed, a deployment measured in gigawatts. QS Compute supplies AI accelerators, GPU servers and rack-scale compute for inference-heavy deployments — contact us for pricing and availability.
Key Benefits
Inference-optimised balance: 192 GB HBM3E at 7.4 TB/s against 4,614 FP8 TFLOPS. Native FP8: first TPU with hardware FP8 support. Pod-scale fabric: 1.2 TB/s ICI per chip across 9,216-chip pods. Proven at scale: in production for frontier model serving.
Applications
Serving frontier-scale language and multimodal models, low-latency inference at low batch sizes, long-context serving, and large-scale AI deployments where memory capacity per chip is the binding constraint.
Request a Quote — GOOGLE IRONWOOD TPU V7 — 192GB HBM3E INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →