Specifications

Vendor

Microsoft (Azure)

Product type

Datacenter AI inference accelerator (XPU)

Process node

TSMC 3 nm

Transistor count

~140 billion

Memory

216 GB HBM3e

Memory bandwidth

7 TB/s

Memory stacks

6x 36 GB, 12-high HBM3e stacks

On-die SRAM

272 MB scratchpad / cluster SRAM

Matrix engine precision

FP8 / FP6 / FP4 native tensor cores

Vector engine precision

BF16 / FP16 / FP32

FP4 throughput

10.15 PFLOPS (vendor specification)

FP8 throughput

5.07 PFLOPS on tensor units

BF16 throughput

1.27 PFLOPS on vector units

Thermal design power

750 W TDP

Primary workload

RL post-training, inference and token generation

Deployment model

Azure first-party fleet

Availability

2026

Overview

The Microsoft Maia 200 is Microsoft's second-generation in-house AI accelerator and a deliberate shift in datacenter silicon design: rather than optimising for training throughput alone, Maia 200 is engineered around the economics of token generation and reinforcement-learning post-training. It pairs native FP8/FP6/FP4 tensor cores with a separate BF16/FP16/FP32 vector engine, so a single device handles both the low-precision matrix work that dominates inference and the higher-precision vector maths that modern RL pipelines require.

Memory capacity is the headline differentiator. At 216 GB of HBM3e across six 12-high stacks with 7 TB/s of bandwidth, Maia 200 comfortably holds large mixture-of-experts weights and long KV caches on-package, while 272 MB of on-die cluster SRAM acts as a scratchpad to keep hot activations off HBM. The 750 W thermal envelope places it in the liquid-cooled rack class alongside current flagship datacenter accelerators.

For buyers assembling Azure-adjacent or on-premises inference capacity, Maia 200 represents the XPU tier that sits between merchant GPUs and fixed-function inference ASICs: general enough to run transformers, MoE and multimodal models, but tuned end-to-end for the cost-per-token metric that hyperscale deployments are judged on.

Key Benefits

216 GB HBM3e at 7 TB/s keeps frontier-class weights and long KV caches resident on a single device. Native FP4/FP8/FP6 tensor cores deliver up to 10.15 PFLOPS at the lowest-precision points where inference economics are decided. 272 MB of on-die SRAM cuts HBM round-trips for hot activations. Separate BF16/FP16/FP32 vector engines make the part viable for RL post-training as well as serving.

Applications

High-volume LLM and multimodal token generation, reinforcement-learning post-training, mixture-of-experts inference, long-context serving with large KV caches, and Azure-scale inference capacity planning.

Request a Quote — MICROSOFT MAIA 200 — 216GB HBM3E DATACENTER AI INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Qualcomm DragonFly AI300 Accelerator Intel Crescent Island Inference GPU Qualcomm AI200 / AI250 Inference Accelerators NVIDIA H200 141GB HBM3e