Specifications
Vendor
Microsoft (Azure)
Product type
Datacenter AI inference accelerator (XPU)
Process node
TSMC 3 nm
Transistor count
~140 billion
Memory
216 GB HBM3e
Memory bandwidth
7 TB/s
Memory stacks
6x 36 GB, 12-high HBM3e stacks
On-die SRAM
272 MB scratchpad / cluster SRAM
Matrix engine precision
FP8 / FP6 / FP4 native tensor cores
Vector engine precision
BF16 / FP16 / FP32
FP4 throughput
10.15 PFLOPS (vendor specification)
FP8 throughput
5.07 PFLOPS on tensor units
BF16 throughput
1.27 PFLOPS on vector units
Thermal design power
750 W TDP
Primary workload
RL post-training, inference and token generation
Deployment model
Azure first-party fleet
Availability
2026
Overview
The Microsoft Maia 200 is Microsoft's second-generation in-house AI accelerator and a deliberate shift in datacenter silicon design: rather than optimising for training throughput alone, Maia 200 is engineered around the economics of token generation and reinforcement-learning post-training. It pairs native FP8/FP6/FP4 tensor cores with a separate BF16/FP16/FP32 vector engine, so a single device handles both the low-precision matrix work that dominates inference and the higher-precision vector maths that modern RL pipelines require.
Memory capacity is the headline differentiator. At 216 GB of HBM3e across six 12-high stacks with 7 TB/s of bandwidth, Maia 200 comfortably holds large mixture-of-experts weights and long KV caches on-package, while 272 MB of on-die cluster SRAM acts as a scratchpad to keep hot activations off HBM. The 750 W thermal envelope places it in the liquid-cooled rack class alongside current flagship datacenter accelerators.
For buyers assembling Azure-adjacent or on-premises inference capacity, Maia 200 represents the XPU tier that sits between merchant GPUs and fixed-function inference ASICs: general enough to run transformers, MoE and multimodal models, but tuned end-to-end for the cost-per-token metric that hyperscale deployments are judged on.
Key Benefits
216 GB HBM3e at 7 TB/s keeps frontier-class weights and long KV caches resident on a single device. Native FP4/FP8/FP6 tensor cores deliver up to 10.15 PFLOPS at the lowest-precision points where inference economics are decided. 272 MB of on-die SRAM cuts HBM round-trips for hot activations. Separate BF16/FP16/FP32 vector engines make the part viable for RL post-training as well as serving.
Applications
High-volume LLM and multimodal token generation, reinforcement-learning post-training, mixture-of-experts inference, long-context serving with large KV caches, and Azure-scale inference capacity planning.
Request a Quote — MICROSOFT MAIA 200 — 216GB HBM3E DATACENTER AI INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →