Hexagon NPU · up to 768GB LPDDR per card · ~70B transistors N3E · 350B-param on one card · AI250 near-memory computing
Inference-optimized datacenter accelerator available as individual chips, PCIe cards, or liquid-cooled server racks.
Next-gen accelerator introducing near-memory computing for a step-change in effective bandwidth and power efficiency.
Rack-level reference system integrating acceleration, memory, interconnect and management software (demonstrated at MWC 2026).
Qualcomm Technologies, Inc.
Cloud AI 200 / AI 250
Datacenter inference accelerator
Hexagon NPU
Up to 768GB LPDDR
per PCIe card
~70B transistors · TSMC N3E
Near-memory computing
>10× effective bandwidth
Chip · PCIe card
Liquid-cooled rack
AI200 — 2026
AI250 — 2027
350B-param model on
a single AI200 card (MWC 2026)
Rack-scale inference
industry-leading TCO
Available
Quote
Qualcomm announced the Cloud AI 200 and Cloud AI 250 on October 27, 2025, marking its push from mobile NPUs into rack-scale datacenter AI inference. Built on Qualcomm's Hexagon NPU heritage — refined across phones and PCs — these inference-optimized accelerators target large language model and multimodal AI serving with industry-leading total cost of ownership, directly challenging NVIDIA and AMD on inference economics.
The AI200 supports up to 768GB of LPDDR memory per PCIe card on roughly 70 billion transistors (TSMC N3E), and is offered as individual chips, PCIe accelerator cards, or fully integrated liquid-cooled server racks. At MWC 2026, Qualcomm demonstrated a 350-billion-parameter model running on a single AI200 card, underscoring its memory-first design for high-capacity inference. Commercial availability is 2026.
The AI250 (2027) introduces near-memory computing for over 10× higher effective bandwidth and reduced power consumption, extending Qualcomm's annual-cadence data-center roadmap. Early lighthouse deployments — including Saudi-based Humain's plans for hundreds of megawatts of Qualcomm-based inference capacity — validate the platform's scale.
Need Qualcomm Cloud AI 200 / 250?
Contact QS Compute for availability, configuration, and volume pricing.
Request Quote