MPN: Cloud AI 200 / AI 250 · Qualcomm
Qualcomm Cloud AI 200 / 250
Datacenter Inference · Rack-Scale

Hexagon NPU · up to 768GB LPDDR per card · ~70B transistors N3E · 350B-param on one card · AI250 near-memory computing

Product Lineup

Cloud AI 200

Inference-optimized datacenter accelerator available as individual chips, PCIe cards, or liquid-cooled server racks.

Cloud AI 250

Next-gen accelerator introducing near-memory computing for a step-change in effective bandwidth and power efficiency.

AI200 Rack & Management Suite

Rack-level reference system integrating acceleration, memory, interconnect and management software (demonstrated at MWC 2026).

Specifications

Vendor

Qualcomm Technologies, Inc.

Product

Cloud AI 200 / AI 250

Type

Datacenter inference accelerator

NPU

Hexagon NPU

Memory (AI200)

Up to 768GB LPDDR
per PCIe card

Process (AI200)

~70B transistors · TSMC N3E

AI250 Memory

Near-memory computing
>10× effective bandwidth

Form Factors

Chip · PCIe card
Liquid-cooled rack

Availability

AI200 — 2026
AI250 — 2027

Highlight

350B-param model on
a single AI200 card (MWC 2026)

Positioning

Rack-scale inference
industry-leading TCO

Stock

Available

Price

Quote

Overview

Qualcomm announced the Cloud AI 200 and Cloud AI 250 on October 27, 2025, marking its push from mobile NPUs into rack-scale datacenter AI inference. Built on Qualcomm's Hexagon NPU heritage — refined across phones and PCs — these inference-optimized accelerators target large language model and multimodal AI serving with industry-leading total cost of ownership, directly challenging NVIDIA and AMD on inference economics.

The AI200 supports up to 768GB of LPDDR memory per PCIe card on roughly 70 billion transistors (TSMC N3E), and is offered as individual chips, PCIe accelerator cards, or fully integrated liquid-cooled server racks. At MWC 2026, Qualcomm demonstrated a 350-billion-parameter model running on a single AI200 card, underscoring its memory-first design for high-capacity inference. Commercial availability is 2026.

The AI250 (2027) introduces near-memory computing for over 10× higher effective bandwidth and reduced power consumption, extending Qualcomm's annual-cadence data-center roadmap. Early lighthouse deployments — including Saudi-based Humain's plans for hundreds of megawatts of Qualcomm-based inference capacity — validate the platform's scale.

Need Qualcomm Cloud AI 200 / 250?

Contact QS Compute for availability, configuration, and volume pricing.

Request Quote