Specifications
Product
Qualcomm Cloud AI 100 Ultra
AI SoC Count
4 per card
AI Cores
16 per SoC — 64 AI cores per card
AI SRAM
144 MB on-card AI SRAM
Process Node
7 nm
Data Types
FP16, INT16, INT8, FP32
Form Factor
PCIe accelerator card
Family
Cloud AI 100 Standard / Pro / Ultra, and Cloud AI 080 series
Software
Qualcomm Cloud AI SDK / platform SDK, compiler and model zoo
Design Intent
High-throughput, high-performance-per-watt generative-AI inference
Deployment
Data-center and edge inference servers
Companion Series
Cloud AI 100 Pro (16 cores) and Cloud AI 100 Standard (14 cores)
Availability
In stock — quote for volume
Overview
The Qualcomm Cloud AI 100 Ultra is the four-SoC member of the Cloud AI 100 accelerator line: 16 AI cores per SoC, 64 cores per card, built on a 7 nm process and carrying 144 MB of on-card AI SRAM. It is designed for inference specifically — not training — and competes on performance per watt and per dollar in large-scale LLM serving.
Qualcomm positions the Ultra tier for the heaviest inference workloads in a deployment, with the Pro and Standard SKUs slotting underneath for lower-density nodes. Because all three share the same Cloud AI SDK and compiler, a model compiled once can be deployed across the tier.
QS Compute supplies Cloud AI 100 accelerator cards alongside the NVIDIA and AMD inference parts and the Tenstorrent Wormhole/Blackhole cards, so customers can benchmark inference cost per token across architectures before standardising.
Key Benefits
64 AI cores across four SoCs give the highest inference throughput in the Cloud AI 100 family. 144 MB AI SRAM keeps weights and activations close to compute. Single SDK across the tier lets one compiled model deploy on Ultra, Pro or Standard cards. 7 nm / inference-first design lowers power per token versus general-purpose GPUs.
Applications
Large-scale LLM and generative-AI inference serving; recommendation and embedding inference at the edge; computer vision and video analytics pipelines; and cost-sensitive inference nodes where tokens-per-watt matters more than training capability.
Request a Quote — QUALCOMM CLOUD AI 100 ULTRA — 64-CORE DATA CENTER AI INFERENCE ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →