Specifications

Product

Qualcomm Cloud AI 100 Ultra

AI SoC Count

4 per card

AI Cores

16 per SoC — 64 AI cores per card

AI SRAM

144 MB on-card AI SRAM

Process Node

7 nm

Data Types

FP16, INT16, INT8, FP32

Form Factor

PCIe accelerator card

Family

Cloud AI 100 Standard / Pro / Ultra, and Cloud AI 080 series

Software

Qualcomm Cloud AI SDK / platform SDK, compiler and model zoo

Design Intent

High-throughput, high-performance-per-watt generative-AI inference

Deployment

Data-center and edge inference servers

Companion Series

Cloud AI 100 Pro (16 cores) and Cloud AI 100 Standard (14 cores)

Availability

In stock — quote for volume

Overview

The Qualcomm Cloud AI 100 Ultra is the four-SoC member of the Cloud AI 100 accelerator line: 16 AI cores per SoC, 64 cores per card, built on a 7 nm process and carrying 144 MB of on-card AI SRAM. It is designed for inference specifically — not training — and competes on performance per watt and per dollar in large-scale LLM serving.

Qualcomm positions the Ultra tier for the heaviest inference workloads in a deployment, with the Pro and Standard SKUs slotting underneath for lower-density nodes. Because all three share the same Cloud AI SDK and compiler, a model compiled once can be deployed across the tier.

QS Compute supplies Cloud AI 100 accelerator cards alongside the NVIDIA and AMD inference parts and the Tenstorrent Wormhole/Blackhole cards, so customers can benchmark inference cost per token across architectures before standardising.

Key Benefits

64 AI cores across four SoCs give the highest inference throughput in the Cloud AI 100 family. 144 MB AI SRAM keeps weights and activations close to compute. Single SDK across the tier lets one compiled model deploy on Ultra, Pro or Standard cards. 7 nm / inference-first design lowers power per token versus general-purpose GPUs.

Applications

Large-scale LLM and generative-AI inference serving; recommendation and embedding inference at the edge; computer vision and video analytics pipelines; and cost-sensitive inference nodes where tokens-per-watt matters more than training capability.

Request a Quote — QUALCOMM CLOUD AI 100 ULTRA — 64-CORE DATA CENTER AI INFERENCE ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA L40S 48GB AMD Instinct MI300A Intel Arc Pro B60 24GB Tenstorrent Blackhole p150