Specifications
Manufacturer
AMD
Model
AMD Instinct MI350P PCIe
Family
Instinct MI350 Series
Architecture
AMD CDNA 4
Lithography
TSMC 3nm / 6nm FinFET
Compute Units
128
Stream Processors
8,192
Matrix Cores
512
Peak Engine Clock
2,200 MHz
Transistor Count
73 billion
Dedicated Memory
144GB HBM3E
Memory Interface
4,096-bit
Peak Memory Bandwidth
4 TB/s (8 Gbps × 4096-bit)
Infinity Cache (LLC)
128 MB
MXFP4 / MXFP6 Matrix
4.6 PFLOPs (structured sparsity)
MXFP8 Matrix
2.3 PFLOPs
BF16 Matrix
1.15 PFLOPs (2.3 PFLOPs sparse)
INT8 Matrix
4.6 POPs (structured sparsity)
Form Factor
Full-height full-length (FHFL) dual-slot PCIe add-in card, 267 mm
Bus Interface
PCIe 5.0 x16 (128 GB/s)
Cooling
Passive
Typical Board Power
600W max; 450W configurable
External Power
12V-2x6
RAS
Full-chip ECC, page retirement, page avoidance
Multi-Card
Up to 8 cards per air-cooled server
OS Support
Linux x86-64
Overview
The AMD Instinct MI350P is the PCIe add-in-card member of the Instinct MI350 series. It carries the same CDNA 4 architecture as the OAM-form MI350X but at half the scale: 128 compute units instead of 256, 512 matrix cores instead of 1,024, 144GB of HBM3E instead of 288GB, and 4TB/s of memory bandwidth instead of 8TB/s. In exchange it fits a standard full-height, full-length, double-slot PCIe 5.0 x16 slot — no OAM baseboard, no specialised interconnect topology — so existing air-cooled servers can be upgraded in place.
Peak matrix throughput is 4.6 PFLOPs at MXFP4 and MXFP6, 2.3 PFLOPs at MXFP8 and 1.15 PFLOPs at BF16 — roughly 40% above the H200 NVL on FP16 and FP8 theoretical compute for a card in the same physical class. The 144GB of HBM3E is the headline for inference: at 4TB/s it holds 70B-class models plus a large KV cache without sharding across multiple cards, and up to eight MI350P cards fit in a single air-cooled chassis for larger deployment.
Full-chip ECC with page retirement and page avoidance covers the memory subsystem, and the card runs on the open AMD ROCm software stack with no licensing fee. Two TBP configurations are offered — 600W standard and a 450W setting for tighter rack power and cooling envelopes. The MI350P pairs with AMD EPYC host processors and is aimed at organisations that want bleeding-edge inference throughput without rebuilding their datacentre around a new form factor.
Key Benefits
The fastest accelerator available in a plain PCIe slot, 144GB of HBM3E for large-model-in-one-card inference, air-cooled so it drops into existing servers, two power settings (600W / 450W), and an open ROCm stack with no per-GPU licensing.
Applications
Enterprise LLM and RAG inference, fine-tuning and LoRA adapters, recommendation and ranking models, computer vision training pipelines, HPC simulation, and mixed GPU fleets where a PCIe card is required rather than OAM.
Request a Quote — AMD INSTINCT MI350P PCIE — 144GB HBM3E ACCELERATOR
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →