Specifications

Manufacturer

AMD

Model

AMD Instinct MI350P PCIe

Family

Instinct MI350 Series

Architecture

AMD CDNA 4

Lithography

TSMC 3nm / 6nm FinFET

Compute Units

128

Stream Processors

8,192

Matrix Cores

512

Peak Engine Clock

2,200 MHz

Transistor Count

73 billion

Dedicated Memory

144GB HBM3E

Memory Interface

4,096-bit

Peak Memory Bandwidth

4 TB/s (8 Gbps × 4096-bit)

Infinity Cache (LLC)

128 MB

MXFP4 / MXFP6 Matrix

4.6 PFLOPs (structured sparsity)

MXFP8 Matrix

2.3 PFLOPs

BF16 Matrix

1.15 PFLOPs (2.3 PFLOPs sparse)

INT8 Matrix

4.6 POPs (structured sparsity)

Form Factor

Full-height full-length (FHFL) dual-slot PCIe add-in card, 267 mm

Bus Interface

PCIe 5.0 x16 (128 GB/s)

Cooling

Passive

Typical Board Power

600W max; 450W configurable

External Power

12V-2x6

RAS

Full-chip ECC, page retirement, page avoidance

Multi-Card

Up to 8 cards per air-cooled server

OS Support

Linux x86-64

Overview

The AMD Instinct MI350P is the PCIe add-in-card member of the Instinct MI350 series. It carries the same CDNA 4 architecture as the OAM-form MI350X but at half the scale: 128 compute units instead of 256, 512 matrix cores instead of 1,024, 144GB of HBM3E instead of 288GB, and 4TB/s of memory bandwidth instead of 8TB/s. In exchange it fits a standard full-height, full-length, double-slot PCIe 5.0 x16 slot — no OAM baseboard, no specialised interconnect topology — so existing air-cooled servers can be upgraded in place.

Peak matrix throughput is 4.6 PFLOPs at MXFP4 and MXFP6, 2.3 PFLOPs at MXFP8 and 1.15 PFLOPs at BF16 — roughly 40% above the H200 NVL on FP16 and FP8 theoretical compute for a card in the same physical class. The 144GB of HBM3E is the headline for inference: at 4TB/s it holds 70B-class models plus a large KV cache without sharding across multiple cards, and up to eight MI350P cards fit in a single air-cooled chassis for larger deployment.

Full-chip ECC with page retirement and page avoidance covers the memory subsystem, and the card runs on the open AMD ROCm software stack with no licensing fee. Two TBP configurations are offered — 600W standard and a 450W setting for tighter rack power and cooling envelopes. The MI350P pairs with AMD EPYC host processors and is aimed at organisations that want bleeding-edge inference throughput without rebuilding their datacentre around a new form factor.

Key Benefits

The fastest accelerator available in a plain PCIe slot, 144GB of HBM3E for large-model-in-one-card inference, air-cooled so it drops into existing servers, two power settings (600W / 450W), and an open ROCm stack with no per-GPU licensing.

Applications

Enterprise LLM and RAG inference, fine-tuning and LoRA adapters, recommendation and ranking models, computer vision training pipelines, HPC simulation, and mixed GPU fleets where a PCIe card is required rather than OAM.

Request a Quote — AMD INSTINCT MI350P PCIE — 144GB HBM3E ACCELERATOR

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

A100 Pcie A100 Sxm4 Adlink Egx Mxm Blackwell