NVIDIA GPU Form Factor Guide 2026 — SXM vs PCIe vs OAM: NVLink, NVSwitch & Which to Buy for AI Training vs Inference

Published: August 25, 2026 | Category: Buying Guide | QSCompute

Two servers with "the same GPU" can behave like entirely different machines. An H100 sold as an SXM module in an NVIDIA HGX baseboard and an H100 sold as a PCIe add-in card are the same silicon, but the SXM build connects to seven other GPUs over 900 GB/s NVLink while the PCIe build is capped at a PCIe 5.0 x16 bus. The form factor — SXM, PCIe, or OAM — decides whether a node can scale to multi-GPU training, whether you can upgrade it later, and whether it needs liquid cooling.

This guide explains the three form factors, why NVLink and NVSwitch exist, and how to pick the right one for training versus inference without paying for interconnect you will never use.

The Three GPU Form Factors, Explained

SXM (Server Module). NVIDIA's socketed, high-power module for data-center GPUs. There is no PCIe slot: the module plugs into a high-speed mezzanine connector on a purpose-built OEM baseboard (the NVIDIA HGX or DGX reference boards). SXM parts carry the highest TDPs (up to 700 W for H100 SXM) and — critically — expose the full NVLink fabric through NVSwitch chips on the baseboard. You cannot buy an SXM module and drop it into a generic server; you buy it pre-integrated in a vendor's HGX/DGX-class system.

PCIe (Add-in Card). The familiar double-wide, full-height card in a standard PCIe slot. PCIe GPUs (L40S, RTX 6000 Ada, RTX PRO 6000, L4, A2, and the H100 PCIe / H100 NVL variants) install into any server or workstation with the right slot and power budget, making them vendor-agnostic, reconfigurable, and resellable. The trade-off is interconnect: multi-GPU traffic has to cross the PCIe bus and the host CPU, and there is no NVSwitch fabric.

OAM (OCP Accelerator Module). The Open Compute Project's answer to SXM — a vendor-neutral mezzanine module standard adopted by AMD Instinct (MI250X, MI300X) and Intel Gaudi. OAM modules sit eight per OCP Universal Baseboard, mirroring the SXM topology but under an open standard, which is why hyperscalers favor it for their own rack designs.

Form FactorInstall MethodTypical TDPMulti-GPU InterconnectExample GPUsBest For
SXMSocketed on OEM baseboard (HGX/DGX)350–700 WNVLink + NVSwitch (all-to-all)H100 SXM, H200 SXM, B200 SXMMulti-GPU training, dense serving
PCIeStandard slot, user-installable70–450 WPCIe bus (through host CPU)L40S, L4, RTX 6000 Ada, RTX PRO 6000, H100 PCIeInference, edge, flexible servers
OAMOCP Universal Baseboard (8 per node)350–750 WInfinity Fabric / proprietaryAMD MI300X, Intel Gaudi 3Hyperscaler, open-standard racks

NVLink & NVSwitch — Why Multi-GPU Needs More Than PCIe

PCIe is a host-centric bus: any GPU-to-GPU transfer must travel through the CPU root complex. A PCIe 5.0 x16 slot delivers roughly 128 GB/s bidirectional (64 GB/s each direction) — and that bandwidth is shared with the NVMe storage and network traffic on the same CPU. NVLink is a direct, point-to-point GPU-to-GPU interconnect that bypasses the host entirely.

The gap has widened with every generation. NVLink 4.0 on Hopper delivers 900 GB/s per GPU — about 14× the bandwidth of a PCIe 4.0 x16 slot. NVSwitch takes this further: a switch chip on the baseboard connects all eight GPUs in an HGX node for all-to-all bandwidth, so any GPU can talk to any other at full speed simultaneously.

NVLink GenerationGPUBandwidth per GPU (bidirectional)
NVLink 1.0P100160 GB/s
NVLink 2.0V100300 GB/s
NVLink 3.0A100600 GB/s
NVLink 4.0H100/H200900 GB/s
NVLink 5.0B2001.8 TB/s

This bandwidth only matters when GPUs are constantly exchanging data. That describes training — tensor parallelism shards a single model across GPUs and streams activations between them on every step — and large-model inference sharded across multiple GPUs. It does not describe the common inference case of one model resident on one GPU, where inter-GPU traffic is near zero and PCIe is more than adequate.

Training vs Inference — The Decision Matrix

The form factor decision reduces to one question: does this workload need multiple GPUs working on the same model at the same time?

WorkloadMulti-GPU CouplingRecommended Form Factor
LLM training (tensor parallel, ≥2 GPUs)High — constant all-to-all trafficSXM + NVSwitch (HGX/DGX)
Fine-tuning (LoRA/QLoRA)ModeratePCIe (2–4 cards) or SXM
Large-model inference (405B sharded)HighSXM or multi-GPU PCIe with NVLink bridge
Single-model inference (≤70B)Low — one GPU per modelPCIe (L40S, RTX PRO 6000, H100 PCIe)
Edge / video / vision inferenceNone — one GPU per nodePCIe (L4, A2, RTX 4000 SFF)
Hyperscaler custom racksHighOAM
SXM is a topology purchase; PCIe is a GPU purchase. If you are buying one or two GPUs to serve models or run a workload per GPU, PCIe is almost always the right economic choice — the same silicon at a lower price, in a server you can reconfigure and resell. If you are building a training node where GPUs must behave as a single large accelerator, the NVSwitch fabric in an SXM system is not optional.

Buying Guidance & Reference Configurations

Single-node inference server (PCIe). Two L40S or RTX 6000 Ada cards in a dual-socket Xeon/EPYC server, each serving an 8–13B model independently. Air-cooled, standard rack chassis, roughly $25,000–$40,000 per node. This is the workhorse for most production inference fleets.

Four-GPU fine-tuning node (PCIe). Four RTX PRO 6000 (96 GB each) in a GPU-optimized chassis. For LoRA/QLoRA fine-tuning of 70B models, PCIe bandwidth is not the bottleneck. Budget $45,000–$60,000.

Eight-GPU training node (SXM). An HGX H100 or HGX B200 baseboard with eight SXM modules and NVSwitch, liquid-cooled, in a dedicated rack. Street pricing for HGX H100 systems runs in the high six figures — a capital decision that only makes sense when tensor parallelism is the workload.

A checklist for procurement teams:

The one-line rule: NVLink is for GPUs that talk to each other on every step. If your GPUs run independent workloads, buy PCIe and spend the difference on more GPUs.

Need help matching GPU form factor to your workload?

QSCompute configures and burn-in tests both PCIe GPU servers (L40S, RTX 6000 Ada, RTX PRO 6000, H100 PCIe) and SXM/HGX training systems, with full documentation and a power/thermal test report.

Contact: +86 137-1464-6179 | sherry@qscompute.com