Published: August 25, 2026 | Category: Buying Guide | QSCompute
Two servers with "the same GPU" can behave like entirely different machines. An H100 sold as an SXM module in an NVIDIA HGX baseboard and an H100 sold as a PCIe add-in card are the same silicon, but the SXM build connects to seven other GPUs over 900 GB/s NVLink while the PCIe build is capped at a PCIe 5.0 x16 bus. The form factor — SXM, PCIe, or OAM — decides whether a node can scale to multi-GPU training, whether you can upgrade it later, and whether it needs liquid cooling.
This guide explains the three form factors, why NVLink and NVSwitch exist, and how to pick the right one for training versus inference without paying for interconnect you will never use.
SXM (Server Module). NVIDIA's socketed, high-power module for data-center GPUs. There is no PCIe slot: the module plugs into a high-speed mezzanine connector on a purpose-built OEM baseboard (the NVIDIA HGX or DGX reference boards). SXM parts carry the highest TDPs (up to 700 W for H100 SXM) and — critically — expose the full NVLink fabric through NVSwitch chips on the baseboard. You cannot buy an SXM module and drop it into a generic server; you buy it pre-integrated in a vendor's HGX/DGX-class system.
PCIe (Add-in Card). The familiar double-wide, full-height card in a standard PCIe slot. PCIe GPUs (L40S, RTX 6000 Ada, RTX PRO 6000, L4, A2, and the H100 PCIe / H100 NVL variants) install into any server or workstation with the right slot and power budget, making them vendor-agnostic, reconfigurable, and resellable. The trade-off is interconnect: multi-GPU traffic has to cross the PCIe bus and the host CPU, and there is no NVSwitch fabric.
OAM (OCP Accelerator Module). The Open Compute Project's answer to SXM — a vendor-neutral mezzanine module standard adopted by AMD Instinct (MI250X, MI300X) and Intel Gaudi. OAM modules sit eight per OCP Universal Baseboard, mirroring the SXM topology but under an open standard, which is why hyperscalers favor it for their own rack designs.
| Form Factor | Install Method | Typical TDP | Multi-GPU Interconnect | Example GPUs | Best For |
|---|---|---|---|---|---|
| SXM | Socketed on OEM baseboard (HGX/DGX) | 350–700 W | NVLink + NVSwitch (all-to-all) | H100 SXM, H200 SXM, B200 SXM | Multi-GPU training, dense serving |
| PCIe | Standard slot, user-installable | 70–450 W | PCIe bus (through host CPU) | L40S, L4, RTX 6000 Ada, RTX PRO 6000, H100 PCIe | Inference, edge, flexible servers |
| OAM | OCP Universal Baseboard (8 per node) | 350–750 W | Infinity Fabric / proprietary | AMD MI300X, Intel Gaudi 3 | Hyperscaler, open-standard racks |
PCIe is a host-centric bus: any GPU-to-GPU transfer must travel through the CPU root complex. A PCIe 5.0 x16 slot delivers roughly 128 GB/s bidirectional (64 GB/s each direction) — and that bandwidth is shared with the NVMe storage and network traffic on the same CPU. NVLink is a direct, point-to-point GPU-to-GPU interconnect that bypasses the host entirely.
The gap has widened with every generation. NVLink 4.0 on Hopper delivers 900 GB/s per GPU — about 14× the bandwidth of a PCIe 4.0 x16 slot. NVSwitch takes this further: a switch chip on the baseboard connects all eight GPUs in an HGX node for all-to-all bandwidth, so any GPU can talk to any other at full speed simultaneously.
| NVLink Generation | GPU | Bandwidth per GPU (bidirectional) |
|---|---|---|
| NVLink 1.0 | P100 | 160 GB/s |
| NVLink 2.0 | V100 | 300 GB/s |
| NVLink 3.0 | A100 | 600 GB/s |
| NVLink 4.0 | H100/H200 | 900 GB/s |
| NVLink 5.0 | B200 | 1.8 TB/s |
This bandwidth only matters when GPUs are constantly exchanging data. That describes training — tensor parallelism shards a single model across GPUs and streams activations between them on every step — and large-model inference sharded across multiple GPUs. It does not describe the common inference case of one model resident on one GPU, where inter-GPU traffic is near zero and PCIe is more than adequate.
The form factor decision reduces to one question: does this workload need multiple GPUs working on the same model at the same time?
| Workload | Multi-GPU Coupling | Recommended Form Factor |
|---|---|---|
| LLM training (tensor parallel, ≥2 GPUs) | High — constant all-to-all traffic | SXM + NVSwitch (HGX/DGX) |
| Fine-tuning (LoRA/QLoRA) | Moderate | PCIe (2–4 cards) or SXM |
| Large-model inference (405B sharded) | High | SXM or multi-GPU PCIe with NVLink bridge |
| Single-model inference (≤70B) | Low — one GPU per model | PCIe (L40S, RTX PRO 6000, H100 PCIe) |
| Edge / video / vision inference | None — one GPU per node | PCIe (L4, A2, RTX 4000 SFF) |
| Hyperscaler custom racks | High | OAM |
Single-node inference server (PCIe). Two L40S or RTX 6000 Ada cards in a dual-socket Xeon/EPYC server, each serving an 8–13B model independently. Air-cooled, standard rack chassis, roughly $25,000–$40,000 per node. This is the workhorse for most production inference fleets.
Four-GPU fine-tuning node (PCIe). Four RTX PRO 6000 (96 GB each) in a GPU-optimized chassis. For LoRA/QLoRA fine-tuning of 70B models, PCIe bandwidth is not the bottleneck. Budget $45,000–$60,000.
Eight-GPU training node (SXM). An HGX H100 or HGX B200 baseboard with eight SXM modules and NVSwitch, liquid-cooled, in a dedicated rack. Street pricing for HGX H100 systems runs in the high six figures — a capital decision that only makes sense when tensor parallelism is the workload.
A checklist for procurement teams:
Need help matching GPU form factor to your workload?
QSCompute configures and burn-in tests both PCIe GPU servers (L40S, RTX 6000 Ada, RTX PRO 6000, H100 PCIe) and SXM/HGX training systems, with full documentation and a power/thermal test report.
Contact: +86 137-1464-6179 | sherry@qscompute.com