Published: September 3, 2026 | Category: Buying Guide | QSCompute
The top-of-rack (ToR) switch has quietly become one of the most expensive and most mistake-prone line items in an AI cluster budget. A 32-GPU H100 buildout needs 32 ports of 400G at the leaf layer before you count spine uplinks — and at roughly $700–1,400 per 400G port, a wrong port count or a wrong buffer profile costs more to fix after the rack is wired than a GPU does. This guide covers the 2026 ToR landscape: what changed in AI traffic, how the three switch tiers compare, a decision matrix by cluster size, and a pre-PO checklist.
Traditional data center traffic is north-south: client to server, mostly small flows. AI training traffic is east-west — every GPU talks to every other GPU every few milliseconds, and the leaf switch is where that conversation starts. A modern 8-GPU H100/H200 server with ConnectX-7 NICs presents eight 400G ports to the network. At 1:1 (no oversubscription), a 32-port 400G ToR therefore serves exactly four such servers; a 64-port 400G switch serves one full rack of eight.
| ToR class | Silicon generation | Typical port layouts | What it connects in 2026 |
|---|---|---|---|
| 100G ToR | 12.8 Tbps (Tomahawk 3/4, NVIDIA Spectrum-3, Trident 4) | 48–128× 100G, up to 32× 400G with breakouts | Inference nodes, small labs, edge data centers, campus leaf |
| 400G ToR | 25.6 Tbps (NVIDIA Spectrum-4, Broadcom Tomahawk 4) | 32–64× 400G | H100/H200 training leaves — the 2026 default for 16–128 GPU clusters |
| 800G ToR | 51.2 Tbps (Tomahawk 5/6, NVIDIA Spectrum-5, deep-buffer routing ASICs) | 64× 800G, 128× 400G | B200/GB200-era leaves and spines, 256+ GPU fabrics, often cold-plate cooled |
Two implications follow. First, never buy the ToR after the servers — the port count is a function of GPU count, NIC speed, and the oversubscription ratio you can tolerate, and it must be decided together. Second, oversubscription policy is the real cost lever: training clusters should be 1:1 from GPU to leaf (NCCL all-reduce is latency- and bandwidth-starved), while inference and data-plane traffic tolerates 2:1 to 4:1 comfortably — which is how you end up with one 100G ToR instead of three 400G ones.
Switch silicon advances in ~25.6 Tbps jumps, and each generation roughly doubles both port speed and per-port cost while improving buffer and telemetry behavior for lossless fabrics. For RoCEv2 — the dominant Ethernet choice for AI, covered in depth in our fabric comparison — the switch must hold enough shared buffer to absorb incast bursts and support PFC/ECN marking without head-of-line blocking.
| 100G ToR (12.8T) | 400G ToR (25.6T) | 800G ToR (51.2T) | |
|---|---|---|---|
| Example platforms | Spectrum-3 SN4600-class, Tomahawk-3/4 boxes | Spectrum-4 SN5600 (64× 400G), Tomahawk-4 leaf | Tomahawk-5/6 and Spectrum-5 boxes, 64× 800G |
| Cut-through latency | ~500–700 ns | ~400–600 ns | ~300–500 ns |
| Shared buffer | 16–32 MB class | 32–64 MB class | 64 MB+ and deep-buffer variants |
| Power draw | 150–350 W | 500–900 W | 1.5–3 kW+ (liquid-cooled options) |
| 2026 street price (per box) | $4k–12k | $20k–45k | $60k–130k+ |
| Street price per port | ~$100–250 | ~$700–1,400 | ~$1,800–3,000+ |
| Best fit | ≤16-GPU inference, edge | 16–128 GPU training | 256+ GPU, GB200-scale |
The 25.6 Tbps tier is where most serious AI buyers land in 2026: 64-port 400G leaves give a full rack of 8-GPU servers at 1:1, the ASICs have mature RoCEv2 offload and streaming telemetry, and per-port pricing has fallen roughly 40% since 2024 as 51.2T silicon displaced them from the flagship slot. The 51.2T tier is real but punishing on optics: 64× 800G means 64 transceivers at $2,000+ each, which is why most 800G deployments appear at spine or superpod level rather than at every rack. A healthy secondary market also exists for 12.8T 100G switches from data-center refreshes — fine for inference fabrics, wrong for anything running NCCL at scale.
| Your setup | Recommended ToR | Rationale |
|---|---|---|
| ≤16 GPUs, inference (L40S, RTX 6000 Ada, Jetson-server mixes) | One 32–48 port 100G ToR | 2:1–4:1 oversubscription is invisible at inference batch loads; lowest capex and power |
| 16–64 GPUs, training (4–8× H100/H200 servers) | 25.6T 400G ToR, 1:1 downlinks | 32× 400G covers 4 nodes; buy a second for the rack pair and run MLAG/EVPN |
| 64–256 GPUs, production training | 400G leaf pair + 400G spine (2-tier CLOS) | Leaf 1:1 to servers, spine 2:1–3:1; RoCEv2 with ECN; validate with NCCL tests |
| 256+ GPUs or NVL72/GB200-class | 800G leaf or InfiniBand NDR fabric | 51.2T silicon, liquid-cooled switches, professional design — optics cost dominates |
Two structural warnings. First, do not oversubscribe the training leaf: with PFC and ECN misconfigured or buffers undersized, a single incast burst throttles the whole rack — and the failure shows up as mysteriously slow NCCL runs, not as switch errors. Second, budget the optics before the box: transceivers and cables frequently exceed the switch hardware cost, and our DAC/AOC guide covers when copper stops being an option (roughly 2 m at 400G).
The ToR switch is the one component you cannot quietly swap after the cluster is live — servers, NICs, and cables all plug into its port map. Do the port math first, pick the tier by your oversubscription tolerance, and validate the RoCEv2 config before the GPUs arrive.
Building an AI cluster and need a validated network design?
QSCompute supplies NVIDIA-certified 100G/400G/800G switches, ConnectX NICs, vendor-coded DACs/AOCs and optics, and burn-in-tested GPU servers. Tell us your GPU count and target ratio — we will return a port map and full BOM within 48 hours.
Contact: +86 137-1464-6179 | info@qscompute.com