Published: August 26, 2026 | Category: Buying Guide | QSCompute
In 2026 the best open-weight models — DeepSeek-V3 and R1, Mixtral 8x7B and 8x22B, Qwen3-MoE, and Llama 4 Maverick — share one architectural trait: they are sparse, built as a Mixture of Experts (MoE) rather than a single dense network. That single design choice quietly inverts the GPU buying calculus. A dense model pays for its intelligence with compute; an MoE model pays with memory.
For procurement teams this matters because the vendor spec that sells dense-model servers — TFLOPS — is close to irrelevant for MoE. The binding constraints are memory capacity (can the GPU hold every expert?) and memory bandwidth (how fast can the active experts stream through?). Get those two numbers right and a mid-range H200 can serve a 46B-parameter MoE model faster than a higher-TFLOPS GPU with less VRAM.
A dense 70B model activates all 70 billion parameters on every token. An MoE model keeps far more parameters in total but routes each token through only a small subset. A lightweight router scores the token and dispatches it to the top-k of N parallel "expert" feed-forward networks. The result: a model with frontier-scale capacity at a fraction of the per-token compute.
| Model | Total Params | Active Params | Experts (active) | FP8 Footprint |
|---|---|---|---|---|
| Mixtral 8x7B | 46.7B | ~12.9B | 8 (2) | ~47 GB |
| Mixtral 8x22B | 141B | ~39B | 8 (2) | ~141 GB |
| Qwen3-235B-A22B | 235B | 22B | 128 (8) | ~235 GB |
| Llama 4 Maverick | 400B | 17B | 128 (1) | ~400 GB |
| DeepSeek-V3 / R1 | 671B | 37B | 256 (8) | ~671 GB |
Token generation is memory-bound, and MoE splits that memory problem into two distinct constraints that point at different hardware:
The practical consequence: for MoE, buy capacity first (enough VRAM to hold all experts without spilling), then bandwidth second (to raise the throughput ceiling). Compute is a distant third. This is the opposite ranking from dense-model server shopping.
Match the GPU to the model's total-parameter footprint, not its active count. Mid-size MoE models (Mixtral 8x7B and 8x22B) are single-GPU feasible on HBM parts; frontier MoE models (Qwen3-235B, DeepSeek-V3) are multi-GPU or B200-class decisions.
| GPU | Memory | Bandwidth | Fits in FP8/INT8 | MoE Verdict |
|---|---|---|---|---|
| B200 SXM | 192 GB HBM3e | 8 TB/s | Mixtral 8x22B, Qwen3-235B (partial) | Best single-GPU for large MoE |
| AMD MI300X | 192 GB HBM3 | 5.3 TB/s | Mixtral 8x22B, Qwen3-235B | Best value for open-model MoE |
| H200 SXM | 141 GB HBM3e | 4.8 TB/s | Mixtral 8x7B, 8x22B | Best single-GPU for mid MoE |
| H100 SXM | 80 GB HBM3 | 3.35 TB/s | Mixtral 8x7B (INT8) | Entry MoE; tight for 8x22B |
| RTX PRO 6000 | 96 GB GDDR7 | 1.79 TB/s | Mixtral 8x7B (INT8) | Workstation MoE serving |
| RTX 6000 Ada | 48 GB GDDR6 ECC | 960 GB/s | Small MoE (INT4) | Limited to compact MoE |
| L40S | 48 GB GDDR6 ECC | 864 GB/s | Small MoE (INT4) | Budget edge MoE |
Frontier MoE models exceed any single GPU. DeepSeek-V3's 671B parameters in FP8 require roughly eight H200s (8 × 141 GB = 1.13 TB) once KV cache and activation buffers are accounted for. Two deployment patterns dominate:
ktransformers and llama.cpp keep the frequently-used experts in GPU VRAM and spill the rest to system DRAM. A single RTX PRO 6000 plus a large DDR5 host can run a 671B MoE at a few tokens/sec — viable for evaluation, not for production serving.Buying checklist for MoE workloads:
Need a GPU node sized for Mixtral, Qwen3-MoE, or DeepSeek-V3?
QSCompute configures and burn-in tests single- and multi-GPU MoE servers — H200, H100, B200, MI300X, and RTX PRO 6000 — matched to your model's total-parameter footprint and throughput target, with the NVLink/NVSwitch or InfiniBand fabric sized for expert parallelism.
Contact: +86 137-1464-6179 | sherry@qscompute.com