Published: August 28, 2026 | Category: Technical | QSCompute
When a B2B buyer specs a GPU server for LLM inference, the first question is almost always "how many GPUs, and which one?" The answer changes dramatically depending on quantization. A 70B model that appears to require two H100s at FP16 can run on a single 48 GB card at 4-bit with a 1–3% quality difference — that is a roughly $20,000 purchasing swing driven entirely by software choices. This guide covers the four formats that matter in 2026 — GPTQ, AWQ, FP8 and NVFP4 — and maps them onto the hardware you actually buy.
LLM inference is bound by VRAM capacity and memory bandwidth, not raw FLOPs. Quantization attacks both: it shrinks weights so more model fits per card, and it cuts bytes-per-token so each card serves tokens faster. For a buyer, the practical result is that the quantized model size — not the published FP16 size — decides the GPU count and generation. Get this wrong and you either over-buy (GPUs that sit half-empty) or under-buy (a model that does not fit in memory at all). Get it right and a quantized deployment routinely cuts hardware cost 2–4× at near-identical output quality.
| Format | Bit width | Method | Calibration data | Native hardware | Main serving stacks |
|---|---|---|---|---|---|
| GPTQ | 2–8 bit (4-bit standard) | Weight-only, layer-wise error minimization | Yes, ~128–512 samples | Ampere and newer (CUDA kernels) | vLLM, TensorRT-LLM, AutoGPTQ, ExLlamaV2 |
| AWQ | 4-bit (3–8 bit configurable) | Activation-aware weight quantization — protects salient weights | Yes, ~128–512 samples | Ampere and newer (AWQ-Marlin kernels) | vLLM, TensorRT-LLM, SGLang |
| FP8 | 8-bit (E4M3/E5M2) | Native floating-point, no re-training | Optional (static scaling) | Hopper (H100/H200), Ada (L40S, RTX 6000 Ada) | TensorRT-LLM, vLLM, PyTorch |
| NVFP4 | 4-bit (FP4 + block scale) | Native floating-point with shared per-block scale | No | Blackwell (B200, RTX PRO 6000), GB10 (DGX Spark) | TensorRT-LLM (Blackwell) |
GPTQ and AWQ are post-training weight-only methods: you run them once on a copy of the model with a small calibration set, and the result is a smaller file you deploy. AWQ generally holds quality slightly better than GPTQ at 4-bit on larger models, at the cost of slightly more complex kernels. FP8 is the lossless production default on Hopper and Ada — no calibration, no quality debate, native tensor-core acceleration. NVFP4 is Blackwell's answer: 4-bit floating point with a per-block scale, roughly 2× the throughput of FP8 on B200-class parts.
Weights-only figures below (add 10–30% for KV cache and activations depending on context length). A 128K-token KV cache on a 70B model can add 10–16 GB at FP16, which is why KV-cache quantization (FP8 or INT8 keys/values) matters in production — most serving stacks now support it.
| Model | FP16 weights | INT8 | INT4 | Typical fit (weights + KV) |
|---|---|---|---|---|
| Llama 3.1 8B | ~16 GB | ~8 GB | ~5 GB | Single 24 GB card (RTX A5000, L4, RTX 4090) |
| Qwen2.5-14B | ~29 GB | ~15 GB | ~8 GB | Single 24 GB card at INT4 |
| Llama 3.1 70B | ~140 GB | ~70 GB | ~40 GB | Single 48 GB card at INT4 (L40S, RTX 6000 Ada, A6000) |
| Llama 3.1 405B | ~810 GB | ~405 GB | ~203 GB | 8× H100/H200 at FP8/INT4, or 4× 96 GB RTX PRO 6000 Blackwell |
| DeepSeek-V3 (671B, MoE) | ~1.3 TB | ~670 GB | ~350 GB | 8× 80 GB H100/H800-class at FP8 (the production reference deployment) |
| GPU generation | Native formats | Notes for buyers |
|---|---|---|
| Ampere (A100, A30, RTX A6000) | INT8 tensor cores | No FP8 — 4-bit runs via GPTQ/AWQ software kernels only; fine for INT8, but FP8 serving is a Hopper/Ada feature |
| Hopper (H100, H200) | INT8 + FP8 | The data-center FP8 workhorse; 2× FP8 vs FP16 tensor throughput |
| Ada (L40S, RTX 6000 Ada, RTX 4000 SFF) | INT8 + FP8 | Same FP8 path in PCIe and fanless-friendly form factors — the popular "budget H100" for quantized serving |
| Blackwell (B200, RTX PRO 6000, DGX Spark GB10) | INT8 + FP8 + FP4/NVFP4 | FP4 doubles throughput again; GB10-class parts ship FP4 TOPS natively |
This table is the real buying argument: if your roadmap is FP8 serving, Ampere cards (however cheap) are the wrong purchase — you pay for software-emulated FP8 or settle for INT8. Conversely, if you plan 4-bit GPTQ/AWQ, Ampere remains perfectly serviceable because those methods run in software on any CUDA GPU.
Spec'ing a GPU server for LLM inference?
QSCompute configures and burn-in tests data-center GPU nodes — H100, H200, B200, L40S, RTX 6000 Ada — with TensorRT-LLM and vLLM stacks pre-tuned for your quantized model. Tell us the model and the quality bar; we'll tell you the fewest GPUs that meet it.
Contact: +86 137-1464-6179 | info@qscompute.com