LLM Quantization for GPU Deployment 2026 — GPTQ vs AWQ vs FP8 vs NVFP4

Published: August 28, 2026 | Category: Technical | QSCompute

When a B2B buyer specs a GPU server for LLM inference, the first question is almost always "how many GPUs, and which one?" The answer changes dramatically depending on quantization. A 70B model that appears to require two H100s at FP16 can run on a single 48 GB card at 4-bit with a 1–3% quality difference — that is a roughly $20,000 purchasing swing driven entirely by software choices. This guide covers the four formats that matter in 2026 — GPTQ, AWQ, FP8 and NVFP4 — and maps them onto the hardware you actually buy.

Why Quantization Is a Hardware Purchase Decision

LLM inference is bound by VRAM capacity and memory bandwidth, not raw FLOPs. Quantization attacks both: it shrinks weights so more model fits per card, and it cuts bytes-per-token so each card serves tokens faster. For a buyer, the practical result is that the quantized model size — not the published FP16 size — decides the GPU count and generation. Get this wrong and you either over-buy (GPUs that sit half-empty) or under-buy (a model that does not fit in memory at all). Get it right and a quantized deployment routinely cuts hardware cost 2–4× at near-identical output quality.

The Four Formats That Matter in 2026

FormatBit widthMethodCalibration dataNative hardwareMain serving stacks
GPTQ2–8 bit (4-bit standard)Weight-only, layer-wise error minimizationYes, ~128–512 samplesAmpere and newer (CUDA kernels)vLLM, TensorRT-LLM, AutoGPTQ, ExLlamaV2
AWQ4-bit (3–8 bit configurable)Activation-aware weight quantization — protects salient weightsYes, ~128–512 samplesAmpere and newer (AWQ-Marlin kernels)vLLM, TensorRT-LLM, SGLang
FP88-bit (E4M3/E5M2)Native floating-point, no re-trainingOptional (static scaling)Hopper (H100/H200), Ada (L40S, RTX 6000 Ada)TensorRT-LLM, vLLM, PyTorch
NVFP44-bit (FP4 + block scale)Native floating-point with shared per-block scaleNoBlackwell (B200, RTX PRO 6000), GB10 (DGX Spark)TensorRT-LLM (Blackwell)

GPTQ and AWQ are post-training weight-only methods: you run them once on a copy of the model with a small calibration set, and the result is a smaller file you deploy. AWQ generally holds quality slightly better than GPTQ at 4-bit on larger models, at the cost of slightly more complex kernels. FP8 is the lossless production default on Hopper and Ada — no calibration, no quality debate, native tensor-core acceleration. NVFP4 is Blackwell's answer: 4-bit floating point with a per-block scale, roughly 2× the throughput of FP8 on B200-class parts.

VRAM Math — What Actually Fits on Which GPU

Weights-only figures below (add 10–30% for KV cache and activations depending on context length). A 128K-token KV cache on a 70B model can add 10–16 GB at FP16, which is why KV-cache quantization (FP8 or INT8 keys/values) matters in production — most serving stacks now support it.

ModelFP16 weightsINT8INT4Typical fit (weights + KV)
Llama 3.1 8B~16 GB~8 GB~5 GBSingle 24 GB card (RTX A5000, L4, RTX 4090)
Qwen2.5-14B~29 GB~15 GB~8 GBSingle 24 GB card at INT4
Llama 3.1 70B~140 GB~70 GB~40 GBSingle 48 GB card at INT4 (L40S, RTX 6000 Ada, A6000)
Llama 3.1 405B~810 GB~405 GB~203 GB8× H100/H200 at FP8/INT4, or 4× 96 GB RTX PRO 6000 Blackwell
DeepSeek-V3 (671B, MoE)~1.3 TB~670 GB~350 GB8× 80 GB H100/H800-class at FP8 (the production reference deployment)
The row to memorize: a 4-bit 70B model fits on a single 48 GB card. Two years ago that workload "required" a multi-GPU H100 node; today a single L40S or RTX 6000 Ada serves it. For 405B-class and MoE giants, 8× H100/H200 remains the workhorse — which is exactly why FP8 (native, lossless) is the default format there, not 4-bit.

Quality and Speed Trade-Offs in Production

Hardware Support — Which GPU Generation Runs What

GPU generationNative formatsNotes for buyers
Ampere (A100, A30, RTX A6000)INT8 tensor coresNo FP8 — 4-bit runs via GPTQ/AWQ software kernels only; fine for INT8, but FP8 serving is a Hopper/Ada feature
Hopper (H100, H200)INT8 + FP8The data-center FP8 workhorse; 2× FP8 vs FP16 tensor throughput
Ada (L40S, RTX 6000 Ada, RTX 4000 SFF)INT8 + FP8Same FP8 path in PCIe and fanless-friendly form factors — the popular "budget H100" for quantized serving
Blackwell (B200, RTX PRO 6000, DGX Spark GB10)INT8 + FP8 + FP4/NVFP4FP4 doubles throughput again; GB10-class parts ship FP4 TOPS natively

This table is the real buying argument: if your roadmap is FP8 serving, Ampere cards (however cheap) are the wrong purchase — you pay for software-emulated FP8 or settle for INT8. Conversely, if you plan 4-bit GPTQ/AWQ, Ampere remains perfectly serviceable because those methods run in software on any CUDA GPU.

Practical Recommendations When You Buy

Spec'ing a GPU server for LLM inference?

QSCompute configures and burn-in tests data-center GPU nodes — H100, H200, B200, L40S, RTX 6000 Ada — with TensorRT-LLM and vLLM stacks pre-tuned for your quantized model. Tell us the model and the quality bar; we'll tell you the fewest GPUs that meet it.

Contact: +86 137-1464-6179 | info@qscompute.com