Published: August 24, 2026 | Category: Buying Guide | QSCompute
Spec sheets lead with TFLOPS, but for AI inference the number that actually determines throughput is the one printed further down the page: memory bandwidth. A large language model generates a single token by streaming the entire set of weights through the GPU on every forward pass. The arithmetic finishes in microseconds; it is the memory subsystem that cannot keep up. That gap is the memory wall — and it, not compute, is why a 4.8 TB/s NVIDIA H200 serves a 70B model nearly twice as fast as a 3.35 TB/s H100 that has almost identical FP16 math throughput.
This guide explains how to read GPU memory specifications, when HBM beats GDDR, how much VRAM different model sizes actually need, and how to match a GPU to your workload without paying for silicon you cannot feed.
Token generation is memory-bound. On every token, the GPU streams every weight plus the growing KV cache through the compute units. For a 70B model in FP16 (~140 GB), a single token reads on the order of 140 GB from memory. The multiply-accumulate work is trivial next to moving that data.
The practical rule of thumb: tokens/sec ≈ memory bandwidth ÷ bytes per token. This is why bandwidth, not TFLOPS, tracks real-world LLM throughput. Over the last five years GPU compute (TFLOPS) has scaled roughly 30× while memory bandwidth has scaled only about 3× — the widening gap is the memory wall, and it makes bandwidth the binding constraint for inference.
| Memory Type | Typical Bandwidth | Found On | Cost Profile |
|---|---|---|---|
| HBM3e | 4.8–8 TB/s per stack | H200, B200, MI300X | Highest, supply-constrained |
| HBM3 | 3.35 TB/s | H100 SXM | High |
| GDDR7 | 960 GB/s–1.79 TB/s | RTX 5090, RTX PRO 6000 | Mid |
| GDDR6X | ~1.0 TB/s | RTX 4090 | Mid |
| GDDR6 (ECC) | 300–960 GB/s | L40S, L4, A2, RTX 6000 Ada | Lowest |
HBM stacks memory dies vertically beside the GPU die, delivering massive bandwidth and capacity at high cost and constrained supply — it is effectively data-center-only. GDDR is off-package, cheaper, more available, but lower-bandwidth. For a 70B+ LLM that must stay resident in VRAM, HBM is effectively mandatory. For edge inference (7–13B models, vision, small generative workloads), GDDR on the L40S, RTX 6000 Ada, or RTX 5090 is the cost-optimal choice.
| GPU | Memory | Capacity | Bandwidth | Best For |
|---|---|---|---|---|
| H200 SXM | HBM3e | 141 GB | 4.8 TB/s | 70B+ LLMs, high-QPS serving |
| H100 SXM | HBM3 | 80 GB | 3.35 TB/s | 70B (quantized), training |
| B200 | HBM3e | 192 GB | 8 TB/s | 405B-class frontier models |
| AMD MI300X | HBM3 | 192 GB | 5.3 TB/s | 70B+ open-model serving |
| RTX PRO 6000 Blackwell | GDDR7 | 96 GB | 1.79 TB/s | 70B (INT8), workstation inference |
| RTX 6000 Ada | GDDR6 ECC | 48 GB | 960 GB/s | 13–34B, vision, fine-tuning |
| L40S | GDDR6 ECC | 48 GB | 864 GB/s | 8–13B serving, video |
| RTX 5090 | GDDR7 | 32 GB | 1.79 TB/s | 8–13B, SDXL, workstation |
| L4 | GDDR6 | 24 GB | 300 GB/s | 7B INT8, video analytics |
The sizing math is simple once you separate weights from the rest. Weights = parameters × bytes-per-parameter (FP16 = 2, INT8/FP8 = 1, FP4/INT4 = 0.5). Then add the KV cache (which grows with context length) and a buffer for activations and framework overhead.
| Model (params) | FP16 Weights | INT8 | FP4/INT4 | Comfortable GPU |
|---|---|---|---|---|
| 3B (Llama 3.2 3B) | 6 GB | 3 GB | 1.5 GB | Jetson AGX Orin (64 GB), L4 |
| 7–8B (Llama 3.1 8B) | 16 GB | 8 GB | 4 GB | L4 (24 GB), RTX 5090 |
| 13B (Llama 2 13B) | 26 GB | 13 GB | 6.5 GB | RTX 6000 Ada (48 GB), L40S |
| 34B | 68 GB | 34 GB | 17 GB | RTX PRO 6000 (96 GB) |
| 70B (Llama 3.1 70B) | 140 GB | 70 GB | 35 GB | H100 (INT8), 2× RTX PRO 6000, H200 |
| 405B (Llama 3.1 405B) | 810 GB | 405 GB | ~200 GB | B200 / multi-H200 |
A practical checklist for procurement teams:
When a model exceeds a single GPU's VRAM you have three levers: a larger-VRAM GPU (H200 for 70B+), multi-GPU with NVLink/NVSwitch for tensor parallelism, or quantization. For most edge and mid-tier inference, a GDDR GPU is the right economic answer; HBM earns its premium only when the model is large enough that offloading would dominate latency.
Need a pre-built GPU inference server with the right memory configuration?
QSCompute configures and burn-in tests HBM and GDDR GPU nodes — H100/H200, L40S, RTX 6000 Ada, and RTX PRO 6000 — matched to your model size and throughput target, with full documentation and a power/thermal test report.
Contact: +86 137-1464-6179 | sherry@qscompute.com