Published: August 30, 2026 | Category: Buying Guide | QSCompute
The three most-requested GPUs on Q3 2026 procurement lists are not three generations apart — they are three different answers to three different jobs. The A100 80GB (Ampere, 2020) is now a used-market value play with a huge HBM2e memory pool. The H100 80GB (Hopper, 2022) is the data-center workhorse for FP8 training and serving. The RTX 6000 Ada (48GB) is the air-cooled professional card that fits in a workstation or a 2U edge server and quietly serves most 8B–70B inference workloads. This guide compares them spec-by-spec and maps each one to the workloads where it is the right purchase in 2026.
| Specification | A100 80GB SXM | H100 80GB SXM | H100 80GB PCIe | RTX 6000 Ada |
|---|---|---|---|---|
| Architecture | Ampere (GA100) | Hopper (GH100) | Hopper (GH100) | Ada Lovelace (AD102) |
| Memory | 80 GB HBM2e | 80 GB HBM3 | 80 GB HBM2e | 48 GB GDDR6 ECC |
| Memory bandwidth | 2.0 TB/s | 3.35 TB/s | 2.0 TB/s | 960 GB/s |
| FP16 compute (dense) | 312 TFLOPS | 989 TFLOPS | 756 TFLOPS | 145.7 TFLOPS |
| FP8 support | No | Yes (1,979 TOPS) | Yes (1,513 TOPS) | Yes (291 TOPS) |
| NVLink | 600 GB/s (SXM) | 900 GB/s (SXM) | None (PCIe 5.0 x16) | Bridge, 2-GPU only |
| TDP | 400 W | 700 W | 350 W | 300 W |
| Typical price, Q3 2026 | $6,000–$9,000 (used/refurb) | $25,000–$35,000 | $20,000–$28,000 | $6,000–$7,500 |
The single most important column is FP8. Ampere has no FP8 path — A100 buyers run INT8 or 4-bit software quantization (GPTQ/AWQ). Hopper and Ada both have native FP8 tensor cores, which roughly doubles inference throughput versus FP16 and is the lossless production default for LLM serving in 2026. Everything else in this comparison follows from that one feature gap.
| Workload | Recommended GPU | Reasoning |
|---|---|---|
| LLM pre-training / dense multi-GPU training | H100 or H200 SXM (HGX node) | NVSwitch all-to-all fabric, FP8, 700 W needs data-center cooling |
| 70B-class fine-tuning (LoRA/QLoRA) | 2× RTX 6000 Ada or 2× L40S | 48 GB holds a 4-bit 70B; PCIe bandwidth is not the bottleneck |
| 70B FP8 inference, production | H100 PCIe or L40S | Native FP8, high tokens/s, no NVSwitch cost |
| 8B–13B inference at scale | RTX 6000 Ada / L40S / A100 | VRAM- and bandwidth-bound; one model per card |
| Diffusion, vision, edge deployment | RTX 6000 Ada | 300 W air-cooled, fits workstations and edge servers |
| Budget or long-context serving (used) | A100 80GB | 80 GB HBM2e at ~$7k is the best $/GB in the market; no FP8, so INT8/4-bit only |
The A100's second life is the used market. A100 80GB cards are out of production, but refurbished units at $6,000–$9,000 are the cheapest large-VRAM pool available — 80 GB of HBM2e at 2.0 TB/s still serves 32B–70B models at 4-bit comfortably. The catch: no FP8, 400 W TDP, and SXM parts need an HGX-era motherboard. For INT8/4-bit serving fleets on a budget, the A100 remains defensible; for anything FP8, it is the wrong purchase regardless of price.
H100 is a power and cooling decision, not just a silicon decision. SXM modules run 700 W — eight of them is an 8 kW node that realistically wants liquid cooling or a very well-ventilated rack. The H100 PCIe at 350 W trades about 25% of the compute for air-coolability and chassis flexibility, which is why it is the more common choice for inference nodes. Budgets that cannot absorb the power and cooling infrastructure should look at L40S or RTX 6000 Ada nodes first.
RTX 6000 Ada is the quiet workhorse. 48 GB of ECC GDDR6, FP8, 300 W — it drops into a workstation or a 2U edge server, runs a 4-bit 70B model by itself, and pairs with a second card via NVLink bridge when you need 96 GB. At $6,000–$7,500 it is frequently the correct answer to "we need to serve LLMs in a facility that is not a data center." The RTX PRO 6000 Blackwell (96 GB) is the natural upgrade path when budgets allow.
Rent before you buy. H100 cloud rental rates have been falling through 2026; for workloads under roughly 60% utilization, renting beats owning. The buy case is strongest for sustained 24/7 serving and for facilities where data egress or security rules make the cloud impractical.
Not sure which GPU your workload actually needs?
QSCompute supplies all four configurations — refurbished A100 nodes, H100/H200 HGX systems, L40S servers, and RTX 6000 Ada workstations — each burn-in tested and pre-loaded with a serving stack tuned to your model. Tell us the workload and the budget; we'll tell you which of these three cards is the right one.
Contact: +86 137-1464-6179 | info@qscompute.com