Published: August 30, 2026 | Category: Technical | QSCompute
"10,000+ tokens/s on Llama 2 70B" sounds impressive — but is that submission comparable to the one from the other vendor? MLPerf is the closest thing AI hardware has to a standardized benchmark, and it is the number quoted in almost every GPU procurement conversation in 2026. The problem is that most buyers read the headline number and skip the methodology that determines whether that number means anything for their workload.
This guide decodes MLPerf Inference results so you can compare H100, H200, B200, L40S and RTX 6000 Ada submissions fairly — and spot the cherry-picking before it costs you a purchase decision.
MLPerf (MLCommons) is a consortium-run benchmark suite. The Inference benchmark measures how fast a system runs a fixed set of models under strict rules. Two divisions exist:
Each MLPerf Inference submission reports up to four scenarios, and each maps to a different real-world deployment:
| Scenario | What It Measures | Real-World Workload | Metric |
|---|---|---|---|
| Offline | Maximum throughput, no latency constraint | Batch jobs: data processing, RAG indexing, offline translation | Samples/sec |
| Server | Throughput under a strict latency SLO (e.g. 99% of queries under 1s) | Cloud API serving with many concurrent users | Samples/sec at SLO |
| SingleStream | Latency of one query at a time | Single-camera inference, sequential edge processing | 90th-percentile latency |
| MultiStream | Batched inference on multiple parallel streams | Multi-camera vision, radar/LiDAR fusion, NVR analytics | Latency per stream batch |
This is why a GPU that "wins" Offline can lose Server: high Offline throughput usually means large batches, and large batches add latency. For interactive applications, the Server scenario number is the one that matters.
The table below shows representative published Closed-division ranges for Llama 2 70B-99 (8-GPU node, Offline and Server scenarios) from recent MLPerf rounds. Individual submissions vary by ±10–15% depending on software stack and system tuning — treat them as ranges, not absolutes.
| GPU (8-GPU node) | Llama2-70B Offline (samples/s) | Llama2-70B Server (samples/s @ 1s SLO) | Memory |
|---|---|---|---|
| NVIDIA H100 SXM | ≈ 11,000 – 13,000 | ≈ 1,900 – 2,300 | 80 GB HBM3 |
| NVIDIA H200 | ≈ 13,000 – 15,500 | ≈ 2,400 – 2,900 | 141 GB HBM3e |
| NVIDIA B200 | ≈ 20,000+ | ≈ 3,500+ | 192 GB HBM3e |
| NVIDIA L40S | ≈ 2,800 – 3,500 | ≈ 600 – 800 | 48 GB GDDR6 |
| NVIDIA RTX 6000 Ada | ≈ 1,800 – 2,300 | ≈ 400 – 550 | 48 GB GDDR6 |
The pattern worth internalizing: the gap between Offline and Server numbers is largest on the fastest hardware, because high-throughput GPUs are starved when latency limits batch size. If your workload is interactive, don't buy on Offline numbers.
| # | Check | Why It Matters |
|---|---|---|
| 1 | Same division (Closed vs Open) | Open allows model changes — not a hardware test |
| 2 | Same model + quality target (99 vs 99.9) | 99.9 submissions use more compute for accuracy |
| 3 | Same scenario | Offline ≠ Server; pick your deployment type |
| 4 | Same GPU count and node count | 8-GPU vs 4-GPU results are not comparable |
| 5 | Same precision path (FP8 vs FP16 vs INT8) | FP8 submissions are faster but may change quality |
| 6 | Power and thermal state disclosed | A 700W vs 350W cap changes results massively |
| 7 | Check the submitter's own notes | Software versions, TensorRT/compiler tunings are listed per submission |
MLPerf measures one thing: inference speed on a fixed model suite. It does not measure cost-per-inference, power-per-inference at your actual utilization, multi-tenant behavior under mixed workloads, or how the system performs on your model (especially custom architectures or LoRA adapters). A 2× MLPerf gap can evaporate — or invert — on a workload the benchmark suite doesn't cover. For that reason, treat MLPerf as a shortlist filter, then validate with your own benchmark on representative hardware before committing.
Ask your vendor for their own benchmark runs on your model — any GPU supplier that can't reproduce a benchmark on your actual workload should be treated with suspicion.
Need GPU benchmarks run on your actual model before you buy?
QSCompute configures and burn-tests NVIDIA GPU servers and edge AI systems — we run your workload on candidate hardware and hand you the numbers.
Contact: +86 137-1464-6179 | info@qscompute.com