How to Read MLPerf Results 2026 — GPU Benchmark Methodology & Buyer's Cheat Sheet

Published: August 30, 2026 | Category: Technical | QSCompute

"10,000+ tokens/s on Llama 2 70B" sounds impressive — but is that submission comparable to the one from the other vendor? MLPerf is the closest thing AI hardware has to a standardized benchmark, and it is the number quoted in almost every GPU procurement conversation in 2026. The problem is that most buyers read the headline number and skip the methodology that determines whether that number means anything for their workload.

This guide decodes MLPerf Inference results so you can compare H100, H200, B200, L40S and RTX 6000 Ada submissions fairly — and spot the cherry-picking before it costs you a purchase decision.

MLPerf 101: What It Actually Measures

MLPerf (MLCommons) is a consortium-run benchmark suite. The Inference benchmark measures how fast a system runs a fixed set of models under strict rules. Two divisions exist:

Rule #1: only compare Closed-division results with Closed-division results, same model, same scenario, same GPU count. An Open-division 8-GPU submission is not "4× faster" than a Closed-division one.

The Four Scenarios, Decoded

Each MLPerf Inference submission reports up to four scenarios, and each maps to a different real-world deployment:

ScenarioWhat It MeasuresReal-World WorkloadMetric
OfflineMaximum throughput, no latency constraintBatch jobs: data processing, RAG indexing, offline translationSamples/sec
ServerThroughput under a strict latency SLO (e.g. 99% of queries under 1s)Cloud API serving with many concurrent usersSamples/sec at SLO
SingleStreamLatency of one query at a timeSingle-camera inference, sequential edge processing90th-percentile latency
MultiStreamBatched inference on multiple parallel streamsMulti-camera vision, radar/LiDAR fusion, NVR analyticsLatency per stream batch

This is why a GPU that "wins" Offline can lose Server: high Offline throughput usually means large batches, and large batches add latency. For interactive applications, the Server scenario number is the one that matters.

Published Results on the GPUs Buyers Actually Quote

The table below shows representative published Closed-division ranges for Llama 2 70B-99 (8-GPU node, Offline and Server scenarios) from recent MLPerf rounds. Individual submissions vary by ±10–15% depending on software stack and system tuning — treat them as ranges, not absolutes.

GPU (8-GPU node)Llama2-70B Offline (samples/s)Llama2-70B Server (samples/s @ 1s SLO)Memory
NVIDIA H100 SXM≈ 11,000 – 13,000≈ 1,900 – 2,30080 GB HBM3
NVIDIA H200≈ 13,000 – 15,500≈ 2,400 – 2,900141 GB HBM3e
NVIDIA B200≈ 20,000+≈ 3,500+192 GB HBM3e
NVIDIA L40S≈ 2,800 – 3,500≈ 600 – 80048 GB GDDR6
NVIDIA RTX 6000 Ada≈ 1,800 – 2,300≈ 400 – 55048 GB GDDR6

The pattern worth internalizing: the gap between Offline and Server numbers is largest on the fastest hardware, because high-throughput GPUs are starved when latency limits batch size. If your workload is interactive, don't buy on Offline numbers.

The 7-Point Fair-Comparison Checklist

#CheckWhy It Matters
1Same division (Closed vs Open)Open allows model changes — not a hardware test
2Same model + quality target (99 vs 99.9)99.9 submissions use more compute for accuracy
3Same scenarioOffline ≠ Server; pick your deployment type
4Same GPU count and node count8-GPU vs 4-GPU results are not comparable
5Same precision path (FP8 vs FP16 vs INT8)FP8 submissions are faster but may change quality
6Power and thermal state disclosedA 700W vs 350W cap changes results massively
7Check the submitter's own notesSoftware versions, TensorRT/compiler tunings are listed per submission

What MLPerf Does Not Tell You

MLPerf measures one thing: inference speed on a fixed model suite. It does not measure cost-per-inference, power-per-inference at your actual utilization, multi-tenant behavior under mixed workloads, or how the system performs on your model (especially custom architectures or LoRA adapters). A 2× MLPerf gap can evaporate — or invert — on a workload the benchmark suite doesn't cover. For that reason, treat MLPerf as a shortlist filter, then validate with your own benchmark on representative hardware before committing.

Using MLPerf in a 2026 GPU Purchase

Ask your vendor for their own benchmark runs on your model — any GPU supplier that can't reproduce a benchmark on your actual workload should be treated with suspicion.

Need GPU benchmarks run on your actual model before you buy?

QSCompute configures and burn-tests NVIDIA GPU servers and edge AI systems — we run your workload on candidate hardware and hand you the numbers.

Contact: +86 137-1464-6179 | info@qscompute.com