Published: August 27, 2026 | Category: Technical | QSCompute
Most AI servers run at full TDP all day — and waste a meaningful fraction of it. On inference workloads, capping a 700 W H100 to 500 W typically costs 3–5% throughput while cutting power draw by ~28%. On a 100-GPU fleet at $0.12/kWh, that's tens of thousands of dollars a year in electricity alone, before counting the smaller PSUs, lighter cooling, and longer component life. This guide covers the three levers — power limits (nvidia-smi -pl), undervolting, and fleet-wide control via DCGM — with real numbers for the GPUs QSCompute ships.
GPUs are most efficient at 60–80% of TDP. Boost clocks scale sub-linearly with power: near the top of the curve, a 10% power increase buys only 2–4% frequency. Because inference is largely memory-bandwidth-bound rather than compute-bound, an H100 or L40S running Llama-class or YOLO-class workloads barely notices a 20–30% power reduction — the SM clocks drop, but memory bandwidth is what carries the throughput.
Training is the opposite: dense GEMMs push SMs to the wall, so the same cap costs more. Rule of thumb: power capping fits inference fleets, 24/7 serving nodes, and multi-tenant GPU boxes; it's a poor fit for throughput-critical training clusters.
| GPU | Default TDP | Practical Cap | Power Saved | Inference Perf Impact |
|---|---|---|---|---|
| H100 SXM (700W) | 700 W | 500 W | ~28% | 3–5% |
| H100 PCIe (350W) | 350 W | 275 W | ~21% | 3–5% |
| L40S (350W) | 350 W | 250 W | ~29% | 4–6% |
| RTX 6000 Ada (300W) | 300 W | 220 W | ~27% | 4–6% |
| RTX 4090 (450W) | 450 W | 300 W | ~33% | 5–8% (undervolt recovers most) |
| A2000 (70W) | 70 W | 55 W | ~21% | 2–4% |
Numbers are representative ranges from production inference fleets (vLLM, TensorRT-LLM, Triton, YOLO pipelines) — validate on your workload, since the mix of memory-bound vs compute-bound kernels moves the impact by a few points.
# Check current limits and supported range nvidia-smi -q -d POWER # Set a persistent power limit (Watts) sudo nvidia-smi -pl 500 # Lock clocks after capping (optional, keeps perf predictable) sudo nvidia-smi -lgc 1500,1900 # Make the limit persist across reboots sudo nvidia-smi --persistence-mode=1
Set the cap before loading the workload — changing it live forces a GPU re-initialization that can stall running jobs. For multi-GPU nodes, apply per-device limits with -i, or better, manage everything through DCGM so limits survive driver updates and reboots.
Power capping tells the GPU how much it may draw; undervolting improves how efficiently it converts that power into clock. On data-center GPUs the accessible form is a locked clock range (nvidia-smi -lgc) paired with a cap — the GPU then runs its target clock at the lowest voltage that sustains it. On workstation/consumer parts (RTX 4090, RTX 6000 Ada), a voltage-frequency curve edit can hold 95%+ performance at 300 W instead of 450 W — the single biggest lever for fanless and compact builds.
For anything beyond a handful of servers, set caps through NVIDIA DCGM (Data Center GPU Manager) rather than ad-hoc nvidia-smi calls:
A practical deployment pattern: run uncapped for one week, collect DCGM power telemetry, then set each node's cap at the 85th percentile of observed power. Most fleets land at 70–80% of TDP and lose almost nothing.
Whether you're running one edge inference box or a 100-GPU fleet, power capping is the cheapest capacity upgrade available — it costs nothing, takes an afternoon, and typically recovers 20–35% of your power bill for under 5% throughput. QSCompute ships every GPU server with power-management pre-configured and can tune caps for your specific workloads.
Want your GPU servers tuned for efficiency before they ship?
QSCompute builds and configures GPU servers and edge AI systems — power capping, thermal validation, and DCGM monitoring included.
Contact: +86 137-1464-6179 | info@qscompute.com