Published: August 28, 2026 | Category: Technical | QSCompute
Every AI team hits the same fork: rent GPU time from a cloud provider, or buy servers and run them yourself. In 2025 the answer was heavily skewed toward renting — H100s were scarce and cloud pricing was the only game in town. In late 2026 the equation has flipped for many workloads: GPU supply has loosened (H200 and Blackwell ramp, plus a wave of L40S/RTX 6000 Ada surplus from data-center churn), on-prem hardware is cheaper per hour than it has ever been, and cloud spot prices are volatile. This guide puts real numbers on both sides so you can compute your break-even, not a generic rule of thumb.
Prices below are representative per-GPU-hour rates across major providers (AWS, Azure, GCP, Lambda, CoreWeave, RunPod, Vast) for on-demand, 1-year reserved, and spot/interruptible tiers, late August 2026. Spot prices swing widely — figures shown are typical midpoints.
| GPU | On-demand / hr | 1-yr reserved / hr | Spot / hr | Notes |
|---|---|---|---|---|
| H100 SXM 80GB | $2.80–$3.50 | $2.00–$2.40 | $1.20–$1.80 | Training workhorse; 8-GPU nodes $22–28/hr |
| H200 141GB | $3.50–$4.50 | $2.60–$3.20 | $1.80–$2.50 | Premiums shrinking as H200 supply normalizes |
| A100 80GB | $1.80–$2.40 | $1.30–$1.60 | $0.80–$1.20 | Aging but still the budget training option |
| L40S 48GB | $1.20–$1.80 | $0.90–$1.20 | $0.50–$0.90 | Best price/performance for 70B-class FP8 inference |
| RTX 4090 24GB | $0.50–$0.80 | $0.35–$0.55 | $0.20–$0.40 | Consumer-class; check ECC and interconnect limits |
| RTX 6000 Ada 48GB | $1.00–$1.50 | $0.75–$1.00 | $0.45–$0.75 | Workstation-class inference; ECC + pro drivers |
These prices exclude egress (data transfer out can add 5–15% for training-heavy jobs), storage (file systems, checkpoints), and orchestration overhead. A realistic all-in cloud cost for an H100 is 15–25% above the raw GPU rate.
Buying is not just the server price. The honest model over a 3-year horizon includes hardware, facility (power + cooling + space), labor, and the capital cost of money. For a typical 8× H100 SXM node:
| Cost component | 3-year total (8× H100 node) | Per-GPU-hour (at 60% util) |
|---|---|---|
| Hardware (node, storage, networking) | ~$280,000–$340,000 | ~$0.85–$1.05 |
| Power & cooling (~11 kW/node avg) | ~$80,000–$110,000 | ~$0.25–$0.35 |
| Facility space (rack, colo or DC share) | ~$25,000–$45,000 | ~$0.08–$0.14 |
| Admin labor (fractional DevOps/ops) | ~$40,000–$70,000 | ~$0.12–$0.22 |
| Total | ~$425,000–$565,000 | ~$1.30–$1.75 |
At 60% utilization an owned H100 costs roughly $1.30–1.75/GPU-hour — cheaper than on-demand ($2.80+) but comparable to 1-year reserved. The moment utilization drops, the math reverses fast: at 25% utilization the same node costs $3.10–4.20/GPU-hour, worse than spot.
| Workload | Typical utilization | Verdict | Rationale |
|---|---|---|---|
| LLM fine-tuning / RL runs | 70–95% while active, bursty | Hybrid | Buy the steady base fleet, rent burst peaks on spot |
| Pre-training large models | 90%+ for months | Buy | This is where on-prem breaks even fastest; also check H200 economics |
| Production inference serving | 40–70% steady | Buy (right-sized) | Predictable traffic → own L40S/RTX 6000 Ada-class; autoscale to cloud for spikes |
| Dev / CI / experimentation | 10–30% | Rent | Spot GPUs or shared dev nodes; never buy for this |
| Batch jobs, sweep grid | Bursty, interruptible OK | Rent spot | Spot discounts of 50–60% make this a no-brainer |
| Edge inference (factory, retail) | Variable | Neither — buy edge | Cloud round-trip latency and data egress kill the economics; see edge GPU guide |
Three things changed the rent-vs-buy calculus this year:
Decided to buy? Get the full picture before you sign.
QSCompute builds and burn-in tests on-prem AI nodes — H100, H200, L40S, RTX 6000 Ada — with storage, networking and power budgeting sized to your real utilization. Tell us your workload and we'll quote the owned fleet that beats your cloud bill.
Contact: +86 137-1464-6179 | info@qscompute.com