Published: September 9, 2026 | Category: Buying Guide | QSCompute
The most common GPU-server question we hear is no longer about training — it is about desktops. Engineering teams want their CAD and 3D applications on thin clients in branch offices or factory floors; OT teams want Windows HMI sessions virtualized instead of bolted to aging PCs beside every machine; IT wants the data to never leave the building. All three converge on the same architecture: a GPU server running virtual desktops, with one NVIDIA license per concurrent user. This guide sizes that server — how many users per GPU, which sharing model, and what the software really costs in 2026.
Every GPU-virtualization decision starts with the sharing model. NVIDIA's official licensing FAQ is blunt: a passthrough deployment is enforced through the EULA rather than per-CCU software, which is why 1:1 passthrough is the license-free default — but it is also the least dense option, and workstation-grade features still need an RTX Virtual Workstation entitlement.
| Sharing model | Typical density | Isolation | Best for | License |
|---|---|---|---|---|
| GPU passthrough / DDA | 1 VM per GPU | Full (dedicated VRAM) | Max performance, GPU compute in VMs, legacy apps | None enforced (EULA); vWS features require vWS |
| vGPU time-slicing (vPC) | 8–24 desktops per 48 GB GPU | Medium (shared SMs, partitioned VRAM) | Knowledge workers, HMI/SCADA clients, office + video | vPC — $50 per CCU / year |
| vGPU time-slicing (vWS) | 3–6 workstations per 48 GB GPU | Medium (shared SMs, partitioned VRAM) | CAD/3D, ISV-certified graphics, data science | vWS — $250 per CCU / year |
| MIG hardware partitioning | Up to 7 slices (A100/H100-class) | High (hardware-isolated) | Multi-tenant compute + inference, not graphics | Compute workloads; no per-CCU vGPU fee |
MIG and vGPU cover different jobs: MIG partitions an A100/H100 for compute tenants, while vGPU multiplexes a graphics card for desktop users — our earlier MIG/vGPU deep-dive covers the node-level detail. The desktop choice is therefore vPC vs vWS, and it is driven almost entirely by the application list, not by the GPU.
Density starts from the vGPU profile: vPC desktops commonly run 2–4 GB profiles and vWS workstations 8–16 GB, so a 48 GB card has a framebuffer ceiling of roughly 24 desktop profiles or 3–6 workstation profiles. NVIDIA's own sizing guidance adds two caveats: concurrent-vGPU limits per GPU, and the fact that time-slicing shares the same SMs — pack 24 users onto one L40S and every user feels the contention at login and in video. A defensible planning number is 8–12 vPC users or 3–4 vWS users per 48 GB Ada GPU for fluid interactive work.
| GPU | VRAM | vPC (2–4 GB profile) | vWS (8–16 GB profile) | 2026 positioning |
|---|---|---|---|---|
| RTX 6000 Ada | 48 GB | 12–24 ceiling; 8–12 practical | 3–6 | 960 GB/s, 300 W — the balanced VDI workhorse |
| L40S | 48 GB | 12–24 ceiling; 8–12 practical | 3–6 | Data-center build, 350 W, same framebuffer math |
| RTX PRO 6000 Blackwell | 96 GB | Up to 48 concurrent vGPUs per NVIDIA | 6–12 | Highest density; 2× VRAM halves the GPU count |
| A16 (legacy density card) | 64 GB | 16–32 | — | Still found in recertified fleets; no FP8, aging NVENC |
Plan 1.5–2× the CPU and RAM you think you need: every vPC desktop wants 2 vCPU and 4 GB, every vWS session 4–8 vCPU and 16–32 GB. A 2-socket server with 1 TB RAM typically hosts 4× 48 GB GPUs — 32–48 vPC users or 12–16 vWS users per 4U chassis.
NVIDIA sells vGPU software per concurrent user (CCU), not per GPU. Current published pricing: vApps $10, vPC $50, and RTX vWS $250 per CCU per year on annual subscription (4-year and 5-year terms discount to $43.75/$40 and $206.25/$180). Perpetual licenses are $20/$100/$450 with $5/$25/$100 annual SUMS. The license often outlives the GPU's depreciation — a 40-engineer pilot with 20 concurrent vWS seats is $5,000/year before a single server is bought, versus $1,000/year if vPC covers them.
| Edition | Annual sub (per CCU) | Perpetual + SUMS/yr | What it entitles |
|---|---|---|---|
| vApps | $10 | $20 + $5 | RDSH-hosted apps, no full desktop |
| vPC | $50 | $100 + $25 | Virtual desktops, office + video, 5K display |
| RTX vWS | $250 | $450 + $100 | ISV-certified 3D/CAD, 8K, compute virtualization |
On the client side, thin clients run $150–500 per seat and a 1080p desktop streams in roughly 1.5–3 Mbps with H.264 — meaning a 1 Gbps uplink comfortably serves 200+ concurrent sessions and the network is rarely the constraint. The constraint is always the same: license cost per seat versus GPU density, which is why the vPC/vWS split deserves a written policy before the PO.
Sizing a VDI or remote-workstation deployment?
QSCompute supplies pre-built GPU VDI servers on L40S, RTX 6000 Ada and RTX PRO 6000 Blackwell — 1U to 8-GPU 4U chassis, validated with VMware, Proxmox and Hyper-V, with license-forecast sheets so the software line item is in the budget before you sign. Send us your user count, application list and concurrency target for a density calculation.
Contact: +86 137-1464-6179 | info@qscompute.com