Published: August 4, 2026 | Category: Technical | QSCompute
A single H100 GPU sitting idle at 35% utilization on a factory edge server is wasted capital. Three adjacent production lines — AOI inspection, predictive maintenance inference, and worker safety monitoring — each need GPU compute, but not a full card. GPU virtualization splits one physical GPU into isolated, guaranteed-performance slices that run concurrent production workloads without crosstalk. For multi-tenant 边缘AI deployments — where uptime, deterministic latency, and workload isolation are non-negotiable — GPU virtualization turns a single edge server into a secure multi-workload inference hub. This guide compares NVIDIA's three GPU virtualization technologies (MIG, vGPU, MPS), their hard isolation guarantees, real performance overhead, and deployment patterns for factory AI.
| Technology | Isolation Type | GPU Support | Granularity | Fault Isolation | Memory QoS | CUDA Version per Tenant |
|---|---|---|---|---|---|---|
| NVIDIA MIG (Multi-Instance GPU) |
Hardware — physically partitioned SMs, memory, L2 cache, and memory controllers | A100, A30, H100, H200, B200 | Fractional GPU slices (e.g., 1g.10gb = 1/7 of A100) | Complete: one instance crash does not affect others | Guaranteed dedicated HBM bandwidth per slice | Independent per MIG instance |
| NVIDIA vGPU (Virtual GPU) |
Software (hypervisor-mediated) — NVIDIA GRID driver with SR-IOV-like scheduling | L40S, A16, A40, RTX 6000 Ada, H100 (non-MIG) | Time-sliced — equal share of GPU time, not dedicated hardware | Partial: driver crash in one VM may affect scheduling for others | Best-effort — no hard HBM bandwidth guarantee | Must match host driver version |
| NVIDIA MPS (Multi-Process Service) |
Software — CUDA context sharing within a single OS instance | All CUDA GPUs | Cooperative — processes share a single CUDA context | None: one process OOM/crash kills all MPS clients | None — HBM is shared pool with no QoS | Must match within MPS server |
MIG is the gold standard for multi-tenant edge AI where workloads cannot afford to interfere with each other. An A100-80GB GPU can be partitioned into up to seven 1g.10gb slices (10 GB HBM2e + 14 SMs each), or combined into larger profiles. An H100-80GB supports similar configurations with HBM3 bandwidth guarantees.
For a three-production-line factory, a single H100-80GB MIG configuration might look like:
The trade-off: MIG only works on datacenter GPUs (A100, H100, H200, B200). L40S, RTX 6000 Ada, and RTX 5090 do NOT support MIG — they fall back to vGPU or MPS.
vGPU is the right choice when your factory edge server runs a hypervisor (VMware ESXi, KVM, Xen) with separate VMs for IT and OT workloads. The NVIDIA vGPU manager driver sits in the hypervisor and time-slices the physical GPU across VMs. Each VM gets its own NVIDIA driver instance, CUDA toolkit version, and container runtime — important when the AOI VM runs an older JetPack SDK and the data analytics VM needs the latest CUDA 12.6.
Key vGPU profiles for L40S (48 GB):
| vGPU Profile | Frame Buffer | Virtual Displays | Max vGPUs per L40S | Best For |
|---|---|---|---|---|
| L40S-48Q | 48 GB | 4 | 1 | Dedicated 1:1 pass-through for a full-card AOI engine |
| L40S-24Q | 24 GB | 4 | 2 | Two concurrent AOI lines sharing one L40S |
| L40S-12Q | 12 GB | 4 | 4 | Light inference + visualization dashboards per VM |
| L40S-8C | 8 GB | 1 | 6 | Compute-only (headless) inference containers |
The vGPU time-slicing scheduler runs at 1 ms granularity with a best-effort QoS policy. Under heavy concurrent load, a 4-way L40S-12Q split delivers ~22% per-vGPU throughput vs bare-metal — acceptable for dashboards, insufficient for latency-critical AOI. For factory workloads with hard latency requirements, pair vGPU with MIG-capable A100/H100 instead.
MPS (Multi-Process Service) is the simplest path to GPU sharing — no hypervisor, no hardware partitioning. It runs entirely within a single Linux OS instance, merging multiple CUDA contexts into a single MPS server context. This eliminates context-switch overhead (typically 20–50 µs per switch on A100) and lets multiple inference containers share the GPU without mutual exclusion.
MPS is ideal for Kubernetes-based edge GPU clusters running Triton Inference Server, where multiple model replicas share one GPU. A Jetson AGX Orin or L40S server running 3–5 inference containers (YOLO + Whisper + Llama) benefits from MPS's zero-context-switch model — but critically, there's no isolation. An OOM kill in the LLM container will cascade to all MPS clients.
| Workload | A100-80GB Bare-Metal | A100 MIG 3g.40gb Slice | Overhead |
|---|---|---|---|
| YOLOv8x inference (fp16, batch=1) | 2,040 fps | 822 fps (scaled: 7/3 × slice-perf = 1,918 fps) | −6% vs scaled bare-metal |
| Llama 3.1 8B (INT8, batch=1) | 1,412 tok/s | 578 tok/s (scaled: 1,349 tok/s) | −4.5% vs scaled bare-metal |
| ResNet-50 training (fp16, batch=256) | 3,120 img/s | 1,265 img/s (scaled: 2,951 img/s) | −5.4% vs scaled bare-metal |
| HBM bandwidth (STREAM triad) | 2,039 GB/s | 870 GB/s (3/7 × theoretical = 874) | −0.5% vs theoretical cap |
MIG overhead is 4–7% relative to a proportionally-scaled bare-metal GPU. That's the cost of hard fault isolation and guaranteed memory bandwidth — a bargain for safety-critical factory deployments.
| Scenario | Best Technology | Recommended GPU | Why |
|---|---|---|---|
| 3 production lines on one edge server — AOI + predictive maintenance + operator LLM | MIG | H100 80GB (3 slices) | Hard fault isolation, guaranteed HBM3 bandwidth per line |
| Factory IT + OT VMs on VMware ESXi — separate Windows/Linux tenants | vGPU | L40S 48GB (4× L40S-12Q) | Hypervisor-native, independent NVIDIA driver per VM |
| Kubernetes Triton Inference Server — 5 YOLO model replicas sharing one GPU | MPS | L40S or RTX 6000 Ada | Zero context-switch overhead, simple container deployment |
| Single-tenant, full-card AOI — no sharing needed | None (bare-metal passthrough) | RTX 6000 Ada or L40S | No overhead, simpler driver stack, lower cost |
We build, burn-in test, and pre-configure GPU edge servers with your choice of MIG/vGPU/MPS partitioning:
Need a multi-tenant GPU edge server with MIG, vGPU, or MPS pre-configured?
We build and burn-in test every configuration — OS, drivers, CUDA, and GPU partitioning are validated before shipping. DDP worldwide.
Contact: +86 137-1464-6179 | sherry@qscompute.com