Published: September 1, 2026 | Category: Technical | QSCompute
Kubernetes is now the default way to orchestrate edge AI — not just in data centers but on Jetson clusters, fanless industrial PCs and compact GPU servers managed with k3s. The moment you schedule a model-serving pod on one of those nodes, you hit the same question: how does the container actually see the GPU? Docker alone cannot pass an NVIDIA GPU into a container, and Kubernetes will happily schedule a pod on a node where it then crashes with a CUDA driver error. Two pieces of open-source software close the gap: the NVIDIA Container Toolkit (runtime-level) and the NVIDIA Kubernetes device plugin (scheduler-level). This guide explains both, how they fit together, and the exact setup for edge nodes — including Jetson-specific caveats.
The Container Toolkit is a runtime hook (the nvidia-container-runtime and a nvidia-ctk CLI) that intercepts container creation, inspects what the container needs, and injects the driver libraries, device nodes and environment variables — NVIDIA_VISIBLE_DEVICES, NVIDIA_DRIVER_CAPABILITIES — into the container at launch. It solves "my container cannot see the GPU." It does nothing for scheduling.
The device plugin is a DaemonSet that runs on every GPU node. It asks nvidia-smi what GPUs exist, registers them with the kubelet as allocatable resources (nvidia.com/gpu), and enforces the limits: nvidia.com/gpu: 1 request in your pod spec. It solves "the scheduler puts two GPU pods on a node with one GPU." Without the toolkit the pod starts blind; without the plugin the scheduler is blind.
curl -s -L https://nvidia.github.io/libnvidia-container/gpgkey | sudo apt-key add - curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list sudo apt update && sudo apt install -y nvidia-container-toolkit sudo nvidia-ctk runtime configure --runtime=docker # or --runtime=containerd sudo systemctl restart docker docker run --rm --gpus all nvidia/cuda:12.4-base-ubuntu22.04 nvidia-smi
The last line is the smoke test: if nvidia-smi prints inside the container, the runtime layer works. Then install the plugin:
kubectl create -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/v0.17.0/nvidia-device-plugin.yml kubectl describe node <node> | grep -A4 "nvidia.com/gpu" # shows allocatable GPUs
On k3s, configure containerd instead of Docker: nvidia-ctk runtime configure --runtime=containerd, then restart k3s. On Jetson (JetPack 5/6), the runtime is preinstalled as nvidia-container-runtime — you only add the device plugin, and your pods must declare runtimeClassName: nvidia (JetPack's containerd config exposes the nvidia runtime class out of the box). An AGX Orin 64GB node therefore advertises exactly one nvidia.com/gpu device; treat each Jetson as a single-GPU node.
| Model | Mechanism | Isolation | Good for |
|---|---|---|---|
| Exclusive (default) | one GPU per pod, hard limit | full | production inference, predictable latency |
| MIG | hardware partitioning (A100/H100) | strong (separate SMs/memory) | multi-tenant edge nodes, 2-7 slices per GPU |
| Time-slicing | plugin config sharing: timeSlicing | weak (software) | bursty dev workloads, cost savings, no isolation guarantee |
| vGPU | licensed vGPU manager (data center) | strong | VDI and enterprise virtualized fleets |
MIG matters at the edge when one node must serve several tenants or models. On an H100 (7 profiles) you can advertise 1g.9gb slices; an A100 80GB offers 1g.10gb, 2g.20gb, 3g.40gb and 7g.40gb profiles. Enable MIG with nvidia-smi -mig 1, then set the plugin's MIG_STRATEGY=single env var so each slice registers as its own nvidia.com/gpu. Note that Ada workstation cards (RTX 6000 Ada, L40S) do not support MIG — if you need partitioning on those, time-slicing is the only option, and the plugin configuration file (--config flag) lets you cap replicas per GPU to keep oversubscription bounded.
nvcr.io containers or dustynv Jetson images; stock CUDA images lack the aarch64 Jetson userspace pieces.nvidia runtime class must be set or the plugin registers devices but pods still fail with "could not select device driver".Once the plugin is healthy, GPU nodes become ordinary schedulable capacity: kubectl get nodes -o json | jq '.items[].status.allocatable' shows nvidia.com/gpu counts, and a test pod with resources.limits["nvidia.com/gpu"]=1 is the final end-to-end check. This is the same pattern whether the node is a $2,000 RTX A2000 edge server or a $3,499 Jetson AGX Thor dev kit — which makes containerized edge AI fleets dramatically easier to manage than the script-based deploys most teams still run in 2026.
Need a GPU edge node that runs Kubernetes out of the box?
QSCompute ships Jetson clusters, RTX A2000/A16 edge servers and L40S nodes pre-configured with the NVIDIA Container Toolkit and device plugin — validated, with a burn-in report. Tell us the model you want to serve and we will deliver a node that schedules GPU pods on day one.
Contact: +86 137-1464-6179 | info@qscompute.com