NVIDIA GPUs in Kubernetes at the Edge 2026 — Container Toolkit & Device Plugin Guide

Published: September 1, 2026 | Category: Technical | QSCompute

Kubernetes is now the default way to orchestrate edge AI — not just in data centers but on Jetson clusters, fanless industrial PCs and compact GPU servers managed with k3s. The moment you schedule a model-serving pod on one of those nodes, you hit the same question: how does the container actually see the GPU? Docker alone cannot pass an NVIDIA GPU into a container, and Kubernetes will happily schedule a pod on a node where it then crashes with a CUDA driver error. Two pieces of open-source software close the gap: the NVIDIA Container Toolkit (runtime-level) and the NVIDIA Kubernetes device plugin (scheduler-level). This guide explains both, how they fit together, and the exact setup for edge nodes — including Jetson-specific caveats.

The two layers, and why you need both

The Container Toolkit is a runtime hook (the nvidia-container-runtime and a nvidia-ctk CLI) that intercepts container creation, inspects what the container needs, and injects the driver libraries, device nodes and environment variables — NVIDIA_VISIBLE_DEVICES, NVIDIA_DRIVER_CAPABILITIES — into the container at launch. It solves "my container cannot see the GPU." It does nothing for scheduling.

The device plugin is a DaemonSet that runs on every GPU node. It asks nvidia-smi what GPUs exist, registers them with the kubelet as allocatable resources (nvidia.com/gpu), and enforces the limits: nvidia.com/gpu: 1 request in your pod spec. It solves "the scheduler puts two GPU pods on a node with one GPU." Without the toolkit the pod starts blind; without the plugin the scheduler is blind.

Install on an edge server (x86, Ubuntu/Debian)

curl -s -L https://nvidia.github.io/libnvidia-container/gpgkey | sudo apt-key add -
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update && sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker   # or --runtime=containerd
sudo systemctl restart docker
docker run --rm --gpus all nvidia/cuda:12.4-base-ubuntu22.04 nvidia-smi

The last line is the smoke test: if nvidia-smi prints inside the container, the runtime layer works. Then install the plugin:

kubectl create -f https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/v0.17.0/nvidia-device-plugin.yml
kubectl describe node <node> | grep -A4 "nvidia.com/gpu"   # shows allocatable GPUs

On k3s, configure containerd instead of Docker: nvidia-ctk runtime configure --runtime=containerd, then restart k3s. On Jetson (JetPack 5/6), the runtime is preinstalled as nvidia-container-runtime — you only add the device plugin, and your pods must declare runtimeClassName: nvidia (JetPack's containerd config exposes the nvidia runtime class out of the box). An AGX Orin 64GB node therefore advertises exactly one nvidia.com/gpu device; treat each Jetson as a single-GPU node.

Sharing models: exclusive, MIG, time-slicing

ModelMechanismIsolationGood for
Exclusive (default)one GPU per pod, hard limitfullproduction inference, predictable latency
MIGhardware partitioning (A100/H100)strong (separate SMs/memory)multi-tenant edge nodes, 2-7 slices per GPU
Time-slicingplugin config sharing: timeSlicingweak (software)bursty dev workloads, cost savings, no isolation guarantee
vGPUlicensed vGPU manager (data center)strongVDI and enterprise virtualized fleets

MIG matters at the edge when one node must serve several tenants or models. On an H100 (7 profiles) you can advertise 1g.9gb slices; an A100 80GB offers 1g.10gb, 2g.20gb, 3g.40gb and 7g.40gb profiles. Enable MIG with nvidia-smi -mig 1, then set the plugin's MIG_STRATEGY=single env var so each slice registers as its own nvidia.com/gpu. Note that Ada workstation cards (RTX 6000 Ada, L40S) do not support MIG — if you need partitioning on those, time-slicing is the only option, and the plugin configuration file (--config flag) lets you cap replicas per GPU to keep oversubscription bounded.

Pitfalls that bite edge deployments

Once the plugin is healthy, GPU nodes become ordinary schedulable capacity: kubectl get nodes -o json | jq '.items[].status.allocatable' shows nvidia.com/gpu counts, and a test pod with resources.limits["nvidia.com/gpu"]=1 is the final end-to-end check. This is the same pattern whether the node is a $2,000 RTX A2000 edge server or a $3,499 Jetson AGX Thor dev kit — which makes containerized edge AI fleets dramatically easier to manage than the script-based deploys most teams still run in 2026.

Need a GPU edge node that runs Kubernetes out of the box?

QSCompute ships Jetson clusters, RTX A2000/A16 edge servers and L40S nodes pre-configured with the NVIDIA Container Toolkit and device plugin — validated, with a burn-in report. Tell us the model you want to serve and we will deliver a node that schedules GPU pods on day one.

Contact: +86 137-1464-6179 | info@qscompute.com