Published: August 29, 2026 | Category: Technical | QSCompute
Xid errors are the NVIDIA driver's early-warning system for dying GPUs. A NVRM: Xid line in dmesg typically appears minutes to hours before the workload crashes — long before DCGM counters or application-level errors make the failure visible. On a 24/7 inference or training server, knowing which Xid code means "monitor", which means "RMA now", and how to catch a degrading GPU early is the difference between a scheduled replacement and a silently corrupted training run.
Every NVIDIA driver kernel module logs hardware and driver faults as numbered Xid error codes to the kernel log (dmesg / journalctl). The number identifies the specific failure mode: a double-bit ECC fault (48), a GPU that fell off the PCIe bus (79), a contained ECC event (94). Xid events surface before failures become visible at the application level, which makes them the fastest early-warning signal available for production GPU fleets.
The severity range is enormous. Xid 94 on Hopper/Ada GPUs is often informational — the hardware ECC caught and corrected a single-bit error and nothing was lost. Xid 79 means the GPU is physically unreachable over PCIe and the node is down. Both appear in the same log format with the same prefix. Misclassifying either one converts a quick automated reset into a 45-minute manual triage session — or keeps a degrading GPU in production until outputs silently corrupt.
| Xid | Error | Severity | First Response |
|---|---|---|---|
| 13 | Graphics engine exception | Low–Med | Note and monitor; usually workload-related |
| 43 | GPU stopped processing | High | Check power delivery, thermals, driver version |
| 45 | Preemptive cleanup | High | GPU reset; escalate if recurring |
| 48 | Double-bit ECC error (uncorrectable) | Critical | RMA if it repeats more than once a week |
| 63 | ECC page retirement / row remapping | Critical | Check retired pages; RMA when count exceeds 32 |
| 64 | ECC page retirement (uncorrectable) | Critical | Same as 63 |
| 74 | NVLink error | High | Check NVLink topology, cabling, firmware |
| 79 | GPU fell off the PCIe bus | Critical | Reseat GPU, try another slot; RMA if it persists |
| 92 | High single-bit ECC error rate | Med–High | Watch the trend; schedule replacement if escalating |
| 94 | Contained ECC error (Hopper/Ada) | Low–Med | Informational — but escalate if rate exceeds ~10/day/GPU |
| 95 | Uncontained ECC error | Critical | Workload outputs may be corrupted — stop and validate |
On newer drivers (R565+, CUDA 12.7), Xid 119/120 are logged as GSP RPC timeouts — errors in the GPU's system-processor firmware rather than memory. A GPU reset usually clears them; recurring events need NVIDIA's GPU Debug Guidelines.
Start with the built-in tools:
# ECC error counters per GPU nvidia-smi -q -d ECC # Retired (remapped) memory pages — the RMA trigger nvidia-smi -q -d RETIRED_PAGES # Live Xid events as they are logged dmesg -T | grep -i xid journalctl -k --since "24 hours ago" | grep -i xid # Per-GPU error query for scripts nvidia-smi --query-gpu=name,ecc.errors.uncorrected.volatile.total --format=csv
Retired pages are the strongest RMA signal. When ECC detects persistent faults, the GPU permanently retires (remaps) those DRAM pages. Check nvidia-smi -q -d RETIRED_PAGES — past 32 retired pages of the 64-page limit, the trend will not reverse and the GPU should be replaced.
Fleet monitoring with DCGM. For anything beyond a handful of servers, run the NVIDIA Data Center GPU Manager (DCGM) with dcgm-exporter into Prometheus/Grafana. Alert on:
ECC availability differs by SKU. H100/H200, L40S, RTX 6000 Ada, A6000 and A2000 all have hardware ECC — you get clean counters and early warning. Consumer GPUs like the RTX 4090 have no ECC: memory faults surface as random application crashes and silent data corruption instead, so run memory stress tests (repeated matrix-heavy kernels) during burn-in and watch Xid 43/13 patterns closely.
Before RMA'ing a GPU, rule out environmental causes — roughly half of first-time Xid events are not the GPU's fault:
nvidia-smi -q -d POWER under load; verify PSU headroom with a power-budget calculation.nvidia-smi -q -d TEMPERATURE and the chassis airflow path.RMA when: Xid 48 repeats more than once in a week; retired pages exceed 32; Xid 79 persists after reseating and a different slot; Xid 95 occurs at all. Collect nvidia-smi -q output, Xid log excerpts, and the GPU serial number (nvidia-smi -q | grep -i serial) before contacting your vendor — it turns a back-and-forth into a same-day replacement.
QSCompute burns in every GPU server before shipping — memory stress, thermal soak, and Xid log review — and can pre-configure DCGM alerting so a failing GPU pages your team instead of silently corrupting a training run.
Buying GPU servers you can't afford to lose to a silent failure?
QSCompute builds and validates GPU servers and edge AI systems — burn-in testing, ECC verification, and DCGM monitoring pre-configured.
Contact: +86 137-1464-6179 | info@qscompute.com