NVIDIA GPU Xid Errors 2026 — ECC Faults, Failure Detection & RMA Decisions

Published: August 29, 2026 | Category: Technical | QSCompute

Xid errors are the NVIDIA driver's early-warning system for dying GPUs. A NVRM: Xid line in dmesg typically appears minutes to hours before the workload crashes — long before DCGM counters or application-level errors make the failure visible. On a 24/7 inference or training server, knowing which Xid code means "monitor", which means "RMA now", and how to catch a degrading GPU early is the difference between a scheduled replacement and a silently corrupted training run.

What Xid Errors Are and Why They Matter

Every NVIDIA driver kernel module logs hardware and driver faults as numbered Xid error codes to the kernel log (dmesg / journalctl). The number identifies the specific failure mode: a double-bit ECC fault (48), a GPU that fell off the PCIe bus (79), a contained ECC event (94). Xid events surface before failures become visible at the application level, which makes them the fastest early-warning signal available for production GPU fleets.

The severity range is enormous. Xid 94 on Hopper/Ada GPUs is often informational — the hardware ECC caught and corrected a single-bit error and nothing was lost. Xid 79 means the GPU is physically unreachable over PCIe and the node is down. Both appear in the same log format with the same prefix. Misclassifying either one converts a quick automated reset into a 45-minute manual triage session — or keeps a degrading GPU in production until outputs silently corrupt.

Xid Error Code Reference

Xid Error Severity First Response
13 Graphics engine exception Low–Med Note and monitor; usually workload-related
43 GPU stopped processing High Check power delivery, thermals, driver version
45 Preemptive cleanup High GPU reset; escalate if recurring
48 Double-bit ECC error (uncorrectable) Critical RMA if it repeats more than once a week
63 ECC page retirement / row remapping Critical Check retired pages; RMA when count exceeds 32
64 ECC page retirement (uncorrectable) Critical Same as 63
74 NVLink error High Check NVLink topology, cabling, firmware
79 GPU fell off the PCIe bus Critical Reseat GPU, try another slot; RMA if it persists
92 High single-bit ECC error rate Med–High Watch the trend; schedule replacement if escalating
94 Contained ECC error (Hopper/Ada) Low–Med Informational — but escalate if rate exceeds ~10/day/GPU
95 Uncontained ECC error Critical Workload outputs may be corrupted — stop and validate

On newer drivers (R565+, CUDA 12.7), Xid 119/120 are logged as GSP RPC timeouts — errors in the GPU's system-processor firmware rather than memory. A GPU reset usually clears them; recurring events need NVIDIA's GPU Debug Guidelines.

Detecting Failures Before the Workload Dies

Start with the built-in tools:

# ECC error counters per GPU
nvidia-smi -q -d ECC

# Retired (remapped) memory pages — the RMA trigger
nvidia-smi -q -d RETIRED_PAGES

# Live Xid events as they are logged
dmesg -T | grep -i xid
journalctl -k --since "24 hours ago" | grep -i xid

# Per-GPU error query for scripts
nvidia-smi --query-gpu=name,ecc.errors.uncorrected.volatile.total --format=csv

Retired pages are the strongest RMA signal. When ECC detects persistent faults, the GPU permanently retires (remaps) those DRAM pages. Check nvidia-smi -q -d RETIRED_PAGES — past 32 retired pages of the 64-page limit, the trend will not reverse and the GPU should be replaced.

Fleet monitoring with DCGM. For anything beyond a handful of servers, run the NVIDIA Data Center GPU Manager (DCGM) with dcgm-exporter into Prometheus/Grafana. Alert on:

ECC availability differs by SKU. H100/H200, L40S, RTX 6000 Ada, A6000 and A2000 all have hardware ECC — you get clean counters and early warning. Consumer GPUs like the RTX 4090 have no ECC: memory faults surface as random application crashes and silent data corruption instead, so run memory stress tests (repeated matrix-heavy kernels) during burn-in and watch Xid 43/13 patterns closely.

Isolate First, RMA Second — The Decision Framework

Before RMA'ing a GPU, rule out environmental causes — roughly half of first-time Xid events are not the GPU's fault:

  1. Power — a sagging PSU or loose 8-pin/12VHPWR connector causes Xid 43/79 patterns. Check nvidia-smi -q -d POWER under load; verify PSU headroom with a power-budget calculation.
  2. Thermals — throttle-then-error sequences point at airflow, not silicon. Check nvidia-smi -q -d TEMPERATURE and the chassis airflow path.
  3. PCIe seating — reseat the card, try a different slot, and check for sag or bracket stress.
  4. Driver/firmware — pin a known-good driver version; error onset right after an upgrade is a driver bug until proven otherwise.

RMA when: Xid 48 repeats more than once in a week; retired pages exceed 32; Xid 79 persists after reseating and a different slot; Xid 95 occurs at all. Collect nvidia-smi -q output, Xid log excerpts, and the GPU serial number (nvidia-smi -q | grep -i serial) before contacting your vendor — it turns a back-and-forth into a same-day replacement.

QSCompute burns in every GPU server before shipping — memory stress, thermal soak, and Xid log review — and can pre-configure DCGM alerting so a failing GPU pages your team instead of silently corrupting a training run.

Buying GPU servers you can't afford to lose to a silent failure?

QSCompute builds and validates GPU servers and edge AI systems — burn-in testing, ECC verification, and DCGM monitoring pre-configured.

Contact: +86 137-1464-6179 | info@qscompute.com