Published: September 17, 2026 | Category: Technical Guide | QSCompute
A GPU only becomes a confidential GPU when three separate things are true at once: the card's memory is encrypted, the CPU platform it sits behind is a trusted execution environment, and you can prove the configuration to an auditor with a signed attestation report. A card that satisfies two of those three is not a confidential GPU — it is a normal accelerator with a security datasheet.
That distinction is now a procurement problem rather than a research question. Healthcare, financial services and defence buyers whose model weights are the product are writing confidential-computing requirements into 2026 tenders. This guide covers what a GPU TEE protects, which hardware supports it, what it costs in throughput, and what to ask before signing a purchase order.
Confidential computing on NVIDIA hardware rests on two mechanisms working together. Inside the GPU, a dedicated encryption engine in the memory controller encrypts high-bandwidth memory with AES-256-GCM; the keys are generated in, and never leave, the GPU's secure enclave. Between the GPU and the CPU, PCIe traffic is encrypted by the CPU's memory encryption engine, so an analyzer sitting on the PCIe slot sees only ciphertext. On multi-GPU NVLink topologies in CC mode, peer traffic is staged through encrypted bounce buffers in unprotected memory.
That combination defines a specific threat model: it protects your data and model weights from the infrastructure — the hypervisor, the cloud operator, a co-tenant, or a technician with a physical probe — but not from your own application code.
| Threat | Covered by GPU CC mode? | Mechanism / caveat |
|---|---|---|
| Cloud operator or hypervisor reading VRAM | Yes | HBM encryption with enclave-held keys |
| Bus-level interception (PCIe analyzer, interposer) | Yes | CPU TEE encrypts PCIe DMA paths |
| Co-tenant on the same physical host | Yes | Confidential VM isolation |
| Model IP exfiltration by the infrastructure operator | Yes | Weights remain ciphertext outside the enclave |
| Physical DRAM or DIMM attacks on the host | Partly | Requires CPU TEE (TDX or SEV-SNP), not the GPU alone |
| Tampered or substituted firmware / VBIOS | Detected, not prevented | Remote attestation flags the mismatch |
| A compromised guest OS inside your own CVM | No | CC protects you from the host, not from yourself |
| Data at rest on disk | No | Use TCG Opal / FIPS SSDs — a separate control |
The last two rows are where most compliance projects go wrong. Confidential computing is an in-use data control. It is complementary to encryption at rest, not a substitute for it.
Confidential Computing is a hardware feature, not a driver setting, and it is generation-gated. Ampere does not have the silicon; Hopper introduced it; Blackwell builds it into the die.
| GPU | Architecture | CC support | Multi-GPU | Requires CPU TEE |
|---|---|---|---|---|
| A100 80GB | Ampere | No | — | — |
| RTX 4090 / RTX 6000 Ada | Ada | No | — | — |
| H100 SXM5 / PCIe / NVL | Hopper | Yes (CC-On, CC-DevTools) | Up to 8-GPU HGX with NVLink encryption | Intel TDX or AMD SEV-SNP |
| H200 SXM / NVL | Hopper | Yes | Up to 8-GPU HGX | Intel TDX or AMD SEV-SNP |
| B200 / B300 HGX | Blackwell | Yes, in silicon | Up to 8 GPUs with NVLink encryption | Intel TDX or AMD SEV-SNP |
| GB200 / GB300 NVL | Grace-Blackwell | Yes, unified-memory CC | Rack scale | CPU-side TDX or SEV-SNP |
| RTX PRO 6000 Blackwell | Blackwell | Yes | Workstation-class | Host TEE |
The CPU side is equally specific: Intel TDX starts at 5th Gen Xeon (Emerald Rapids) and AMD SEV-SNP at EPYC Genoa, both with compliant memory populations. A CC-capable GPU on an older host yields a confidential VM that cannot attest the GPU path — the most common silent failure in these builds.
The consequence for procurement is blunt: a 2026 confidential-AI requirement invalidates every Ampere card. The A100 remains good value for INT8 and 4-bit serving, but it cannot join a confidential workload at any price. Firmware is the other friction point — CC-enabled VBIOS images are frequently requested separately from the vendor, so raise the requirement at purchase order time, not after delivery.
The honest answer is that CC mode is close to free for compute-bound work and noticeably not free for interconnect-bound work.
| Performance primitive | Impact in CC mode |
|---|---|
| Raw tensor-core compute | At par — compute engines execute normally |
| HBM memory bandwidth | At par — encryption runs on the memory-controller path |
| CPU↔GPU interconnect (PCIe DMA) | Capped by CPU encryption throughput (~4 GB/s on Hopper-era CC) |
| Multi-GPU NVLink peer bandwidth | Reduced — encrypted bounce buffers in unprotected memory |
| Driver metadata, command buffers, sync primitives | Extra encrypt/decrypt overhead per exchange |
| Profiling counters | Unavailable in CC-On; use CC-DevTools in a lab only |
In workload terms, Hopper-generation confidential inference and training typically land in the 5–15% overhead band, with the H200's improved encryption engines at the lower end. On Blackwell the penalty has narrowed further: published HGX B300 benchmarks on a Qwen 3.5 397B FP8 model show roughly −1% to −8% relative throughput across concurrency levels and sequence lengths, once the serving stack is CC-aware.
Two planning rules follow. If your bottleneck is host-to-device movement or NVLink all-to-all, CC costs real capacity — budget the extra card. If you run single-GPU inference with the model resident in VRAM, CC is nearly transparent and the decision should be driven by compliance, not percentage points.
Encryption without attestation is an unverifiable claim. NVIDIA's attestation stack issues a signed report containing the device identity, VBIOS and firmware measurements, and the current security settings. An attestation agent inside the confidential VM verifies that report against expected values before sensitive work is allowed to run; if the firmware was swapped or CC silently disabled, verification fails and the workload refuses to start.
For regulated buyers this is the deliverable. HIPAA, PCI-DSS, ITAR and GDPR Article 32 all require demonstrable technical controls; an attestation report is evidence that the compute environment was genuine and untampered at run time, which a signed security policy document is not. It also unlocks a genuinely new architecture: a workload where neither the cloud operator nor a collaborating partner ever holds plaintext. In multi-tenant settings this is what lets two competitors train on pooled data without either side — or the operator — seeing the other's inputs.
QSCompute supplies CC-capable H100, H200 and B200/B300 systems on Intel TDX and AMD SEV-SNP platforms, with CC-enabled VBIOS configured before shipment, TCG Opal industrial SSDs for the at-rest layer, and per-unit burn-in and attestation reports. Tell us your compliance regime and target model; we return a specification with the attestation path documented end to end.
Specifying confidential AI hardware for a regulated workload?
Send us your compliance regime, target model and throughput SLA — our engineers return a platform shortlist with CC mode, CPU TEE, VBIOS and attestation path documented per configuration.
Contact: +86 137-1464-6179 | info@qscompute.com