Confidential Computing for AI 2026:
GPU TEEs, Attestation & Hardware Selection

Published: September 17, 2026 | Category: Technical Guide | QSCompute

A GPU only becomes a confidential GPU when three separate things are true at once: the card's memory is encrypted, the CPU platform it sits behind is a trusted execution environment, and you can prove the configuration to an auditor with a signed attestation report. A card that satisfies two of those three is not a confidential GPU — it is a normal accelerator with a security datasheet.

That distinction is now a procurement problem rather than a research question. Healthcare, financial services and defence buyers whose model weights are the product are writing confidential-computing requirements into 2026 tenders. This guide covers what a GPU TEE protects, which hardware supports it, what it costs in throughput, and what to ask before signing a purchase order.

What a GPU TEE Actually Protects

Confidential computing on NVIDIA hardware rests on two mechanisms working together. Inside the GPU, a dedicated encryption engine in the memory controller encrypts high-bandwidth memory with AES-256-GCM; the keys are generated in, and never leave, the GPU's secure enclave. Between the GPU and the CPU, PCIe traffic is encrypted by the CPU's memory encryption engine, so an analyzer sitting on the PCIe slot sees only ciphertext. On multi-GPU NVLink topologies in CC mode, peer traffic is staged through encrypted bounce buffers in unprotected memory.

That combination defines a specific threat model: it protects your data and model weights from the infrastructure — the hypervisor, the cloud operator, a co-tenant, or a technician with a physical probe — but not from your own application code.

ThreatCovered by GPU CC mode?Mechanism / caveat
Cloud operator or hypervisor reading VRAMYesHBM encryption with enclave-held keys
Bus-level interception (PCIe analyzer, interposer)YesCPU TEE encrypts PCIe DMA paths
Co-tenant on the same physical hostYesConfidential VM isolation
Model IP exfiltration by the infrastructure operatorYesWeights remain ciphertext outside the enclave
Physical DRAM or DIMM attacks on the hostPartlyRequires CPU TEE (TDX or SEV-SNP), not the GPU alone
Tampered or substituted firmware / VBIOSDetected, not preventedRemote attestation flags the mismatch
A compromised guest OS inside your own CVMNoCC protects you from the host, not from yourself
Data at rest on diskNoUse TCG Opal / FIPS SSDs — a separate control

The last two rows are where most compliance projects go wrong. Confidential computing is an in-use data control. It is complementary to encryption at rest, not a substitute for it.

Which Hardware Actually Supports Confidential AI

Confidential Computing is a hardware feature, not a driver setting, and it is generation-gated. Ampere does not have the silicon; Hopper introduced it; Blackwell builds it into the die.

GPUArchitectureCC supportMulti-GPURequires CPU TEE
A100 80GBAmpereNo——
RTX 4090 / RTX 6000 AdaAdaNo——
H100 SXM5 / PCIe / NVLHopperYes (CC-On, CC-DevTools)Up to 8-GPU HGX with NVLink encryptionIntel TDX or AMD SEV-SNP
H200 SXM / NVLHopperYesUp to 8-GPU HGXIntel TDX or AMD SEV-SNP
B200 / B300 HGXBlackwellYes, in siliconUp to 8 GPUs with NVLink encryptionIntel TDX or AMD SEV-SNP
GB200 / GB300 NVLGrace-BlackwellYes, unified-memory CCRack scaleCPU-side TDX or SEV-SNP
RTX PRO 6000 BlackwellBlackwellYesWorkstation-classHost TEE

The CPU side is equally specific: Intel TDX starts at 5th Gen Xeon (Emerald Rapids) and AMD SEV-SNP at EPYC Genoa, both with compliant memory populations. A CC-capable GPU on an older host yields a confidential VM that cannot attest the GPU path — the most common silent failure in these builds.

The consequence for procurement is blunt: a 2026 confidential-AI requirement invalidates every Ampere card. The A100 remains good value for INT8 and 4-bit serving, but it cannot join a confidential workload at any price. Firmware is the other friction point — CC-enabled VBIOS images are frequently requested separately from the vendor, so raise the requirement at purchase order time, not after delivery.

The Performance Cost, Measured

The honest answer is that CC mode is close to free for compute-bound work and noticeably not free for interconnect-bound work.

Performance primitiveImpact in CC mode
Raw tensor-core computeAt par — compute engines execute normally
HBM memory bandwidthAt par — encryption runs on the memory-controller path
CPU↔GPU interconnect (PCIe DMA)Capped by CPU encryption throughput (~4 GB/s on Hopper-era CC)
Multi-GPU NVLink peer bandwidthReduced — encrypted bounce buffers in unprotected memory
Driver metadata, command buffers, sync primitivesExtra encrypt/decrypt overhead per exchange
Profiling countersUnavailable in CC-On; use CC-DevTools in a lab only

In workload terms, Hopper-generation confidential inference and training typically land in the 5–15% overhead band, with the H200's improved encryption engines at the lower end. On Blackwell the penalty has narrowed further: published HGX B300 benchmarks on a Qwen 3.5 397B FP8 model show roughly −1% to −8% relative throughput across concurrency levels and sequence lengths, once the serving stack is CC-aware.

Two planning rules follow. If your bottleneck is host-to-device movement or NVLink all-to-all, CC costs real capacity — budget the extra card. If you run single-GPU inference with the model resident in VRAM, CC is nearly transparent and the decision should be driven by compliance, not percentage points.

Attestation Is the Part That Makes It Auditable

Encryption without attestation is an unverifiable claim. NVIDIA's attestation stack issues a signed report containing the device identity, VBIOS and firmware measurements, and the current security settings. An attestation agent inside the confidential VM verifies that report against expected values before sensitive work is allowed to run; if the firmware was swapped or CC silently disabled, verification fails and the workload refuses to start.

For regulated buyers this is the deliverable. HIPAA, PCI-DSS, ITAR and GDPR Article 32 all require demonstrable technical controls; an attestation report is evidence that the compute environment was genuine and untampered at run time, which a signed security policy document is not. It also unlocks a genuinely new architecture: a workload where neither the cloud operator nor a collaborating partner ever holds plaintext. In multi-tenant settings this is what lets two competitors train on pooled data without either side — or the operator — seeing the other's inputs.

Specification Checklist for Confidential AI

  1. Verify the exact SKU against the Secure AI Compatibility Matrix. Confirm GPU model, VBIOS version, driver branch and CC mode combination — not just the card family.
  2. Order CC-enabled VBIOS and firmware with the hardware. CC firmware is often requested separately from the vendor; raise it at PO time.
  3. Match the CPU generation. 5th Gen Xeon or newer with TDX, or EPYC Genoa or newer with SEV-SNP, plus a compliant memory population.
  4. Choose the topology deliberately. Single-GPU PCIe for inference; 8-GPU HGX with NVLink encryption only if tensor parallelism justifies the peer-bandwidth cost.
  5. Budget throughput headroom. Plan for 5–15% on Hopper and roughly 1–8% on Blackwell, or add a GPU to hold the SLA.
  6. Require an attestation verifier in your stack. NVIDIA Attestation SDK plus RIM service, integrated into deployment — not bolted on at audit time.
  7. Pin driver, CUDA and VBIOS versions. CC support is sensitive to specific combinations; version drift breaks attestation.
  8. Keep CC-DevTools out of production. Profiling mode disables the protections — lab only, never on a live tenant.
  9. Pair with encryption at rest. TCG Opal or FIPS-validated industrial SSDs close the at-rest gap that CC does not cover.
  10. Demand attestation reports from the exact units shipped. A burn-in report plus a passing attestation per GPU is the artifact your auditor will request.

QSCompute supplies CC-capable H100, H200 and B200/B300 systems on Intel TDX and AMD SEV-SNP platforms, with CC-enabled VBIOS configured before shipment, TCG Opal industrial SSDs for the at-rest layer, and per-unit burn-in and attestation reports. Tell us your compliance regime and target model; we return a specification with the attestation path documented end to end.

Specifying confidential AI hardware for a regulated workload?

Send us your compliance regime, target model and throughput SLA — our engineers return a platform shortlist with CC mode, CPU TEE, VBIOS and attestation path documented per configuration.

Contact: +86 137-1464-6179 | info@qscompute.com