Published: September 20, 2026 | Category: Technical Guide | QSCompute
When an AI product leaves the factory, the model leaves with it. Weights, architecture, calibration tables and the inference logic behind years of training and domain data now sit inside a box the customer physically owns, powers up, and can open with a screwdriver. A competitor buying one unit and dumping its storage is the cheapest route to a cloned product — cheaper than building the model, and usually cheaper than reverse-engineering the board.
This is a different problem from data confidentiality. Confidential computing protects the customer’s data from the infrastructure operator; here the asset is the model itself, and the adversary is the party holding your hardware. The two overlap in the silicon they use, but the design target is inverted: you are protecting yourself from the person you sold the device to.
Name the adversary before choosing hardware, because each arrives with different access.
Five attacks follow from that list, roughly in order of effort: dumping the flash or NVMe; using a live JTAG or UART debug port; scraping RAM on a running device; patching the licence branch out of the binary; and side-channel analysis. Only the first three are defeated by hardware you can buy, which is why they get the budget.
It is worth being blunt about why file obfuscation alone fails. Any wrapper must turn ciphertext back into plaintext weights before the GPU can multiply against them. If the model runs, its plaintext exists in RAM for part of every inference — the attacker only has to be present at the right moment. The controls that hold are the ones that keep the key out of reach, not the ones that hide the ciphertext. That is the difference between deterrence and a security property.
Five layers, each with a different trust anchor:
| Attack | Primary defence | Hardware required | BOM impact (illustrative) |
|---|---|---|---|
| Flash/NVMe dump on a bench | Encrypted model store, key sealed to TPM | TPM 2.0 or Opal 2.0 SSD | $5–15, or a drive-class uplift |
| Debug-port access | Fused JTAG/UART, production OTP | SoC OTP programming in manufacturing | Engineering time, near-zero unit cost |
| RAM scrape on a live device | GPU confidential-compute mode, encrypted HBM | Hopper/Blackwell-class GPU or enclave-capable SoC | Large — new silicon tier |
| Licence-branch patching | Measured boot, signed image, remote attestation | Secure boot plus a provisioning service | Server-side service cost |
| Cloned unit re-provisioned | Per-device identity and attestation-gated model delivery | Discrete secure element per unit | $1–3 per unit plus provisioning |
Most teams should pick a posture deliberately rather than defaulting to whichever feature the SoC happens to expose.
| Posture | Stops | Hardware | Ops burden | Residual exposure |
|---|---|---|---|---|
| A — Obfuscation only | Casual inspection | None | None | Trivial to defeat; weights readable at runtime |
| B — Signed boot + encrypted store, online key | Bench dump, tampered image | TPM 2.0 or Opal SSD | Key management, signing infrastructure | Weights and key both live in RAM during inference |
| C — B plus secure element and attestation-gated provisioning | Bench dump, cloned units, unauthenticated model delivery | Secure element + provisioning service | Manufacturing key injection, fleet PKI | Runtime scrape still possible on a hostile device |
| D — C plus GPU TEE / enclave inference | Memory scraping on an owned device | Confidential-compute-capable GPU or SoC | Heaviest: attestation policy, CC-mode performance budget | Reduced to side-channel and supply-chain attack |
Choose by the ratio of model value to unit volume. A $5,000 inspection appliance running a model that cost $2M to build justifies posture D. A $200 sensor node whose model anyone could retrain does not. The common mistake is buying posture B for a genuinely valuable model: signed boot and an encrypted store with an online key still leave the weights in plaintext memory at every inference, which is exactly where the attacker was going to look.
| Platform | Secure boot | Runtime isolation | Storage encryption | Best fit |
|---|---|---|---|---|
| Jetson Orin / Thor (Arm) | Secure Boot + OP-TEE | TrustZone, CC-capable on newer parts | Opal or LUKS + TPM | Edge appliances, wide-temperature |
| x86 industrial PC | Intel Boot Guard / AMD PSP | Intel TDX, AMD SEV-SNP | Opal or LUKS + TPM 2.0 | Multi-model, GPU-attached lines |
| Discrete data-centre GPU | Holds host chain | GPU TEE with encrypted HBM | Host-side | On-prem inference servers |
| Arm SoC (RK3588-class) | Vendor secure boot, varies | TrustZone, limited | LUKS + discrete secure element | Cost-sensitive nodes |
Two procurement realities matter more than the matrix. First, capability is per-SKU, not per-brand: secure boot and TrustZone are available on Jetson Orin and Thor, but debug fusing and per-device key injection are contract-manufacturing steps that must be specified in writing, not assumed. Second, the standards now carry dates. NIST SP 800-193 covers firmware resiliency, IEC 62443-4-2 sets component-level security requirements, FIPS 140-3 governs the crypto module itself, and the EU Cyber Resilience Act’s Article 14 vulnerability-reporting duty went live on 11 September 2026, with Annex I conformity obligations following on 11 December 2027.
QSCompute supplies the hardware side of secure inference deployment — Jetson Orin and Thor modules with secure boot, TPM-equipped fanless industrial PCs, confidential-compute-capable GPUs, and TCG Opal industrial NVMe storage for encrypted model stores.
Shipping a model on hardware you do not control?
Send us your model value, unit volume and threat model — our engineers return a secure-inference BOM: SoC secure boot, TPM or secure element, encrypted storage and the provisioning path.
Contact: +86 137-1464-6179 | info@qscompute.com