Mechanism · TEE remote attestation for AI workloads
Evidence & limits
On this page
R3In production for showing which software ran to a party that distrusts the operator holding the hardware
Tinfoil relies on the attestation of commercial GPU confidential computing in production, but the independent evaluations left critical flaws open.
Assessed use: showing which software ran to a party that distrusts the operator holding the hardware
Rubric assessment
- R1 met: designs with stated claims and assumptions are published for audits, property attestation and policy enforcement 8 9 11.
- R2 met through Tinfoil's model-identity chain. It is an open-source production deployment on NVIDIA H100, H200 or B200 GPUs with AMD SEV-SNP or Intel TDX (provider-reported) 18 20 21. Attestable Audits and PAL*M add end-to-end results against stated adversaries 8 9.
- R3 met through reliance by another party. Tinfoil, which did not build the TEEs, relies on the attestation in production: its service checks each GPU's attestation at boot and does not start if the check fails 19 20. Apple Private Cloud Compute is another production deployment (provider-reported) 24. As context, GPU confidential computing is a documented product feature 1, and NVIDIA reports that Azure's confidential H100 virtual machines became generally available in 2024 22.
- R4 not met. The independent public evaluations of the underlying TEEs (TEE.fail, DDRop, Battering RAM, WireTap, RMPocalypse and Fabricked) all found critical flaws, and the four that need physical access remain open. TEE.fail used physical access, root privileges and under $1000 of equipment to extract a CPU's Intel provisioning certification key, which anchors SGX and TDX attestation, and forge TDX attestations. Paired with relayed H100 attestations, the forgeries let a workload outside TEE protection pass both checks 3. DDRop forged TDX attestation reports on an up-to-date platform with a DDR5 interposer costing under $200 33. On DDR4 servers, Battering RAM and WireTap forged SGX attestations with interposers costing under $50 and under $1000, and Battering RAM also broke SEV-SNP attestation 4 5. These physical attacks defeat the main claim when the prover controls the hardware. NVIDIA lists sophisticated physical attacks as out of scope 1. Intel and AMD treat interposer and other physical attacks on memory as out of scope, according to the researchers 3 4 5 33. RMPocalypse and Fabricked forged SEV-SNP attestations from malicious host software, with no physical access 6 34. AMD reports firmware fixes for both 7 35. An independent analysis of NVIDIA's GPU confidential computing found residual metadata and timing leaks but reported no attestation break 36. Trail of Bits' pre-launch audit of WhatsApp's deployment found high-severity implementation flaws that Meta fixed, and notes that SEV-SNP does not fully protect against advanced physical attacks 29 30. The breaks do not invalidate the R2 and R3 evidence, which concerns working implementations and production use.
- Attestation that survives an attacker who physically holds the hardware, for example through memory integrity and freshness protection or tamper-responsive enclosures, confirmed by independent red-teaming.
- GPU attestation cryptographically bound to the specific confidential VM it serves.
- Coverage of whole-chip and multi-node activity, beyond a single deployment.
- Roots of trust and key provenance that rival parties accept, beyond one vendor's certificate authority.
- Measurement of runtime configuration as well as launch state.
Assessed 2026-09-25 against rubric v1.1.
Mechanism properties
| Threat model | Semi-trusted prover |
|---|---|
| Adversarial evaluation | Independent red-team |
| Hardware needed | Existing features |
| Prover cooperation | Required |
| Confidentiality | Preserving |
Evidence
- NVIDIA documents confidential computing for Hopper and Blackwell GPUs 1.
- Tinfoil reports support for H100, H200 and B200 GPUs 18 and a production inference deployment 20. Its model-identity tool is open source 21.
- PAL*M reports under 11% overhead for common operations on Intel TDX with an H100. For inference attestation across three models, the added time was 3.8–11.4% of total run time for multi-turn sessions and 45.5–66.4% for single prompts. Its code is "to be released after peer review" 9.
- Attestable Audits ran its prototype on CPU-only AWS Nitro Enclaves with a 4-bit Llama-3.1-8B. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost. The authors expect a production version on confidential-computing GPUs to have "an overhead as small as 5×" 8.
- GuardAIn reports under 0.1% inference overhead for Llama variants on a Huawei Ascend 910A 12.
- FLI and Mithril call their SGX prototype "not necessarily deployable as is", because of performance and hardware attacks that need mitigation 16.
- Public clouds sell confidential GPU virtual machines. NVIDIA reports that Azure's NCC H100 v5 confidential VMs became generally available in two regions in September 2024 22. Google reports that its Private AI Compute, launched in November 2025, uses remote attestation and encryption to connect users' devices to Gemini models in a sealed environment on its TPUs 23.
- Apple Private Cloud Compute sends each request only to servers that attest to software listed in a public transparency log, according to Apple 24. In June 2026 Apple announced an extension to Google Cloud on NVIDIA confidential computing and Intel TDX 25. An independent researcher reports that a node in Apple's research environment passed its attestation check with tampered configuration files, until Apple fixed the bug that allowed the tampering 26.
- Meta reports that WhatsApp's Private Processing runs AI requests in confidential VMs on AMD SEV-SNP with NVIDIA Hopper GPUs, and that clients check each attestation against a third-party transparency log 28. In a pre-launch audit, Trail of Bits reports finding 28 issues, eight of them high severity, including configuration data loaded outside the measurement, attestations with no freshness guarantee and GPU attestation that was not verified 29 30. Meta fixed 16 and partly addressed four before launch 29.
- Anthropic, with Pattern Labs, sketches a confidential inference design in which a small model loader with an attested boot decrypts data only inside a trusted environment before passing it to the accelerator. It aims to protect both user data and model weights, and Anthropic calls the work early 27.
Limitations
- A physical attacker can forge attestations. In TEE.fail, independent researchers with physical access, root privileges and equipment costing under $1000 extracted a per-CPU Intel key that certifies attestation keys from an up-to-date machine, and forged TDX attestations. They paired the forgeries with genuine H100 attestations relayed from rented hardware, and a workload outside TEE protection passed both checks 3. A second team forged TDX attestation reports with an active DDR5 interposer costing under $200 33. On DDR4 servers, Battering RAM and WireTap forged SGX attestations with interposers costing under $50 and under $1000, and Battering RAM also broke AMD SEV-SNP attestation 4 5.
- A software attacker has forged attestations too. In RMPocalypse, a malicious hypervisor faked SEV-SNP attestation on Zen 3, Zen 4 and Zen 5 processors without physical access 6. Fabricked did the same on Zen 5 from the hypervisor and UEFI firmware, by misconfiguring the processor's interconnect 34. AMD reports firmware fixes for both 7 35.
- Side channels and other attacks by the host remain. PAL*M and Attestable Audits cite earlier side-channel, interrupt-injection and memory-aliasing attacks on CPU TEEs 8 9, and new ones such as StackWarp continue to appear 32. An independent analysis of NVIDIA's GPU confidential computing found metadata and timing leaks in unprotected shared memory 36. NVIDIA disables performance counters in confidential mode because they could provide an avenue for side-channel attacks 1. Gloria Z notes that counters have leaked secrets from TEEs 11.
- Coverage of large systems is incomplete. Hopper does not encrypt NVLink 1. NVIDIA's early-access release lacked multi-node support in mid-2025 17, and its April 2026 release notes list confidential modes only for GPUs passed through to a confidential VM, with no multi-node mode 31. O'Gara et al. ask what changes would be needed to use TEEs for verifying AI training across multi-node, multi-GPU systems at scale 13.
- The root of trust rests with a few vendors (AMD, Intel and NVIDIA) 11.
Known flaws
- DDR5 memory-bus interposers forge Intel TDX attestations and break SEV-SNP protections (TEE.fail, DDRop)
- DDR4 memory-bus interposers forge SGX and SEV-SNP attestation (Battering RAM, WireTap)
- Software-only forgery of SEV-SNP attestation (RMPocalypse, Fabricked)
- H100 attestation not bound to a specific confidential VM
- Side channels and other attacks by the host on CPU and GPU TEEs
- Attestation covers launch state, and measurements can be incomplete
- Deployment-level attestation does not cover the whole chip
- Root of trust concentrated in a few hardware vendors
Blockers
Vendor threat models exclude sophisticated physical attacks, but in international verification the prover holds the hardware.
Negative claims such as "no undeclared training" need chip-wide accounting of all workloads, which attestation does not provide.
Multi-GPU and multi-node coverage is incomplete, because Hopper leaves NVLink traffic unencrypted and NVIDIA's April 2026 release notes list no multi-node confidential mode.
Rival parties have not agreed on trust roots and key provenance they would accept.
CPU-only enclaves are costly for large models, because in the Attestable Audits prototype CPU inference cost 21.7 times as much per token as GPU inference and the enclave roughly doubled the CPU cost.