Implementation · Attestable Audits
Evidence & limits
On this page
R2Demonstrated for showing users that the model answering them is the audited one
ReadinessLow confidence
One developer paper reports end-to-end results on commercial hardware, but there is no public code and no independent reproduction, so confidence is low.
Assessed use: showing users that the model answering them is the audited one
Rubric assessment
- R1 met: the protocol, goals and adversaries are published 1.
- R2 met through reproducible published end-to-end results. The paper specifies the protocol and its algorithms, the enclave and baseline instances, the models and the benchmarks. It reports results for Llama-3.1-8B on AWS Nitro Enclaves, which are commercially available hardware, against a stated adversary 1. The reported evaluation covers the audit step. The paper links no code, which the rubric does not require for this route.
- R3 not met: no party other than the developers is documented relying on it for a verification decision, and the prototype is CPU-only research code rather than a production-grade system.
- R4 not met: as of September 2026 no independent security evaluation has been published.
Gaps to the next level
- Reliance by a party other than the developers for a verification decision, or a production-grade, available implementation, for example on GPU confidential computing at realistic model scale.
- An independent public security evaluation (audit, red-team or peer-reviewed analysis).
Assessed 2026-09-25 against rubric v1.1.
Evidence
The authors report the following results:
- Setup. The prototype ran on CPU-only AWS Nitro Enclaves with a 4-bit Llama-3.1-8B-Instruct and three benchmarks: MMLU, XSum and ToxicChat. The reported evaluation covers the audit step 1.
- Throughput. The enclave produced 1.84 tokens per second, against 202 on an NVIDIA L40S GPU baseline 1.
- Cost. The enclave cost $5.80 per 100K tokens, against $2.61 on a plain CPU instance with the same number of cores and $0.12 on the GPU baseline. The authors put the cost of CPU inference at 21.7 times that of GPU inference. They attribute the roughly twofold enclave cost to needing a larger instance, or giving up cores 1.
- Accuracy. Zero-shot MMLU accuracy was 51.4% at 4-bit, which the authors describe as similar to the unquantized model on the GPU baseline 1.
- Loading. Copying a model into the enclave took at most 2 minutes 1.
- Outlook. The authors expect a production version on confidential-computing GPUs to have "an overhead as small as 5×" 1.
- Code. The paper links no code repository 1.
Limitations
- Vendor trust. The approach assumes the TEE vendor is trusted 1.
- Attacks on confidential VMs. The authors cite memory aliasing (BadRAM), ciphertext side channels (CIPHERLEAKS) and malicious interrupts (Heckler), and propose revoking affected images when such attacks are found 1. Physical attacks on the underlying TEEs, such as memory-bus interposition, are discussed under TEE remote attestation for AI workloads.
- Exfiltration. The authors name prompt-based model exfiltration during user interaction as a residual gap 1.
- Model size. Memory limits forced 4-bit quantization, which lowered MMLU accuracy 1.
- Multi-GPU. The authors wrote that H100 confidential computing had no multi-GPU support 1. NVIDIA's August 2025 white paper describes an eight-GPU protected-PCIe mode on Hopper 2.
Known flaws
Blockers
The prototype needs porting to GPU confidential computing to handle larger models; the authors expect an overhead as small as 5 times there.
As of September 2026 no code has been released for the prototype.