Mechanism · Confidential multi-party verification

Evidence & limits

On this page

R2Demonstrated for audits or evaluations of a private model that reveal neither party's inputs

A 2026 pilot evaluated a proprietary model on a generally available cloud enclave service, but the evaluation workflows are pilots or research prototypes, and no party is documented as relying on their results for a verification decision.

Assessed use: audits or evaluations of a private model that reveal neither party's inputs

Rubric assessment

  • R1 met: designs that state what is verified and what is trusted are published for TEE workflows 1 2 and zero-knowledge audits 3.
  • R2 met. In the double-blind pilot, AVERI evaluated Gemini 2.5 Flash Lite against private benchmark prompts in a Google Cloud enclave with an NVIDIA H100, with the prompts kept from Google DeepMind and the weights from the evaluators (participant-reported) 9. Cove has an open-source reference implementation on Intel TDX via Phala Cloud's dstack, demonstrated end to end on a benchmark workflow 1 2. ZkAudit reports peer-reviewed end-to-end audits of MobileNet image classifiers and a recommender model 3. Attestable Audits ran safety benchmarks on an 8-billion-parameter model in AWS Nitro Enclaves 4.
  • R3 not met for this use. Google reports that Confidential Space, which releases each data owner's data only to an attested workload that meets the owner's conditions, is generally available, on H100 GPUs since April 2026 11 12. That release step is TEE attestation, assessed on its own record. The evaluation workflow that ran on it, OpenMined's PySyft, is documented only as a pilot 9. Pour Demain, an outside auditor, ran interpretability evaluations of a 744-billion-parameter open-weights model on Tinfoil's production confidential-computing platform, with its governance layer shown for single sessions only 10. No party is documented as relying on such a result for a verification decision.
  • R4 not met: no independent audit or red-team of Confidential Space for this use, Cove, Attestable Audits or ZkAudit has been published. An independent analysis of one monitoring tool's published evaluation found that its one-bit verdicts leak information 7. Attackers with physical access to the memory bus have forged the Intel TDX attestation the pilot relied on 13 14.

Confidence is medium. The R2 evidence includes peer-reviewed results 3 and a pilot on a proprietary model 9. Whether a general-purpose enclave service can carry this use to R3 is a judgment call.

Gaps to the next level
  • A production-grade, available confidential evaluation or audit workflow, or reliance by a party other than its developer on such a result for a verification decision.
  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a confidential multi-party verification system, such as Confidential Space with PySyft or Cove, that leaves no critical flaw open.
  • Confidential evaluation on multi-node enclave clusters at the scale of the largest frontier models, or zero-knowledge audits at that scale.

Assessed 2026-09-28 against rubric v1.1.

Mechanism properties

Threat modelSemi-trusted prover
Adversarial evaluationAnalysis
Hardware neededExisting features
Prover cooperationRequired
ConfidentialityPreserving

Evidence

  • Double-blind pilot. Google DeepMind, OpenMined, AVERI, MLCommons and the Singapore AI Safety Institute report evaluating Gemini 2.5 Flash Lite against private AILuminate benchmark prompts in Google Cloud's Confidential Space, on one NVIDIA H100 with Intel TDX, using OpenMined's PySyft 9. Each party checked the enclave's attestation before uploading its assets, the prompts stayed hidden from Google DeepMind and the weights from the evaluators, and AVERI staff decrypted and scored the outputs 9. The participants note that not all model code could be inspected or allowlisted, and that the guest operating system builds are not independently reproducible 9. The report's figure of under 5% compute overhead is cited from an NVIDIA blog, not measured in the pilot 9.
  • Pour Demain on Tinfoil. Pour Demain reports running five interpretability workloads on GLM-5.1, a 744-billion-parameter open-weights model, on Tinfoil's production confidential-computing platform, using Intel TDX with eight H200 GPUs 10. Raw tensors stayed inside the enclave, and only bounded, aggregated, signed and budget-capped exports left it 10. It measured 5–7% added wall time from confidential computing, rising to 33–38% with interpretability instrumentation, and it demonstrated the governance layer for single sessions only 10.
  • Cove. Its authors show how its primitives express three applications: capability-attested inference, attested confidential benchmarks and bilateral capability verification 1. They report an open-source reference implementation on Intel TDX via Phala Cloud's dstack, demonstrated end to end on the benchmark workflow only 1 2.
  • Attestable Audits. The authors ran MMLU, XSum and ToxicChat on a 4-bit Llama-3.1-8B model in CPU-only AWS Nitro Enclaves 4. They report that CPU inference cost 21.7 times as much per token as GPU inference and ran about 100 times slower 4.
  • ZkAudit. Peer-reviewed at ICML 2024 3. The authors audited MobileNet v2 image classifiers on three datasets, with accuracy 0.5–0.7 percentage points below full precision, and a small recommender whose error matched full precision 3.
  • Auditor-in-a-Box. A reference implementation runs in Tinfoil confidential virtual machines 6. Its authors state that user data and plan execution in the demo are not actually secure, and that it has not been stress-tested by a counterparty 6.

Limitations

  • Verdict leakage. Abdelghafar and Kulp found that one-bit reports can reveal sensitive attributes 7. They propose designing the evidence itself to limit this 7.
  • Physical attacks on TEEs. Researchers interposing on the memory bus extracted a per-CPU Intel attestation key and forged TDX attestations 14, and a second team forged TDX attestation reports with an active interposer 13. Other attacks forged AMD SEV-SNP attestations, one through a DDR4 interposer and one from software alone on platforms without AMD's fix 15 16.
  • Stack trust. Compromise of the Docker daemon, host kernel or TEE stack breaks Cove's guarantees 2.
  • Scale. Frontier model inference typically needs the resources of several GPUs 8. Pour Demain's evaluation used one server with eight H200 GPUs, and the double-blind pilot one H100 9 10. ZkAudit was shown on image classifiers and a recommender model, not frontier-scale language models 3.
  • Process. Plan negotiation, false positives and appeals remain open problems 6.

For attesting that a declared safeguard ran on a single service, see Safeguard attestation.

Known flaws

Blockers

Search

Full search page