# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-08. Interactive version: https://trustbutveri.fyi/explorer/?mechanisms=M-0016,M-0012&ready=R2

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's readiness and open flaws. Readiness levels R0 to R4 describe one record's public evidence for its assessed use and are never combined. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and flaws) and https://trustbutveri.fyi/about/readiness/ (readiness levels).

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

- **Minimum readiness: R2 Demonstrated.** Keeps mechanisms whose readiness level is at least this one. A level describes the public evidence for a mechanism's stated use, not its cost or feasibility. R3 can still have open critical flaws.

15 of 25 mechanisms on the map pass these filters.

## Overview

One row per mechanism, read from its record. Open flaws: critical / significant / minor. The last three columns are the editors' reading of what the verifier sees.

| Mechanism | Readiness | Prover | Attack testing | Hardware | Open flaws | Weights | Inputs and outputs | Training data |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Timed challenge-response and memory-occupation challenges | R2 | Adversarial | Analysis | None | 0 / 1 / 1 | not involved | not involved | not involved |
| Model identity attestation | R3 | Semi-trusted | Independent red-team | Existing features | 1 / 2 / 0 | depends | depends | not involved |

## Claims

No claims chosen.

## Mechanisms

### Timed challenge-response and memory-occupation challenges

A verifier sends unpredictable questions that a device can answer in time only if it holds specified data, or dedicates specified resources, locally. ([Timed challenge-response and memory-occupation challenges](https://trustbutveri.fyi/mechanisms/timed-challenge-response/))

- Assessment: mechanism family.
- Readiness: R2 Demonstrated, assessed for detecting whether a GPU is doing other work.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: none. Prover cooperation: required. Attack testing: analysis. Category: Cryptographic & computational.
- What the verifier sees: model weights not involved; inputs and outputs not involved; training data not involved. Uses verifier-chosen challenges; it does not handle model data.

### Model identity attestation

Establishes that responses come from a specific, committed set of model weights, using enclave measurements or recomputation of sampled outputs. ([Model identity attestation](https://trustbutveri.fyi/mechanisms/model-identity-attestation/))

- Assessment: mechanism family.
- Readiness: R3 In production, assessed for showing users that a service runs the declared model weights.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: required. Attack testing: independent red-team. Category: Cryptographic & computational.
- What the verifier sees: model weights depends; inputs and outputs depends; training data not involved. The enclave route shows only hashes; the recomputation route gives the verifier the weights and the sampled requests and responses.


## Properties

**In production**

- Model identity attestation: R3 In production, assessed for showing users that a service runs the declared model weights

**Built for an adversarial prover**

- Timed challenge-response and memory-occupation challenges

**No new hardware needed**

- Timed challenge-response and memory-occupation challenges
- Model identity attestation

**Flaws since mitigated**

- Launch-state attestation does not by itself cover weights loaded later (in Model identity attestation) [11][23]


## Attack testing

Published attempts to break a system, including those that found failures. Testing history does not establish that open flaws are resolved.

**Testing history**

- Timed challenge-response and memory-occupation challenges: Analysis
- Model identity attestation: Independent red-team


## Limits

**Open critical flaws**

- Underlying attestation can be forged or relayed (demonstrated attack, in Model identity attestation; https://trustbutveri.fyi/mechanisms/model-identity-attestation/#flaw-1) [12][14][18][20][21][22]. Inherited finding. Critical for the enclave route against an operator with physical access to affected hardware, or control of an unpatched SEV-SNP hypervisor. It does not apply to the recomputation route. PAL*M excludes physical attacks, and Tinfoil acknowledges this boundary. The enclave route inherits the platform-specific TEE attestation failures. Intel TDX forgery and H100 relay were demonstrated with physical access and host control. Battering RAM defeated AMD SEV-SNP attestation on DDR4 servers; RMPocalypse did so from malicious host software on platforms without AMD's fixes. These demonstrate failures of the trust roots, not of each model-commitment protocol. Related finding: https://trustbutveri.fyi/mechanisms/tee-remote-attestation/#flaw-1.

  Response: The TEE.fail authors report that physical interposer attacks are outside Intel's and AMD's threat models. AMD reports fixes for RMPocalypse.

  Related mechanism: Hardware-enabled guarantees (flexHEG) and guarantee processors (R1, not in the proposal). A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.

**Open significant flaws**

- Remote memory narrows the timing margin (theoretical argument, in Timed challenge-response and memory-occupation challenges; https://trustbutveri.fyi/mechanisms/timed-challenge-response/#flaw-2) [1]. Data-centre remote memory access returns in about 1–2 µs, against about 70–200 ns for local DRAM. The MIRI overview says verification of memory saturation depends on ruling out remote access by latency or physical disconnection. It adds that pre-staging data is ruled out only by unpredictable, capacity-filling challenges.

  Related mechanism: Bandwidth limits and compartmentalization (R2, not in the proposal). Physical disconnection is proposed to exclude remote memory between the separated groups during a challenge. It depends on the isolation boundary being enforced.
- For private models, a user can confirm consistency but not content (open question, in Model identity attestation; https://trustbutveri.fyi/mechanisms/model-identity-attestation/#flaw-3) [11][24]. When weights are not published, users can check that the same root hash is served each time, but not what the model is. Pairing the hash with an attested evaluation, as in Attestable Audits, is one proposed remedy.
- Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel (demonstrated attack, in Model identity attestation; https://trustbutveri.fyi/mechanisms/model-identity-attestation/#flaw-4) [13][19]. The recomputation variant assumes that every input, output and seed is logged correctly, and that the attacker can neither predict nor manipulate which messages are sampled for verification. Legitimate nondeterminism concentrates at a few token positions, and slow leaks within the tolerated slack remain possible. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token, reducing the exfiltration slowdown from 146–254 times under benign prompts to 60–118 times. The attack targets the exfiltration bound, not the check that outputs match the declared model.

  Related mechanism: Network taps and certifiers (R1, not in the proposal). Taps are proposed to copy and hash traffic on the monitored links, reducing reliance on the prover's own log. This still depends on the monitored boundary and trusted capture.

  Related mechanism: Deterministic and bit-exact inference (R3, not in the proposal). Bit-exact inference would remove the numerical tolerance that leaves this channel.

**Open minor flaws**

- Error rates not quantified (open question, in Timed challenge-response and memory-occupation challenges; https://trustbutveri.fyi/mechanisms/timed-challenge-response/#flaw-3) [3]. Monfared et al. show separable timing distributions but do not define thresholds or statistical tests, so false-positive and false-negative rates are not quantified.


## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open flaw or a dependency. Pointers, not recommendations: each brings its own readiness level and flaws, and none is claimed to close a flaw.

- **Deterministic and bit-exact inference** (R3 In production, assessed for reproducing open-model inference from receipts in Gensyn's information-market service)
  - Bears on the open significant flaw "Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel" in Model identity attestation. Bit-exact inference would remove the numerical tolerance that leaves this channel.
  - Model identity attestation waits on it: Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling.
- **Bandwidth limits and compartmentalization** (R2 Demonstrated, assessed for monitoring inter-node traffic with operator-run software on four GPUs)
  - Bears on the open significant flaw "Remote memory narrows the timing margin" in Timed challenge-response and memory-occupation challenges. Physical disconnection is proposed to exclude remote memory between the separated groups during a challenge. It depends on the isolation boundary being enforced.
  - Timed challenge-response and memory-occupation challenges waits on it: Outside help, such as remote memory, must be excluded during challenges.
- **TEE remote attestation for AI workloads** (R3 In production, assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Model identity attestation waits on it: Attestation that resists physical attackers, for the enclave variant.
- **Hardware-enabled guarantees (flexHEG) and guarantee processors** (R1 Proposed, assessed for checking and enforcing training-compute limits on chips, against adversaries up to states). Excluded by the filters: readiness R1
  - Bears on the open critical flaw "Underlying attestation can be forged or relayed" in Model identity attestation. A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.
- **Network taps and certifiers** (R1 Proposed, assessed for committing a complete record of cluster traffic, so declared inference can be checked). Excluded by the filters: readiness R1
  - Bears on the open significant flaw "Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel" in Model identity attestation. Taps are proposed to copy and hash traffic on the monitored links, reducing reliance on the prover's own log. This still depends on the monitored boundary and trusted capture.


## Dependencies

**Blockers**

- Timed challenge-response and memory-occupation challenges: No network-level memory challenge across data-centre servers has been demonstrated. (adversarial validation) [1]
- Timed challenge-response and memory-occupation challenges: Challenges that fill memory displace workloads; filling a pod's volatile memory takes tens of minutes and SSDs take hours. (performance & compatibility) [1][3]
- Timed challenge-response and memory-occupation challenges: Outside help, such as remote memory, must be excluded during challenges. (coverage & hidden compute; waits on Bandwidth limits and compartmentalization) [1]
- Model identity attestation: Attestation that resists physical attackers, for the enclave variant. (hardware trust; waits on TEE remote attestation for AI workloads) [18]
- Model identity attestation: Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling. (protocol soundness; waits on Deterministic and bit-exact inference) [13]
- Model identity attestation: The recomputation variant needs the verifier to hold the declared weights. (access & governance) [13]


## What the verifier sees

- Model weights: shown by none; depends on the design for Model identity attestation; hidden by none; not involved in Timed challenge-response and memory-occupation challenges; unspecified for none.
- Inputs and outputs: shown by none; depends on the design for Model identity attestation; hidden by none; not involved in Timed challenge-response and memory-occupation challenges; unspecified for none.
- Training data: shown by none; depends on the design for none; hidden by none; not involved in Timed challenge-response and memory-occupation challenges and Model identity attestation; unspecified for none.

## Implementations

- Timed challenge-response and memory-occupation challenges: [Data-centre memory challenging](https://trustbutveri.fyi/implementations/data-centre-memory-challenging/) (R1, proposed architecture); [GPU contention probes](https://trustbutveri.fyi/implementations/gpu-contention-probes/) (R2, research prototype); [Low-trust AI compute verification system overview](https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/) (R1, proposed architecture); [SAGE](https://trustbutveri.fyi/implementations/sage-gpu-attestation/) (R2, research prototype); [VRAM-residency challenge](https://trustbutveri.fyi/implementations/vram-residency-challenge/) (R2, research prototype)
- Model identity attestation: [Attestable Audits](https://trustbutveri.fyi/implementations/attestable-audits/) (R2, research prototype); [PAL*M](https://trustbutveri.fyi/implementations/palm/) (R2, research prototype); [Tinfoil model identity (Modelwrap)](https://trustbutveri.fyi/implementations/tinfoil-model-identity/) (R3, product)

## Sources

1. A System Overview for Near-Term, Low-Trust AI Compute Verification, N. Cankaya (2026). https://intelligence.org/wp-content/uploads/2026/06/A-system-overview-for-near-term-low-trust-AI-compute-verification.pdf
2. Verification Plan, R. Dean (2026). https://ai-2040.com/supplements/verification-plan
3. Timing and Memory Telemetry on GPUs for AI Governance, S. K. Monfared et al. (2026). https://arxiv.org/abs/2602.09369
4. SAGE: Software-based Attestation for GPU Execution, A. Ivanov et al. (2023). https://www.usenix.org/conference/atc23/presentation/ivanov
5. SWATT: SoftWare-based ATTestation for Embedded Devices, A. Seshadri et al. (2004). https://netsec.ethz.ch/publications/papers/swatt.pdf
6. Proofs of Space, S. Dziembowski et al. (2015). https://eprint.iacr.org/2013/796
7. Secure Code Update for Embedded Devices via Proofs of Secure Erasure, D. Perito & G. Tsudik (2010). https://link.springer.com/chapter/10.1007/978-3-642-15497-3_39
8. Software-Based Memory Erasure with Relaxed Isolation Requirements, S. Bursuc et al. (2024). https://ieeexplore.ieee.org/document/10664348/
9. On the Difficulty of Software-Based Attestation of Embedded Devices, C. Castelluccia et al. (2009). https://s3.eurecom.fr/docs/ccs09_Castelluccia.pdf
10. Refutation of "On the Difficulty of Software-Based Attestation of Embedded Devices", A. Perrig & L. van Doorn (2010). https://netsec.ethz.ch/publications/papers/perrig-ccs-refutation.pdf
11. How Tinfoil Proves Exactly What Model Is Running, Tinfoil Team (2026). https://tinfoil.sh/blog/2026-02-03-proving-model-identity
12. PAL*M: Property Attestation for Large Generative Models, P. Chantasantitam et al. (2026). https://arxiv.org/abs/2601.16199
13. Verifying LLM Inference to Detect Model Weight Exfiltration, R. Rinberg et al. (2025). https://arxiv.org/abs/2511.02620
14. A primer on secure enclaves, Tinfoil (2026). https://docs.tinfoil.sh/verification/secure-enclave-primer
15. Backend infrastructure, Tinfoil (2026). https://docs.tinfoil.sh/verification/attestation-architecture
16. How verification works in Tinfoil, Tinfoil (2026). https://docs.tinfoil.sh/verification/verification-in-tinfoil
17. modelwrap: Reproducible dm-verity read-only image of Huggingface models, Tinfoil (2026). https://github.com/tinfoilsh/modelwrap
18. TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition, J. Chuang et al. (2026). https://tee.fail/
19. Adversarial Entropy Inflation Against Gumbel-Based Inference Verification, N. Kezins (2026). https://arxiv.org/abs/2608.23375
20. Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing, J. De Meulemeester et al. (2026). https://batteringram.eu/
21. RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP, B. Schlüter & S. Shinde (2025). https://rmpocalypse.github.io/
22. SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020), AMD (2025). https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3020.html
23. On TEEs for Privacy-Preserving Monitoring in AI Governance, Gloria Z (2026). https://techgov.intelligence.org/blog/on-tees-for-privacy-preserving-monitoring-in-ai-governance
24. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments, C. Schnabl et al. (2025). https://arxiv.org/abs/2506.23706
