# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-09. Interactive version: https://trustbutveri.fyi/explorer/?mechanisms=M-0001,M-0012&implementations=M-0001:I-0002

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's development status, security evidence and findings. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and findings) and https://trustbutveri.fyi/about/readiness/ (development status).

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Overview

One row per mechanism, read from its record. Open failures: critical / significant / minor. The last three columns are the editors' reading of what the verifier sees. Findings are grouped as known failures, scope limitations and open questions. Only known failures count as failures. Counts are an inventory of published findings, not a risk score.

| Mechanism | Development | Security evidence | Prover | Attack testing | Hardware | Open failures | Weights | Inputs and outputs | Training data |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Sampled inference recomputation / DiFR (Divergence From Reference) | Research demonstration | Published security analysis | Adversarial | Analysis | None | 0 / 1 / 1 | unspecified | unspecified | unspecified |
| Hardware-attested weight binding | Operational use | Published attack testing | Semi-trusted | Independent red-team | Existing features | 1 / 0 / 0 | hidden | unspecified | not involved |

## Claims

No claims chosen.

## Mechanisms

### Sampled inference recomputation

DiFR checks that an inference provider ran its declared model by comparing output tokens or activations with a trusted re-run using the same random seed. ([Sampled inference recomputation](https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/))

- Assessment: selected implementation [DiFR (Divergence From Reference)](https://trustbutveri.fyi/implementations/difr/).
- Development: Research demonstration (legacy code R2), assessed for checking that outputs match the declared model, precision and sampling settings.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: adversarial prover. Hardware: none. Prover cooperation: required. Attack testing: analysis. Category: Cryptographic & computational.
- What the verifier sees: model weights unspecified; inputs and outputs unspecified; training data unspecified. This Explorer has no asset-specific exposure assessment for this implementation. Check its source and deployment assumptions.

### Hardware-attested weight binding

Checks that a hardware enclave serves committed model weights, by attesting software that tests the weights against a hash commitment when they are read. ([Hardware-attested weight binding](https://trustbutveri.fyi/mechanisms/model-identity-attestation/))

- Assessment: mechanism family.
- Development: Operational use (legacy code R3), assessed for hardware-attested weight binding showing users that a service runs its committed weights.
- Security evidence: Published attack testing. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: required. Attack testing: independent red-team. Category: Cryptographic & computational.
- What the verifier sees: model weights hidden; inputs and outputs unspecified; training data not involved. Verifiers check a hash commitment to the weights, carried in a hardware attestation, and need no access to the weights themselves. The mechanism does not specify whether prompts and outputs are disclosed to a verifier. It involves no training data.


## Properties

**Operational use**

- Hardware-attested weight binding: Operational use (legacy code R3), assessed for hardware-attested weight binding showing users that a service runs its committed weights

**Built for an adversarial prover**

- Sampled inference recomputation

**No new hardware needed**

- Sampled inference recomputation
- Hardware-attested weight binding

**Failures since mitigated**

- Launch-state attestation does not by itself cover weights loaded later (in Hardware-attested weight binding) [12][22]


## Attack testing

Attack testing records published testing for this use. It does not by itself show independent review, a formal proof or that a deployed system is secure.

**Testing history**

- Sampled inference recomputation / DiFR (Divergence From Reference): Analysis
- Hardware-attested weight binding: Independent red-team


## Limits

**Open critical failures**

- Underlying attestation can be forged or relayed (known failure, demonstrated attack, in Hardware-attested weight binding; https://trustbutveri.fyi/mechanisms/model-identity-attestation/evidence/flaws/1/) [13][14][18][19][20][21]. Inherited finding. Critical for weight binding against an operator with physical access to affected hardware, or with control of the hypervisor on an AMD SEV-SNP platform without AMD's fixes. PAL*M excludes physical attacks, and Tinfoil acknowledges this boundary. The enclave route inherits the platform-specific TEE attestation failures. Intel TDX forgery and H100 relay were demonstrated with physical access and host control. Battering RAM defeated AMD SEV-SNP attestation on DDR4 servers; RMPocalypse did so from malicious host software on platforms without AMD's fixes. These demonstrate failures of the trust roots, not of each model-commitment protocol. Related finding: https://trustbutveri.fyi/mechanisms/tee-remote-attestation/evidence/flaws/1/.

  Response: The TEE.fail authors report that physical interposer attacks are outside Intel's and AMD's threat models. AMD reports fixes for RMPocalypse.

  Related mechanism: Hardware-enabled guarantees (flexHEG) and guarantee processors (R1, not in the proposal). A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.

**Open significant failures**

- Statistical tolerance leaves a covert channel (known failure, demonstrated attack, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/1/) [2][7][9]. Statistical schemes can put an upper bound on an adversary's covert bandwidth, but cannot close it. In the companion exfiltration study, the detector cut exfiltratable information to under 0.5% under benign prompt traffic. It did not cut it to zero. Its authors called the channel's size under adversarial prompts an open empirical question. An independent study attacked that detector, which uses the same Gumbel-margin statistic. An adversary who controls the prompts roughly doubled the bits leaked per token. Across six models, this cut the slowdown from 146–254 times under benign prompts to 60–118 times. The attack widens the exfiltration bound. It does not target DiFR's check that outputs match the declared configuration.

**Family finding context**

- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Tolerance for numerical noise leaves a covert channel (known failure, demonstrated attack, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/1/) [2][7][9]. Schemes that accept approximate matches can put an upper bound on an adversary's covert bandwidth, but they cannot close the channel. The weight-exfiltration detector cut exfiltratable information to under 0.5%, not to zero, on a 30-billion-parameter mixture-of-experts model under benign prompt traffic. Its authors called the channel's size under adversarial prompts an open empirical question. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token. Across six models, that cut the slowdown from 146–254 times under benign prompts to 60–118 times. The attack widens the exfiltration bound. It does not target the check that outputs match the declared model.

  Related mechanism: Deterministic and bit-exact inference (R3, not in the proposal). Bit-exact inference would remove the numerical tolerance if exact replay can be deployed with the required weights and configuration.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Only recorded traffic is checked (scope limitation, theoretical argument, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/2/) [2][10]. Recomputation checks that recorded, declared workloads are correct. It cannot show that the record is complete. The published schemes do not cover hidden workloads run on the same compute, or substituted work. Rinberg et al. say their exfiltration-detection scheme cannot stand alone.

  Related mechanism: Network taps and certifiers (R1, not in the proposal). Taps copy and hash all traffic on the monitored links, which bears on whether the traffic record is complete. They do not show what else ran on the same chips.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Some inference optimizations are not covered (known failure, theoretical argument, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/3/) [1][11]. TOPLOC's authors state that it cannot detect speculative decoding in which a cheaper model does the decoding. They did not test whether it distinguishes types of key-value (KV) cache compression. DiFR was evaluated only on sampling from a single model. Its authors sketch an extension to one speculative-decoding algorithm but do not test it.
- Context for DiFR (Divergence From Reference). Findings from the mechanism family appear here as context. They apply to an implementation only when its own record lists them, under the conditions stated there. Mixed hardware widens the honest baseline (known failure, open question, in Sampled inference recomputation; https://trustbutveri.fyi/mechanisms/sampled-inference-recomputation/evidence/flaws/4/) [1]. When honest reference runs span different GPU types, the spread of benign scores grows. In DiFR's tests on Qwen3-30B-A3B, pooling A100 and H200 runs left Token-DiFR unable to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy separated them. Matched provider and verifier environments, or pooling that weights rare large deviations, restored detection.

**Open minor failures**

- Mixed hardware widens the honest baseline (known failure, open question, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/2/) [1]. For Qwen3-30B-A3B, pooling honest runs across A100 and H200 GPUs and parallelism setups broadened the honest score distribution. Token-DiFR then failed to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy did. The authors report that matched provider and verifier environments, or pooling that weights rare large deviations, restore detection.

**Scope limitations**

- For private models, a user can confirm consistency but not content (scope limitation, open question, in Hardware-attested weight binding; https://trustbutveri.fyi/mechanisms/model-identity-attestation/evidence/flaws/3/) [12][23]. When weights are not published, users can check that the same root hash is served each time, but not what the model is. Pairing the hash with an attested evaluation, as in Attestable Audits, is one proposed remedy.

**Open questions**

- Speculative decoding and multi-model sampling not evaluated (open question, open question, in DiFR (Divergence From Reference); https://trustbutveri.fyi/implementations/difr/evidence/flaws/3/) [1]. The algorithms and experiments cover sampling from a single LLM. Speculative decoding was not evaluated. The authors sketch an extension to one speculative-decoding algorithm, without experiments. They note that other variants would need modified verification and extra metadata.


## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open failure or a dependency. Pointers, not recommendations: each brings its own readiness level and findings, and none is claimed to close a failure.

- **Hardware-enabled guarantees (flexHEG) and guarantee processors** (Proposed (legacy code R1), assessed for checking and enforcing training-compute limits on chips, against adversaries up to states)
  - Bears on the open critical failure "Underlying attestation can be forged or relayed" in Hardware-attested weight binding. A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.
- **TEE remote attestation for AI workloads** (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Hardware-attested weight binding waits on it: Attestation that resists physical attackers, for the enclave variant.


## Dependencies

**Missing prerequisites**

- TEE remote attestation for AI workloads (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware), needed by Hardware-attested weight binding

**Blockers**

- Sampled inference recomputation: The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models. (privacy & leakage) [1]
- Sampled inference recomputation: The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences. (performance & compatibility) [1][4]
- Sampled inference recomputation: No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic. (adversarial validation) [6][7]
- Hardware-attested weight binding: Attestation that resists physical attackers, for the enclave variant. (hardware trust; waits on TEE remote attestation for AI workloads) [18]


## What the verifier sees

- Model weights: shown by none; depends on the design for none; hidden by Hardware-attested weight binding; not involved in none; unspecified for Sampled inference recomputation.
- Inputs and outputs: shown by none; depends on the design for none; hidden by none; not involved in none; unspecified for Sampled inference recomputation and Hardware-attested weight binding.
- Training data: shown by none; depends on the design for none; hidden by none; not involved in Hardware-attested weight binding; unspecified for Sampled inference recomputation.

## Implementations

- Sampled inference recomputation: [AI 2040 inference-only verification stack](https://trustbutveri.fyi/implementations/ai-2040-inference-only-verification-plan/) (R1, proposed architecture); [DiFR (Divergence From Reference)](https://trustbutveri.fyi/implementations/difr/) (R2, research prototype); [Low-trust AI compute verification system overview](https://trustbutveri.fyi/implementations/low-trust-compute-verification-system-overview/) (R1, proposed architecture); [SASH confidential network logger](https://trustbutveri.fyi/implementations/sash-confidential-network-logger/) (R1, research prototype); [TOPLOC](https://trustbutveri.fyi/implementations/toploc/) (R3, open-source project)
- Hardware-attested weight binding: [Attestable Audits](https://trustbutveri.fyi/implementations/attestable-audits/) (R2, research prototype); [PAL*M](https://trustbutveri.fyi/implementations/palm/) (R2, research prototype); [Tinfoil model identity (Modelwrap)](https://trustbutveri.fyi/implementations/tinfoil-model-identity/) (R3, product)

## Sources

1. DiFR: Inference Verification Despite Nondeterminism, A. Karvonen et al. (2025). https://arxiv.org/abs/2511.20621
2. Verifying LLM Inference to Detect Model Weight Exfiltration, R. Rinberg et al. (2025). https://arxiv.org/abs/2511.02620
3. adamkarvonen/difr (GitHub repository), A. Karvonen (2025). https://github.com/adamkarvonen/difr
4. Scaling Recomputation Inference Verification, Amodo Design (2026). https://amododesign.com/notes/2026-09-02-scaling-recomputation-inference-verification/
5. Amodo-Design/Inference-Recomputation-Prototype (GitHub repository), Amodo Design (2026). https://github.com/Amodo-Design/Inference-Recomputation-Prototype
6. AI 2040 Plan A — Verification SITREP, Amodo Design (2026). https://amododesign.com/ai-verification/plan-a-sitrep/
7. Adversarial Entropy Inflation Against Gumbel-Based Inference Verification, N. Kezins (2026). https://arxiv.org/abs/2608.23375
8. An Inference Verification Prototype — Stage 1, Amodo Design (2026). https://amododesign.com/notes/2026-06-29-inference-verification-prototype/
9. Bit-Exact AI Inference Verification Without Performance Tradeoffs, N. Cankaya (2026). https://arxiv.org/abs/2606.00279
10. Example Schemes for Verifying High-Stakes AI Agreements, Amodo Design (2026). https://amododesign.com/notes/2026-06-23-verification-algorithms/
11. TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference, J. M. Ong et al. (2025). https://proceedings.mlr.press/v267/ong25a.html
12. How Tinfoil Proves Exactly What Model Is Running, Tinfoil Team (2026). https://tinfoil.sh/blog/2026-02-03-proving-model-identity
13. PAL*M: Property Attestation for Large Generative Models, P. Chantasantitam et al. (2026). https://arxiv.org/abs/2601.16199
14. A primer on secure enclaves, Tinfoil (2026). https://docs.tinfoil.sh/verification/secure-enclave-primer
15. Backend infrastructure, Tinfoil (2026). https://docs.tinfoil.sh/verification/attestation-architecture
16. How verification works in Tinfoil, Tinfoil (2026). https://docs.tinfoil.sh/verification/verification-in-tinfoil
17. modelwrap: Reproducible dm-verity read-only image of Huggingface models, Tinfoil (2026). https://github.com/tinfoilsh/modelwrap
18. TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition, J. Chuang et al. (2026). https://tee.fail/
19. Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing, J. De Meulemeester et al. (2026). https://batteringram.eu/
20. RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP, B. Schlüter & S. Shinde (2025). https://rmpocalypse.github.io/
21. SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020), AMD (2025). https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3020.html
22. On TEEs for Privacy-Preserving Monitoring in AI Governance, Gloria Z (2026). https://techgov.intelligence.org/blog/on-tees-for-privacy-preserving-monitoring-in-ai-governance
23. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments, C. Schnabl et al. (2025). https://arxiv.org/abs/2506.23706
