# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-09. Interactive version: https://trustbutveri.fyi/explorer/?mechanisms=M-0023,M-0019,M-0012&tested=analysis

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's development status, security evidence and findings. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and findings) and https://trustbutveri.fyi/about/readiness/ (development status).

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

- **Attack testing: Published analysis.** Keeps mechanisms whose strongest published attack testing is at least this. The strongest published attempt to break the mechanism for its verification use: a security analysis, red-teaming by its developers or collaborators, or a red team independent of them.

24 of 25 mechanisms on the map pass these filters.

## Overview

One row per mechanism, read from its record. Open failures: critical / significant / minor. The last three columns are the editors' reading of what the verifier sees. Findings are grouped as known failures, scope limitations and open questions. Only known failures count as failures. Counts are an inventory of published findings, not a risk score.

| Mechanism | Development | Security evidence | Prover | Attack testing | Hardware | Open failures | Weights | Inputs and outputs | Training data |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Safeguard attestation | Research demonstration | Published security analysis | Semi-trusted | Analysis | Existing features | 0 / 2 / 0 | depends | depends | not involved |
| Chip registries and manufacturing records | Proposed | Published security analysis | Semi-trusted | Analysis | Existing features | 0 / 1 / 0 | not involved | not involved | not involved |
| Hardware-attested weight binding | Operational use | Published attack testing | Semi-trusted | Independent red-team | Existing features | 1 / 0 / 0 | hidden | unspecified | not involved |

## Claims

No claims chosen.

## Mechanisms

### Safeguard attestation

Hardware-signed evidence that an AI service sent a given response through its declared safeguard path, such as a wrapper that calls a guardrail classifier. ([Safeguard attestation](https://trustbutveri.fyi/mechanisms/safeguard-attestation/))

- Assessment: mechanism family.
- Development: Research demonstration (legacy code R2), assessed for attesting that a declared safeguard mediated a service's responses.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: required. Attack testing: analysis. Category: Cryptographic & computational.
- What the verifier sees: model weights depends; inputs and outputs depends; training data not involved. The enclave route signs hashes of the safeguard, request and response; a low-trust design has the verifier re-run and screen sampled requests itself.

### Chip registries and manufacturing records

Recording each AI chip's identity and owner from the fab onwards, and cryptographically fixing manufacturing records, so that chips can be accounted for later. ([Chip registries and manufacturing records](https://trustbutveri.fyi/mechanisms/chip-registries-and-manufacturing-records/))

- Assessment: mechanism family.
- Development: Proposed (legacy code R1), assessed for a checkable record of which chips were made and who declared owning them.
- Security evidence: Published security analysis. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: required. Attack testing: analysis. Category: Compute accounting & provenance.
- What the verifier sees: model weights not involved; inputs and outputs not involved; training data not involved. Records chip identities and owners; it does not handle model data.

### Hardware-attested weight binding

Checks that a hardware enclave serves committed model weights, by attesting software that tests the weights against a hash commitment when they are read. ([Hardware-attested weight binding](https://trustbutveri.fyi/mechanisms/model-identity-attestation/))

- Assessment: mechanism family.
- Development: Operational use (legacy code R3), assessed for hardware-attested weight binding showing users that a service runs its committed weights.
- Security evidence: Published attack testing. Independent evaluation: unassessed. Formal proof: unassessed. Deployment assurance: unassessed.
- Claims in this proposal: none of them.
- Threat model: semi-trusted prover. Hardware: existing features. Prover cooperation: required. Attack testing: independent red-team. Category: Cryptographic & computational.
- What the verifier sees: model weights hidden; inputs and outputs unspecified; training data not involved. Verifiers check a hash commitment to the weights, carried in a hardware attestation, and need no access to the weights themselves. The mechanism does not specify whether prompts and outputs are disclosed to a verifier. It involves no training data.


## Properties

**Operational use**

- Hardware-attested weight binding: Operational use (legacy code R3), assessed for hardware-attested weight binding showing users that a service runs its committed weights

**No new hardware needed**

- Safeguard attestation
- Chip registries and manufacturing records
- Hardware-attested weight binding

**Failures since mitigated**

- Launch-state attestation does not by itself cover weights loaded later (in Hardware-attested weight binding) [9][19]


## Attack testing

Attack testing records published testing for this use. It does not by itself show independent review, a formal proof or that a deployed system is secure.

**Testing history**

- Safeguard attestation: Analysis
- Chip registries and manufacturing records: Analysis
- Hardware-attested weight binding: Independent red-team


## Limits

**Open critical failures**

- Underlying attestation can be forged or relayed (known failure, demonstrated attack, in Hardware-attested weight binding; https://trustbutveri.fyi/mechanisms/model-identity-attestation/evidence/flaws/1/) [4][7][10][11][12][20]. Inherited finding. Critical for weight binding against an operator with physical access to affected hardware, or with control of the hypervisor on an AMD SEV-SNP platform without AMD's fixes. PAL*M excludes physical attacks, and Tinfoil acknowledges this boundary. The enclave route inherits the platform-specific TEE attestation failures. Intel TDX forgery and H100 relay were demonstrated with physical access and host control. Battering RAM defeated AMD SEV-SNP attestation on DDR4 servers; RMPocalypse did so from malicious host software on platforms without AMD's fixes. These demonstrate failures of the trust roots, not of each model-commitment protocol. Related finding: https://trustbutveri.fyi/mechanisms/tee-remote-attestation/evidence/flaws/1/.

  Response: The TEE.fail authors report that physical interposer attacks are outside Intel's and AMD's threat models. AMD reports fixes for RMPocalypse.

  Related mechanism: Hardware-enabled guarantees (flexHEG) and guarantee processors (R1, not in the proposal). A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.

**Open significant failures**

- Components outside the attested boundary (known failure, theoretical argument, in Safeguard attestation; https://trustbutveri.fyi/mechanisms/safeguard-attestation/evidence/flaws/4/) [1][2]. In the proof-of-guardrail experiments, the guardrail model and the agent's backend model were both reached through external APIs, and the authors leave the decision to trust those APIs to the verifier. The measured wrapper must also have no vulnerability that lets the unmeasured agent bypass the guardrail, for example by executing arbitrary commands inside the enclave. The code's README states that the enclave does not currently restrict the agent's arbitrary command execution, which could be used to bypass guardrails.
- Memory-bus interposition extracts attestation keys and forges attestations (known failure, demonstrated attack, in Safeguard attestation; https://trustbutveri.fyi/mechanisms/safeguard-attestation/evidence/flaws/5/) [1][4][7][8][9][10][11][12][13]. Inherited finding. Applies to variants using the affected Intel or AMD trust roots. PAL*M excludes physical attacks. A TDX-backed safeguard claim against a physical host attacker would be defeated, but these studies do not demonstrate a break of the AWS Nitro proof-of-guardrail prototype or of verifier-side recomputation. The TEE findings cover DDR5 attacks on Intel TDX, the H100 relay demonstration, DDR4 attacks on AMD SEV-SNP, and software-only SEV-SNP forgery before AMD's fixes. These are inherited hardware limits; a governance analysis explains why physical access matters in a treaty setting. Related finding: https://trustbutveri.fyi/mechanisms/tee-remote-attestation/evidence/flaws/1/.

  Response: Intel and AMD place the physical attack class outside their threat models, according to the researchers. AMD reports firmware fixes for RMPocalypse.

  Related mechanism: Hardware-enabled guarantees (flexHEG) and guarantee processors (R1, not in the proposal). A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.
- Documents and serial numbers can be forged (known failure, theoretical argument, in Chip registries and manufacturing records; https://trustbutveri.fyi/mechanisms/chip-registries-and-manufacturing-records/evidence/flaws/2/) [15]. Avellar and Grunewald note that export documents can be forged, that companies can hide information behind obscure corporate structures, and that it may be possible to forge serial numbers on chips and racks. They recommend cryptographic attestation of a powered-on chip as an extra check.

**Scope limitations**

- Attestation shows a safeguard ran, not that it is effective (scope limitation, theoretical argument, in Safeguard attestation; https://trustbutveri.fyi/mechanisms/safeguard-attestation/evidence/flaws/1/) [1]. Proof of guardrail ensures that the guardrail executed, but the guardrail can still err or be jailbroken. Because the guardrail must be open source, a malicious developer can attack it with jailbreaks while still presenting a valid proof. In the authors' evaluation, Llama Guard 3 reached an F1 score of 0.56 on the unsafe class of the ToxicChat dataset. The authors state that proof of guardrail should not be interpreted or advertised as proof of safety.
- Selective attestation leaves traffic uncovered (scope limitation, theoretical argument, in Safeguard attestation; https://trustbutveri.fyi/mechanisms/safeguard-attestation/evidence/flaws/2/) [1][4][9]. Attestations are issued per response. In the prototype, the agent offers them when it receives high-stakes questions, so nothing shows that unattested traffic went through the same path. PAL*M's authors note that a prover could cherry-pick favourable executions, and suggest verifier-published nonces or requesting only session-level proofs. A governance analysis notes that auditors also need assurance that all activity is accounted for, since a host could start a second confidential virtual machine that bypasses monitoring.

  Related mechanism: On-chip telemetry from timing, memory and performance counters (R2, not in the proposal). On-chip counters are a proposed route to evidence about everything a chip runs, including a second virtual machine that skips the safeguard.
- Measurements may omit behaviour-relevant configuration or runtime changes (scope limitation, theoretical argument, in Safeguard attestation; https://trustbutveri.fyi/mechanisms/safeguard-attestation/evidence/flaws/3/) [9]. Every component that influences inference behaviour must be covered by the launch measurement, including feature flags, environment variables and invocation arguments. A launch measurement also does not show that a program keeps running as measured if the kernel is later compromised.
- Records cover only chips that were recorded (scope limitation, theoretical argument, in Chip registries and manufacturing records; https://trustbutveri.fyi/mechanisms/chip-registries-and-manufacturing-records/evidence/flaws/1/) [16][18]. A registry or commitment accounts only for chips entered into it. Cankaya asks how a verifier would know it had found all chips, or how much "dark compute" remains, and notes that a fraudulent original record would mean unregistered chips had been made in advance. Halstead and Larsen propose reconstructing earlier production by auditing upstream suppliers.

  Related mechanism: Remote detection of data centres (R1, not in the proposal). Looks for large data centres that were never declared, which a registry cannot show.
- Insiders could alter records before they are fixed (scope limitation, theoretical argument, in Chip registries and manufacturing records; https://trustbutveri.fyi/mechanisms/chip-registries-and-manufacturing-records/evidence/flaws/3/) [16]. Cankaya argues that insiders who can photograph process secrets could also tamper with production records. A commitment makes changes after publication detectable, but it cannot show that the records were accurate when committed.
- For private models, a user can confirm consistency but not content (scope limitation, open question, in Hardware-attested weight binding; https://trustbutveri.fyi/mechanisms/model-identity-attestation/evidence/flaws/3/) [19][24]. When weights are not published, users can check that the same root hash is served each time, but not what the model is. Pairing the hash with an attested evaluation, as in Attestable Audits, is one proposed remedy.

**Not yet demonstrated**

- Chip registries and manufacturing records: Proposed (legacy code R1), assessed for a checkable record of which chips were made and who declared owning them


## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open failure or a dependency. Pointers, not recommendations: each brings its own readiness level and findings, and none is claimed to close a failure.

- **Hardware-enabled guarantees (flexHEG) and guarantee processors** (Proposed (legacy code R1), assessed for checking and enforcing training-compute limits on chips, against adversaries up to states)
  - Bears on the open critical failure "Underlying attestation can be forged or relayed" in Hardware-attested weight binding. A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.
  - Bears on the open significant failure "Memory-bus interposition extracts attestation keys and forges attestations" in Safeguard attestation. A tamper-protected enclosure around the chip is the proposed answer when the party that holds the hardware may attack it physically.
- **TEE remote attestation for AI workloads** (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Safeguard attestation waits on it: Frontier model inference typically needs several GPUs, GPU confidential computing is less mature than CPU support, and CPU inference, which an enclave prototype had to use, ran about 100 times slower than GPU inference.
  - Safeguard attestation waits on it: Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.
  - Hardware-attested weight binding waits on it: Attestation that resists physical attackers, for the enclave variant.


## Dependencies

**Missing prerequisites**

- TEE remote attestation for AI workloads (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware), needed by Safeguard attestation and Hardware-attested weight binding

**Shared foundations**

- TEE remote attestation for AI workloads, relied on by Safeguard attestation and Hardware-attested weight binding

**Blockers**

- Safeguard attestation: No published design shows that all of a provider's traffic passes through the attested safeguard path; current evidence covers individual attested responses. (coverage & hidden compute) [1][9]
- Safeguard attestation: Frontier model inference typically needs several GPUs, GPU confidential computing is less mature than CPU support, and CPU inference, which an enclave prototype had to use, ran about 100 times slower than GPU inference. (performance & compatibility; waits on TEE remote attestation for AI workloads) [9][24]
- Safeguard attestation: Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack. (hardware trust; waits on TEE remote attestation for AI workloads) [7][9]
- Safeguard attestation: Safeguard evidence must be bound to the model actually served, which depends on model-identity attestation. (evidence binding; waits on Hardware-attested weight binding) [19][24]
- Safeguard attestation: No independent red-team or audit of a safeguard-attestation system has been published, and the available prototypes are described by their authors as proofs of concept that have not been stress-tested by a counterparty. (adversarial validation) [2][6]
- Chip registries and manufacturing records: No implementation of an AI chip registry has been publicly reported, and covering re-exports would need cooperation from re-exporters and foreign governments that may not be feasible everywhere. (access & governance) [14][15]
- Chip registries and manufacturing records: Linking records to physical chips needs hard-to-spoof unique IDs and inspections. (hardware trust) [14][15][16]
- Chip registries and manufacturing records: Chips produced before a registry starts must be reconstructed from supplier records. (coverage & hidden compute) [16][18]
- Hardware-attested weight binding: Attestation that resists physical attackers, for the enclave variant. (hardware trust; waits on TEE remote attestation for AI workloads) [7]


## What the verifier sees

- Model weights: shown by none; depends on the design for Safeguard attestation; hidden by Hardware-attested weight binding; not involved in Chip registries and manufacturing records; unspecified for none.
- Inputs and outputs: shown by none; depends on the design for Safeguard attestation; hidden by none; not involved in Chip registries and manufacturing records; unspecified for Hardware-attested weight binding.
- Training data: shown by none; depends on the design for none; hidden by none; not involved in Safeguard attestation, Chip registries and manufacturing records and Hardware-attested weight binding; unspecified for none.

## Implementations

- Safeguard attestation: none on the map
- Chip registries and manufacturing records: none on the map
- Hardware-attested weight binding: [Attestable Audits](https://trustbutveri.fyi/implementations/attestable-audits/) (R2, research prototype); [PAL*M](https://trustbutveri.fyi/implementations/palm/) (R2, research prototype); [Tinfoil model identity (Modelwrap)](https://trustbutveri.fyi/implementations/tinfoil-model-identity/) (R3, product)

## Sources

1. Proof-of-Guardrail in AI Agents and What (Not) to Trust from It, X. Jin et al. (2026). https://arxiv.org/abs/2603.05786
2. Verifiable-ClawGuard: proof-of-guardrail reference code, SaharaLabsAI (2026). https://github.com/SaharaLabsAI/Verifiable-ClawGuard
3. Safety Without Compromising on Privacy, D. McCann-Sayles et al. (2026). https://tinfoil.sh/blog/2026-09-14-safety-without-compromising-privacy
4. PAL*M: Property Attestation for Large Generative Models, P. Chantasantitam et al. (2026). https://arxiv.org/abs/2601.16199
5. Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute, B. Penchas et al. (2026). https://icml.cc/virtual/2026/78630
6. Auditor-in-a-Box: Tools for Third-Party Auditing, R. Rinberg & B. Penchas (2026). https://www.lesswrong.com/posts/uWYk7MM9hAf9GEbGe/auditor-in-a-box-tools-for-third-party-auditing
7. TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition, J. Chuang et al. (2026). https://tee.fail/
8. DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes, J. De Meulemeester et al. (2026). https://ddropattack.eu/
9. On TEEs for Privacy-Preserving Monitoring in AI Governance, Gloria Z (2026). https://techgov.intelligence.org/blog/on-tees-for-privacy-preserving-monitoring-in-ai-governance
10. Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing, J. De Meulemeester et al. (2026). https://batteringram.eu/
11. RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP, B. Schlüter & S. Shinde (2025). https://rmpocalypse.github.io/
12. SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020), AMD (2025). https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3020.html
13. A System Overview for Near-Term, Low-Trust AI Compute Verification, N. Cankaya (2026). https://intelligence.org/wp-content/uploads/2026/06/A-system-overview-for-near-term-low-trust-AI-compute-verification.pdf
14. Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment, M. Baker et al. (2025). https://www.rand.org/pubs/working_papers/WRA4077-1.html
15. Near-Term Verification Methods for AI Chip Exports, B. Avellar & E. Grunewald (2026). https://www.iaps.ai/research/near-term-verification-methods-for-ai-chip-exports
16. TSMC most definitely has a golden record of all AI chips it made, N. Cankaya (2025). https://nacicankaya.substack.com/p/tsmc-most-definitely-has-a-golden
17. Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification, S. Ansari (2026). https://arxiv.org/abs/2604.04712
18. Covert AI Projects, B. Halstead & T. Larsen (2026). https://ai-2040.com/supplements/covert-ai-projects
19. How Tinfoil Proves Exactly What Model Is Running, Tinfoil Team (2026). https://tinfoil.sh/blog/2026-02-03-proving-model-identity
20. A primer on secure enclaves, Tinfoil (2026). https://docs.tinfoil.sh/verification/secure-enclave-primer
21. Backend infrastructure, Tinfoil (2026). https://docs.tinfoil.sh/verification/attestation-architecture
22. How verification works in Tinfoil, Tinfoil (2026). https://docs.tinfoil.sh/verification/verification-in-tinfoil
23. modelwrap: Reproducible dm-verity read-only image of Huggingface models, Tinfoil (2026). https://github.com/tinfoilsh/modelwrap
24. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments, C. Schnabl et al. (2025). https://arxiv.org/abs/2506.23706
