# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-08. Interactive version: https://trustbutveri.fyi/explorer/?hide=weights,io&goal=G-0005

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's readiness and open flaws. Readiness levels R0 to R4 describe one record's public evidence for its assessed use and are never combined. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and flaws) and https://trustbutveri.fyi/about/readiness/ (readiness levels).

## Goal

The proposal is for the goal "Prevent weight theft" (https://trustbutveri.fyi/goals/prevent-weight-theft/): Keep the weights of capable AI models from being copied out of the facilities that hold them. The links from the goal to claims are the editors' judgment. Direct: The claim states part of what the goal requires. The goal cannot be verified without it. Supporting: Verifying the claim makes a direct claim easier to check or a violation less useful. The goal could be verified without it.

- Direct: Model weights have not left the facility (not in this proposal). This claim is the goal in a form a verifier can check: no copy of the specified weights has left the facility. (Sources: S. Nevo et al. 2024; A. Scher & L. Thiergart 2025.)
- Supporting: Communication between compute groups is bounded (not in this proposal). If only a set amount of data can leave a data centre, an adversary cannot steal more than that amount. (Sources: R. Rinberg et al. 2026.)

Outside this map, the goal also needs:

- Insider and physical security. Nevo and colleagues group their 38 attack vectors into nine categories, which include unauthorised physical access, supply chain attacks and human intelligence. Access controls, insider threat programmes and physical security are outside this map's records. (Sources: S. Nevo et al. 2024.)

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

- **Keep hidden from the verifier: model weights, inputs and outputs.** Removes mechanisms that show the asset to the verifier. Conditional or unspecified exposure stays with a note and needs checking against the privacy requirement. Model weights: the checked model's parameters. Inputs and outputs: the requests a deployed model serves and its responses. Training data: what a model was trained on. Each mechanism's exposure is the editors' reading of its record: shown, depends on the design (kept, with a note), hidden, not involved, or unspecified for a selected implementation. Code and configuration are not covered yet.

23 of 25 mechanisms on the map pass these filters.

## Claims

No claims chosen.

## Mechanisms

No mechanisms chosen.
