# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-09. Interactive version: https://trustbutveri.fyi/explorer/?claims=C-0005&goal=G-0004

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's development status, security evidence and findings. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and findings) and https://trustbutveri.fyi/about/readiness/ (development status).

## Goal

The proposal is for the goal "Prevent catastrophic misuse" (https://trustbutveri.fyi/goals/prevent-catastrophic-misuse/): Keep capable AI models from helping anyone carry out catastrophic attacks, such as biological or chemical ones. The links from the goal to claims are the editors' judgment. Direct: The claim states part of what the goal requires. The goal cannot be verified without it. Supporting: Verifying the claim makes a direct claim easier to check or a violation less useful. The goal could be verified without it.

- Direct: Declared safeguards were applied during inference (not in this proposal). Baker and colleagues give filtering some inputs and running oversight checks on outputs as deployment mitigations. Cankaya's proposed system screens sampled workloads for outputs free of blacklisted use. (Sources: M. Baker et al. 2025; N. Cankaya 2026.)
- Direct: Model weights have not left the facility (not in this proposal). Nevo and colleagues write that an attacker who has a model's weights can abuse the model without restrictions or monitoring. (Sources: S. Nevo et al. 2024.)
- Supporting: The declared model is the one being served. Safeguards are specified and checked for one model. They say little if a different model serves the requests. (Editors' reasoning.)
- Supporting: Communication between compute groups is bounded (not in this proposal). A limit on the total data that can leave a facility caps how much of a model's weights can be stolen. (Sources: R. Rinberg et al. 2026.)

Outside this map, the goal also needs:

- Capability evaluations. Baker and colleagues describe mitigations as proportionate to evaluated risks. They leave improving model evaluations as a separate unsolved problem. (Sources: M. Baker et al. 2025.)
- User identity. Cankaya's example rule separates whitelisted users from others. This map has no records for checking who a user is. (Sources: N. Cankaya 2026.)

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Claims

### 1. The declared model is the one being served

Outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared, not from a substitute. ([The declared model is the one being served](https://trustbutveri.fyi/claims/declared-model-is-served/))

Status: unaddressed. No mechanism in the proposal addresses it.


## Mechanisms

No mechanisms chosen.

## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open failure or a dependency. Pointers, not recommendations: each brings its own readiness level and findings, and none is claimed to close a failure.

- **Hardware-attested weight binding** (Operational use (legacy code R3), assessed for hardware-attested weight binding showing users that a service runs its committed weights)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Sampled inference recomputation** (Operational use (legacy code R3), assessed for checking untrusted workers' activations against the declared model, prompt and precision)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **TEE remote attestation for AI workloads** (Operational use (legacy code R3), assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Confidential multi-party verification** (Research demonstration (legacy code R2), assessed for audits or evaluations of a private model that reveal neither party's inputs)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Zero-knowledge proofs of inference** (Research demonstration (legacy code R2), assessed for proving a language model's output follows from committed weights, against a cheating prover)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.

