# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-08. Interactive version: https://trustbutveri.fyi/explorer/?claims=C-0006,C-0005&goal=G-0004

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's readiness and open flaws. Readiness levels R0 to R4 describe one record's public evidence for its assessed use and are never combined. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and flaws) and https://trustbutveri.fyi/about/readiness/ (readiness levels).

## Goal

The proposal is for the goal "Prevent catastrophic misuse" (https://trustbutveri.fyi/goals/prevent-catastrophic-misuse/): Keep capable AI models from helping anyone carry out catastrophic attacks, such as biological or chemical ones. The links from the goal to claims are the editors' judgment. Direct: The claim states part of what the goal requires. The goal cannot be verified without it. Supporting: Verifying the claim makes a direct claim easier to check or a violation less useful. The goal could be verified without it.

- Direct: Declared safeguards were applied during inference. Baker and colleagues give filtering some inputs and running oversight checks on outputs as deployment mitigations. Cankaya's proposed system screens sampled workloads for outputs free of blacklisted use. (Sources: M. Baker et al. 2025; N. Cankaya 2026.)
- Direct: Model weights have not left the facility (not in this proposal). Nevo and colleagues write that an attacker who has a model's weights can abuse the model without restrictions or monitoring. (Sources: S. Nevo et al. 2024.)
- Supporting: The declared model is the one being served. Safeguards are specified and checked for one model. They say little if a different model serves the requests. (Editors' reasoning.)
- Supporting: Communication between compute groups is bounded (not in this proposal). A limit on the total data that can leave a facility caps how much of a model's weights can be stolen. (Sources: R. Rinberg et al. 2026.)

Outside this map, the goal also needs:

- Capability evaluations. Baker and colleagues describe mitigations as proportionate to evaluated risks. They leave improving model evaluations as a separate unsolved problem. (Sources: M. Baker et al. 2025.)
- User identity. Cankaya's example rule separates whitelisted users from others. This map has no records for checking who a user is. (Sources: N. Cankaya 2026.)

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Claims

### 1. Declared safeguards were applied during inference

Specified safety measures, such as input filters, output checks or monitoring, actually ran on the requests a deployed model served. ([Declared safeguards were applied during inference](https://trustbutveri.fyi/claims/safeguards-were-applied/))

Status: unaddressed. No mechanism in the proposal addresses it.

### 2. The declared model is the one being served

Outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared, not from a substitute. ([The declared model is the one being served](https://trustbutveri.fyi/claims/declared-model-is-served/))

Status: unaddressed. No mechanism in the proposal addresses it.


## Mechanisms

No mechanisms chosen.

## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open flaw or a dependency. Pointers, not recommendations: each brings its own readiness level and flaws, and none is claimed to close a flaw.

- **Deterministic and bit-exact inference** (R3 In production, assessed for reproducing open-model inference from receipts in Gensyn's information-market service)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Model identity attestation** (R3 In production, assessed for showing users that a service runs the declared model weights)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Sampled inference recomputation** (R3 In production, assessed for checking untrusted workers' activations against the declared model, prompt and precision)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **TEE remote attestation for AI workloads** (R3 In production, assessed for showing which software ran to a party that distrusts the operator holding the hardware)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Confidential multi-party verification** (R2 Demonstrated, assessed for audits or evaluations of a private model that reveal neither party's inputs)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.
- **Safeguard attestation** (R2 Demonstrated, assessed for attesting that a declared safeguard mediated a service's responses)
  - Aimed at the claim "Declared safeguards were applied during inference", which is unaddressed.
- **Zero-knowledge proofs of inference** (R2 Demonstrated, assessed for proving a language model's output follows from committed weights, against a cheating prover)
  - Aimed at the claim "The declared model is the one being served", which is unaddressed.

