# AI verification proposal

A proposal built with the Proposal Explorer of the AI Verification Tech Map (https://trustbutveri.fyi/), from its records of 2026-10-08. Interactive version: https://trustbutveri.fyi/explorer/?claims=C-0006,C-0009,C-0008&goal=G-0004

How to read it: a claim is something one party wants to verify about another's AI hardware or software. A mechanism is a general technique for verifying claims; it is "aimed at" a claim when that is its direct purpose, and "supporting" when it contributes without being aimed at it. A claim is addressed when a mechanism in the proposal is aimed at it and is not excluded by the filters; addressed does not mean verified, so check that mechanism's readiness and open flaws. Readiness levels R0 to R4 describe one record's public evidence for its assessed use and are never combined. Definitions: https://trustbutveri.fyi/about/methodology/ (roles, properties and flaws) and https://trustbutveri.fyi/about/readiness/ (readiness levels).

## Goal

The proposal is for the goal "Prevent catastrophic misuse" (https://trustbutveri.fyi/goals/prevent-catastrophic-misuse/): Keep capable AI models from helping anyone carry out catastrophic attacks, such as biological or chemical ones. The links from the goal to claims are the editors' judgment. Direct: The claim states part of what the goal requires. The goal cannot be verified without it. Supporting: Verifying the claim makes a direct claim easier to check or a violation less useful. The goal could be verified without it.

- Direct: Declared safeguards were applied during inference. Baker and colleagues give filtering some inputs and running oversight checks on outputs as deployment mitigations. Cankaya's proposed system screens sampled workloads for outputs free of blacklisted use. (Sources: M. Baker et al. 2025; N. Cankaya 2026.)
- Direct: Model weights have not left the facility. Nevo and colleagues write that an attacker who has a model's weights can abuse the model without restrictions or monitoring. (Sources: S. Nevo et al. 2024.)
- Supporting: The declared model is the one being served (not in this proposal). Safeguards are specified and checked for one model. They say little if a different model serves the requests. (Editors' reasoning.)
- Supporting: Communication between compute groups is bounded. A limit on the total data that can leave a facility caps how much of a model's weights can be stolen. (Sources: R. Rinberg et al. 2026.)

Outside this map, the goal also needs:

- Capability evaluations. Baker and colleagues describe mitigations as proportionate to evaluated risks. They leave improving model evaluations as a separate unsolved problem. (Sources: M. Baker et al. 2025.)
- User identity. Cankaya's example rule separates whitelisted users from others. This map has no records for checking who a user is. (Sources: N. Cankaya 2026.)

## Filters

Filters apply to mechanisms only and describe the setting the proposal is for.

None set. Every mechanism on the map was available.

## Claims

### 1. Declared safeguards were applied during inference

Specified safety measures, such as input filters, output checks or monitoring, actually ran on the requests a deployed model served. ([Declared safeguards were applied during inference](https://trustbutveri.fyi/claims/safeguards-were-applied/))

Status: unaddressed. No mechanism in the proposal addresses it.

### 2. Model weights have not left the facility

No copy of specified model weights has left a designated facility through networks, physical media or other channels. ([Model weights have not left the facility](https://trustbutveri.fyi/claims/weights-have-not-left/))

Status: unaddressed. No mechanism in the proposal addresses it.

### 3. Communication between compute groups is bounded

Data flowing between specified groups of chips, or out of a facility, stays below a declared rate. ([Communication between compute groups is bounded](https://trustbutveri.fyi/claims/bandwidth-is-bounded/))

Status: unaddressed. No mechanism in the proposal addresses it.


## Mechanisms

No mechanisms chosen.

## Possible additions

Mechanisms on the map, not in the proposal, that the records connect to an unaddressed or partly addressed claim, an open flaw or a dependency. Pointers, not recommendations: each brings its own readiness level and flaws, and none is claimed to close a flaw.

- **Bandwidth limits and compartmentalization** (R2 Demonstrated, assessed for monitoring inter-node traffic with operator-run software on four GPUs)
  - Aimed at the claim "Communication between compute groups is bounded", which is unaddressed.
- **Bounding unexplained information in outputs** (R2 Demonstrated, assessed for bounding how much hidden information can leave in checked inference outputs)
  - Aimed at the claim "Model weights have not left the facility", which is unaddressed.
- **Safeguard attestation** (R2 Demonstrated, assessed for attesting that a declared safeguard mediated a service's responses)
  - Aimed at the claim "Declared safeguards were applied during inference", which is unaddressed.
- **Side-channel suppression for isolated facilities** (R1 Proposed, assessed for bounding physical covert channels out of a verified enclosure)
  - Aimed at the claim "Communication between compute groups is bounded", which is unaddressed.

