Deploy only evaluated models

Draft

Under this goal, a powerful AI model is widely deployed only after its risks have been evaluated and judged manageable.

Baker and colleagues give this as something states might need to verify, with risks to international security as the risks in question 1. Cankaya's demonstrative ruleset requires AI systems to be whitelisted before deployment, including internal deployment 2. Verifying the goal means showing that the declared evaluation ran and that the model being served is the one evaluated.

This map covers evidence of evaluation execution. It does not assess whether an evaluation adequately measures risk.

On this page

Proposals

  • Baker and colleagues list hypothetical rules on AI 1. Two apply here. Model evaluations regularly test frontier models for risks to international security 1. Deployment mitigations are risk mitigation practices for large-scale deployment, such as filtering some kinds of inputs and running oversight checks on outputs 1. The first goal of their framework is to verify that declared uses of large-scale AI compute are compliant: declared accurately, and with the required properties 1.
  • Cankaya gives a demonstrative ruleset for a low-trust verification system 2. AI systems are reported during and after training, and whitelisted before deployment 2. In the proposed system, a reconstructed workload is screened for whether its model is on the whitelist 2.
  • Scher and Thiergart take verifying the authenticity of model evaluations as one of three example policy goals 3.

Claims

What would have to be verified to check this goal. The two levels are the editors' judgment. How goals link to claims

Direct

  • A reported evaluation needs evidence that the declared procedure ran on the model and data it names. Scher and Thiergart include evaluation authenticity among their example verification goals. 3
  • An evaluation speaks only for the model that was evaluated. Cankaya's proposed system screens workloads for whether their model is on the whitelist, and the first subgoal of Baker and colleagues is to verify that the claimed deployment took place. 1 2

Supporting

  • Baker and colleagues describe deployment mitigations, such as filtering some inputs and running oversight checks on outputs, as proportionate to evaluated risks. A judgment that risks are manageable can depend on them. 1
  • Undeclared compute could be used for violations, such as serving a model that was never evaluated. The second goal of Baker and colleagues' framework is to verify that there are no undeclared uses of large-scale AI compute. 1
  • Nevo and colleagues write that an attacker who has a model's weights can abuse the model without restrictions or monitoring. 4

Start a proposal from these claims →

Outside this map

What the goal also needs that no record on this map covers.

  • Evaluation qualityBaker and colleagues leave specifying rules well, including improving model evaluations, as a separate unsolved problem. Evidence that an evaluation ran does not determine whether its tests adequately measure risk or whether its result justifies deployment. 1

Sources

  1. BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: wide deployment only after risks to international security are evaluated and deemed manageable; model evaluations and deployment mitigations among hypothetical rules; goals 1 and 2 of the framework; rule specification, including improving model evaluations, out of scope · abstract; §2.1, Table 3; §2.3; §3.2
  2. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: demonstrative ruleset: reporting during and after training, whitelisting before deployment including internal deployment; workloads screened for a whitelisted model · §2a; §3.2.2
  3. BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: verifying the authenticity of model evaluations as an example policy goal · Verifying the authenticity of model evaluations
  4. BS. Nevo et al. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. RAND Corporation. Source recordSupports: an attacker with a model's weights can abuse it without restrictions or monitoring · main report, pp. 2-3

Search

Full search page