Prevent catastrophic misuse

Draft

The goal is to keep capable AI models from helping people carry out attacks with catastrophic effects.

The 2025 International AI Safety Report found that general-purpose AI systems had shown some ability to give instructions and troubleshooting guidance for reproducing known biological and chemical weapons 1. It added that real-world attempts to develop such weapons still needed substantial additional resources and expertise 1. Baker and colleagues write that continued progress in capabilities could enable catastrophic misuse, for example in biological and cyber attacks 2.

Proposals act on what a deployed model will do 2 and on protecting its weights from theft 4. Verifying them means showing that required safeguards ran, and that the weights have not left the facilities that hold them.

On this page

Proposals

  • Cankaya's demonstrative ruleset includes a blacklist of illicit uses 3. One example is aiding users who are not whitelisted in high-risk dual-use areas such as chemical, biological, radiological and nuclear (CBRN) work 3.
  • Baker and colleagues include deployment mitigations among their hypothetical rules, such as filtering some kinds of inputs and running oversight checks on outputs 2.
  • Nevo and colleagues write that protecting frontier models from theft and misuse will become more important as the models become more capable 4. They define five security levels, each set by how capable an attacker a system can withstand 4.

Claims

What would have to be verified to check this goal. The two levels are the editors' judgment. How goals link to claims

Direct

Supporting

Start a proposal from these claims →

Outside this map

What the goal also needs that no record on this map covers.

  • Capability evaluationsBaker and colleagues describe mitigations as proportionate to evaluated risks. They leave improving model evaluations as a separate unsolved problem. 2
  • User identityCankaya's example rule separates whitelisted users from others. This map has no records for checking who a user is. 3

Sources

  1. BY. Bengio et al. (2025). International AI Safety Report. International AI Safety Report. Source recordSupports: general-purpose AI systems giving instructions and troubleshooting guidance for known biological and chemical weapons; real-world attempts still need substantial resources and expertise · Executive Summary, Section 2 (Risks), Biological and chemical attacks; §2.1.4
  2. BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: catastrophic misuse, for example in biological and cyber attacks; deployment mitigations among hypothetical rules, proportionate to evaluated risks; improving model evaluations out of scope · §1; §2.1, Table 3; §2.3
  3. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: blacklist of illicit uses, including aiding non-whitelisted users in high-risk dual-use areas such as CBRN; sampled workloads screened for blacklisted use · §2a; §3.2.2
  4. BS. Nevo et al. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. RAND Corporation. Source recordSupports: protecting frontier models from theft and misuse grows in importance with capability; five security levels; an attacker with the weights can abuse the model without restrictions or monitoring · abstract; ch. 6; main report, pp. 2-3
  5. BR. Rinberg et al. (2026). Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains. arXiv. Source recordSupports: egress limits cap what can be stolen · §5.1

Search

Full search page