{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "G-0004",
    "slug": "prevent-catastrophic-misuse",
    "title": "Prevent catastrophic misuse",
    "aliases": [
      "CBRN misuse",
      "biological and chemical attacks",
      "misuse safeguards"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-03",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": [
        "claude-review"
      ]
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "Keep capable AI models from helping anyone carry out catastrophic attacks, such as biological or chemical ones.",
    "summary": "The goal is to keep capable AI models from helping people carry out attacks with catastrophic effects. The 2025 International AI Safety Report found that general-purpose AI systems had shown some ability to give instructions and troubleshooting guidance for reproducing known biological and chemical weapons [[S-0054]]. It added that real-world attempts to develop such weapons still needed substantial additional resources and expertise [[S-0054]]. Baker and colleagues write that continued progress in capabilities could enable catastrophic misuse, for example in biological and cyber attacks [[S-0002]]. Proposals act on what a deployed model will do [[S-0002]] and on protecting its weights from theft [[S-1610]]. Verifying them means showing that required safeguards ran, and that the weights have not left the facilities that hold them.",
    "claims": [
      {
        "id": "C-0006",
        "title": "Declared safeguards were applied during inference",
        "url": "https://trustbutveri.fyi/claims/safeguards-were-applied/",
        "assessment": true,
        "relevance": "direct",
        "note": "Baker and colleagues give filtering some inputs and running oversight checks on outputs as deployment mitigations. Cankaya's proposed system screens sampled workloads for outputs free of blacklisted use.",
        "sources": [
          "S-0002",
          "S-0018"
        ],
        "editorial": false
      },
      {
        "id": "C-0009",
        "title": "Model weights have not left the facility",
        "url": "https://trustbutveri.fyi/claims/weights-have-not-left/",
        "assessment": true,
        "relevance": "direct",
        "note": "Nevo and colleagues write that an attacker who has a model's weights can abuse the model without restrictions or monitoring.",
        "sources": [
          "S-1610"
        ],
        "editorial": false
      },
      {
        "id": "C-0005",
        "title": "The declared model is the one being served",
        "url": "https://trustbutveri.fyi/claims/declared-model-is-served/",
        "assessment": true,
        "relevance": "supporting",
        "note": "Safeguards are specified and checked for one model. They say little if a different model serves the requests.",
        "sources": [],
        "editorial": true
      },
      {
        "id": "C-0008",
        "title": "Communication between compute groups is bounded",
        "url": "https://trustbutveri.fyi/claims/bandwidth-is-bounded/",
        "assessment": true,
        "relevance": "supporting",
        "note": "A limit on the total data that can leave a facility caps how much of a model's weights can be stolen.",
        "sources": [
          "S-1508"
        ],
        "editorial": false
      }
    ],
    "outside": [
      {
        "label": "capability evaluations",
        "text": "Baker and colleagues describe mitigations as proportionate to evaluated risks. They leave improving model evaluations as a separate unsolved problem.",
        "sources": [
          "S-0002"
        ],
        "editorial": false
      },
      {
        "label": "user identity",
        "text": "Cankaya's example rule separates whitelisted users from others. This map has no records for checking who a user is.",
        "sources": [
          "S-0018"
        ],
        "editorial": false
      }
    ],
    "sources": [
      {
        "source": "S-0054",
        "supports": "general-purpose AI systems giving instructions and troubleshooting guidance for known biological and chemical weapons; real-world attempts still need substantial resources and expertise",
        "locator": "Executive Summary, Section 2 (Risks), Biological and chemical attacks; §2.1.4"
      },
      {
        "source": "S-0002",
        "supports": "catastrophic misuse, for example in biological and cyber attacks; deployment mitigations among hypothetical rules, proportionate to evaluated risks; improving model evaluations out of scope",
        "locator": "§1; §2.1, Table 3; §2.3"
      },
      {
        "source": "S-0018",
        "supports": "blacklist of illicit uses, including aiding non-whitelisted users in high-risk dual-use areas such as CBRN; sampled workloads screened for blacklisted use",
        "locator": "§2a; §3.2.2"
      },
      {
        "source": "S-1610",
        "supports": "protecting frontier models from theft and misuse grows in importance with capability; five security levels; an attacker with the weights can abuse the model without restrictions or monitoring",
        "locator": "abstract; ch. 6; main report, pp. 2-3"
      },
      {
        "source": "S-1508",
        "supports": "egress limits cap what can be stolen",
        "locator": "§5.1"
      }
    ],
    "order": 4,
    "type": "goal",
    "url": "https://trustbutveri.fyi/goals/prevent-catastrophic-misuse/",
    "source_file": "content/goals/prevent-catastrophic-misuse.md",
    "flags_all": [],
    "body_markdown": "## Proposals\n\n- **Cankaya's** demonstrative ruleset includes a blacklist of illicit uses [[S-0018]]. One example is aiding users who are not whitelisted in high-risk dual-use areas such as chemical, biological, radiological and nuclear (CBRN) work [[S-0018]].\n- **Baker and colleagues** include deployment mitigations among their hypothetical rules, such as filtering some kinds of inputs and running oversight checks on outputs [[S-0002]].\n- **Nevo and colleagues** write that protecting frontier models from theft and misuse will become more important as the models become more capable [[S-1610]]. They define five security levels, each set by how capable an attacker a system can withstand [[S-1610]].",
    "body_text": "Proposals - Cankaya's demonstrative ruleset includes a blacklist of illicit uses [S-0018]. One example is aiding users who are not whitelisted in high-risk dual-use areas such as chemical, biological, radiological and nuclear (CBRN) work [S-0018]. - Baker and colleagues include deployment mitigations among their hypothetical rules, such as filtering some kinds of inputs and running oversight checks on outputs [S-0002]. - Nevo and colleagues write that protecting frontier models from theft and misuse will become more important as the models become more capable [S-1610]. They define five security levels, each set by how capable an attacker a system can withstand [S-1610].",
    "referenced_by": []
  }
}