{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "G-0005",
    "slug": "prevent-weight-theft",
    "title": "Prevent weight theft",
    "aliases": [
      "weight security",
      "weight exfiltration",
      "model theft"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-03",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": [
        "claude-review"
      ]
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "Keep the weights of capable AI models from being copied out of the facilities that hold them.",
    "summary": "The goal is to keep the weights of capable AI models inside the facilities that are meant to hold them. Nevo and colleagues write that protecting frontier models from theft and misuse will become more important as the models become more capable [[S-1610]]. They identify 38 meaningfully distinct attack vectors [[S-1610]]. Scher and Thiergart describe an approach to monitoring AI inference under international agreements: strong security keeps weights from leaving a data centre, and that data centre is then monitored closely [[S-0005]]. Verifying the goal means showing that no copy of the weights left by any channel.",
    "claims": [
      {
        "id": "C-0009",
        "title": "Model weights have not left the facility",
        "url": "https://trustbutveri.fyi/claims/weights-have-not-left/",
        "assessment": true,
        "relevance": "direct",
        "note": "This claim is the goal in a form a verifier can check: no copy of the specified weights has left the facility.",
        "sources": [
          "S-1610",
          "S-0005"
        ],
        "editorial": false
      },
      {
        "id": "C-0008",
        "title": "Communication between compute groups is bounded",
        "url": "https://trustbutveri.fyi/claims/bandwidth-is-bounded/",
        "assessment": true,
        "relevance": "supporting",
        "note": "If only a set amount of data can leave a data centre, an adversary cannot steal more than that amount.",
        "sources": [
          "S-1508"
        ],
        "editorial": false
      }
    ],
    "outside": [
      {
        "label": "insider and physical security",
        "text": "Nevo and colleagues group their 38 attack vectors into nine categories, which include unauthorised physical access, supply chain attacks and human intelligence. Access controls, insider threat programmes and physical security are outside this map's records.",
        "sources": [
          "S-1610"
        ],
        "editorial": false
      }
    ],
    "sources": [
      {
        "source": "S-1610",
        "supports": "importance of protecting frontier weights; 38 attack vectors in nine categories; five security levels; attackers up to nation-state operations; recommendations",
        "locator": "abstract; key takeaways; recommendations; ch. 5, Table 5.1"
      },
      {
        "source": "S-0005",
        "supports": "strong security to keep weights in a data centre, plus close monitoring, as a way to monitor inference",
        "locator": "executive summary, Verifying various policy goals"
      },
      {
        "source": "S-1508",
        "supports": "egress limits cap what can be stolen",
        "locator": "§5.1"
      }
    ],
    "order": 5,
    "type": "goal",
    "url": "https://trustbutveri.fyi/goals/prevent-weight-theft/",
    "source_file": "content/goals/prevent-weight-theft.md",
    "flags_all": [],
    "body_markdown": "## Proposals\n\n- **Nevo and colleagues** identify 38 attack vectors and define five security levels [[S-1610]]. The attackers they consider range from opportunistic criminals to highly resourced nation-state operations [[S-1610]]. Their recommendations include centralising copies of the weights, reducing the number of people with access, hardening interfaces against exfiltration and running insider threat programmes [[S-1610]].\n- **Scher and Thiergart** describe one approach to monitoring all inference of a deployed model: strong security prevents model weights from leaving a data centre, and that data centre is then monitored closely [[S-0005]].\n- **Rinberg and colleagues** note that egress limits cap how much can be stolen [[S-1508]]. If only 10 GB leaves a data centre, an adversary cannot steal more than 10 GB [[S-1508]].",
    "body_text": "Proposals - Nevo and colleagues identify 38 attack vectors and define five security levels [S-1610]. The attackers they consider range from opportunistic criminals to highly resourced nation-state operations [S-1610]. Their recommendations include centralising copies of the weights, reducing the number of people with access, hardening interfaces against exfiltration and running insider threat programmes [S-1610]. - Scher and Thiergart describe one approach to monitoring all inference of a deployed model: strong security prevents model weights from leaving a data centre, and that data centre is then monitored closely [S-0005]. - Rinberg and colleagues note that egress limits cap how much can be stolen [S-1508]. If only 10 GB leaves a data centre, an adversary cannot steal more than 10 GB [S-1508].",
    "referenced_by": []
  }
}