Prevent weight theft
DraftThe goal is to keep the weights of capable AI models inside the facilities that are meant to hold them.
Nevo and colleagues write that protecting frontier models from theft and misuse will become more important as the models become more capable 1. They identify 38 meaningfully distinct attack vectors 1.
Scher and Thiergart describe an approach to monitoring AI inference under international agreements: strong security keeps weights from leaving a data centre, and that data centre is then monitored closely 2. Verifying the goal means showing that no copy of the weights left by any channel.
On this page
Proposals
- Nevo and colleagues identify 38 attack vectors and define five security levels 1. The attackers they consider range from opportunistic criminals to highly resourced nation-state operations 1. Their recommendations include centralising copies of the weights, reducing the number of people with access, hardening interfaces against exfiltration and running insider threat programmes 1.
- Scher and Thiergart describe one approach to monitoring all inference of a deployed model: strong security prevents model weights from leaving a data centre, and that data centre is then monitored closely 2.
- Rinberg and colleagues note that egress limits cap how much can be stolen 3. If only 10 GB leaves a data centre, an adversary cannot steal more than 10 GB 3.
Claims
What would have to be verified to check this goal. The two levels are the editors' judgment. How goals link to claims
Direct
- This claim is the goal in a form a verifier can check: no copy of the specified weights has left the facility. 1 2
Supporting
- If only a set amount of data can leave a data centre, an adversary cannot steal more than that amount. 3
Outside this map
What the goal also needs that no record on this map covers.
- Insider and physical securityNevo and colleagues group their 38 attack vectors into nine categories, which include unauthorised physical access, supply chain attacks and human intelligence. Access controls, insider threat programmes and physical security are outside this map's records. 1
Sources
- BS. Nevo et al. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. RAND Corporation. Source recordSupports: importance of protecting frontier weights; 38 attack vectors in nine categories; five security levels; attackers up to nation-state operations; recommendations · abstract; key takeaways; recommendations; ch. 5, Table 5.1
- BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: strong security to keep weights in a data centre, plus close monitoring, as a way to monitor inference · executive summary, Verifying various policy goals
- BR. Rinberg et al. (2026). Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains. arXiv. Source recordSupports: egress limits cap what can be stolen · §5.1