Model weights have not left the facility
On this page
Sources
- BS. Nevo et al. (2024). Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models. RAND Corporation. Source recordSupports: importance of protecting frontier weights; 38 attack vectors; five security levels; adversaries up to nation-states; comprehensive defences · summary
- BR. Rinberg et al. (2025). Verifying LLM Inference to Detect Model Weight Exfiltration. arXiv. Source recordSupports: steganographic exfiltration via inference server responses; security game; fixed-seed re-run and token-plausibility check; code release; <0.5% exfiltratable information at <0.01% FPR on MoE-Qwen-30B under benign prompts; >200x adversary slowdown · §1 Contributions; §5–7
- BN. Kezins (2026). Adversarial Entropy Inflation Against Gumbel-Based Inference Verification. arXiv. Source recordSupports: prompt-controlling adversary roughly doubles bits leaked per token; slowdown falls to 60–118x · abstract; results
- BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: strong security to keep weights in a data centre plus close monitoring · data-centre security discussion
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: verifier aims to exfiltrate prover secrets; only commitments leave the facility; egress explainable by ingress; memory wiping; side-channel target · threat model; architecture; open problems
- CN. Cankaya (2026). Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses. MIRI Technical Governance Team. Source recordSupports: physical side channels can bypass network monitoring; defences · channels of concern; defences
- BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: confidentiality as protecting models, data and code from theft; whistleblower and interview layers · §1; §4
- CN. Cankaya (2026). The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use. The Datacenter Lie Detector. Source recordSupports: active taps scrubbing headers against covert channels on front-end links · frontend vs backend
- BR. Rinberg et al. (2026). Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains. arXiv. Source recordSupports: egress limits cap what can be stolen · §5.1
- BN. Cankaya (2026). Bit-Exact AI Inference Verification Without Performance Tradeoffs. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: approximate output matching leaves degrees of freedom, including steganography, that covert adversaries can exploit · abstract
- BS. F. Comer et al. (2026). Highly Secure Inference Data Centers: A Vertically Integrated Strategy for Security Engineering. RAND Corporation (Research Report RR-A4827-1). Source recordSupports: secure inference data centre design against state-backed attackers; no external verification path described · Summary; ch. 1
- BN. Cankaya et al. (2026). Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors. arXiv. Source recordSupports: residual covert egress of about 40 Mbit/s for a 200k-GPU inference cluster at about 0.1 bits per token after replay checks · §5.2