Mechanism · Bounding unexplained information in outputs

Sources

On this page
  1. BJ. Petrie & Y. Mühlhäuser (2026). Verifying AI Compute by Bounding Unexplained Information Exfiltration. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: architecture: isolation, interlock, commitments, challenge-based prediction; principle; three confidentiality options; stage of work · abstract (read via the ICML 2026 virtual poster page; the OpenReview PDF was not reachable)
  2. BR. Rinberg et al. (2025). Verifying LLM Inference to Detect Model Weight Exfiltration. arXiv. Source recordSupports: fixed-seed sampling likelihood; theorems; threat model and assumptions; results; code release; side channels out of scope · abstract; §4; §5 (Theorems 5.1-5.3); §6; Appendix E
  3. BN. Kezins (2026). Adversarial Entropy Inflation Against Gumbel-Based Inference Verification. arXiv. Source recordSupports: entropy-inflation attack, results and recommended mitigation; applicability to unexplained-information architectures · abstract; introduction; method; conclusion
  4. BR. Rinberg et al. (2026). Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains. arXiv. Source recordSupports: compressibility of LLM text; egress limiting rationale; dual-use note · abstract; §5.1; §5.3
  5. CR. Greenblatt (2024). Preventing model exfiltration with upload limits. AI Alignment Forum. Source recordSupports: upload limits with compression against a weaker model; below 1 bit per token; assumptions; hidden-distillation route; author's uncertainty and probability estimate · whole post
  6. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: sanitized egress and one-bit fault leakage; side-channel suppression target; exact-replay metadata · §4.3.3; §5.2.2; §5.3.1
  7. BN. Cankaya et al. (2026). Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors. arXiv. Source recordSupports: residual covert egress of about 40 Mbit/s for a 200k-GPU inference cluster at about 0.1 bits per token after replay checks · §5.2

Search

Full search page