Declared safeguards were applied during inference

On this page

Sources

  1. BM. Baker et al. (2025). Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment. RAND Corporation. Source recordSupports: deployment mitigations specified by inputs and outputs (filters, oversight checks); difficulty of choosing technical rules; whistleblower and interview layers · Table 3; §4
  2. BA. R. Wasil et al. (2024). Verification methods for international AI agreements. arXiv. Source recordSupports: AI developer inspections to check authorised code and implementation of evaluations and safeguards; software can be quickly modified or hidden · Access-dependent methods; Table 1
  3. BM. Brundage et al. (2026). Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies. arXiv. Source recordSupports: third-party verification of developers' safety and security claims with deep, secure access; AI Assurance Levels · abstract
  4. CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: TEEs for policy-adherent inference; launch measurement must cover all components; attestation-key holder can forge reports · main argument; Limitations
  5. BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: TEE-based verifiable benchmarks keeping model IP and datasets confidential; prototype with Llama-3.1 · abstract
  6. BP. Chantasantitam et al. (2026). PAL*M: Property Attestation for Large Generative Models. arXiv. Source recordSupports: property attestation across training and inference on TDX + H100; formal model in Tamarin · abstract
  7. BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: output cross-checks; inspector agents vulnerable to prompt injection · architecture; open problems
  8. BX. Jin et al. (2026). Proof-of-Guardrail in AI Agents and What (Not) to Trust from It. arXiv. Source recordSupports: proof-of-guardrail protocol, prototype on AWS Nitro Enclaves with an external guardrail API, tamper tests, opt-in attestation, jailbreak risk · abstract; §3; §4.1; Table 1; Appendix A
  9. BB. Penchas et al. (2026). Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: monitoring scoped to a jointly signed plan and executed in an attested TEE · abstract
  10. CR. Rinberg & B. Penchas (2026). Auditor-in-a-Box: Tools for Third-Party Auditing. LessWrong. Source recordSupports: Auditor-in-a-Box demo of plan-scoped monitoring; authors state user data and plan execution are not actually secure and the demo has not been stress-tested by a counterparty · whole post

Search

Full search page