Mechanism · Safeguard attestation
Sources
On this page
- BX. Jin et al. (2026). Proof-of-Guardrail in AI Agents and What (Not) to Trust from It. arXiv. Source recordSupports: problem statement; protocol; threat model and trust assumptions; implementation and overheads; tamper tests; guardrail accuracy on ToxicChat; jailbreak risk; proof-of-safety caveat; external APIs; wrapper-bypass risk; selective attestation · abstract; §3; §4.1; Tables 1-3; Appendix A
- BSaharaLabsAI (2026). Verifiable-ClawGuard: proof-of-guardrail reference code. GitHub. Source recordSupports: public code; proof-of-concept status; stated limitation on agent command execution · README, including Limitations
- CD. McCann-Sayles et al. (2026). Safety Without Compromising on Privacy. Tinfoil blog. Source recordSupports: Tinfoil's enclave-run safeguard pipeline: models, what leaves the enclave, open-source and attested code, rollout status · whole post (updated 2026-09-16)
- BP. Chantasantitam et al. (2026). PAL*M: Property Attestation for Large Generative Models. arXiv. Source recordSupports: inference property definitions; TDX+H100 implementation; overheads and baseline; threat model; cherry-picking discussion · abstract; §3.2; §4.3.4 (Defs. 7-8); §4.4; Table 6; Appendix A
- BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: inference protocol linking model, audit result, prompt and response; CPU-only enclave prototype; CPU versus GPU cost and slowdown · §3 (Inference protocol); §5; Table 2
- BB. Penchas et al. (2026). Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute. ICML 2026 Workshop on Technical AI Governance Research. Source recordSupports: verifiably-scoped monitoring protocol · abstract
- CR. Rinberg & B. Penchas (2026). Auditor-in-a-Box: Tools for Third-Party Auditing. LessWrong. Source recordSupports: reference implementation in Tinfoil confidential VMs; stated limitations · reference implementation; limitations
- CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: policy adherence as a verification property; measurement completeness; runtime state; second-CVM completeness gap; vendor root of trust; treaty threat model; GPU TEE maturity and multi-GPU inference · deployment integrity; hardware auditability; resource accounting; physical attack surface
- BM. Brundage et al. (2026). Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies. arXiv. Source recordSupports: configuration drift, such as swapping safety classifiers or relaxing filter thresholds · §5.2
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: verifier-side compliance screening of re-executed records · §3.2.2; §5.2.3
- BR. Levin et al. (2025). Has My System Prompt Been Used? Large Language Model Prompt Membership Inference. arXiv. Source recordSupports: black-box statistical test for system-prompt use; its prompt-protection setting · abstract; §3.2
- AJ. Chuang et al. (2026). TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition. 2026 IEEE Symposium on Security and Privacy (SP). Source recordSupports: physical key extraction from TDX and signing-key extraction in SEV-SNP; forged attestations against NVIDIA GPU confidential computing; cost; vendor acknowledgement and positions · project site summary; paper abstract and disclosure
- AJ. De Meulemeester et al. (2026). DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes. 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26). Source recordSupports: DDRop forges attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer · site summary
- AJ. De Meulemeester et al. (2026). Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing. 47th IEEE Symposium on Security and Privacy (S&P 2026). Source recordSupports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer
- AB. Schlüter & S. Shinde (2025). RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP. 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25). Source recordSupports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor
- BAMD (2025). SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020). AMD product security bulletin. Source recordSupports: AMD firmware fixes for RMPocalypse (CVE-2025-0033)
- CTinfoil Team (2026). How Tinfoil Proves Exactly What Model Is Running. Tinfoil. Source recordSupports: binding model weights to enclave attestation · whole post