The declared evaluation was run
On this page
Sources
- BC. Schnabl et al. (2025). Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments. ICML 2025 Workshop on Technical AI Governance. Source recordSupports: audit attestation binds model, audit code and data, and result; separate inference protocol · §3; Algorithms 2–3
- BP. Chantasantitam et al. (2026). PAL*M: Property Attestation for Large Generative Models. arXiv. Source recordSupports: proof of evaluation binds model, tokenizer, evaluation data and metric; hardware and threat model · §3.2; §4.3.3; Table 5
- Bcovehub (2026). Cove: Compositional Multi-Party Confidential Workflows for Verifiable AI Governance (reference implementation). GitHub. Source recordSupports: Cove certificates, reviewed node manifests and trust assumptions · docs/internal/architecture.md; docs/internal/security_model.md
- BA. Trask et al. (2026). Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing. Google DeepMind. Source recordSupports: double-blind evaluation workflow and pilot; accepted proprietary code; guest OS and cloud verification assumptions · §2.5; §3; §4