Organization · Nonprofit
Machine Intelligence Research Institute
A nonprofit focused on preventing human extinction from artificial superintelligence; its Technical Governance Team publishes designs and analyses for verifying AI agreements.
intelligence.org · also called MIRI
MIRI describes itself as "a nonprofit focused on preventing human extinction from artificial superintelligence". Its publications on verifying AI agreements, several of them from its Technical Governance Team, include:
- Low-trust verification system. Cankaya's system overview combines network taps, sampled recomputation, memory challenges and facility monitoring in one reference architecture for near-term, low-trust compute verification 1. See Low-trust AI compute verification system overview, Network taps and certifiers, Sampled inference recomputation and Timed challenge-response and memory-occupation challenges.
- Side-channel suppression. A post on suppressing side channels in an untrusted data centre with retrofitted defences sets out channel classes, attenuation targets and costs 2; see Side-channel suppression for isolated facilities.
- TEEs for monitoring. A post written in MIRI's Technical Governance Fellowship analyses trusted execution environments for privacy-preserving monitoring 3. It contrasts the usual confidential-computing adversary, a dishonest operator, with a state that has physical access to data centres 3. See TEE remote attestation for AI workloads, Safeguard attestation and Confidential multi-party verification.
- Mechanisms survey. Scher and Thiergart survey mechanisms for verifying international agreements about AI development, with a detailed analysis of interconnect bandwidth limits 4.
- Draft agreement. A draft international agreement would prohibit concentrations of more than 16 H100-equivalents outside monitored facilities and would track new chip production 5; see Compute stock is at most a declared amount.
Implementations
- A retrofittable reference design in which network taps commit to all facility traffic, and air-gapped, independently sourced checkers later re-run randomly challenged records.
Mechanisms
Mechanisms this organization has designed, built, evaluated or supplied.
- Lets mutually distrusting parties run an agreed check over private models or records inside attested enclaves or zero-knowledge proofs, revealing only the result.
- Establishes that responses come from a specific, committed set of model weights, using enclave measurements or recomputation of sampled outputs.
- Uses timing, memory-residency and performance-counter signals measured on AI accelerators as evidence about which workloads they are running.
- Hardware-signed evidence that an AI service ran its declared safeguards, such as a guardrail classifier or monitor, when producing a given response.
- A verifier re-runs a random sample of an AI provider's logged queries on a trusted copy of the declared model and checks the outputs match.
- Trusted execution environments (TEEs) in CPUs and GPUs sign reports of loaded software, so a remote party can check which code ran an AI workload.
- Capping or removing the network links between groups of accelerators, so that serving models still works but large training runs become impractically slow.
- Overwriting all of a device's memory in a way a verifier can check, so that nothing from earlier, undeclared work survives the wipe.
- Devices on a cluster's network links that copy and hash all traffic, so a verifier can later check sampled records against declared work.
- Shielding, filtering, jamming and inspecting an AI facility so that no hidden physical channel can bypass the checks placed on its official links.
- A verifier sends unpredictable questions that a device can answer in time only if it holds specified data, or dedicates specified resources, locally.
Publications
Sources this organization authored or published.
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. RecordCited by Bandwidth limits and compartmentalization; Bounding unexplained information in outputs; Deterministic and bit-exact inference; Memory wiping and proofs of secure erasure; Network taps and certifiers; Proofs of useful work and resource exhaustion; Safeguard attestation; Sampled inference recomputation; Side-channel suppression for isolated facilities; Tamper evidence for verifier devices; Timed challenge-response and memory-occupation challenges; Zero-knowledge proofs of inference; Low-trust AI compute verification system overview; zkLLM; Communication between compute groups is bounded; The declared model is the one being served; This compute runs inference, not training; Declared safeguards were applied during inference; A training run stayed within declared limits; Model weights or data have not left the facility; Compartmentalization; Cryptographic commitment; Evidence binding; FLOP accounting; Inference and training workloads; Interconnect bandwidth; Network tap; Numerical nondeterminism; Positive and negative claims; Proof of space; Recomputation; Sampling and assurance; Side channel; Threat model; Undeclared compute; Verifier; Weight exfiltration; Machine Intelligence Research Institute
- CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. RecordCited by Confidential multi-party verification; Model identity attestation; On-chip telemetry from timing, memory and performance counters; Safeguard attestation; TEE remote attestation for AI workloads; The declared model is the one being served; There is no undeclared relevant compute; Declared safeguards were applied during inference; Root of trust; Side channel; Trusted execution environment (TEE); Machine Intelligence Research Institute
- CN. Cankaya (2026). Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses. MIRI Technical Governance Team. RecordCited by Side-channel suppression for isolated facilities; Communication between compute groups is bounded; Model weights or data have not left the facility; Compartmentalization; Side channel; Weight exfiltration; Machine Intelligence Research Institute
- BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. RecordCited by Chips are where they are declared to be; Compute stock is at most a declared amount; This compute runs inference, not training; There is no undeclared relevant compute; A training run stayed within declared limits; FLOP accounting; Machine Intelligence Research Institute
- BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. RecordCited by Proofs of useful work and resource exhaustion; Communication between compute groups is bounded; Chips are where they are declared to be; Compute stock is at most a declared amount; Declared hardware is idle or shut down; This compute runs inference, not training; There is no undeclared relevant compute; Model weights or data have not left the facility; Compartmentalization; Evidence binding; Inference and training workloads; Interconnect bandwidth; Positive and negative claims; Proof of (useful) work; Remote attestation; Side channel; Undeclared compute; Machine Intelligence Research Institute
Sources
- BN. Cankaya (2026). A System Overview for Near-Term, Low-Trust AI Compute Verification. Machine Intelligence Research Institute. Source recordSupports: low-trust compute verification system overview
- CN. Cankaya (2026). Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses. MIRI Technical Governance Team. Source recordSupports: side-channel suppression with retrofitted defences
- CGloria Z (2026). On TEEs for Privacy-Preserving Monitoring in AI Governance. MIRI Technical Governance Team. Source recordSupports: analysis of TEEs for privacy-preserving monitoring (Technical Governance Fellowship)
- BA. Scher & L. Thiergart (2025). Mechanisms to Verify International Agreements About AI Development. arXiv. Source recordSupports: survey of mechanisms to verify international agreements; interconnect bandwidth limits
- BA. Scher et al. (2025). An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. Machine Intelligence Research Institute. Source recordSupports: draft international agreement: monitored facilities and tracking of chip production