Mechanism · Timed challenge-response and memory-occupation challenges
Evidence & limits
On this page
R2Demonstrated for detecting whether a GPU is doing other work
R2 for detecting whether a GPU is doing other work: published experiments on T4 and H100 GPUs show that challenge response times reveal co-running models and data residency, but only on single GPUs, and no challenge that bounds free memory across servers has been shown.
Assessed use: detecting whether a GPU is doing other work
Rubric assessment
- R1 met: the MIRI overview describes memory challenges for verifying the presence of information and the absence of free memory, with timing figures and assumptions 1. The AI 2040 plan names memory-challenge verification as a possible direction 2. The underlying primitives (timed attestation, proofs of space and proofs of secure erasure) are peer-reviewed 5 8 9 11.
- R2 met through reproducible published results. Monfared et al. describe four verifier-issued challenge probes with their parameters and sample counts, under a threat model in which host and GPU firmware "may be modified, virtualized, or colluding" 3. On T4 and H100 GPUs, solve times rise when language models run alongside, and a VRAM-residency challenge separates data in GPU memory from data in host memory by more than 350 ms 3. No code is linked, which the rubric does not require. SAGE shows timed software attestation on A100 GPUs for trusted execution, not for detecting other work 4. The level rests on GPU contention probes and VRAM-residency challenge, the implementations that report these results. Data-centre memory challenging and Low-trust AI compute verification system overview are proposed designs without results.
- R3 not met: no production-grade challenge tool is publicly documented, and no source reports a party other than the authors relying on such challenges for a verification decision. For bounding spare memory (This compute runs inference, not training), the MIRI overview states that, to its author's knowledge, a network-level timing probe of memory contents between servers "has not yet been demonstrated" 1.
Confidence is medium: the results come from one group's single-GPU experiments without quantified error rates.
- Production-grade challenge tooling, or use by a party other than the developers for a verification decision.
- A network-level challenge that bounds free memory across accelerator servers, with public code or measurements described in enough detail to repeat.
- Quantified false-positive and false-negative rates under adversarial conditions.
- Evaluation against known attack classes on timed attestation, such as compression and relocation.
Assessed 2026-09-25 against rubric v1.1.
Mechanism properties
| Threat model | Adversarial prover |
|---|---|
| Adversarial evaluation | Analysis |
| Hardware needed | None |
| Prover cooperation | Required |
| Confidentiality | Preserving |
Evidence
- Contention probes. Monfared et al. report that proof-of-work-style and verifiable-delay challenges on a T4, and matrix-multiplication challenges on an H100, take longer when language models run alongside, and that larger models cause larger delays 3. See GPU contention probes.
- GPU memory residency. On an H100 with a 60 GB challenge dataset, Monfared et al. report that the gap between memory-resident and host-resident responses "exceeds 350 ms, making them trivial to distinguish" 3. See VRAM-residency challenge.
- GPU attestation. SAGE, a peer-reviewed software-based attestation mechanism for A100 GPUs, is reported by its authors to be "already practical today" for trustworthy execution without special hardware support 4. See SAGE.
- Memory wiping. Amodo's wiping design includes a timed challenge phase with an assumed 1 ms round trip, but its July 2026 analysis left the challenge-phase calculations for later 10.
- Across servers. The MIRI overview states that, to its author's knowledge, distinguishing memory contents between servers with a network-level timing probe "has not yet been demonstrated" 1. See Data-centre memory challenging.
Limitations
The limits differ by scheme:
- Disputed embedded-scheme attacks. Castelluccia et al. implemented attacks based on a return-oriented rootkit and on code compression, together with specific attacks on SWATT and ICE-based schemes 6. They conclude that secure time-based attestation is "very difficult, if not impossible, to design correctly" 6. Perrig and van Doorn dispute the rootkit and SWATT attacks' applicability to the original schemes, while accepting the ICE attack 7. Perito and Tsudik cite such weaknesses as motivation for proofs of secure erasure 11.
- Coverage. Castelluccia et al. argue that "all memories (RAM, ROM, EEPROM) have to be attested" 6.
- Overhead. The VRAM-residency test "intentionally incurs substantial memory overhead" 3, and filling a pod's volatile memory takes tens of minutes 1.
- Unquantified error rates. Monfared et al. do not define thresholds or statistical tests 3.
Known flaws
Blockers
No network-level memory challenge across data-centre servers has been demonstrated.
Challenges that fill memory displace workloads; filling a pod's volatile memory takes tens of minutes and SSDs take hours.
Outside help, such as remote memory, must be excluded during challenges.