Implementation · VRAM-residency challenge · Draft

Evidence & limits

On this page

R2Demonstrated for detecting whether verifier-supplied data remains in a single GPU's memory

A published experiment on an H100 separates data held in GPU memory from data held in host memory by more than 350 ms, on one GPU and without a detection threshold.

Assessed use: detecting whether verifier-supplied data remains in a single GPU's memory

Rubric assessment

  • R1 met: the paper describes the protocol, the claim it supports and a threat model in which host and GPU firmware "may be modified, virtualized, or colluding" 1.
  • R2 met through reproducible published results. The paper gives the hash configuration, the challenge schedule and the dataset size, and reports the timing gap on an H100 1. The hardware is realistic and the adversary is stated. No code is linked, which the rubric does not require.
  • R3 not met: no production-grade tool is publicly documented, and no source reports a party other than the authors relying on the test for a verification decision.

Confidence is medium: the result comes from one group and one GPU, without quantified error rates.

Gaps to the next level
  • Production-grade tooling, or use by a party other than the authors for a verification decision.
  • Detection thresholds with measured false-positive and false-negative rates.
  • A test across servers, where the verifier is not on the same host as the GPU.

Assessed 2026-10-05 against rubric v1.1.

Evidence

  • H100. With 60 GB of challenge data, the gap between responses from GPU memory and responses from pinned host memory "exceeds 350 ms, making them trivial to distinguish" 1.
  • Overhead. The authors report negligible power overhead and negligible throughput loss for this test 1.

Limitations

  • Memory cost. The test "intentionally incurs substantial memory overhead" 1.
  • Verifier data only. The challenge runs over a block that the verifier supplied, not over the prover's own data 1.
  • No thresholds. The paper leaves hardware-specific thresholds to future work and does not quantify error rates 1.
  • Single GPU. The experiments ran on single T4 and H100 GPUs 1. A design for challenges across data-centre servers is covered in Data-centre memory challenging.

Known flaws

Blockers

Search

Full search page