Implementation · GPU contention probes · Draft
Evidence & limits
On this page
R2Demonstrated for detecting another workload running on the same GPU
ReadinessMedium confidence
Published experiments on T4 and H100 GPUs show that solve times shift when language models run alongside, but only on single GPUs and without detection thresholds.
Assessed use: detecting another workload running on the same GPU
Rubric assessment
- R1 met: the paper describes the three probes, what each measures and a threat model in which host and GPU firmware "may be modified, virtualized, or colluding" 1.
- R2 met through reproducible published results. The paper gives the probe parameters and reports solve-time distributions on T4 and H100 GPUs with and without co-running language models 1. The hardware is realistic and the adversary is stated. No code is linked, which the rubric does not require.
- R3 not met: no production-grade tool is publicly documented, and no source reports a party other than the authors relying on the probes for a verification decision.
Confidence is medium: the results come from one group, on single GPUs, without quantified error rates.
Gaps to the next level
- Production-grade probe tooling, or use by a party other than the authors for a verification decision.
- Detection thresholds with measured false-positive and false-negative rates.
- Results on multi-GPU servers and against an operator who tries to hide a workload.
Assessed 2026-10-05 against rubric v1.1.
Evidence
- T4. With the hash-search and delay-function probes, solve times rose when language models ran on the same GPU, and larger models caused larger delays 1.
- H100. The matrix-multiplication probe slowed in the same way under co-running models 1.
- Overhead. The authors report that the three probes add power draw, occupy GPU memory and reduce throughput when run continuously 1.
Limitations
- No thresholds. The paper leaves hardware-specific thresholds to future work and does not quantify error rates 1.
- No proof of execution. The authors state that the measurements are "not designed to deliver cryptographic proof of correct execution" 1.
- No device identity. The probes do not distinguish individual GPUs of the same model 1.
- Leakage. The authors note that response timing may correlate with the GPU's background workloads 1.
Known flaws
Blockers
Continuous probes add power draw, occupy GPU memory and reduce inference throughput.
No thresholds or statistical tests define when a timing shift counts as a detection.