GPU contention probes
GPU contention probes are three timed puzzles from a 2026 study by Monfared and colleagues.
A verifier sends a puzzle, the GPU solves it, and the verifier times the answer. One puzzle is a memory-hard hash search, one is a chain of sequential squarings, and one is a large matrix multiplication on the tensor cores. Each loads a different part of the GPU, so another workload on the same device slows the answer.
On T4 and H100 GPUs, solve times rose when language models ran alongside, and larger models caused larger delays. The probes need no trusted hardware, and the study considers a host and GPU that may be modified. The authors define no detection thresholds and do not quantify error rates.
The probes do not identify which GPU answered. Running them continuously costs power and throughput. The paper links no code.
Published experiments on T4 and H100 GPUs show that solve times shift when language models run alongside, but only on single GPUs and without detection thresholds.
Assessed use: detecting another workload running on the same GPU
On this page
What it is
GPU contention probes are three timed puzzles described by Monfared, Ganji, Tajik and Holcomb in a 2026 preprint 1. They are one application of timed challenge-response. A verifier sends a puzzle to a GPU and measures how long the answer takes 1. The paper's fourth probe checks memory instead of compute and is covered in VRAM-residency challenge.
How it works
Each puzzle loads a different part of the GPU 1:
- Hash search. The GPU searches for an input whose memory-hard hash falls below a target. Solve time reflects parallel effort 1.
- Delay function. The GPU runs chains of sequential modular squarings. Solve time reflects sequential execution 1.
- Matrix multiplication. The GPU multiplies large pseudorandom matrices on its tensor cores. Solve time reflects tensor-core throughput 1.
A workload that shares the GPU competes for the same units, so its presence shifts the distribution of solve times 1.