{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "I-0018",
    "slug": "gpu-contention-probes",
    "title": "GPU contention probes",
    "aliases": [
      "GPU timing probes",
      "Proof-of-work, delay-function and matrix-multiplication probes"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-05",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": []
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "Three timed puzzles for GPUs whose solve times lengthen when another workload shares the device, showing a verifier that the GPU is busy.",
    "summary": "GPU contention probes are three timed puzzles from a 2026 study by Monfared and colleagues. A verifier sends a puzzle, the GPU solves it, and the verifier times the answer. One puzzle is a memory-hard hash search, one is a chain of sequential squarings, and one is a large matrix multiplication on the tensor cores. Each loads a different part of the GPU, so another workload on the same device slows the answer. On T4 and H100 GPUs, solve times rose when language models ran alongside, and larger models caused larger delays. The probes need no trusted hardware, and the study considers a host and GPU that may be modified. The authors define no detection thresholds and do not quantify error rates. The probes do not identify which GPU answered. Running them continuously costs power and throughput. The paper links no code.",
    "technical": "- **Hash search.** A custom CUDA implementation of Argon2id runs with 1 pass, 1 lane and 1 MiB of working memory per instance. The difficulty targets about one valid hash per 2^24 trials [[S-0033]].\n- **Delay function.** The reported runs use 256 parallel instances of repeated modular squaring, each of 2^20 steps [[S-0033]].\n- **Matrix multiplication.** The square matrices have dimension 32,768, with FP16 multiplication and FP32 accumulation. Results are checked with Freivalds' algorithm over 5 rounds [[S-0033]].\n- **Co-running models.** The contention experiments used TinyLlama-1.1B at about 2 GB, Qwen2.5-7B at about 9 GB and Llama-2 in FP16 at about 12 GB [[S-0033]].",
    "category": "cryptographic-computational",
    "secondary_categories": [],
    "verifies": [
      {
        "claim": "C-0003",
        "role": "primary",
        "note": "Solve times lengthen when another workload runs on the GPU (S-0033)."
      }
    ],
    "threat_model": "adversarial",
    "adversarial_evaluation": "analysis",
    "hardware_requirement": "none",
    "prover_cooperation": "required",
    "confidentiality": "partial",
    "depends_on": [],
    "readiness": {
      "assessment": true,
      "level": "R2",
      "scope": "detecting another workload running on the same GPU",
      "rubric_version": "1.1",
      "rationale": "Published experiments on T4 and H100 GPUs show that solve times shift when language models run alongside, but only on single GPUs and without detection thresholds.\n\n- **R1** met: the paper describes the three probes, what each measures and a threat model in which host and GPU firmware \"may be modified, virtualized, or colluding\" [[S-0033]].\n- **R2** met through reproducible published results. The paper gives the probe parameters and reports solve-time distributions on T4 and H100 GPUs with and without co-running language models [[S-0033]]. The hardware is realistic and the adversary is stated. No code is linked, which the rubric does not require.\n- **R3** not met: no production-grade tool is publicly documented, and no source reports a party other than the authors relying on the probes for a verification decision.\n\nConfidence is medium: the results come from one group, on single GPUs, without quantified error rates.",
      "evidence": [
        "S-0033"
      ],
      "next_level_gaps": [
        "Production-grade probe tooling, or use by a party other than the authors for a verification decision.",
        "Detection thresholds with measured false-positive and false-negative rates.",
        "Results on multi-GPU servers and against an operator who tries to hide a workload."
      ],
      "confidence": "medium",
      "assessed_by": [
        "ai-draft"
      ],
      "assessed_on": "2026-10-05",
      "status": "current",
      "dispute": null
    },
    "flaws": [
      {
        "assessment": true,
        "title": "Error rates not quantified",
        "kind": "open-question",
        "severity": "significant",
        "status": "open",
        "description": "Monfared et al. show timing distributions that shift under contention, but leave hardware-specific thresholds to future work and state that false-positive and false-negative rates are not quantified.",
        "sources": [
          "S-0033"
        ],
        "response": null
      },
      {
        "assessment": true,
        "title": "Answers are not tied to one GPU",
        "kind": "theoretical-argument",
        "severity": "significant",
        "status": "open",
        "description": "The paper's floating-point fingerprint characterises a GPU model. The authors state that it does not distinguish individual GPUs, so a probe answer does not show which device of that model produced it.",
        "sources": [
          "S-0033"
        ],
        "response": null
      }
    ],
    "blockers": [
      {
        "text": "Continuous probes add power draw, occupy GPU memory and reduce inference throughput.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-0033"
        ]
      },
      {
        "text": "No thresholds or statistical tests define when a timing shift counts as a detection.",
        "theme": "adversarial-validation",
        "blocked_by": null,
        "sources": [
          "S-0033"
        ]
      }
    ],
    "challenge_themes": [
      "adversarial-validation",
      "performance-compatibility",
      "evidence-binding"
    ],
    "organizations": [],
    "people": [],
    "sources": [
      {
        "source": "S-0033",
        "supports": "probe designs and parameters; threat model; T4 and H100 contention results; overheads; limitations",
        "locator": "§3; §4; §6.1–§6.4; §7"
      }
    ],
    "concepts": [
      "K-0001",
      "K-0002",
      "K-0011",
      "K-0018"
    ],
    "kind": "research-prototype",
    "developer": [],
    "realises": [
      "M-0016"
    ],
    "type": "implementation",
    "url": "https://trustbutveri.fyi/implementations/gpu-contention-probes/",
    "source_file": "content/implementations/gpu-contention-probes.md",
    "flags_all": [],
    "body_markdown": "## What it is\n\nGPU contention probes are three timed puzzles described by Monfared, Ganji, Tajik and Holcomb in a 2026 preprint [[S-0033]]. They are one application of [[M-0016|timed challenge-response]]. A verifier sends a puzzle to a GPU and measures how long the answer takes [[S-0033]]. The paper's fourth probe checks memory instead of compute and is covered in [[I-0019]].\n\n## How it works\n\nEach puzzle loads a different part of the GPU [[S-0033]]:\n\n- **Hash search.** The GPU searches for an input whose memory-hard hash falls below a target. Solve time reflects parallel effort [[S-0033]].\n- **Delay function.** The GPU runs chains of sequential modular squarings. Solve time reflects sequential execution [[S-0033]].\n- **Matrix multiplication.** The GPU multiplies large pseudorandom matrices on its tensor cores. Solve time reflects tensor-core throughput [[S-0033]].\n\nA workload that shares the GPU competes for the same units, so its presence shifts the distribution of solve times [[S-0033]].\n\n## Evidence\n\n- **T4.** With the hash-search and delay-function probes, solve times rose when language models ran on the same GPU, and larger models caused larger delays [[S-0033]].\n- **H100.** The matrix-multiplication probe slowed in the same way under co-running models [[S-0033]].\n- **Overhead.** The authors report that the three probes add power draw, occupy GPU memory and reduce throughput when run continuously [[S-0033]].\n\n## Limitations\n\n- **No thresholds.** The paper leaves hardware-specific thresholds to future work and does not quantify error rates [[S-0033]].\n- **No proof of execution.** The authors state that the measurements are \"not designed to deliver cryptographic proof of correct execution\" [[S-0033]].\n- **No device identity.** The probes do not distinguish individual GPUs of the same model [[S-0033]].\n- **Leakage.** The authors note that response timing may correlate with the GPU's background workloads [[S-0033]].",
    "body_text": "What it is GPU contention probes are three timed puzzles described by Monfared, Ganji, Tajik and Holcomb in a 2026 preprint [S-0033]. They are one application of timed challenge-response. A verifier sends a puzzle to a GPU and measures how long the answer takes [S-0033]. The paper's fourth probe checks memory instead of compute and is covered in VRAM-residency challenge. How it works Each puzzle loads a different part of the GPU [S-0033]: - Hash search. The GPU searches for an input whose memory-hard hash falls below a target. Solve time reflects parallel effort [S-0033]. - Delay function. The GPU runs chains of sequential modular squarings. Solve time reflects sequential execution [S-0033]. - Matrix multiplication. The GPU multiplies large pseudorandom matrices on its tensor cores. Solve time reflects tensor-core throughput [S-0033]. A workload that shares the GPU competes for the same units, so its presence shifts the distribution of solve times [S-0033]. Evidence - T4. With the hash-search and delay-function probes, solve times rose when language models ran on the same GPU, and larger models caused larger delays [S-0033]. - H100. The matrix-multiplication probe slowed in the same way under co-running models [S-0033]. - Overhead. The authors report that the three probes add power draw, occupy GPU memory and reduce throughput when run continuously [S-0033]. Limitations - No thresholds. The paper leaves hardware-specific thresholds to future work and does not quantify error rates [S-0033]. - No proof of execution. The authors state that the measurements are \"not designed to deliver cryptographic proof of correct execution\" [S-0033]. - No device identity. The probes do not distinguish individual GPUs of the same model [S-0033]. - Leakage. The authors note that response timing may correlate with the GPU's background workloads [S-0033].",
    "referenced_by": [
      {
        "id": "M-0016",
        "title": "Timed challenge-response and memory-occupation challenges",
        "url": "https://trustbutveri.fyi/mechanisms/timed-challenge-response/"
      },
      {
        "id": "I-0019",
        "title": "VRAM-residency challenge",
        "url": "https://trustbutveri.fyi/implementations/vram-residency-challenge/"
      }
    ]
  }
}