{
  "schema_version": "1.4.0",
  "rubric_version": "1.1",
  "license": "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)",
  "record": {
    "id": "I-0019",
    "slug": "vram-residency-challenge",
    "title": "VRAM-residency challenge",
    "aliases": [
      "GPU memory residency test",
      "HBM residency challenge"
    ],
    "status": "draft",
    "last_reviewed": "2026-10-05",
    "review_interval_days": 90,
    "steward": null,
    "provenance": {
      "drafted_by": "ai",
      "reviewed_by": []
    },
    "risk_flags": [],
    "flags": [],
    "one_liner": "A timed challenge that shows whether a block of verifier-supplied data is still held in a GPU's own memory.",
    "summary": "The VRAM-residency challenge is a timed test of whether data is held in a GPU's on-board memory. It comes from a 2026 study by Monfared and colleagues. The verifier first loads a large block of challenge data into GPU memory. At random times it sends fresh nonces, and the GPU must run a keyed, memory-hard hash over the stored block and return a digest. If the block has been moved to host memory, each access crosses the PCIe link and the answer arrives late. On an H100 with 60 GB of challenge data, the gap between the two cases exceeded 350 ms. The test adds little power or throughput cost, but by design the challenge data occupies a large part of GPU memory. The study covers a single GPU. It defines no thresholds and reports no error rates. The paper links no code.",
    "technical": "- **Hash.** The test uses Argon2id in keyed mode, with a nonce-keyed BLAKE2b masking pass, at 1 pass, 1 lane and 1 MiB of working memory per instance [[S-0033]].\n- **Answer.** The outputs of the parallel instances are concatenated and reduced to one digest with a fast hash such as SHA-256 [[S-0033]].\n- **Timing.** The wait before each challenge is drawn from Uniform(0, T_max). In the reported experiment, challenge times were drawn uniformly over 120 s [[S-0033]].",
    "category": "cryptographic-computational",
    "secondary_categories": [],
    "verifies": [
      {
        "claim": "C-0003",
        "role": "primary",
        "note": "A timely answer shows that the challenge data still occupies GPU memory (S-0033)."
      }
    ],
    "threat_model": "adversarial",
    "adversarial_evaluation": "analysis",
    "hardware_requirement": "none",
    "prover_cooperation": "required",
    "confidentiality": "preserving",
    "depends_on": [],
    "readiness": {
      "assessment": true,
      "level": "R2",
      "scope": "detecting whether verifier-supplied data remains in a single GPU's memory",
      "rubric_version": "1.1",
      "rationale": "A published experiment on an H100 separates data held in GPU memory from data held in host memory by more than 350 ms, on one GPU and without a detection threshold.\n\n- **R1** met: the paper describes the protocol, the claim it supports and a threat model in which host and GPU firmware \"may be modified, virtualized, or colluding\" [[S-0033]].\n- **R2** met through reproducible published results. The paper gives the hash configuration, the challenge schedule and the dataset size, and reports the timing gap on an H100 [[S-0033]]. The hardware is realistic and the adversary is stated. No code is linked, which the rubric does not require.\n- **R3** not met: no production-grade tool is publicly documented, and no source reports a party other than the authors relying on the test for a verification decision.\n\nConfidence is medium: the result comes from one group and one GPU, without quantified error rates.",
      "evidence": [
        "S-0033"
      ],
      "next_level_gaps": [
        "Production-grade tooling, or use by a party other than the authors for a verification decision.",
        "Detection thresholds with measured false-positive and false-negative rates.",
        "A test across servers, where the verifier is not on the same host as the GPU."
      ],
      "confidence": "medium",
      "assessed_by": [
        "ai-draft"
      ],
      "assessed_on": "2026-10-05",
      "status": "current",
      "dispute": null
    },
    "flaws": [
      {
        "assessment": true,
        "title": "Error rates not quantified",
        "kind": "open-question",
        "severity": "minor",
        "status": "open",
        "description": "Monfared et al. report a timing gap of more than 350 ms but leave hardware-specific thresholds to future work and state that false-positive and false-negative rates are not quantified.",
        "sources": [
          "S-0033"
        ],
        "response": null
      },
      {
        "assessment": true,
        "title": "Answers are not tied to one GPU",
        "kind": "theoretical-argument",
        "severity": "significant",
        "status": "open",
        "description": "The paper's floating-point fingerprint characterises a GPU model. The authors state that it does not distinguish individual GPUs, so a timely answer does not show which device of that model held the data.",
        "sources": [
          "S-0033"
        ],
        "response": null
      }
    ],
    "blockers": [
      {
        "text": "The challenge data occupies a large part of GPU memory for as long as the test runs.",
        "theme": "performance-compatibility",
        "blocked_by": null,
        "sources": [
          "S-0033"
        ]
      },
      {
        "text": "No test across servers has been reported, and the MIRI overview lists network-level probing of memory contents as undemonstrated.",
        "theme": "adversarial-validation",
        "blocked_by": null,
        "sources": [
          "S-0033",
          "S-0018"
        ]
      }
    ],
    "challenge_themes": [
      "capacity-bounds",
      "adversarial-validation",
      "performance-compatibility",
      "evidence-binding"
    ],
    "organizations": [],
    "people": [],
    "sources": [
      {
        "source": "S-0033",
        "supports": "residency protocol; hash configuration; challenge schedule; H100 result; overheads; limitations",
        "locator": "§3; §5; §6.3–§6.4; §7"
      },
      {
        "source": "S-0018",
        "supports": "network-level probing of memory contents between servers not yet demonstrated",
        "locator": "§5.1.2"
      }
    ],
    "concepts": [
      "K-0001",
      "K-0002",
      "K-0012",
      "K-0018"
    ],
    "kind": "research-prototype",
    "developer": [],
    "realises": [
      "M-0016"
    ],
    "type": "implementation",
    "url": "https://trustbutveri.fyi/implementations/vram-residency-challenge/",
    "source_file": "content/implementations/vram-residency-challenge.md",
    "flags_all": [],
    "body_markdown": "## What it is\n\nThe VRAM-residency challenge is a test of whether data sits in a GPU's own memory, described by Monfared, Ganji, Tajik and Holcomb in a 2026 preprint [[S-0033]]. The paper uses VRAM and HBM, for high-bandwidth memory, to mean the GPU's on-board memory [[S-0033]]. The test is one application of [[M-0016|timed challenge-response]]. The same paper's compute probes are covered in [[I-0018]].\n\n## How it works\n\n1. **Load.** Before measurement starts, a large block of incompressible challenge data is stored in GPU memory [[S-0033]].\n2. **Challenge.** At random times the verifier sends fresh nonces [[S-0033]].\n3. **Compute.** The GPU runs a keyed, memory-hard hash over the stored block. Each memory access depends on the result of the one before, so the work is bound by memory bandwidth and cannot be computed ahead of time [[S-0033]].\n4. **Answer.** The GPU returns one digest. The verifier checks that it is correct and records the response time [[S-0033]].\n\nIf the block has been moved out of GPU memory, each access has to cross the PCIe link from host memory, and the answer arrives late [[S-0033]].\n\n## Evidence\n\n- **H100.** With 60 GB of challenge data, the gap between responses from GPU memory and responses from pinned host memory \"exceeds 350 ms, making them trivial to distinguish\" [[S-0033]].\n- **Overhead.** The authors report negligible power overhead and negligible throughput loss for this test [[S-0033]].\n\n## Limitations\n\n- **Memory cost.** The test \"intentionally incurs substantial memory overhead\" [[S-0033]].\n- **Verifier data only.** The challenge runs over a block that the verifier supplied, not over the prover's own data [[S-0033]].\n- **No thresholds.** The paper leaves hardware-specific thresholds to future work and does not quantify error rates [[S-0033]].\n- **Single GPU.** The experiments ran on single T4 and H100 GPUs [[S-0033]]. A design for challenges across data-centre servers is covered in [[I-0020]].",
    "body_text": "What it is The VRAM-residency challenge is a test of whether data sits in a GPU's own memory, described by Monfared, Ganji, Tajik and Holcomb in a 2026 preprint [S-0033]. The paper uses VRAM and HBM, for high-bandwidth memory, to mean the GPU's on-board memory [S-0033]. The test is one application of timed challenge-response. The same paper's compute probes are covered in GPU contention probes. How it works 1. Load. Before measurement starts, a large block of incompressible challenge data is stored in GPU memory [S-0033]. 2. Challenge. At random times the verifier sends fresh nonces [S-0033]. 3. Compute. The GPU runs a keyed, memory-hard hash over the stored block. Each memory access depends on the result of the one before, so the work is bound by memory bandwidth and cannot be computed ahead of time [S-0033]. 4. Answer. The GPU returns one digest. The verifier checks that it is correct and records the response time [S-0033]. If the block has been moved out of GPU memory, each access has to cross the PCIe link from host memory, and the answer arrives late [S-0033]. Evidence - H100. With 60 GB of challenge data, the gap between responses from GPU memory and responses from pinned host memory \"exceeds 350 ms, making them trivial to distinguish\" [S-0033]. - Overhead. The authors report negligible power overhead and negligible throughput loss for this test [S-0033]. Limitations - Memory cost. The test \"intentionally incurs substantial memory overhead\" [S-0033]. - Verifier data only. The challenge runs over a block that the verifier supplied, not over the prover's own data [S-0033]. - No thresholds. The paper leaves hardware-specific thresholds to future work and does not quantify error rates [S-0033]. - Single GPU. The experiments ran on single T4 and H100 GPUs [S-0033]. A design for challenges across data-centre servers is covered in Data-centre memory challenging.",
    "referenced_by": [
      {
        "id": "M-0016",
        "title": "Timed challenge-response and memory-occupation challenges",
        "url": "https://trustbutveri.fyi/mechanisms/timed-challenge-response/"
      },
      {
        "id": "I-0020",
        "title": "Data-centre memory challenging",
        "url": "https://trustbutveri.fyi/implementations/data-centre-memory-challenging/"
      },
      {
        "id": "I-0018",
        "title": "GPU contention probes",
        "url": "https://trustbutveri.fyi/implementations/gpu-contention-probes/"
      }
    ]
  }
}