Mechanism · Sampled inference recomputation

Evidence & limits

On this page

R3In production for checking untrusted workers' activations against the declared model, prompt and precision

TOPLOC has been used in production by its developer, but no independent security audit or red-team of TOPLOC has been published.

Assessed use: checking untrusted workers' activations against the declared model, prompt and precision

Rubric assessment

  • R1 met: the design, the claim it verifies and its trust assumptions are published. They include a formal security game for steganographic weight exfiltration 1.
  • R2 met through DiFR and TOPLOC. DiFR has MIT-licensed code with a vLLM integration 10. It has reproducible results on A100 and H200 GPUs with models of 8 to 30 billion parameters 2. TOPLOC is an installable package 8 with peer-reviewed results 7. Rinberg et al. evaluate the method against a stated adversary who controls the inference server 1.
  • R3 met through TOPLOC: it is publicly available 8, and Prime Intellect, its developer, reports using it in two 2025 runs to accept or reject work from untrusted nodes 9 13. No other party is documented relying on sampled recomputation for a verification decision. This checks forward-pass consistency. Token-level checks of sampling settings through DiFR remain at R2 2.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of the consistency check has been published. Amodo lists recomputation red-teaming as not started 15. Amodo's tests of Token-DiFR are independent of its authors but used two fine-tuned models, which by Amodo's account cover only a fraction of possible attacks 14. The one independent attack study targets the weight-exfiltration bound built on this mechanism 16. That bound is a supporting use, assessed under Bounding unexplained information in outputs.

Confidence is low: R3 rests on the developer's own reports of TOPLOC use, the latest from July 2025 13.

Gaps to the next level
  • An independent audit, red-team or peer-reviewed security analysis of a recomputation consistency check against an adaptive adversary.

Assessed 2026-10-08 against rubric v1.1.

Mechanism properties

Threat modelAdversarial prover
Adversarial evaluationAnalysis
Hardware neededNone
Prover cooperationRequired
ConfidentialityPartial

Evidence

  • DiFR detected 4-bit quantization with an AUC above 0.999 within 300 output tokens in its authors' tests 2. They tested Llama 3.1 8B, Qwen3-8B and Qwen3-30B-A3B on A100 and H200 GPUs, with and without tensor parallelism 2. The code is public 10.
  • TOPLOC had no false positives or negatives in the evaluations reported in its peer-reviewed paper. The hardware tests used A100 and RTX 4090 GPUs 7. Prime Intellect reports using TOPLOC to verify work from untrusted inference workers in a decentralized training run. It evicted nodes whose work failed the checks 9. In a later data-generation run on 1,253 GPUs, it reports a false-positive rate of 0.000925% (37 slashes) over 4 million samples 13.
  • Rinberg et al. tested weight-exfiltration detection on open-weight models of 3 to 30 billion parameters 1. On a 30-billion-parameter mixture-of-experts model under benign prompt traffic, they report that their detector cut exfiltratable information to under 0.5% at a false-positive rate below 0.01%. That slows an adversary more than 200-fold 1. An independent study found that choosing the prompts roughly doubles the leakage per token and cuts the slowdown to 60–118 times 16.
  • Amodo built prototypes around the DiFR library 12 14. The second version adds a traffic-capturing proxy and a ledger on Kubernetes 11 12. It tested models of 1.5 to 120 billion parameters on H100 and H200 GPUs and reported a 2–8 times performance advantage for the verifier 11. Amodo also reports that TOPLOC and DiFR are well described and are now being tested on relevant hardware and models 15.

Limitations

  • Tolerance margins leave a small covert channel 1 5. An adversary who controls the prompts can widen it 16.
  • In DiFR's tests on Qwen3-30B-A3B, pooling A100 and H200 reference runs hid the two smallest tested deviations from Token-DiFR at the target false-positive rate, until matched environments or tail-weighted pooling were used 2.
  • DiFR's authors show that tuning the sampling temperature fools a naive cross-entropy check. Token-DiFR still detects the change 2.
  • TOPLOC's authors state it cannot detect speculative decoding with a cheaper model 7. Amodo expects recomputation algorithms to need updating as model architectures and inference techniques change 15.
  • TOPLOC's authors name spoofed last-layer activations and deliberately "unstable" prompts as possible attacks, which remain untested 7.
  • A mismatch does not reveal whether evasion, a hardware fault or an evaluation error caused it 4.
  • Amodo rates recomputation red-teaming as not started, and security of the recomputation server as not on track 15.

Zero-knowledge proofs of inference are an alternative. They need deterministic computation as a precondition 5.

Known flaws

Blockers

  • In tap-based retrofit designs, recording all inference traffic needs network taps and recomputation servers that can ingest it, in the worst case one recomputation-server network interface per inference front-end interface.

  • In retrofit designs, the recomputation server must sit inside the prover's data centre, possibly under the prover's physical control, and still be protected from a compromised provider, which Amodo rates 'not on track'.

  • No independent red-team of a recomputation consistency check has been published (the one independent attack study targets the weight-exfiltration bound), and Amodo rates recomputation red-teaming 'not started'.

  • Tolerance-based checks need calibration on trusted hardware and exact knowledge of the provider's sampling procedure, and in one prototype a sampling-implementation mismatch produced large spurious differences.

  • The verifier needs the model weights, so checking a closed-weights model requires a trusted, confidential recomputation environment, which the retrofit designs place inside the prover's facility.

Search

Full search page