Implementation · DiFR (Divergence From Reference)

Evidence & limits

On this page

R2Demonstrated for checking that outputs match the declared model, precision and sampling settings

The code is public and the results reproduce on data-centre GPUs, but only its developers rely on it and no one has independently evaluated its consistency check.

Assessed use: checking that outputs match the declared model, precision and sampling settings

Rubric assessment

  • R1 met: the paper states the verification claim, the specification the provider must follow and the trust assumptions 1. The assumptions are a trusted reference, calibration on trusted hardware and synchronized seeds. A companion paper embeds the method in a formal security game 2.
  • R2 met: public MIT-licensed code with a vLLM integration exists 3. Results are reproducible on A100 and H200 GPUs with models of 8 to 30 billion parameters 1. A separate team built it into a prototype and tested models of up to 120 billion parameters 4 5.
  • R3 not met: Amodo's prototype is research code, not production software 5, and no party is documented relying on DiFR for a verification decision.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of DiFR's consistency check has been published. Amodo lists recomputation red-teaming as not started 8. Amodo's own tests are independent of the authors but used two fine-tuned models, which by Amodo's account cover only a fraction of possible attacks 6. The one independent attack targets a weight-exfiltration detector built on the same Gumbel-margin statistic 9. That attack bears on the supporting exfiltration use, assessed under Bounding unexplained information in outputs. It does not bear on the primary use.
Gaps to the next level
  • Reliance by a party other than the developers on DiFR for a verification decision, or a production-grade release.
  • An independent public security evaluation (audit, red-team or peer-reviewed analysis) against adaptive adversaries.

Assessed 2026-09-25 against rubric v1.1.

Evidence

  • The authors tested Llama 3.1 8B-Instruct, Qwen3-8B and Qwen3-30B-A3B on 2,000 UltraChat prompts 1. The four inference configurations were H200 with four-way tensor parallelism, A100 with and without it, and H200 without it running Hugging Face 1.
  • The faults tested were FP8 key-value cache quantization, 4-bit model quantization, a temperature of 1.1 instead of 1.0, an incorrect seed, and a simulated sampling bug that picks uniformly among the top-k tokens 1% of the time 1.
  • The authors report that Token-DiFR detects 4-bit quantization with an AUC above 0.999 within 300 output tokens 1. Activation-DiFR does so within 2 tokens, while cutting communication by 25–75% relative to existing methods 1.
  • The authors report that Activation-DiFR Pareto-dominates TOPLOC in communication cost against detection accuracy when detecting FP8 KV-cache quantization 1. They also cite a report that a distributional test called RUT failed to reliably separate 8-bit quantized models. Token-DiFR detects such differences within a few thousand tokens 1.
  • The companion paper's detector limited exfiltratable information to under 0.5% at a false-positive rate below 0.01% on a mixture-of-experts Qwen 30B model, under benign prompt traffic 2. An independent study found that choosing the prompts roughly doubles the leakage per token 9.
  • In a small case study in November 2025, the authors spot-checked five public Llama 3.1 8B API deployments with greedy queries. Scores varied substantially across them. The authors read higher scores as evidence of a different effective specification, such as a different chat template, not as a direct measure of model quality 1.
  • Amodo Design, a separate team, built the DiFR library into recomputation prototypes 5 6. Its first prototype, on eight H100 GPUs, judged each request on its own. The best threshold separated the pass rates of the honest model and two models fine-tuned to append text by about 25 percentage points, and most missed cases were answers shorter than five tokens 6. A later asynchronous version was tested on models of 1.5 to 120 billion parameters on H100 and H200 GPUs 4.

Limitations

  • The methods need the model weights 1. Providers can check their own traffic and users can check open-weights models, but the methods cannot verify closed-weights providers 1.
  • Deployment requires knowing the provider's sampling procedure 1. In Amodo's prototype, vLLM v0.25.1 used a different random-number implementation from the one DiFR expected. This produced large logit differences until Amodo forced the older model runner 4.
  • For Qwen3-30B-A3B, pooling A100 and H200 reference runs left Token-DiFR unable to separate a temperature of 1.1 or a simulated top-2 sampling bug from honest runs at the target false-positive rate. Matched environments, or pooling that weights rare large deviations, restored detection 1.
  • Only single-model sampling was studied. Speculative decoding was not evaluated 1.
  • A statistical tolerance can bound covert bandwidth but cannot close it 7.
  • Adversarial testing is limited. The authors show that simple adversarial temperature tuning defeats a cross-entropy baseline, while Token-DiFR still detects the change 1. An independent study reports that attacker-chosen prompts roughly double the leakage allowed by a weight-exfiltration detector built on the same Gumbel-margin statistic. This cut the detector's slowdown to 60–118 times 9. Amodo rates red-teaming of recomputation schemes as not started 8.

Known flaws

Blockers

  • The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models.

  • The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences.

  • No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic.

Search

Full search page