Mechanism · Bounding unexplained information in outputs
Evidence & limits
On this page
R2Demonstrated for bounding how much hidden information can leave in checked inference outputs
R2 through the inference-output instance, which has public code and an independent attack study; the facility-level architecture is still a proposal.
Assessed use: bounding how much hidden information can leave in checked inference outputs
Rubric assessment
- R1 met: Petrie and Mühlhäuser publish an architecture, the claim it verifies and its setting, in which neither party trusts the other's hardware 1.
- R2 met through the instance for inference outputs. Rinberg et al. bound the covert information in LLM responses beyond what honest sampling from the declared model explains. They publish code and report results on a 30-billion-parameter mixture-of-experts model against a stated adversary who controls the inference server 2. The facility-level architecture remains at R1: its paper outlines protocol details, attacks and prototyping plans, not results 1.
- R3 not met: no party other than the developers relies on the bound for a verification decision, and no production-grade system is available.
- R4 not met. Its evaluation criterion holds for the inference instance only: an independent researcher attacked it in practice, and the flaw shown widens the bound rather than defeating it, so the evaluation left no critical flaw open 3. The facility-level design has not been independently evaluated.
Confidence is low. The R2 evidence covers token outputs of a single inference service, not the general bound over all facility outputs 1, and the demonstrated bound degrades when the attacker controls prompts 3.
- A prototype of the facility-level architecture (interlock, commitments, challenge-based prediction) with published results.
- A bound that holds against an adversary who controls the prompt distribution, for example with entropy-calibrated tolerances, evaluated independently.
- Reliance by a party other than the developer on an unexplained-information bound for a verification decision, or a production-grade deployment.
Assessed 2026-09-25 against rubric v1.1.
Evidence
- Inference outputs. Rinberg et al. tested models from 3 to 30 billion parameters, including two mixture-of-experts models, and published code 2. On the 30-billion-parameter mixture-of-experts model, under benign prompt traffic, they report under 0.5% exfiltratable information at a false-positive rate below 0.01%, a slowdown of more than 200 times 2.
- Independent attack. Kezins tested six instruction-tuned models of 1 to 32 billion parameters over three seeds 3. Under benign prompts the slowdown was 146–254 times 3. Engineered prompts roughly doubled leakage per token and reduced the slowdown to 60–118 times 3.
- Compressibility. In a question-asking protocol, ten yes-or-no answers from a larger model recovered 23% to 72% of the capability gap between a small and a large model on standard benchmarks, at compression ratios of 0.0006 to 0.004 4.
- Facility-level architecture. Petrie and Mühlhäuser publish it as protocol details, attack vectors and prototyping plans 1.
Limitations
- Entropy inflation. An attacker who controls prompts widens the set of tokens the verifier tolerates 3. Kezins argues that designs built on the same unexplained-information bound inherit this weakness, and recommends calibrating tolerances against local token entropy 3.
- Noise tolerance. The covert channel lives in the tolerance allowed for GPU nondeterminism 3. Bit-exact replay needs the hardware model, weights, parallelism layout, software versions and batch size 6; see Deterministic and bit-exact inference.
- Residual physical bandwidth. One design argues that suppressing covert physical bandwidth below kilobits per second is much more achievable than aiming for zero 6. See Side-channel suppression for isolated facilities and Bandwidth limits and compartmentalization.
- Uncertain volumes. The upload-limit author is uncertain about the ratio of generated tokens to model parameters 5. In 2024 he estimated about a 25% chance that the approach can cheaply ensure that weights cannot be taken without bypassing physical security 5.
Known flaws
Blockers
The prover's compute must be isolated so that all traffic passes through the verifier's interlock; any unmonitored path voids the bound.
Physical side channels need separate suppression, and one design treats a low residual bandwidth, rather than zero, as the realistic target.
Tolerance for numerical nondeterminism sets the size of the residual channel; bit-exact replay would remove it but needs full hardware and software metadata.
Recomputation over confidential weights and inputs needs a protected setting: prover recomputation in a verifier-controlled enclosure, verifier recomputation in a prover-controlled enclosure, or zero-knowledge proofs.
No prototype of the facility-level architecture exists to red-team.