Mechanism · Bounding unexplained information in outputs
Technical detail
On this page
- Fixed-seed sampling likelihood. The verifier knows the seed and recomputes the next-token distribution for each logged position, then scores whether each emitted token is plausible under honest sampling. Gumbel-Max and inverse-probability-transform estimators are given, and three theorems bound the rate at which an adversary can embed information while passing the check 2.
- Reported result. On MoE-Qwen-30B under benign prompt traffic, exfiltratable information fell to under 0.5% at a false-positive rate below 0.01%, a slowdown of more than 200 times 2.
- Attack result. Across six instruction-tuned models of 1 to 32 billion parameters and three seeds, the strongest prompt attack raised leakage from 0.067–0.119 to 0.145–0.286 bits per token and cut the slowdown from 146–254 times to 60–118 times, because the verifier's admissible-token set grows with the model's output entropy 3.
- Compression bounds. LLM-based arithmetic coding with domain-adapted LoRA adapters is 2 times better than the base model; succinct rewrites reach compression ratios of about 0.03; and a question-asking protocol reaches 0.0006–0.004 4. An earlier proposal reports compression below 1 bit per token by arithmetic coding against a weaker reference model, in an experiment with fine-tuned Pythia models 5.