Implementation · TOPLOC

Technical detail

On this page
  • The prover commits to its activations every 32 generated tokens. It takes the top-k values of the last hidden layer, with k = 128 in the main configuration. It encodes their indices and values as a polynomial over an integer field, with a modulus chosen to be injective on the index set 1. The result is k two-byte coefficients. For Llama 3.1-8B-Instruct that is 258 bytes per 32 tokens, against 262 KB for storing the embeddings directly 1.
  • The verifier decodes the proof and recomputes the top-k values with a prefill pass. It counts exponent mismatches and computes the mean and median mantissa differences. Validation succeeds if all three are below their thresholds. For bf16 the thresholds are 38, 10 and 8 1.
  • The hardware tests used 1× A100, 1× RTX 4090 and 2× RTX 4090 GPUs, with FlashAttention 2, PyTorch SDPA and FlexAttention. The authors read the activations through a vLLM hook 1.
  • Prime Intellect reports that validation is up to 100 times faster than the original inference 3 4. In SYNTHETIC-2 it reports a median verification cost, averaged across models, 25 times lower than re-running the inference 6. It reports that proof generation cut tokens-per-second throughput by about 1% in INTELLECT-2 4.
  • Prime Intellect reports that TOPLOC v2 adds reproducible Gumbel noise for categorical sampling, so that verifiers can check token sampling. Version 2 also extends the scheme to pipeline-parallel inference. A pass at the final pipeline stage accepts all stages, and a failure triggers a stage-by-stage replay that finds the first faulty node 5 6.
  • The package is published on PyPI as toploc. The latest tag is v0.1.6 2.

Search

Full search page