Mechanism · Deterministic and bit-exact inference
Technical detail
On this page
Three routes lead to exact results.
Invariance. Kernels fix the reduction order for each output element regardless of batch size. Thinking Machines made RMSNorm, matrix multiplication and attention batch-invariant, the last with a fixed split size for the key-value dimension rather than a fixed number of splits 8. vLLM exposes this behind VLLM_BATCH_INVARIANT=1 on NVIDIA GPUs of compute capability 8.0 or higher and on Intel XPUs, in beta 12. SGLang integrated batch-invariant attention for its FlashInfer, FlashAttention 3 and Triton backends 11. LLM-42 instead decodes on a non-deterministic fast path and replays candidate tokens under a fixed-shape reduction schedule, rolling back any that are inconsistent 13.
Record and replay. Stock engines are deterministic but not invariant. Outputs are bitwise reproducible if the verifier knows the hardware model, the exact deployed weights, the parallelism topology (separately for prefill and decode), software versions including custom kernels, and the batch size of each forward pass 1 3. Of these, only batch size changes during serving, and it costs one extra integer per forward pass to record 1. A software emulator reproduces the rounding of other GPU models by modelling tensor-core accumulation and kernel-specific reduction trees 1. Hawkeye reproduces tensor-core matrix multiplication exactly on a CPU for Ampere, Hopper and Ada Lovelace GPUs in FP16, BF16 and FP8 9.
Reproducible operators. Gensyn's RepOps fixes the order of floating-point operations across hardware, so providers and a referee can reproduce the same result 16.
With exact replay, verification is pass/fail, and the chance of catching at least one false output in k samples is 1 − (1 − p)^k for a false-output rate p 1.