Implementation · Batch-invariant inference kernels (Thinking Machines)

Evidence & limits

On this page

R2Demonstrated for exact recomputation of served outputs by a verifier, with a cooperating provider

Public kernels made a 235-billion-parameter model's outputs identical across 1,000 runs and underpin the deterministic modes of two production engines, but no verification system is documented using them.

Assessed use: exact recomputation of served outputs by a verifier, with a cooperating provider

Rubric assessment

  • R1 met: Thinking Machines sets out the design and the property it gives, outputs that do not depend on batch size 1. The condition for identical outputs, a fixed model, inference implementation and device, is stated in the verification literature 6.
  • R2 met: the code is public under an MIT licence, with a deterministic vLLM example 2. With the kernels, 1,000 temperature-zero completions from Qwen3-235B were identical, against 80 unique completions without them 1.
  • R3 not met for this use. The kernels have reached production engines: SGLang builds its deterministic mode on them 3, and vLLM's developers state that its batch-invariant mode, labelled beta, is based on the same work 4 5. The kernels were built for reproducibility, such as on-policy reinforcement learning 1. Gensyn's REE and Eigen Labs' EigenAI, which build verification services on exact replay, use their own kernels 7 8. No party is documented relying on these kernels for a verification decision.
  • R4 not met: no independent security evaluation has been published.

Confidence is low, because whether an engine's deterministic mode counts as production-grade for verification is a judgment call.

Gaps to the next level
  • A production-grade verification stack that uses these kernels for exact-match checks, or a party other than the developers relying on such checks for a verification decision.

Assessed 2026-09-25 against rubric v1.1.

Evidence

  • For 1,000 temperature-zero completions of one prompt on Qwen3-235B, Thinking Machines reports 80 unique outputs with default kernels and one with batch-invariant kernels 1.
  • The library's vLLM example gave 18 unique samples out of 1,000 completions without the upstream vLLM change, and one with it 2.
  • SGLang integrated the mean, log-softmax and matrix-multiplication kernels and wrote batch-invariant attention kernels for several backends. It reports an average slowdown of 34.35% on its FlashInfer and FlashAttention 3 backends, against the 61.5% in the original post 3.
  • vLLM documents its batch-invariant mode as beta and lists tested model families including DeepSeek V3, Qwen3, Llama 3.1 and GPT-OSS 4.

Limitations

  • Outputs are identical only while the model, inference implementation and device stay fixed. For varied stacks and GPU types, DiFR's authors expect statistical verification to remain necessary 6.
  • In Thinking Machines' Qwen3-8B test, vLLM's default took 26 s, the unoptimised deterministic build 55 s and the build with an improved attention kernel 42 s 1.
  • vLLM's mode is in beta, and its tracking issue lists open work on performance, AMD hardware and speculative decoding 4 5.
  • The post's motivation is reproducibility, including on-policy reinforcement learning, not verification against a cheating provider 1.

Blockers

  • Batch invariance costs throughput: on Qwen3-8B the improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average slowdown of 34.35% on its FlashInfer and FlashAttention 3 backends.

  • Outputs are identical only while the model, inference implementation and device stay fixed, so provider and verifier must run the same stack.

Search

Full search page