Mechanism · Deterministic and bit-exact inference

Evidence & limits

On this page

R3In production for reproducing open-model inference from receipts in Gensyn's information-market service

Gensyn reports using exact replay in production so that anyone can check markets it settles with REE, but the production evidence comes from Gensyn, and no independent security evaluation of bit-exact verification has been published.

Assessed use: reproducing open-model inference from receipts in Gensyn's information-market service

Rubric assessment

  • R1 met: the verification claim, a covert-adversary threat model and the information a verifier needs are published 1.
  • R2 met: public code predicts dense LLM blocks bit for bit on A100, L40, L40S and H100 GPUs running unmodified vLLM and Hugging Face engines, against a stated adversary 1. Batch-invariant modes are public in vLLM (beta) and SGLang 11 12, and batch invariance was shown on a 235-billion-parameter model 8. Exact CPU reproduction of GPU matrix multiplication is peer-reviewed 9.
  • R3 met through Gensyn's REE, on its developer's account. Gensyn reports that Delphi, its information-market app, is live on its mainnet, and that markets settled by open models inside REE produce receipts that anyone can re-run to verify the answer 21. Its service documentation also describes the REE judge receipts and mainnet deployment 23. REE is publicly available, and its README lists v0.8.0, released on 5 October 2026 18. Gensyn does not label its releases alpha or beta 22. Eigen Labs' service does not count towards R3, because Eigen Labs launched it as a mainnet alpha 20. No party other than a developer is documented relying on exact replay for a verification decision. The evidence for replay of stock serving engines remains at R2 and covers the emulator's tested dense model blocks 1.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of bit-exact verification has been published.

Confidence is low: R3 rests on Gensyn's account of its own service.

Gaps to the next level
  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open.

Assessed 2026-10-08 against rubric v1.1.

Mechanism properties

Threat modelAdversarial prover
Adversarial evaluationAnalysis
Hardware neededNone
Prover cooperationRequired
ConfidentialityPartial

Evidence

  • Nondeterminism measured. For 1,000 temperature-0 completions of one prompt on Qwen3-235B-A22B, Thinking Machines reports 80 unique outputs with default kernels 8. With batch-invariant kernels, all 1,000 were identical 8.
  • Bit-exact emulation. On Qwen3 4B blocks, the emulator reports zero BF16 differences for feed-forward blocks on A100, L40, L40S and H100 GPUs 1. It reports zero differences out of 71 million elements for FlashAttention-2 at 4,000 tokens 1. The paper received a best-paper award at the ICML 2026 TAIGR workshop, and its code is public 1.
  • Matrix multiplication on CPU. Hawkeye, peer-reviewed at MLSys 2026, reports 100% success replicating 4096 × 4096 matrix multiplications on Ampere, Hopper and Lovelace GPUs 9.
  • Engines. vLLM documents its batch-invariant mode 12. SGLang reports an average slowdown of 34.35% for its deterministic mode on FlashInfer and FlashAttention 3 backends 11.
  • Without determinism. In DiFR's tests with synchronized seeds, over 98% of tokens already match exactly between provider and verifier 4.
  • Batch-invariant kernels. Thinking Machines' MIT-licensed kernels underpin SGLang's deterministic mode, and vLLM's developers state that its batch-invariant mode is based on the same work 11 14 15.
  • Verde and RepOps. Gensyn fixes the order of floating-point operations so that honest compute providers get bitwise-identical results, and settles their disagreements by re-running a single operation 16. It reports running Verde and RepOps in production 17, and settling markets in its Delphi app with REE, its reproducible runtime, whose receipts anyone can re-run 21.
  • EigenAI. Eigen Labs reports a deterministic engine built on llama.cpp with its own matrix-multiplication and reduction kernels. Its outputs were bitwise identical across 10,000 runs on one GPU model, including across hosts, and never matched between A100 and H100 GPUs 19. An optimistic protocol has a committee re-execute challenged outputs inside TEEs 19. Eigen Labs launched the service on mainnet in September 2025, as an alpha whose stake was not yet exposed to slashing 20.

Limitations

  • Throughput. In Thinking Machines' test on Qwen3-8B, vLLM's default took 26 s, the unoptimized deterministic build 55 s, and the build with an improved attention kernel 42 s 8.
  • Coverage. The emulator does not yet cover mixture-of-experts inference, non-NVIDIA GPUs, the proprietary nvjet kernel family on Hopper, or training 1. Hawkeye covers matrix multiplication only. Attention and convolutions need further reverse engineering 9. vLLM's batch-invariant mode is in beta, and its tracking issue lists open work on AMD hardware, NVFP4 and speculative decoding 12 15.
  • Residual nondeterminism. Some integer de-quantization kernels use atomic additions and remain truly nondeterministic 1.
  • Disclosure. Replay requires exact weights and configuration details 1.
  • Maturity. Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack, and red-teaming of recomputation schemes, as not started 7.

Known flaws

Blockers

  • Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.

  • Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.

  • Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.

  • Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.

Search

Full search page