Thinking Machines Lab
An AI research and product company; developer of batch-invariant kernels that make language-model outputs independent of batch size.
Thinking Machines Lab describes itself as "an artificial intelligence research and product company" 1. Its work relevant here is Batch-invariant inference kernels (Thinking Machines):
- In September 2025 it published batch-invariant kernels for LLM inference, arguing that varying batch sizes are the main reason LLM endpoints give nondeterministic outputs 2. Its MIT-licensed batch_invariant_ops library replaces several PyTorch operations with batch-invariant versions 3. See deterministic and bit-exact inference.
- It reports that, with the kernels, 1,000 temperature-zero completions from Qwen3-235B were identical 2. SGLang built its deterministic mode on the kernels 4, and vLLM's developers state that its batch-invariant mode is based on the same work 5.
On this page
Implementations
Implementations this organization develops.
- Open-source kernels from Thinking Machines Lab that make LLM outputs independent of batch size, adopted in vLLM and SGLang to give reproducible inference.
Publications
Sources this organization authored or published.
- BThinking Machines Lab. Thinking Machines Lab. Source recordCited by Thinking Machines Lab
- CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordCited by Deterministic and bit-exact inference; Batch-invariant inference kernels (Thinking Machines); The declared model is the one being served; Numerical nondeterminism; Thinking Machines Lab
- BThinking Machines Lab (2025). thinking-machines-lab/batch_invariant_ops (GitHub repository). GitHub. Source recordCited by Deterministic and bit-exact inference; Batch-invariant inference kernels (Thinking Machines); Thinking Machines Lab
Sources
- BThinking Machines Lab. Thinking Machines Lab. Source recordSupports: self-description · homepage
- CH. He & Thinking Machines Lab (2025). Defeating Nondeterminism in LLM Inference. Thinking Machines Lab: Connectionism. Source recordSupports: batch size as the main cause of nondeterminism; batch-invariant kernels; Qwen3-235B result (provider-reported)
- BThinking Machines Lab (2025). thinking-machines-lab/batch_invariant_ops (GitHub repository). GitHub. Source recordSupports: MIT-licensed batch_invariant_ops library · README
- CThe SGLang Team (2025). Towards Deterministic Inference in SGLang and Reproducible RL Training. LMSYS Org blog. Source recordSupports: SGLang's deterministic mode built on the kernels
- CvLLM project contributors (2025). [Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433). GitHub (vllm-project/vllm issues). Source recordSupports: vLLM developers' statement that batch invariance is based on the post