How Text Is Generated

Why temperature zero is still not deterministic

Setting temperature to zero removes randomness from the choosing, not from the arithmetic — and the arithmetic changes with how many other people are using the server.

On this page 7
  1. Why the order changes at all
  2. Why a tiny difference becomes a different answer
  3. The other things that get blamed
  4. Can it be fixed?
  5. What to do about it in practice
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Temperature zero makes the model always pick its top choice. It does not make the calculation that produced those choices come out the same every time.

Think about splitting a restaurant bill. Three friends add up the same items in a different order, each rounding to the nearest rupee as they go. They get slightly different totals, and nobody made a mistake.

Rounding at different points changes the answer. That is all this lesson is about.

A computer rounds too. It stores each number with limited room, so it rounds constantly. Change the order it adds things up and the last few digits move.

Why the order changes at all

A graphics card does not add a long list of numbers one at a time. It splits the list into chunks, adds the chunks in parallel, then adds the results together.

How many chunks it uses depends on how much work there is in total. And how much work there is depends on how many other people's requests are being processed alongside yours.

So your answer depends on the size of the crowd you were batched with. That is not a bug in the model. It is a property of how fast arithmetic is done.

Why a tiny difference becomes a different answer

Most of the time it does not matter. A difference in the seventh decimal place does not change which word wins.

But occasionally two words are almost exactly tied. The model genuinely has no preference. Then the seventh decimal place is the whole decision.

Once one word changes, everything after it changes, because the model reads its own output. One flipped coin at word forty and the rest of the paragraph is different.

   run 1:  ... the report was submitted on Monday and approved.
   run 2:  ... the report was submitted on Monday and accepted.
                                                      ^ a near-tie broke the other way

The other things that get blamed

People usually blame the wrong causes. It is worth knowing the real list.

Not the sampling. Temperature zero already removed that.

Not the seed. A seed controls random choices. There are none left at temperature zero.

Not the model changing under you. Providers do update models, and that is a real cause. This also happens with a fixed model on your own machine.

It is the batch. The crowd your request travelled with changed the arithmetic.

Can it be fixed?

Yes, and it costs speed.

Researchers showed in 2025 how to remove the effect. Rewrite the core calculations so they always split the work the same way, whatever the crowd size. The published work demonstrated identical results across a thousand repeated runs.

The price is that the calculation can no longer choose the fastest split for the current load. Reported slowdowns are meaningful but not ruinous. Several serving systems now offer this as an option you turn on when you need it.

What to do about it in practice

Do not build anything that requires two runs to match exactly. Cache the answer you got rather than expecting to regenerate it.

For tests, check properties rather than exact strings. Did it return valid JSON? Does the field have the right value? Not: does it equal this exact paragraph.

Remember this

  • Temperature zero removes random choosing, not rounding differences in the arithmetic.
  • The batch you shared the server with changes the last few digits of the calculation.
  • A near-tie turns a tiny difference into a different word, and then a different paragraph.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

Four experiments, from pure arithmetic up to a real model. The model is distilgpt2, about 340 MB, CPU-runnable. The download progress bar is not shown in the output.

The whole story in one file

nondeterminism.py
import torch

# 1. Floating-point addition is not associative.
a, b, c = torch.tensor([1e8]), torch.tensor([-1e8]), torch.tensor([1.0])
print("(a + b) + c =", ((a + b) + c).item())
print("a + (b + c) =", (a + (b + c)).item())

# 2. So the ORDER a sum is split into changes the answer.
torch.manual_seed(0)
x = torch.randn(100_000)
print("\nsum left to right :", f"{x.sum(dtype=torch.float64).item():.12f}")
print("sum in float32    :", f"{x.sum().item():.12f}")
print("sum of 4 chunks   :", f"{sum(c.sum() for c in x.chunk(4)).item():.12f}")
print("sum of 8 chunks   :", f"{sum(c.sum() for c in x.chunk(8)).item():.12f}")

# 3. A GPU or a BLAS library picks the chunking based on how much work there is,
#    and how much work there is depends on the batch. Same row, different batch:
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2").eval()
ids = tok("The monsoon reached Nagpur three days late this year and the",
          return_tensors="pt").input_ids
with torch.no_grad():
    solo = model(ids).logits[0, -1]

print("\nsame prompt, different batch size (temperature is not involved at all):")
for B in (2, 4, 16, 64):
    filler = torch.randint(0, 50257, (B - 1, ids.shape[1]))
    with torch.no_grad():
        batched = model(torch.cat([ids, filler], 0)).logits[0, -1]
    print(f"   batch {B:>3}: bit-identical={torch.equal(solo, batched)}"
          f"   largest logit difference={(solo - batched).abs().max():.2e}"
          f"   argmax {solo.argmax().item()} vs {batched.argmax().item()}")

# 4. A difference that small only matters when the top two tokens are close.
print("\nhow close do the top two get? 200 greedy steps of the same model:")
seq = tok("In the last five years, the way people book train tickets in India has",
          return_tensors="pt").input_ids
near32 = ties_bf16 = 0
for _ in range(200):
    with torch.no_grad():
        lg = model(seq).logits[0, -1]
    top2 = lg.topk(2).values
    near32 += (top2[0] - top2[1]).item() < 1e-4
    bf = lg.to(torch.bfloat16).float().topk(2).values      # what a real server computes in
    ties_bf16 += (bf[0] - bf[1]).item() == 0.0
    seq = torch.cat([seq, lg.argmax().view(1, 1)], dim=1)
print(f"   float32: steps with a top-two gap under 1e-4 : {near32}")
print(f"   bfloat16: steps with an EXACT tie for first  : {ties_bf16}")
print("   bfloat16 rounding:",
      torch.tensor([20.0, 20.05, 20.1, 20.2]).to(torch.bfloat16).float().tolist())
Output
(a + b) + c = 1.0
a + (b + c) = 0.0

sum left to right : -244.559763742878
sum in float32    : -244.559738159180
sum of 4 chunks   : -244.559738159180
sum of 8 chunks   : -244.559722900391

same prompt, different batch size (temperature is not involved at all):
   batch   2: bit-identical=False   largest logit difference=5.34e-05   argmax 1181 vs 1181
   batch   4: bit-identical=False   largest logit difference=5.34e-05   argmax 1181 vs 1181
   batch  16: bit-identical=False   largest logit difference=5.34e-05   argmax 1181 vs 1181
   batch  64: bit-identical=False   largest logit difference=6.10e-05   argmax 1181 vs 1181

The first two sections are exact and reproduce anywhere. The largest logit difference values are machine-specific — they depend on your BLAS library, your CPU's vector width and your PyTorch build. What reproduces everywhere is the word False. Do not copy those numbers into a document as if they were properties of the model.

The final section on a machine with a different BLAS may differ by a step or two:

Output
how close do the top two get? 200 greedy steps of the same model:
   float32: steps with a top-two gap under 1e-4 : 0
   bfloat16: steps with an EXACT tie for first  : 2
   bfloat16 rounding: [20.0, 20.0, 20.125, 20.25]

Reading the output

(a + b) + c = 1.0 and a + (b + c) = 0.0. Same three numbers, same operation, different grouping, different answer. Floating-point addition is not associative, and everything else on this page follows from that one line. 1e8 + 1.0 rounds back to 1e8 because a float32 has no room for both magnitudes at once.

Four chunks and eight chunks disagree in the eighth digit. No parallelism yet, no GPU, no model. Splitting a sum differently is enough. A GPU kernel splits sums by tile size and tile size is chosen from the problem shape.

The same prompt in a batch of 64 gives different logits from the same prompt alone. No sampling, no seed, no temperature. Only the shape of the matrix multiplication changed, because the other rows in the batch changed how the work was tiled. This is the mechanism behind almost every "why did the API give me a different answer at temperature zero" report.

The argmax survived every time here, and that is honest rather than convenient. A difference of 6e-05 flips a decision only when the top two logits are within 6e-05 of each other. On this small model in float32, across 200 greedy steps, that never happened.

And then bfloat16 changes the picture completely. Twice in 200 steps, the top two logits were exactly equal after rounding to bfloat16. At that point the winner is decided by tie-breaking order, which is an implementation detail of the kernel. Production servers run in bfloat16 or fp8, not float32, and the last line shows why: 20.05 and 20.0 are the same number in bfloat16. The grid is about a thousand times coarser, so near-ties become exact ties.

What is really going on inside the kernel

Three separate sources, worth naming individually.

Reduction order in matrix multiplies. A GPU splits the shared dimension across threads and combines partial sums. The split depends on the tile size chosen for the current shape, and the shape depends on the batch.

Atomic accumulation. Some kernels accumulate with atomic adds whose completion order is genuinely non-deterministic between runs, even with everything else fixed.

Mixture-of-experts routing. In an MoE model, expert assignment is computed per batch. A token's expert can change with its batch-mates, which changes the computation applied to it. This one is not floating-point noise at all — it is a different function being evaluated.

To that add the ordinary suspects: a different GPU model, a different CUDA or cuBLAS version, a different attention backend, tensor-parallel degree, and whether CUDA graphs are on.

What to do

Reproducibility on your own machine. Fix the seed, pin every library version, use a fixed batch size, and disable non-deterministic kernels:

python
torch.use_deterministic_algorithms(True)
torch.backends.cudnn.deterministic = True

This gets you run-to-run repeatability on identical hardware with identical shapes. It does not survive a different GPU, and it does not survive a shared server.

Reproducibility on a server. Batch-invariant kernels, discussed below. vLLM and SGLang both expose deterministic modes built on this work.

Everything else. Cache outputs rather than regenerating them. Test properties rather than exact strings. Store the generated text alongside the request if you will need to show it again.

Common mistakes

Setting a seed and expecting determinism at temperature zero. There is nothing left for the seed to control.

Reporting an evaluation to four decimal places from a single run. Run it three times. The spread you see is the precision you actually have, and it is often larger than the difference you are claiming.

Blaming the provider for silently changing the model. They do sometimes. This effect happens on your own hardware with fixed weights, so rule it out first.

Assuming greedy is deterministic because it has no random number generator. Greedy is an argmax over numbers that are themselves not reproducible.

Try it yourself

Change x = torch.randn(100_000) to torch.randn(10) and watch every chunking agree. The effect scales with the length of the reduction, which is exactly why it shows up in large models and never in unit tests. Then run the batch experiment twice in the same process and confirm each individual result is stable. The variation is across batch shapes, not across repeated calls with the same shape.

What to learn next

Researcher — Mathematics and papers.

The mechanism

Floating-point addition is commutative but not associative. For $a, b, c$ in IEEE-754,

$$ \mathrm{fl}\big(\mathrm{fl}(a + b) + c\big) \ne \mathrm{fl}\big(a + \mathrm{fl}(b + c)\big) $$

in general, with relative error bounded by the unit roundoff $u$ ($2^{-24}$ for float32, $2^{-8}$ for bfloat16). Any parallel reduction fixes a summation tree, and different trees give different results.

A kernel's reduction tree is chosen from the problem shape: tile sizes, split-K factor, number of thread blocks, whether a persistent or a stream-K schedule is used. Autotuners pick these from the shape at hand. Since the shape includes the batch dimension, the reduction tree is a function of the batch.

Batch invariance

He and the Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, September 2025, identify this as the dominant cause of nondeterminism in production serving. Not atomics, and not GPU nondeterminism in the abstract.

The key framing: individual kernels are usually run-to-run deterministic, meaning the same input in the same shape gives the same output every time. They are not batch-invariant, meaning the same row gives a different output depending on what else is in the batch. A server batches your request with whatever traffic happened to arrive, and that traffic is outside your control. So your result is effectively nondeterministic, even though every kernel is individually reproducible.

The fix is to require a fixed reduction order independent of batch size, for RMSNorm, matrix multiplication and attention. They demonstrate bit-identical outputs across 1,000 repeated completions of the same prompt on Qwen3-8B under vLLM, where the unmodified stack produced many distinct completions.

The costs are structural rather than incidental. A batch-invariant matmul cannot switch to a split-K strategy when the batch is small. It gives up throughput exactly where the tuned kernel gains most. Attention must use a fixed split size along the KV dimension rather than one chosen from the current sequence length. The LMSYS team reports reducing overhead to roughly 34 percent by combining these kernels with CUDA graphs. That work targets deterministic inference for reproducible reinforcement-learning training.

That last application is the strongest argument for caring. On-policy RL requires the sampler and the trainer to agree on the log-probabilities of the sampled tokens. When they disagree because of numerical differences, the algorithm is silently off-policy, and the resulting instability is easy to misattribute to the learning algorithm.

Sources of divergence, ranked

  1. Batch composition. Reduction order changes with shape.
  2. MoE routing. Expert selection can depend on batch-level statistics. A different function, not a rounding difference.
  3. Atomics. Non-deterministic completion order in kernels that use atomicAdd. Avoidable in principle.
  4. Kernel selection. cuBLAS and cuDNN heuristics pick algorithms by shape, and by hardware.
  5. Tensor parallelism. All-reduce combination order across ranks.
  6. Chunked prefill and prefix caching. The same prefix computed in one pass or several gives different KV values, so a cache hit and a cache miss can produce different logits. This one surprises people, because the cache is meant to be transparent.
  7. Speculative decoding. Exactly lossless in distribution for the sampling variant, and bit-identical for the greedy variant only if the verifier's numerics match the non-speculative path, which they need not.

Amplification

Divergence compounds rather than adding up. Let $\delta$ be the numerical perturbation to a logit vector and $g$ the gap between the top two logits. Under greedy decoding, the argmax flips when $\delta > g/2$ roughly. The probability of a flip at any single step is small, but over $T$ steps

$$ P(\text{some divergence}) = 1 - \prod_{t=1}^{T} \big(1 - P(\text{flip at } t)\big) $$

and $T$ is in the hundreds or thousands. Once one token differs, the sequences are conditioned on different prefixes and every subsequent distribution differs. There is no recovery.

The empirical result above sharpens this. In float32 the gap distribution rarely reaches $10^{-4}$, so flips are uncommon. In bfloat16, exact ties occurred at 1 percent of steps in a 200-step run of a small model. A tie is decided by tie-breaking convention rather than by the model. The precision of production inference is therefore a first-order factor in how reproducible it is, and low-precision quantised serving is the least reproducible of all.

Consequences for evaluation

  • A benchmark score from a single run has error bars you did not measure. Report the spread across at least three runs, and treat differences smaller than that spread as noise.
  • A regression test asserting an exact output string will flake, and the flake rate will look like a bug in something else.
  • Reward-model scores and LLM-as-judge verdicts inherit this variance twice over — once from the model under test and once from the judge.
  • Bisecting a quality regression across library versions is unreliable, because a version bump changes kernels and therefore changes outputs regardless of any real regression.

Sources

What to learn next