AI glossary

Logits

In one sentence Logits are the raw, unnormalised scores a model produces for each possible output, before they are converted into probabilities.

By Updated

Logits are the raw scores a model assigns to every possible answer, before any conversion into probabilities.

Think of judges' raw scorecards at a dance competition, before results are announced. One judge writes 8.1, 4.5, 7.9 — informative numbers, but not yet "who won" and not yet percentages. Logits are those scorecards: one real number per option, unbounded, possibly negative, summing to nothing in particular. The softmax function is the announcer that converts them into proper probabilities.

For an LLM, the picture is concrete: at each step the network's final layer emits one logit per vocabulary entry — say 128,000 numbers. A logit of 12.3 for " the" and 3.1 for " zebra" means the model considers " the" vastly more plausible as the next token. The gap between logits, not their absolute values, is what carries meaning.

logits:  " the" 12.3   " a" 10.9   " zebra" 3.1
softmax:  0.78          0.19        0.00002

Why the two-stage design matters practically: nearly every generation control operates on the logits before softmax. Temperature divides them, banned-word filters set them to negative infinity, and structured-output constraints mask out invalid tokens. Training losses also consume logits directly — PyTorch's CrossEntropyLoss expects raw logits, and feeding it softmaxed probabilities is a classic silent bug.

Where to go next