Softmax
In one sentence Softmax converts a list of raw scores into probabilities that are positive and sum to one, exaggerating the gaps between scores.
Updated
Softmax turns any list of raw scores into a proper probability distribution: every value positive, all of them summing to 1.
It is the conversion from marks to vote share. Three candidates score 9, 7 and 2 in a poll of experts; softmax announces "72%, 26%, 2%". Two things happened: the numbers became shares of a whole, and the gaps stretched — 9 versus 7 became a three-to-one ratio. That stretching is deliberate and useful: softmax rewards confidence.
Mechanically, softmax exponentiates each score and divides by the sum of all the exponentials: softmax(zᵢ) = e^zᵢ ÷ Σⱼ e^zⱼ, where zᵢ is the i-th score (logit) and the sum runs over all options. The exponential is why gaps amplify — each extra point of logit multiplies the resulting probability by a constant factor.
It appears in the two most important places in modern AI. At the output of every classifier and LLM, converting logits over classes or tokens into probabilities to sample from. And inside attention, converting raw relevance scores into weights that sum to one — deciding how much each token contributes.
One practical dial: dividing the scores by a temperature before softmax controls the stretching. Low temperature sharpens the distribution toward the top score; high temperature flattens it toward uniform.
Where to go next
- Full lesson: Activation functions
- Related terms: logits, temperature, attention, activation-function