How Text Is Generated

Min-p and typical sampling

Two truncation methods that decide the shortlist from how confident the model is at that exact step, instead of from a fixed count or a fixed share of probability.

On this page 7
  1. Why the older methods needed replacing
  2. How min-p works
  3. How typical sampling works
  4. Where you have already seen this
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Min-p keeps every word that is at least a certain fraction as good as the model's favourite. Typical sampling keeps the words that are about as surprising as the model expects this step to be.

Think about buying mangoes from a cart. You spot the best one, then you keep any mango that is roughly as good as that one. You do not decide in advance to buy five. You do not keep buying until a fixed weight is filled.

Your shortlist changes with the cart. A great cart gives you many good mangoes. A poor cart gives you one, and you take one.

That is min-p. The shortlist is defined relative to the best option available right now.

Why the older methods needed replacing

The previous lesson showed two problems.

Top-k always keeps the same number of words, whether the model is certain or lost. Top-p keeps a fixed share of confidence, which sounds adaptive but breaks when you turn up the creativity dial.

Here is the awkward interaction. Turning up creativity flattens the model's whole list, so every word looks a bit more possible. Top-p then reaches much further down to collect its fixed share. It scoops up junk on the way.

So the two knobs fight each other. Raising creativity secretly widens the shortlist, and users blame the model.

How min-p works

Min-p looks at the model's top word, then sets a bar as a fraction of that word's confidence.

If the top word carries most of the confidence, the bar is high and almost nothing else clears it. If the top word carries little, the bar is low and many words survive.

   model is CERTAIN                    model is UNSURE
   top word carries most of it         nothing stands out

   bar set relative to the top:        bar set relative to the top:
   only the top word clears it         nine words clear it

   -> one safe choice                  -> nine live options

The bar moves with the model. That is the whole idea, and it is why the setting stays sensible when you change the creativity dial.

How typical sampling works

Typical sampling asks a stranger question, and it is worth reading twice.

At every step the model has an expectation of how surprising the next word should be. In a boring stretch of text, the next word should not be surprising at all. In an unpredictable stretch, some surprise is normal.

Typical sampling keeps the words whose surprise is closest to that expectation. It cuts words that are too obvious, and words that are too shocking.

Yes, it cuts the safest word sometimes. That is deliberate. Text made only of the safest choices reads as flat and repetitive, which is a real and well-documented failure.

Where you have already seen this

  • Local model tools with a min_p slider next to top_p.
  • Roleplay and story-writing communities recommending high creativity with min-p turned on.
  • A model that becomes incoherent above a certain creativity setting with top-p, and stays readable with min-p.

The honest part

Min-p was published with strong claims, and a later paper re-examined those claims and disagreed with several of them. Both papers are worth reading.

What is not in dispute is the mechanism. A confidence-relative bar behaves differently from a fixed share. You can see the difference for yourself in the code below.

Treat the recommended values as starting points to test, not as findings. That is good advice for every sampling setting on this page.

Remember this

  • Min-p sets the bar as a fraction of the model's best word, so the shortlist follows the model's confidence.
  • Typical sampling keeps words that are about as surprising as this step should be, cutting both extremes.
  • Both were built because top-p's shortlist balloons when you raise the creativity setting.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Both methods are short. Seeing them side by side on the same two distributions is worth more than any description.

The two situations that matter

adaptive_truncation.py
import numpy as np

VOCAB = ["chai", "coffee", "water", "juice", "milk",
         "soup", "petrol", "sand", "regret", "Tuesday"]

# Two honest situations, written as probabilities and turned back into logits.
CERTAIN = np.log([0.880, 0.040, 0.025, 0.018, 0.012, 0.009, 0.007, 0.005, 0.003, 0.001])
UNSURE  = np.log([0.200, 0.180, 0.150, 0.120, 0.100, 0.080, 0.070, 0.050, 0.030, 0.020])


def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()


def kept(logits, mask):
    return [w for w, k in zip(VOCAB, mask) if k]


def top_p(logits, thresh):
    order = np.argsort(-logits)
    p = softmax(logits)[order]
    keep = np.zeros(len(logits), bool)
    keep[order[np.cumsum(p) - p <= thresh]] = True     # include the crossing token
    return keep


def min_p(logits, m):
    """Keep every token at least m times as likely as the best token."""
    p = softmax(logits)
    return p >= m * p.max()


def typical(logits, mass):
    """Keep tokens whose surprise is closest to this step's average surprise."""
    logp = logits - logits.max() - np.log(np.exp(logits - logits.max()).sum())
    p = np.exp(logp)
    entropy = -(p * logp).sum()
    distance = np.abs(-logp - entropy)
    order = np.argsort(distance)                       # most typical first
    last = min(int((np.cumsum(p[order]) < mass).sum()), len(logits) - 1)
    keep = np.zeros(len(logits), bool)
    keep[order[: last + 1]] = True
    return keep


for name, L in (("CERTAIN", CERTAIN), ("UNSURE", UNSURE)):
    p = softmax(L)
    print(f"{name}: top probability {p.max():.3f}, entropy {-(p * np.log(p)).sum():.2f} nats")
    print(f"  top-p   0.90 -> {len(kept(L, top_p(L, 0.90))):>2} tokens {kept(L, top_p(L, 0.90))}")
    print(f"  min-p   0.10 -> {len(kept(L, min_p(L, 0.10))):>2} tokens {kept(L, min_p(L, 0.10))}")
    print(f"  typical 0.90 -> {len(kept(L, typical(L, 0.90))):>2} tokens {kept(L, typical(L, 0.90))}")
    print()

# A more realistic step: 5 sensible words and 1,000 tokens of noise.
rng = np.random.default_rng(3)
big = np.concatenate([np.array([6.0, 5.4, 4.9, 4.5, 4.2]), rng.normal(0.0, 1.0, 1000)])

print("1,005-token vocabulary: 5 sensible words, 1,000 pieces of noise")
print(f"{'temperature':>12}{'top-p 0.95':>13}{'min-p 0.05':>13}{'mass on the 5 real words':>28}")
for T in (0.7, 1.0, 1.5, 2.0, 3.0):
    z = big / T
    print(f"{T:>12.1f}{int(top_p(z, 0.95).sum()):>13}{int(min_p(z, 0.05).sum()):>13}"
          f"{softmax(z)[:5].sum():>27.3f}")
Output
CERTAIN: top probability 0.880, entropy 0.59 nats
  top-p   0.90 ->  2 tokens ['chai', 'coffee']
  min-p   0.10 ->  1 tokens ['chai']
  typical 0.90 ->  2 tokens ['chai', 'coffee']

UNSURE: top probability 0.200, entropy 2.12 nats
  top-p   0.90 ->  8 tokens ['chai', 'coffee', 'water', 'juice', 'milk', 'soup', 'petrol', 'sand']
  min-p   0.10 ->  9 tokens ['chai', 'coffee', 'water', 'juice', 'milk', 'soup', 'petrol', 'sand', 'regret']
  typical 0.90 ->  7 tokens ['chai', 'coffee', 'water', 'juice', 'milk', 'soup', 'petrol']

1,005-token vocabulary: 5 sensible words, 1,000 pieces of noise
 temperature   top-p 0.95   min-p 0.05    mass on the 5 real words
         0.7          271            5                      0.768
         1.0          672            7                      0.349
         1.5          827           75                      0.107
         2.0          873          518                      0.053
         3.0          908         1002                      0.025

Reading the output, including the part that argues against min-p

On the CERTAIN step, min-p kept one token and top-p kept two. The second token carried four percent. Min-p at 0.10 requires eight point eight percent, so it was cut. That is the point: when the model is sure, be decisive.

On the UNSURE step, min-p kept nine and top-p kept eight. Min-p relaxed and top-p tightened, on the same pair of settings. The direction reversed because the two rules measure different things.

The big table is the real argument. With a thousand junk tokens present, top-p 0.95 admits 271 of them at temperature 0.7 and 672 at temperature 1.0. Min-p admits 5 and 7. Over a long generation, that difference is the difference between readable output and occasional nonsense.

And the last two rows are the counter-argument, stated honestly. At temperature 2.0 and 3.0 the five real words hold five percent and two point five percent of the mass. Min-p relaxes to 518 and 1002 tokens. It has not saved anything; it collapses too, one step later than top-p.

That is worth internalising. Min-p buys you roughly one extra unit of temperature headroom. It does not make arbitrary temperature safe, and any claim that it does is overselling. The critical re-analysis of the min-p paper makes essentially this point with proper experiments.

Checking against the reference implementations

verify.py
import numpy as np, torch
from transformers import MinPLogitsWarper, TypicalLogitsWarper

CERTAIN = np.log([0.880, 0.040, 0.025, 0.018, 0.012, 0.009, 0.007, 0.005, 0.003, 0.001])


def min_p(logits, m):
    p = np.exp(logits - logits.max())
    p /= p.sum()
    return p >= m * p.max()


def typical(logits, mass):
    logp = logits - logits.max() - np.log(np.exp(logits - logits.max()).sum())
    p = np.exp(logp)
    distance = np.abs(-logp - (-(p * logp).sum()))
    order = np.argsort(distance)
    last = min(int((np.cumsum(p[order]) < mass).sum()), len(logits) - 1)
    keep = np.zeros(len(logits), bool)
    keep[order[: last + 1]] = True
    return keep


ids = torch.zeros((1, 1), dtype=torch.long)
t = torch.tensor(CERTAIN, dtype=torch.float32).unsqueeze(0)

hf_minp = MinPLogitsWarper(0.10, min_tokens_to_keep=1)(ids, t.clone())[0].numpy()
hf_typ = TypicalLogitsWarper(0.90, min_tokens_to_keep=1)(ids, t.clone())[0].numpy()

print("min-p matches HuggingFace  :", bool((np.isfinite(hf_minp) == min_p(CERTAIN, 0.10)).all()))
print("typical matches HuggingFace:", bool((np.isfinite(hf_typ) == typical(CERTAIN, 0.90)).all()))
Output
min-p matches HuggingFace  : True
typical matches HuggingFace: True

Standalone, so it needs nothing from the block above. Written against transformers 5.6.2 and torch 2.5.1.

Using them for real

python
# HuggingFace Transformers
out = model.generate(**inputs, do_sample=True, temperature=1.2, min_p=0.05, top_p=1.0)

# vLLM
from vllm import SamplingParams
params = SamplingParams(temperature=1.2, min_p=0.05, top_p=1.0)

Note top_p=1.0 in both. Leaving top-p at 0.9 alongside min-p means top-p does the cutting first and min-p rarely binds, and you end up tuning a parameter that has no effect. Pick one truncation rule and switch the others off.

Order of application in transformers is temperature, top-k, top-p, then min-p. Since min-p runs last it can only cut further, never restore.

Common mistakes

Stacking every filter at once. top_k=40, top_p=0.9, min_p=0.05, typical_p=0.95 is four rules fighting. The tightest one wins on every step and the other three are noise in your config file.

Copying min-p values between models. A value tuned on one model family is a guess on another. Vocabulary size and calibration both move the right threshold.

Assuming min-p is always better. For extraction, classification, code and tool arguments, the correct setting is greedy decoding. No truncation rule beats not sampling when there is one right answer.

Reading typical_p as a probability cutoff. It is a mass budget over tokens ranked by closeness to the entropy, not by probability. That ranking is why typical sampling can drop the top token, which surprises people reading their logs.

Try it yourself

Add a third row to the first block: a distribution where two tokens tie at 0.45 and the rest share 0.10. Predict what min-p at 0.10 keeps before you run it. Then sweep min-p from 0.01 to 0.30 on UNSURE and plot the survivor count. The curve's steepness is the sensitivity you are signing up for.

What to learn next

Researcher — Mathematics and papers.

Min-p

Nguyen et al., Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs, ICLR 2025 oral (arxiv.org/abs/2407.01082) define the kept set as

$$ V_{\text{min-p}} = \left{ x \in V : P(x) \ge p_{\text{base}} \cdot \max_{y \in V} P(y) \right} $$

with $p_{\text{base}} \in (0, 1]$, typically $0.02$ to $0.1$.

The scale-relative threshold is the whole design. Under temperature scaling $z \mapsto z/T$, the ratio $P(x)/\max_y P(y)$ transforms as

$$ \frac{P_T(x)}{\max_y P_T(y)} = \exp!\left(\frac{z_x - z_{\max}}{T}\right) = \left(\frac{P_1(x)}{\max_y P_1(y)}\right)^{1/T} $$

so the surviving set is a monotone function of $T$, with a far gentler slope than top-p's. Top-p's cutoff depends on the cumulative mass, and therefore on the entire tail. The empirical table above shows both the benefit and its limit.

The dispute, which is instructive

The original paper reported gains on GPQA, GSM8K and AlpacaEval Creative Writing across Mistral and Llama 3 at 1B to 123B, especially at high temperature.

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models, 2025 (arxiv.org/abs/2506.13681) re-examines all four lines of evidence. It reports that the benchmark gains do not reproduce under matched baselines. It finds confounds in the human evaluation, and argues that community adoption was mis-cited as endorsement.

Two things to take from this beyond the specific question. First, decoding-method papers are unusually easy to over-claim, because a baseline evaluated at one hyperparameter setting and a proposed method tuned across many is not a comparison. Second, a method can be mechanically sound and still not be an improvement on the benchmarks used to argue for it. Min-p is now widely implemented — in transformers, vLLM, SGLang and llama.cpp — which is a fact about adoption, not about quality.

Locally typical sampling

Meister et al., Locally Typical Sampling, TACL 2023 (arxiv.org/abs/2202.00666) start from information theory rather than from tail-cutting.

The conditional entropy at step $t$ is

$$ H_t = -\sum_{x \in V} P(x \mid x_{<t}) \log P(x \mid x_{<t}) $$

and a token's information content is $-\log P(x \mid x_{<t})$. The typical set is tokens whose information content lies near $H_t$:

$$ V_{\text{typ}} = \arg\min_{V' \subseteq V} \sum_{x \in V'} \Big| {-\log P(x)} - H_t \Big| \quad \text{subject to} \quad \sum_{x \in V'} P(x) \ge \tau $$

Implementations approximate this greedily: sort by $\big|{-\log P(x)} - H_t\big|$ ascending, take a prefix until cumulative mass reaches $\tau$.

The motivating hypothesis is that human language is efficient in Shannon's sense — speakers spread information evenly rather than always choosing the most predictable word. Text that is too predictable is as unhuman as text that is too surprising, which is the observation Holtzman et al. made empirically about beam search.

Top-h, the newest member

Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation, NeurIPS 2025 (arxiv.org/abs/2509.02510) formulates truncation as an entropy-constrained optimisation. Tokens are added in descending probability order while cumulative entropy stays under $\tau = \alpha H$, with $\alpha \in (0, 1]$ the single parameter. It is available as top_h in transformers 5.x.

Choosing between them

There is no decisive published comparison, and the min-p dispute is a warning against trusting one. Defensible practice:

  • Greedy for anything with a correct answer: extraction, classification, tool arguments, code completion where determinism matters more than variety.
  • Top-p at 0.9 to 0.95 with $T \le 1$ as the boring default. Enormously deployed, well understood.
  • Min-p at 0.02 to 0.1 with $T > 1$ when you deliberately want a hot distribution — creative writing, generating diverse candidates for reranking.
  • Typical at 0.2 to 0.95 where repetitiveness is the specific complaint. It targets that failure directly.

And evaluate on your own task. Sampling settings interact with a model's calibration, which is a property of its post-training, not of the decoding rule.

Papers

What to learn next