The HuggingFace Stack

Controlling generate()

generate() writes text one piece at a time, and its arguments decide how each next piece is chosen, how repetition is punished, and when to stop.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

generate() writes text one word-piece at a time, and its settings control how the next piece gets picked and when the writing stops.

Try this game on your phone: type a few words, then keep pressing the middle keyboard suggestion. Within a line or two, the sentence starts circling — "the best thing about the best thing about...". That is what happens when a machine always picks its single safest next word.

The model behind generate() faces the same choice at every step. Its settings decide whether it plays safe, rolls dice, or drafts several continuations and keeps the best.

Why it exists

The model itself only ever produces one thing: scores for every possible next piece. Turning endless scores into finished text needs a policy — pick how? stop when? Those policy decisions are not in the model. They live in generate(), which is why the same model can sound robotic or creative depending on arguments alone.

How it works

prompt → model → scores for every possible next piece
                     ↓  choose one:
          safest (greedy) / dice roll (sampling) / several drafts (beams)
                     ↓
        append the piece, feed everything back, repeat
        ... until the length limit, or an end mark appears

Three levers matter most:

  • Greedy: always the top score. Reliable, and prone to the circling you saw.
  • Sampling: roll dice over the good options. Temperature sets how wild the dice are.
  • Length: how many new pieces to allow before stopping.

A real example you have seen

Every chatbot regenerate button. Same question, different answer each press — because sampling rolls fresh dice. A calculator app would be broken if "regenerate" changed the answer; a chat app would be creepy without it. The lever is the same do_sample idea.

Remember this

  • The model outputs scores; generate()'s settings turn scores into text.
  • Greedy picking loops; sampling varies; beams draft and compare.
  • Nothing about generation is "in" the model — same weights, different settings, different voice.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install transformers torch

Tested with transformers 5.6. distilgpt2 is a ~356 MB download, cached after the first run. Everything here runs on CPU.

Three decoding policies, one model

decoding.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, set_seed

tok = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2")

prompt = "The best street food in Mumbai is"
inputs = tok(prompt, return_tensors="pt")

greedy = model.generate(**inputs, max_new_tokens=15, do_sample=False,
                        pad_token_id=tok.eos_token_id)
print("greedy :", tok.decode(greedy[0], skip_special_tokens=True).strip())

set_seed(0)
sampled = model.generate(**inputs, max_new_tokens=15, do_sample=True,
                         temperature=0.9, top_p=0.9,
                         pad_token_id=tok.eos_token_id)
print("sampled:", tok.decode(sampled[0], skip_special_tokens=True).strip())

beams = model.generate(**inputs, max_new_tokens=15, num_beams=4,
                       repetition_penalty=1.3, pad_token_id=tok.eos_token_id)
print("beams  :", tok.decode(beams[0], skip_special_tokens=True).strip())
Output
greedy : The best street food in Mumbai is the best street food in Mumbai.
sampled: The best street food in Mumbai is by giving you access to a menu with options such as burgers, sandwiches,
beams  : The best street food in Mumbai is now available on the market.

The greedy line is stable for a given model and version. The sampled line will differ on your machine — even with the seed, sampling results vary across library versions and hardware. The behaviour is what reproduces: greedy loops, sampling varies.

The walkthrough

Greedy walked in a circle, live. It echoed the prompt back at itself, then spent its remaining tokens emitting newlines — which the .strip() trims away. This is not a bug. It is the documented failure mode of always taking the top token, and the reason chat products sample.

temperature does nothing without do_sample=True. The single most common generation bug. Sampling knobs — temperature, top_p, top_k — apply only to dice rolls; with greedy decoding they are ignored, and recent versions warn about the unused arguments.

max_new_tokens counts only new pieces. Its older sibling max_length counts the prompt too, which breaks the moment prompts vary in length. Prefer max_new_tokens always.

pad_token_id=tok.eos_token_id pre-empts a warning: GPT-2 has no pad token, and generate wants one defined for housekeeping.

Beams with repetition_penalty drafted four continuations in parallel and kept the best-scoring one, while taxing any token already used. Short, clean, and duller — beams optimise for safe, and shine in translation and summarisation more than open-ended chat.

Batched generation needs left padding. Continuation starts at the last position of each row. Right-side pads put blanks there; set tok.padding_side = "left" before batch generating, so real text ends every row.

Common mistakes

Tuning temperature with sampling off. Covered above; it silently changes nothing at all.

Expecting the demo model to be smart. distilgpt2 is 82M parameters of 2019-era completion — it suggests burgers in Mumbai. The decoding mechanics transfer unchanged to modern instruction-tuned models; the wisdom does not come from generate().

Forgetting skip_special_tokens=True. Decoded text arrives salted with end-of-text markers you did not want to show users.

Cutting answers off at max_new_tokens. A hard stop mid-sentence usually means the limit fired before the end token. Raise the limit, or check that the model's end token matches eos_token_id in its generation config — a frequent issue with fine-tuned models whose chat template changed the stop token.

Try it yourself

Run the sampled version five times without set_seed and watch it change. Then set temperature=0.2 and run five more — count how often the outputs repeat themselves. Finish with temperature=1.8 and decide for yourself where the dial stops being creative and starts being drunk.

What to learn next

Researcher — Mathematics and papers.

The decoding problem

A causal LM defines p(x_{t+1} | x_{≤t}); generation seeks a continuation, and the objective is contested. Exact MAP decoding is both intractable and — more damningly — undesirable: Holtzman et al. (2020), The curious case of neural text degeneration, show high-likelihood strings are repetitive and low-entropy relative to human text, and document the greedy/beam repetition attractor observed in the demo. The finding that beam search's benefit inverts between closed-ended (MT) and open-ended tasks predates it (Fan et al., 2018).

The standard samplers

With logits z and temperature T, sampling uses softmax(z/T); T → 0 recovers greedy, T > 1 flattens. Truncated variants restrict support before renormalising:

  • top-k (Fan et al., 2018): keep the k highest.
  • nucleus / top-p (Holtzman et al., 2020): keep the smallest set with cumulative mass ≥ p — adaptive to the distribution's entropy, which fixed k is not.
  • repetition penalty (Keskar et al., 2019, CTRL): divide (or scale) logits of already-generated tokens by α > 1. Effective and crude — it cannot distinguish topical reuse from degeneration.
  • Later refinements: typical sampling (Meister et al., 2022), contrastive decoding/search (Li et al., 2022; Su et al., 2022).

Order of operations matters: HuggingFace applies processors as a pipeline (penalties → temperature → truncation), configurable via LogitsProcessorList, and all defaults live in the model's GenerationConfig — checkpoint authors ship their preferred decoding, which your kwargs override per call.

Why seeded sampling is not portable

set_seed fixes the pseudo-random stream, but the probabilities being sampled differ across environments: floating-point reduction order differs by kernel choice, hardware, and batch shape, so logits differ in low-order bits, and one flipped multinomial draw diverges the entire suffix. Bitwise reproducibility requires pinning versions, hardware, and deterministic kernels — see random seeds and reproducibility. Treat sampled outputs as distributions, and evaluate accordingly.

Cost structure

Autoregressive decoding with a KV cache costs one forward pass per emitted token, with attention reading a cache that grows linearly — memory O(L·d·layers) per sequence. Beam search multiplies compute and cache by beam width b. Throughput at serving time is a batching-and-cache engineering problem, which is vLLM's reason to exist.

What to learn next