Post-training and Alignment

Instruction tuning

A freshly pretrained model continues your text instead of answering it, and instruction tuning is the short training phase that teaches it to reply.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The surprising part
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A freshly trained model continues your text. Instruction tuning teaches it to answer you instead.

The analogy you have already lived

Type "how to make" into a search box and look at the grey suggestions underneath. You get "how to make paneer", "how to make a resume", "how to make money".

Not one of them is an answer. They are all continuations — more of the same kind of text.

A model straight out of pretraining behaves exactly like that box. You ask it a question and it writes a more elaborate version of your question.

Why it exists

Pretraining teaches one skill: guess what text comes next. Nothing in that task ever asks the model to be helpful.

So a pretrained model does three annoying things.

It ignores your instruction. Ask for a definition and it writes a list of related headings. On the internet, that is what usually follows such a line.

It does not stop. Text on the internet keeps going. So does the model, often repeating itself forever.

It has no idea who is speaking. There is no notion of "you asked, I answer". It is one undifferentiated stream of text.

Instruction tuning fixes all three, and it is short. Pretraining reads trillions of words. Instruction tuning reads a few tens of thousands of examples.

How it works

You collect examples of the behaviour you want. Each one is a request and a good reply.

   request:  "Explain photosynthesis in one line."
   reply:    "Plants use sunlight to turn water and carbon dioxide into food."

   request:  "List three Indian rivers."
   reply:    "The Ganga, the Godavari and the Narmada."

Then you keep training the model on those pairs, using the same guess-the-next-word method as before. Nothing about the method changes. Only the data changes.

There is one extra trick. You wrap every example in markers saying who is speaking:

   <user>   Explain photosynthesis in one line.   <end>
   <robot>  Plants use sunlight to ...            <end>

Those markers are what teach the model to take turns and to stop. The <end> marker is the model learning where a reply finishes.

The surprising part

This works with astonishingly little data. One famous result got good behaviour from a thousand carefully written examples.

The reason is that the model already knows how to explain photosynthesis. Pretraining gave it that. Instruction tuning is not adding knowledge. It is teaching the model which of its many abilities to use, and in what shape.

Think of a talented cook who has never worked in a restaurant. They do not need cooking lessons. They need to learn to read a ticket and plate the dish.

Where you have already seen this

  • The difference between search suggestions and an AI answer box.
  • A new employee who knows the subject but not the office's way of doing things.
  • Any "Instruct" or "Chat" model on Hugging Face, sitting beside its "Base" version.

Remember this

  • Pretraining teaches text continuation, not helpfulness.
  • Instruction tuning is more of the same training, on request-and-reply examples.
  • It adds almost no knowledge. It selects and shapes what is already there.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch transformers

The first example downloads GPT-2, about 550 MB. The second downloads only a tokenizer, a few MB.

What a base model actually does

base_model_behaviour.py
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tok = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2").eval()


def greedy(prompt, n=40):
    enc = tok(prompt, return_tensors="pt")
    with torch.no_grad():
        out = model.generate(**enc, max_new_tokens=n, do_sample=False,
                             pad_token_id=tok.eos_token_id)
    return tok.decode(out[0, enc.input_ids.shape[1]:])


for p in ["Write a one-line definition of photosynthesis.",
          "What is the capital of France?",
          "List three Indian rivers."]:
    print(f"\nPROMPT: {p}")
    print(f"GPT-2 : {greedy(p)!r}")
Output
PROMPT: Write a one-line definition of photosynthesis.
GPT-2 : '\n\nThe following is a list of the most common photosynthetic plants in the world.\n\n1. Chrysanthemum\n\nThe Chrysanthemum is a perennial plant that grows'

PROMPT: What is the capital of France?
GPT-2 : '\n\nThe capital of France is Paris.\n\nThe capital of France is Paris.\n\nThe capital of France is Paris.\n\nThe capital of France is Paris.\n\nThe capital'

PROMPT: List three Indian rivers.
GPT-2 : '\n\nThe first is the Kudankulam, which is the largest river in the country. It flows through the state of Karnataka, and is the largest river in the country. It'

Greedy decoding with do_sample=False, so this reproduces exactly. Written against transformers 5.6.2 and PyTorch 2.5.1, CPU.

Every failure mode is visible in those three outputs

Prompt 1: the instruction was ignored. GPT-2 did not write a definition. It wrote the heading-plus-listicle that usually follows a sentence like that on a web page. It is doing its job perfectly and its job is not what you wanted.

Prompt 2: correct, then stuck in a loop. It knows the answer. It also has no concept of "the reply is finished", so it keeps producing plausible next text forever. Instruction tuning teaches an end-of-turn token; without one, only your max_new_tokens stops it.

Prompt 3: confidently wrong. Kudankulam is a nuclear power station, not a river. Instruction tuning does not fix this. A tuned model states the same kind of falsehood in a tidier format. Hold on to that: format and truth are separate problems. See hallucination.

The chat template is the actual interface

Instruction-tuned models were trained on one exact string layout. Use anything else and you are off-distribution.

chat_templates.py
from transformers import AutoTokenizer

msgs = [
    {"role": "system", "content": "You are terse."},
    {"role": "user", "content": "Capital of France?"},
    {"role": "assistant", "content": "Paris."},
]

for name in ["Qwen/Qwen3-0.6B", "HuggingFaceTB/SmolLM2-135M-Instruct"]:
    tok = AutoTokenizer.from_pretrained(name)
    print(f"\n===== {name} =====")
    print(repr(tok.apply_chat_template(msgs, tokenize=False)))
    print("--- for generation (assistant turn left open) ---")
    print(repr(tok.apply_chat_template(msgs[:2], tokenize=False, add_generation_prompt=True)))
Output
===== Qwen/Qwen3-0.6B =====
'<|im_start|>system\nYou are terse.<|im_end|>\n<|im_start|>user\nCapital of France?<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nParis.<|im_end|>\n'
--- for generation (assistant turn left open) ---
'<|im_start|>system\nYou are terse.<|im_end|>\n<|im_start|>user\nCapital of France?<|im_end|>\n<|im_start|>assistant\n'

===== HuggingFaceTB/SmolLM2-135M-Instruct =====
'<|im_start|>system\nYou are terse.<|im_end|>\n<|im_start|>user\nCapital of France?<|im_end|>\n<|im_start|>assistant\nParis.<|im_end|>\n'
--- for generation (assistant turn left open) ---
'<|im_start|>system\nYou are terse.<|im_end|>\n<|im_start|>user\nCapital of France?<|im_end|>\n<|im_start|>assistant\n'

Two things to notice, both of which bite people.

add_generation_prompt=True is not optional at inference. It appends the opening of the assistant turn. Leave it out and the model has to guess that it is its turn, which it often gets wrong.

Qwen3 injects an empty <think></think> block into the assistant turn. That is a model-specific detail of a reasoning-capable template, not a general convention. It is exactly why you must render the template from the model's own tokenizer instead of hand-writing <|im_start|>. More on that block in thinking tokens and reasoning models.

The training loop itself

Instruction tuning is ordinary causal language-model training on templated text. In TRL (v1.12.0) the whole thing is:

python
from datasets import load_dataset
from trl import SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen3-0.6B",
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()

No output block: this needs a GPU and downloads gigabytes, and its logs vary by hardware. What matters is that there is no new loss function and no new algorithm on that line. The next lesson, formatting an SFT dataset, covers the part that does need care.

Common mistakes

Hand-writing the chat template. Copying <|im_start|>user\n... from a blog post and getting a newline wrong degrades quality silently. Always call apply_chat_template.

Fine-tuning a base model and evaluating it with an instruct template. Base models have no special tokens for roles. The template will insert tokens the model has never seen.

Expecting instruction tuning to add knowledge. It reshapes behaviour. To add facts, use RAG or continued pretraining on the new domain.

Training on too much mediocre data. Beyond a few tens of thousands of good examples, more mediocre examples make the model worse. Quality dominates quantity here more than anywhere else in the stack.

Forgetting the end-of-turn token in your training data. Then the model never learns to stop, and you have reproduced GPT-2's second failure in a tuned model.

Try it yourself

Run the first script again with "Q: What is the capital of France?\nA:" as the prompt. GPT-2 handles it far better, because that pattern is common in its training data. That is in-context format matching, and it is the poor cousin of instruction tuning — the same idea, done through the prompt instead of the weights.

What to learn next

Researcher — Mathematics and papers.

The lineage

Instruction tuning began as a multi-task generalisation result, not an alignment technique.

Wei et al., 2022 (FLAN: Finetuned Language Models Are Zero-Shot Learners) converted 62 existing NLP datasets into natural-language instruction templates and fine-tuned a 137B model on them. Held-out task performance improved sharply, establishing that instruction-format supervision transfers across tasks. Sanh et al., 2022 (T0) reached the same conclusion independently with a prompt-sourced multi-task mixture.

Ouyang et al., 2022 (InstructGPT) reframed it: not "generalise to new NLP tasks", but "follow arbitrary human requests". Their pipeline — SFT on human demonstrations, then a reward model, then PPO — became the template for the whole field. Note the ordering: SFT comes first and everything after it is a refinement of the SFT policy.

Chung et al., 2022 (Scaling Instruction-Finetuned Language Models) scaled the task count to 1,836 and showed gains continue with more tasks, and that including chain-of-thought data is what preserves reasoning ability.

How little data is needed

Zhou et al., 2023 (LIMA: Less Is More for Alignment) fine-tuned LLaMA-65B on 1,000 carefully curated prompt–response pairs, with no RLHF at all, and reported outputs preferred or equal to GPT-4's in 43% of a human comparison. Their Superficial Alignment Hypothesis states it plainly: a model's knowledge and capabilities are learned almost entirely in pretraining, and alignment teaches which subdistribution of formats to use when interacting with users.

The hypothesis is contested at the margins — RLHF and reasoning-focused post-training demonstrably add capability on verifiable tasks, which pure format selection cannot explain. But for instruction following specifically, the evidence for it is strong, and the practical consequence is reliable: 1,000 excellent examples beat 100,000 mediocre ones.

Corroborating results: Zhou et al., 2023 (AlpaGasus) filtered 52k Alpaca examples down to 9k by GPT-4 quality rating and got a better model; Xu et al., 2023 (WizardLM, Evol-Instruct) generated complexity-escalated instructions and beat larger hand-built sets.

The synthetic-data pipeline

Wang et al., 2023 (Self-Instruct) bootstrapped 52k instructions from 175 seed tasks using the model itself, filtering by ROUGE-L similarity against existing instructions. Alpaca applied it to GPT-3.5 outputs. Nearly every open instruction dataset since is a descendant.

The known failure mode is distribution collapse: naive self-generation narrows diversity with each round. Mitigations that work are complexity escalation (Evol-Instruct), explicit taxonomy coverage, and mixing in human data. Generating more of the same does not work.

Loss masking

The convention is to compute loss only on assistant tokens, masking prompt tokens with -100. The rationale — do not spend capacity modelling the user's distribution — is sound and is demonstrated in the next lesson.

The evidence that it improves results is weaker than the convention suggests. Shi et al., 2024 (Instruction Tuning With Loss Over Instructions) find that including the instruction in the loss helps in exactly the regimes where SFT is most fragile: few examples, and long instructions relative to short completions. They report it acting as a regulariser against overfitting. Treat masking as the sensible default and as a hyperparameter, not as a law.

What instruction tuning costs

The alignment tax is the capability drop on standard benchmarks after alignment training. Ouyang et al. observed it on InstructGPT and reduced it by mixing pretraining gradients into the RLHF phase (PPO-ptx). Chung et al. found instruction tuning with chain-of-thought data avoids the reasoning-specific portion of the tax.

The related and larger risk is catastrophic forgetting when tuning on a narrow domain. The standard mitigations are a replay mixture, parameter-efficient methods, and a low learning rate — typically 1e-5 to 2e-5 for full fine-tuning, roughly two orders of magnitude below pretraining's peak.

Papers

What to learn next