Generative AI

Fine-tuning

Fine-tuning continues training an already-trained model on your own examples, which changes how it behaves — and is the wrong tool for most problems beginners reach for it with.

On this page 8
  1. Why it exists
  2. How it works
  3. The most important paragraph on this page
  4. When NOT to fine-tune
  5. What it costs you, honestly
  6. Where you have already seen it
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Fine-tuning means taking a model that already works and training it a little more on your own examples.

Think about hiring a cook who has run a kitchen for fifteen years. You do not teach them to hold a knife or judge heat. You spend a week teaching them your five signature dishes and how your regulars like them.

That week is fine-tuning. The skill was already there. You shaped it to your kitchen.

The model arrives knowing grammar, facts and reasoning. You give it a few hundred examples of the job you actually want done, and it settles into that job.

Why it exists

Training a language model from nothing is enormously expensive. It takes huge datasets, months of computing time, and budgets in the tens of millions of dollars. Almost nobody can do it, and almost nobody needs to.

Fine-tuning starts from someone else's finished model. Your part might take a few hours on one rented graphics card.

The catch is what it can and cannot change.

How it works

  a trained model  +  your 500 examples  →  more training, gently
                                                    |
                                                    ↓
                                         a model with the same
                                         knowledge, new habits

Each example is a pair: an input, and the reply you wanted. The model reads the input, produces its own reply, and is corrected towards yours. Repeat a few thousand times and the habit sticks.

Nothing is added to the model. Every one of its numbers shifts a little.

The most important paragraph on this page

Fine-tuning teaches behaviour reliably. It teaches facts badly.

Feed it a hundred examples of your company's tone and it will pick up the tone. Feed it a hundred facts about your products and something else happens. You get text that sounds like your product documents, with the details quietly wrong.

Facts are not stored in a model the way rows are stored in a table. They are smeared across billions of numbers. Adding one fact by nudging those numbers is unreliable, and you cannot check whether it landed.

For facts, use RAG — look the fact up and hand it to the model. It is cheaper, it is checkable, and updating it takes a second.

When NOT to fine-tune

Most beginners reach for fine-tuning far too early. Walk down this list, and stop at the first step that works.

1. Write a clearer prompt. Say the format, the tone, the length and what to avoid. Most "the model won't behave" problems die here. See prompt engineering.

2. Put examples in the prompt. Three or four worked examples inside the request teach format remarkably well, and you can change them in a second.

3. Give it the documents. If the problem is that the model does not know something, that is a retrieval problem, not a training problem.

4. Break the task into steps. One prompt doing five things badly is often five prompts doing one thing well.

Only then consider fine-tuning. It earns its place when one of these is true.

  • The behaviour you want is hard to describe, but easy to demonstrate.
  • You have hundreds of good examples already.
  • The task is narrow and will not change much.
  • Your prompt has grown so long that shortening it saves real money.

What it costs you, honestly

You now own a model. That is a heavier thing than it sounds.

  • Someone must collect, clean and check the examples. This is most of the work, and it is dull.
  • Every time the base model is updated, you decide whether to redo the work.
  • You need a way to tell whether the new model is actually better, not only different.
  • A fine-tuned model can get worse at everything you did not train on. That has a name, and the developer tab shows it happening in a few seconds of code.

Where you have already seen it

  • A support assistant that always replies in your company's fixed format, with the ticket ID first.
  • Speech-to-text tuned on Indian accents, because the general model kept mangling names.
  • A coding assistant taught your team's internal libraries and style.
  • Almost every chat model you use is itself a fine-tune. The base model only continues text. Being helpful and answering questions was fine-tuned in.

Remember this

  • Fine-tuning continues training an existing model on your examples.
  • It changes behaviour and style well. It adds facts badly — use RAG for those.
  • Try a better prompt, then examples, then retrieval. Fine-tune last, not first.

What to learn next

  • LoRA — fine-tuning by training a small add-on, which is how it is done in practice.
  • What is RAG? — the right tool when the problem is missing facts, not missing behaviour.
  • Hallucination — why teaching new facts by fine-tuning tends to make things worse.

Developer — Code and libraries.

Setup

bash
pip install torch

Fine-tuning a real language model needs a multi-gigabyte download and a GPU. This page does not ask you for either.

Instead you will build a very small character-level model, pretrain it, then fine-tune it — and watch both the good and the bad effects with your own eyes. It runs in about fifteen seconds on a laptop CPU, and everything happens in one file with no downloads.

Watching a fine-tune happen

tiny_finetune.py
import torch
import torch.nn as nn

torch.manual_seed(0)

# The "pretraining" corpus: ordinary sentences.
BASE = [
    "the train leaves at six in the morning",
    "she opened the window and looked outside",
    "we walked to the market after the rain",
    "he wrote a letter to his brother in delhi",
    "the tea was too hot to drink",
    "birds were singing in the old mango tree",
]

# The "fine-tuning" corpus: one narrow format we want the model to copy.
TARGET = [
    "ticket: refund | status: open | owner: asha",
    "ticket: delay | status: open | owner: ravi",
    "ticket: damage | status: closed | owner: asha",
    "ticket: refund | status: closed | owner: ravi",
]

CHARS = sorted(set("".join(BASE + TARGET)))
STOI = {c: i for i, c in enumerate(CHARS)}
ITOS = {i: c for c, i in STOI.items()}

def encode(s):
    return torch.tensor([[STOI[c] for c in s]])

class CharModel(nn.Module):
    def __init__(self):
        super().__init__()
        self.embed = nn.Embedding(len(CHARS), 24)
        self.gru = nn.GRU(24, 96, batch_first=True)
        self.out = nn.Linear(96, len(CHARS))
    def forward(self, ids, h=None):
        z, h = self.gru(self.embed(ids), h)
        return self.out(z), h

model = CharModel()
lossfn = nn.CrossEntropyLoss()

def corpus_loss(corpus):
    """Average cross-entropy per character. Lower means the text looks more familiar."""
    model.eval()
    with torch.no_grad():
        total = 0.0
        for line in corpus:
            ids = encode(line)
            logits, _ = model(ids[:, :-1])
            total += lossfn(logits[0], ids[0, 1:]).item()
    return total / len(corpus)

def train(corpus, epochs, lr):
    opt = torch.optim.Adam(model.parameters(), lr=lr)
    model.train()
    for _ in range(epochs):
        for line in corpus:
            ids = encode(line)
            opt.zero_grad()
            logits, _ = model(ids[:, :-1])
            loss = lossfn(logits[0], ids[0, 1:])      # predict each character from the ones before it
            loss.backward()
            opt.step()

def sample(prompt, n=40):
    """Greedy continuation: always take the most likely next character."""
    model.eval()
    ids = encode(prompt)
    out = prompt
    h = None
    with torch.no_grad():
        logits, h = model(ids, h)
        for _ in range(n):
            nxt = int(logits[0, -1].argmax())
            out += ITOS[nxt]
            logits, h = model(torch.tensor([[nxt]]), h)
    return out

print("=== after pretraining on ordinary sentences ===")
train(BASE, epochs=120, lr=0.005)
print(f"loss on ordinary sentences : {corpus_loss(BASE):.3f}")
print(f"loss on the ticket format  : {corpus_loss(TARGET):.3f}")
print("sample:", repr(sample("the ")))

print()
print("=== after fine-tuning on 4 ticket lines ===")
train(TARGET, epochs=60, lr=0.002)
print(f"loss on ordinary sentences : {corpus_loss(BASE):.3f}")
print(f"loss on the ticket format  : {corpus_loss(TARGET):.3f}")
print("sample:", repr(sample("the ")))
print("sample:", repr(sample("ticket: ")))
Output
=== after pretraining on ordinary sentences ===
loss on ordinary sentences : 0.009
loss on the ticket format  : 7.296
sample: 'the tea was too hot to drinked to the market'

=== after fine-tuning on 4 ticket lines ===
loss on ordinary sentences : 1.275
loss on the ticket format  : 0.067
sample: 'the ravind | status: closed | owner: ashaves'
sample: 'ticket: delay | status: open | owner: ravind | s'

Exact numbers can shift slightly between PyTorch versions, because initialisation differs. The shape of the result does not.

Everything important is in those eight lines

The fine-tune worked. Loss on the ticket format fell from 7.296 to 0.067. Given the prompt ticket: , the model now produces the format cleanly. Four examples were enough to install a habit.

And it broke the model. Loss on ordinary sentences rose from 0.009 to 1.275. Look at the sample: 'the ravind | status: closed | owner: ashaves'. Given the word "the", it now slides into ticket syntax. It has forgotten how to write a normal sentence.

This is catastrophic forgetting: training on a narrow new task damages performance on everything the model could previously do. It is not an artefact of this toy. It is the central risk of fine-tuning, and it happens on real models too, more subtly and therefore more dangerously.

You did not see it happen here by reading about it. You measured it, on data the model was no longer being trained on. That measurement is the job. A fine-tune without a held-back evaluation set is a change of unknown sign.

Four things that reduce the damage

  • Lower the learning rate. The fine-tune above used 0.002 against 0.005 for pretraining. Real fine-tunes typically drop it by ten to a hundred times. Too high, and you overwrite instead of adjust.
  • Fewer epochs. Two or three passes over your data is normal. Ten is usually memorisation.
  • Mix in general data. Blending some general examples into your fine-tuning set — called replay — keeps the old ability alive. Try adding two BASE lines to TARGET and rerunning.
  • Freeze most of the model. Train a small add-on instead of every weight. That is LoRA, and it is what almost everyone actually uses.

What a real fine-tuning dataset looks like

Every major toolchain wants JSONL: one JSON object per line, each a short conversation.

train.jsonl
{"messages":[{"role":"system","content":"You classify support tickets. Reply with one word."},{"role":"user","content":"My payment failed but money was deducted."},{"role":"assistant","content":"refund"}]}
{"messages":[{"role":"system","content":"You classify support tickets. Reply with one word."},{"role":"user","content":"The courier has not arrived in nine days."},{"role":"assistant","content":"delay"}]}
{"messages":[{"role":"system","content":"You classify support tickets. Reply with one word."},{"role":"user","content":"The screen was cracked when I opened the box."},{"role":"assistant","content":"damage"}]}

Three rules that matter more than the model you pick:

  1. Every example must be an answer you would be happy to ship. The model copies your data, including its mistakes. One inconsistent label teaches inconsistency.
  2. The system message must match production exactly. Train with one and serve without it and you have changed the task.
  3. Hold back 10 to 20 percent and never train on it. That held-out set is the only evidence you will have that the fine-tune helped.

Fifty excellent examples beat five thousand scraped ones. This is measured, not folklore — see the LIMA result in the researcher tab.

Common mistakes

Fine-tuning to add knowledge. The single most common error. Facts land unreliably and cannot be audited or updated. Use RAG.

No evaluation set. Without one you are comparing vibes. Write twenty test questions with known good answers before you train, and score the base model on them first.

Training on the prompt tokens. In supervised fine-tuning you compute loss on the assistant's reply, not the user's question. Getting this wrong teaches the model to generate questions. Most libraries handle it, but check the masking setting.

Too many epochs on a small set. Training loss near zero on 200 examples means memorisation. Watch held-out loss and stop when it turns upward.

Changing the base model and the data at once. Then you cannot tell which change caused the result. Move one thing at a time.

Try it yourself

Change TARGET to include two of the BASE sentences alongside the four ticket lines. Rerun and compare the loss on ordinary sentences. That is replay, in the smallest form that demonstrates it.

Then set the fine-tune learning rate to 0.02 and watch the forgetting get dramatically worse. Learning rate is the main dial between "adjusted" and "overwritten".

What to learn next

  • LoRA — fine-tuning by training a small add-on, which is how it is done in practice.
  • What is RAG? — the right tool when the problem is missing facts, not missing behaviour.
  • Hallucination — why teaching new facts by fine-tuning tends to make things worse.

Researcher — Mathematics and papers.

The objective

Supervised fine-tuning is ordinary maximum likelihood on a conditional distribution, with the loss masked to the response:

L(θ) = − Σ_{i}  Σ_{t ∈ response_i}  log p_θ ( y_t | y_{<t}, prompt_i )

θ are the model parameters, prompt_i the input tokens of example i, y_t the t-th token of the target response. Tokens in the prompt receive no loss. The pretraining objective is identical apart from the mask covering everything, which is why SFT is best understood as a continuation of pretraining on a narrow, curated distribution rather than a different algorithm.

Memory arithmetic for full fine-tuning

Per parameter, with Adam and mixed precision:

bf16 weights                    2 bytes
bf16 gradients                  2 bytes
fp32 master weights             4 bytes
fp32 Adam first moment          4 bytes
fp32 Adam second moment         4 bytes
                               ---------
                               16 bytes / parameter

A 7B model therefore needs about 112 GB before a single activation is stored, which exceeds any single accelerator commonly available. Activations add O(layers × batch × sequence × hidden) on top. This arithmetic, not accuracy, is the reason parameter-efficient methods dominate practice — see LoRA.

Catastrophic forgetting

McCloskey and Cohen (1989) named the phenomenon in connectionist networks: sequential training on task B degrades task A, because the parameters encoding A are freely reused for B. It has never been solved, only managed.

Three families of mitigation, in increasing order of practical use:

  • Regularisation towards the old parameters. Elastic Weight Consolidation (Kirkpatrick et al., 2017) adds Σ_i F_i (θ_i − θ*_i)², where θ* are the pretrained parameters and F_i the diagonal Fisher information estimating how much parameter i mattered to the old task. Principled, and rarely used at LLM scale because estimating F is expensive.
  • Replay. Mixing a fraction of general-purpose data into the fine-tuning mixture. Crude, cheap, effective; 5 to 30 percent is the usual range.
  • Constrained capacity. Restrict the update to a low-rank or otherwise small subspace. Biderman et al. (2024), LoRA Learns Less and Forgets Less, quantify the trade directly: LoRA underperforms full fine-tuning on target-domain gains, and preserves base-model capability substantially better.

Data quality dominates data quantity

Zhou et al. (2023), LIMA: Less Is More for Alignment, fine-tuned LLaMA-65B on 1,000 carefully curated prompt–response pairs with no RLHF, and reached responses preferred to or equal with GPT-4's in 43% of a human comparison. Their superficial alignment hypothesis is the load-bearing claim: knowledge and capability are acquired almost entirely during pretraining, and alignment tuning mainly selects which subdistribution of formats and styles the model emits.

This reframes what SFT is for. If capability is already present, adding examples past the point where the format is unambiguous buys little and risks forgetting. It also explains the persistent failure of SFT as a knowledge-injection mechanism: you are selecting a style, not writing to a store.

Gekhman et al. (2024) sharpen this with a measured mechanism: fine-tuning examples containing facts absent from pretraining are learned slowly, and as the model fits them it becomes measurably more prone to hallucinating on other questions. Teaching new facts by SFT trains the behaviour of asserting unsupported claims.

Beyond supervised fine-tuning

RLHF (Ouyang et al., 2022) trains a reward model on human pairwise preferences, then optimises the policy with PPO under a KL penalty against the SFT reference:

max_θ  E [ r_φ(x, y) ]  −  β · D_KL( π_θ(y|x) ‖ π_ref(y|x) )

r_φ is the learned reward, π_ref the frozen SFT model, β the strength of the anchor. The KL term is what stops reward hacking from destroying fluency.

DPO (Rafailov et al., 2023) shows this objective has a closed-form optimal policy, allowing the reward model to be eliminated and preferences to be optimised with a simple classification loss on preferred and rejected pairs. Far cheaper and far more stable; it is the default starting point for preference tuning now. Later variants (IPO, KTO, ORPO) adjust the loss shape or drop the need for paired data.

The practical pipeline is SFT for format and task, then preference optimisation for the qualities that are easier to compare than to demonstrate.

Hyperparameters that actually matter

  • Learning rate: 1e-5 to 2e-5 for full fine-tuning, 1e-4 to 3e-4 for LoRA adapters. The most common cause of a destroyed model is a learning rate borrowed from a pretraining recipe.
  • Epochs: 1 to 3. Held-out loss usually turns upward before training loss looks suspicious.
  • Warmup and cosine decay: standard, and worth keeping; a cold start at full learning rate on a converged model is needlessly destructive.
  • Sequence packing: concatenating short examples to fill the context improves throughput, but requires correct attention masking or examples bleed into each other.

Evaluation

Fine-tuning changes behaviour globally, so evaluate globally.

  • A task set for the thing you trained for, scored automatically where possible.
  • A regression set of general capability, to detect forgetting. Even a small held-out slice of general instruction data catches most damage.
  • A safety set, because narrow fine-tuning is known to weaken alignment. Qi et al. (2023) demonstrated that fine-tuning on a small number of benign examples measurably degrades safety behaviour in aligned models.

Papers

What to learn next

  • LoRA — fine-tuning by training a small add-on, which is how it is done in practice.
  • What is RAG? — the right tool when the problem is missing facts, not missing behaviour.
  • Hallucination — why teaching new facts by fine-tuning tends to make things worse.

What to learn next

These follow on from what you just read.

  • Generative AI

    LoRA

    LoRA fine-tunes a large model by freezing it and training a small add-on beside it, which cuts the memory cost enormously and lets you swap behaviours like plug-in packs.

  • Generative AI

    AI agents

    An AI agent is a language model placed in a loop where it can use tools, look at the result, and decide what to do next.

  • Generative AI

    Diffusion models

    A diffusion model makes an image by starting from pure noise and removing a little of it at a time, having first learned what noise looks like by adding it to real pictures on purpose.