How it works
How fine-tuning works
Fine-tuning takes a model that already knows language, shows it your examples, and nudges its weights gently until it behaves your way without forgetting what it knew.
- 4 min read
- Updated
On this page 7
Training a large model from scratch costs millions of dollars. Fine-tuning skips almost all of it: you take a model that already speaks the language, and teach it your specific job — your tone, your format, your domain. Like hiring an experienced doctor and giving them two weeks of hospital-specific induction, instead of running a new medical school. Here is the pipeline.
The pipeline at a glance
pretrained model (already knows language)
|
v
[1. collect examples] input -> desired output pairs
|
v
[2. format the data] consistent prompt/response template
|
v
[3. train gently] small learning rate, few passes
| (often via LoRA adapters)
v
[4. evaluate] held-out examples it never trained on
|
v
[5. deploy] base model + your small deltaStage 1 — collect example pairs
Fine-tuning data is pairs: an input, and the output you wish the model had given. A support question and the ideal reply. A contract clause and the correct risk summary. Hundreds of clean pairs often beat tens of thousands of sloppy ones, because the model copies whatever it sees — including your mistakes, at scale.
One decision comes first, though: do you need fine-tuning at all? If the model lacks knowledge, retrieval (RAG) is usually the better fix. Fine-tuning shines at changing behaviour — tone, format, and reliably following a house style.
Stage 2 — format the data
Every example is pressed into one consistent template — the same structure the model will see in production, with the same markers separating instruction from response. Inconsistent formatting is the most boring and most common reason fine-tunes disappoint. The model learns the template as part of the task.
Stage 3 — train, gently
Now the actual training. Show a batch of examples, and measure how far the model's output is from the target. Then nudge every weight slightly in the direction that reduces the error. This is ordinary gradient descent — the same mechanism that trained the base model.
The critical word is gently. The learning rate — the size of each nudge — is kept tiny, and the data gets only a few passes. Push too hard and the model overwrites what it knew before, a failure with its own name: catastrophic forgetting. Your legal-summary model suddenly writes broken English.
Most fine-tuning today does not touch the original weights at all. LoRA — low-rank adaptation — freezes the base model and trains small add-on matrices beside it, under 1% of the total parameters. The result trains on one modest GPU and ships as a file of a few hundred megabytes instead of a full model copy. You can keep one base model and swap task-specific adapters like lenses on a camera.
Stage 4 — evaluate on unseen examples
Before deploying, test on held-out pairs the model never trained on. Training loss falling means little by itself — a model can memorise your examples and generalise nothing, which is overfitting. Compare outputs side by side against the base model, and check the model still handles general requests, to catch forgetting early.
Stage 5 — deploy
Ship the merged model, or the base model plus your adapter. From the outside nothing changed: same API, same speed. The difference is that the behaviour you used to beg for in a long prompt is now baked into the weights — which also makes every future request cheaper and shorter.
Chat models like ChatGPT went through this same machinery at industrial scale: instruction tuning is fine-tuning on question–answer pairs, followed by preference tuning (RLHF) on human judgements.
Which lessons teach each stage
- The whole technique in depth: Fine-tuning
- Stage 3, the adapter trick: LoRA
- The nudging mechanism itself: Gradient descent
- The preference-tuning stage after this one: RLHF