How it works

How model training works

Training is one loop run millions of times — guess, measure the error, trace blame backwards through the network, nudge every weight, repeat until the guesses get good.

On this page 8
  1. The pipeline at a glance
  2. Stage 0 — data, split honestly
  3. Stage 1 — forward: the model guesses
  4. Stage 2 — loss: score the miss
  5. Stage 3 — backward: assign the blame
  6. Stage 4 — update: nudge every weight
  7. Stage 5 — validate, and stop at the right time
  8. Which lessons teach each stage

Every model you have heard of — ChatGPT, the camera's face unlock, the fraud alert on your UPI app — was made by the same loop. Guess, measure the miss, adjust, repeat. Millions of times. This page walks through one turn of that loop, and what surrounds it.

The pipeline at a glance

 [0. data]  examples with right answers, split into train / validation
        |
        v
 +--------------------------------------------------+
 |              THE TRAINING LOOP                   |
 |                                                  |
 |  [1. forward]   model makes a prediction         |
 |       |                                          |
 |  [2. loss]      one number: how wrong was it?    |
 |       |                                          |
 |  [3. backward]  trace each weight's share of     |
 |       |         the blame (backpropagation)      |
 |  [4. update]    nudge every weight a tiny step   |
 |       |                                          |
 |       +---> next batch, back to 1                |
 +--------------------------------------------------+
        |
        v
 [5. validate]  check on unseen data; stop at the right time

Stage 0 — data, split honestly

Training needs examples with known answers: photos labelled "cat", sentences with their next word, transactions marked fraud or fine. Before anything runs, the data is split. Most becomes the training set, the part the model learns from. A slice is held back as the validation set, which the model never trains on — the mock exam with fresh questions.

Data is fed in batches — a few dozen examples at a time. Like rolling several chapatis before firing the tawa: one at a time wastes the stove, all at once will not fit.

Stage 1 — forward: the model guesses

A batch flows through the network, input to output, and the model produces predictions. At the start these are garbage, and that is fine. The model's weights — the millions of adjustable numbers inside it — begin random. Training exists to make them non-random.

Stage 2 — loss: score the miss

A loss function compares predictions to the right answers and boils the difference down to one number. High means badly wrong, near zero means close. This single number is the entire training signal. The model is not told what the right answer was, only how wrong it scored — like a hot-and-cold game where every guess gets back one word: colder, warmer.

Stage 3 — backward: assign the blame

Here is the famous part. Backpropagation works backwards from the loss, layer by layer, computing for every single weight: if this weight rose a touch, would the loss rise or fall, and how sharply? That per-weight answer is its gradient — a direction and a strength for the nudge to come.

This is calculus doing detective work: one number of blame at the output, apportioned fairly among millions of contributors.

Stage 4 — update: nudge every weight

Every weight takes a small step against its gradient — downhill, toward less error. This is gradient descent. How big a step is set by the learning rate, and it is the touchiest dial in training. Too large, and the model overshoots the valley and the loss explodes. Too small, and training crawls for weeks. Think of walking to a doorway in the dark: big confident strides, then smaller careful ones.

Then the next batch loads, and the loop turns again. One full pass over the training data is an epoch; training runs for several.

Stage 5 — validate, and stop at the right time

Falling training loss is not the goal — it can fall for a bad reason. Given enough epochs, the model starts memorising the training examples instead of learning the pattern. That is overfitting: the student who memorised last year's paper and fails this year's.

The validation set exposes it. Training loss keeps improving while validation loss stalls, then worsens. The standard response is early stopping: keep the checkpoint from where validation was best, and end training there. Passing the mock exam, not reciting the textbook, was always the goal.

Which lessons teach each stage