AI glossary

Scaling laws

In one sentence Scaling laws are the observed formulas relating model size, data and compute to performance, letting labs forecast a model before spending millions training it.

By Updated

Scaling laws are empirical formulas showing that a language model's loss falls smoothly and predictably as you increase model size, training data, and compute.

Civil engineers can promise what a bridge will bear before pouring concrete, because materials follow known formulas. For decades, machine learning had no such promise — you trained and found out. Scaling laws changed that: plot loss against model size, data size, or compute on log-log paper, and the points fall on strikingly straight lines. The straight lines extrapolate, so a lab can train a family of small models, fit the curve, and forecast a 1,000× larger run before committing the millions.

Two results anchor the field. Kaplan et al. (OpenAI, 2020) established the smooth power-law behaviour and justified racing to bigger models. Hoffmann et al. (DeepMind, 2022 — the "Chinchilla" paper) corrected the balance: for a fixed compute budget, models were being built too big and fed too little, with the compute-optimal recipe near 20 training tokens per parameter. Modern practice deliberately "overtrains" past that ratio on far more data, because a smaller model trained longer is cheaper to serve forever after.

Two honest caveats. The laws predict next-token loss (equivalently perplexity), and specific downstream skills can improve in jumps even while loss falls smoothly — the debate around "emergent abilities". And each law describes one training recipe; better data or architecture shifts the whole curve, which is where much current competition lives.

Where to go next