Pretraining
In one sentence Pretraining is the massive first stage of building an LLM — predicting the next token across trillions of words until general capability emerges.
Updated
Pretraining is the first, enormous stage of training a model: months of predicting the next token across a huge slice of human text, before any task-specific teaching.
It is the childhood-and-education phase of the model's life. A doctor's training does not start with medicine; it starts with two decades of general education — language, mathematics, how the world works. Only on top of that foundation does the specialisation land. Pretraining is those two decades, compressed: the model reads trillions of tokens of web text, books and code, with one unchanging exercise — guess the next token, check, adjust.
That one exercise is self-supervised-learning, and its magic is what accurate guessing requires. To continue "the second-largest planet in the solar system is", the model needs astronomy. To continue dialogue, it needs some model of how people speak and think. Grammar, facts, style and a surprising amount of reasoning get absorbed as side effects of one prediction game — with no labelled data, which is what makes internet scale possible.
The scale is the defining feature: thousands of GPUs, months of continuous training, budgets in the tens or hundreds of millions of dollars, with outcomes forecast by scaling-laws. The output — a "base model" — is a foundation-model: capable, general, and unhelpful in raw form. The far cheaper later stages (instruction-tuning, RLHF) shape it into an assistant. Roughly: pretraining installs the knowledge; post-training installs the manners.
Where to go next
- Full lesson: How LLMs work
- Related terms: self-supervised-learning, fine-tuning, foundation-model, scaling-laws