AI glossary

Checkpoint

In one sentence A checkpoint is a saved snapshot of a model's weights at a moment in training — the file you resume from, evaluate, or ship.

By Updated

A checkpoint is a snapshot of a model's parameters saved to disk at some point during or after training.

It is the save file of a long video game. A power cut at hour 40 of a 60-hour game is a tragedy only if you never saved. Training runs are long games played on flaky hardware — large LLM runs expect node failures — so training loops save regularly, and a crash costs hours, not weeks. That resume-ability alone would justify the practice.

But checkpoints do more than insure. Saving at intervals gives you the model at many moments in time, and the best model is often not the last one — validation performance peaks and then decays into overfitting, so the standard pattern keeps "best so far" and "latest" side by side, and early-stopping is implemented as "restore the best checkpoint". A full training checkpoint holds more than weights: optimizer state (Adam's running averages), the learning-rate schedule position, and the data position — everything needed to continue as if never interrupted. A deployment checkpoint strips down to weights alone, at a fraction of the size.

The word has also become the unit of distribution: "download the checkpoint from Hugging Face" means the trained weights file. Formats you will meet: PyTorch's .pt/.bin (pickle-based — capable of executing code on load, so treat unknown files with caution) and the now-standard safetensors, which loads safely and quickly. A published base model is a checkpoint others fine-tune from — one more reason the save file, not the training run, is the field's currency.

Where to go next