Hyperparameter
In one sentence A hyperparameter is a setting you choose before training starts, unlike a parameter, which the model learns during training.
Updated
A hyperparameter is a setting you pick before training begins — as opposed to a parameter, which the model works out for itself while training.
Baking separates the two cleanly. You choose the oven temperature and the baking time before the tray goes in. What happens to the batter inside — how it rises, how the crust forms — is not something you set, it is the result. Temperature and time are hyperparameters. The cake is the learned parameters.
The distinction matters because the two are found in completely different ways. Parameters are found by gradient descent, automatically, millions at a time. Hyperparameters are found by trying a value, training, measuring on held-out data, and trying another.
The ones you will actually set
| Hyperparameter | What it controls | Typical starting point |
|---|---|---|
| Learning rate | Step size per update | 0.001 with Adam |
| Batch size | Examples per update | 32 or 64, or whatever fits in memory |
| Epochs | Passes over the data | Use early stopping instead of a fixed number |
| Layers and width | Model capacity | Start small, grow if underfitting |
| Dropout | Randomly switching off neurons to fight overfitting | 0.1 to 0.5 |
| Weight decay | Penalty on large weights | 0.01 |
Learning rate is worth more of your attention than the rest combined. It is the setting most likely to make the difference between a model that trains and a model that produces nonsense.
For searching, random search beats grid search on the same budget in most real problems, because only a few hyperparameters matter much and random sampling tries more distinct values of those. Tools like Optuna go further with Bayesian search. Whichever you use, tune against a validation set, never the test set — otherwise your final number is a measure of how hard you searched, not how well the model works.
Where to go next
- Full lesson: Model evaluation
- Related terms: parameter, batch-size, epoch, gradient-descent