Parameter
In one sentence A parameter is one number inside a model that training adjusts — the model's whole learned knowledge is nothing but billions of these numbers.
Updated
A parameter is a single number inside a model that training is allowed to change. Everything a model knows is stored in these numbers and nowhere else.
Picture the mixing desk in a recording studio, with rows of sliders. Each slider changes how much of one instrument reaches the final track. Nobody labels a slider "make it sound good" — the good sound is what emerges when all of them are in the right positions. Training is the long process of moving every slider a tiny amount at a time until the output is right, and a modern model has billions of sliders.
Each parameter is either a weight, which scales an incoming signal, or a bias, which shifts it. Both are found by gradient descent, never chosen by you. Settings you choose yourself, like learning rate or batch size, are hyperparameters — a different thing with a confusingly similar name.
Why the count is printed on every model
When a model is called 7B, that is 7 billion parameters, and the number tells you what hardware you need.
7 billion parameters, weights only
float32 4 bytes each → 28 GB
float16 2 bytes each → 14 GB ← common default for inference
8-bit 1 byte each → 7 GB
4-bit 0.5 byte each → 3.5 GB ← runs on a consumer GPUAdd roughly 1 to 3 GB on top for activations and the KV cache during generation. That arithmetic is the whole reason quantization exists, and it explains at a glance why a 70B model will not load on a 16 GB card in half precision.
More parameters generally means more capability and always means more memory, more electricity and more latency. It does not automatically mean better for your task — a well fine-tuned small model regularly beats a large general one on a narrow job, at a fraction of the cost.
Where to go next
- Full lesson: What is a neural network?
- Related terms: hyperparameter, quantization, tensor, llm