AI glossary

Quantization

In one sentence Quantization stores a model's numbers with fewer bits, so it needs less memory and runs faster, at the cost of a small drop in quality.

By Updated

Quantization stores each of a model's numbers using fewer bits, which shrinks the model and speeds it up, in exchange for a small loss of precision.

Saving a photograph as a smaller JPEG is the same trade. Some fine detail goes, the file drops from 8 MB to 800 KB, and on a phone screen nobody notices. A model's weights carry the same kind of redundancy: storing each one to eight decimal places is far more precision than the model's answers actually need.

Normally a weight is a 16-bit or 32-bit floating point number. Quantization maps ranges of those values onto a much smaller set: 256 levels for 8-bit, 16 levels for 4-bit. It also records a scale factor per group of weights, so the original range can be approximated on the way back.

What you get for it

Precision7B model sizeQualityTypical use
float16~14 GBReferenceTraining, high-end serving
8-bit~7 GBVery close to referenceSafe default for serving
4-bit~3.5-4 GBSlightly weaker, usually acceptableRunning locally on one GPU

Memory is not the only gain. Moving weights from memory to the compute units is the slowest part of generating a token, so halving the bytes often nearly halves that wait. This is why a 4-bit model can feel faster as well as smaller.

The honest caveats: quality loss is small at 8-bit and real at 4-bit. It shows up first in long reasoning chains and less common languages, rather than in short answers. Test on your own task instead of trusting a benchmark. Very small models suffer more than large ones, because they have less redundancy to spare. And formats matter. GGUF is what llama.cpp and Ollama use on CPU or mixed hardware, while GPTQ and AWQ are GPU-oriented. bitsandbytes gives you 4-bit and 8-bit loading directly inside Hugging Face Transformers, which is also what QLoRA builds on.

Where to go next