Loading a model in 4-bit
Quantisation stores each weight in four bits instead of sixteen, cutting a model's memory to roughly a quarter so it fits on hardware that could not hold it before.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Four-bit loading stores every number in a model with far less precision, so the whole thing takes about a quarter of the memory.
Think of writing down a phone bill. The exact amount is ₹843.6217, but you write ₹844 and nothing goes wrong. You lost precision you were never using. Do that across a whole ledger and the book gets much thinner.
A model is millions of numbers. Most of them do not need full precision to be useful. Quantisation is the word for rounding them down on purpose.
Why it exists
Models grew faster than graphics cards did. A model needing 28 GB of memory does not run on a 12 GB card. It does not run slowly — it refuses to start.
The previous approach was to park spare layers in slower memory. That works, and it crawls, because data has to travel for every answer.
Rounding is the other answer. Instead of moving the model somewhere else, make the model smaller. A quarter of the size fits where the original could not, and it stays on the fast chip where it belongs.
How it works
original weight 0.4812736… 16 bits each
↓ round to one of 16 allowed values
stored weight "value #11" 4 bits each
memory: 538 MB (full) → 269 MB (half) → 110 MB (4-bit)Something is lost. The rounded model gives slightly different answers from the original, and on hard tasks a little worse. For most work that trade is worth taking — a model that runs imperfectly beats a model that does not run.
A real example you have seen
Music streaming. The studio recording is enormous; the file on your phone is a fraction of that, with detail thrown away that most ears never notice. You accept a tiny quality loss to fit a thousand songs on a phone.
Remember this
- Quantisation stores each of a model's numbers in fewer bits.
- Four bits instead of sixteen means roughly a quarter of the memory.
- Accuracy drops a little. Fitting on your hardware at all is usually worth it.
What to learn next
- QLoRA: fine-tuning a large model on one GPU — training on top of a 4-bit base.
- Quantization in practice — the same idea aimed at phones and edge devices.
- Model compression — the other levers: pruning and distillation.
Developer — Code and libraries.
Setup
pip install transformers torch accelerate bitsandbytesTested with transformers 5.6, accelerate 1.13 and bitsandbytes 0.50. The model is HuggingFaceTB/SmolLM2-135M — about 272 MB of files, small enough to download on a slow connection.
Read this before running the second script. bitsandbytes 4-bit loading needs an NVIDIA GPU. On a CPU-only laptop it raises an error rather than falling back. The first script below runs anywhere; the second one needs a CUDA machine.
The arithmetic, on any machine
import torch
from transformers import AutoModelForCausalLM
NAME = "HuggingFaceTB/SmolLM2-135M"
model = AutoModelForCausalLM.from_pretrained(NAME, dtype=torch.float32)
params = sum(p.numel() for p in model.parameters())
print(f"parameters : {params:,}")
print(f"fp32 footprint : {round(model.get_memory_footprint() / 1e6)} MB")
for bits, name in [(16, "fp16 / bf16"), (8, "int8"), (4, "4-bit")]:
print(f"{name:15}: {round(params * bits / 8 / 1e6)} MB of weights (before overheads)")parameters : 134,515,008 fp32 footprint : 538 MB fp16 / bf16 : 269 MB of weights (before overheads) int8 : 135 MB of weights (before overheads) 4-bit : 67 MB of weights (before overheads)
Do this arithmetic for any model before downloading it. Parameters × bytes-per-weight is the floor, and inference needs perhaps 20% more for activations and the attention cache.
The real thing, on a CUDA GPU
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
NAME = "HuggingFaceTB/SmolLM2-135M"
half = AutoModelForCausalLM.from_pretrained(NAME, dtype=torch.float16, device_map="cuda:0")
print("fp16 footprint :", round(half.get_memory_footprint() / 1e6), "MB")
del half; torch.cuda.empty_cache()
cfg = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # the 4-bit format designed for normally-distributed weights
bnb_4bit_use_double_quant=True, # quantise the quantisation constants too
bnb_4bit_compute_dtype=torch.bfloat16, # maths still happens in 16-bit
)
q = AutoModelForCausalLM.from_pretrained(NAME, quantization_config=cfg, device_map="cuda:0")
print("4-bit footprint:", round(q.get_memory_footprint() / 1e6), "MB")
print("attention layer:", type(q.model.layers[0].self_attn.q_proj).__name__)
print("output head :", type(q.lm_head).__name__)
tok = AutoTokenizer.from_pretrained(NAME)
ids = tok("The best street food in Mumbai is", return_tensors="pt").to(q.device)
out = q.generate(**ids, max_new_tokens=12, do_sample=False, pad_token_id=tok.eos_token_id)
print("sample :", tok.decode(out[0], skip_special_tokens=True))fp16 footprint : 269 MB 4-bit footprint: 110 MB attention layer: Linear4bit output head : Linear sample : The best street food in Mumbai is the popular street food in Mumbai, the most popular street food
These numbers are from an NVIDIA RTX A6000. Footprints should match closely on any CUDA card; the generated sentence is greedy and stable, but a different bitsandbytes build can shift it.
The walkthrough
269 down to 110, not 67. The arithmetic promised 67 MB of weights, so where did the extra 43 MB come from? Look at the last two printed types. Attention layers became Linear4bit, but lm_head is still a plain Linear. The output head is left in 16-bit on purpose — it is the most precision-sensitive layer, and quantising it costs quality for little gain. Embeddings stay full width too. Expect a real reduction near 2.5× on small models, closer to the promised 4× on large ones where the transformer blocks dominate.
nf4 is not ordinary rounding. NormalFloat4 places its 16 allowed values at the quantiles of a normal distribution, because trained weights are roughly normally distributed. Even spacing wastes most of its levels on ranges that hold almost no weights. Use "nf4" unless you have measured "fp4" being better for your case, which is rare.
bnb_4bit_compute_dtype is the one people leave wrong. Weights are stored in 4 bits and computed in 16. Leave this at its float32 default and every matrix multiply runs in full precision, which is markedly slower for no quality gain. Set bfloat16 on Ampere-generation cards and newer, float16 on older ones.
Double quantisation is a small free win. Quantisation stores a scaling constant per block of weights, and those constants are themselves numbers worth compressing. bnb_4bit_use_double_quant=True saves roughly 0.4 bits per parameter — around 3% of the model, at no measurable quality cost.
A 4-bit model still generates. The output above is not a great sentence, because a 135M-parameter model is not a great model. But the loop ran, on a fraction of the memory. That is the whole claim.
Common mistakes
Expecting exactly a quarter. Unquantised layers, activations and the attention cache all take room. Budget from measurement, not from arithmetic — and remember the cache grows with your prompt length.
Calling .to("cuda") on a quantised model. It is already placed by device_map, and moving it raises an error about quantised weights. Pick device_map and leave placement alone, as in loading models bigger than your GPU.
Trying to save it with save_pretrained and reload elsewhere. A bitsandbytes 4-bit model can be serialised, but it reloads only where bitsandbytes runs. For portable quantised weights, the GPTQ and AWQ formats exist and ship pre-quantised on the Hub.
Assuming quantisation is free accuracy-wise. On short factual answers the drop is usually small. On long chains of reasoning, errors compound. Measure your own task before and after — see model evaluation.
Reaching for 4-bit when the model already fits. A quantised model is often slower per token than the same model in fp16, because weights are unpacked back to 16 bits before every multiply. Quantise to make something possible, not to make it fast — quantization in practice covers when the speed argument does hold.
Try it yourself
Run memory_math.py for a model you actually want: read its parameter count from the Hub page and work out the 4-bit floor before downloading a single byte. Then, on a CUDA machine, load SmolLM2-135M twice — bnb_4bit_compute_dtype=torch.float32 and torch.bfloat16 — and time 50 generated tokens each way.
What to learn next
- QLoRA: fine-tuning a large model on one GPU — training on top of a 4-bit base.
- Quantization in practice — the same idea aimed at phones and edge devices.
- Model compression — the other levers: pruning and distillation.
Researcher — Mathematics and papers.
Block-wise quantisation, stated exactly
Weights are split into contiguous blocks of size B (64 for 4-bit in bitsandbytes). For each block, an absolute-maximum scale c = max|w| is stored, and each weight is mapped to the nearest of 2⁴ codebook entries after dividing by c. Dequantisation is ŵ = c · Q[i]. Block-wise scaling — rather than one scale per tensor — is what bounds the damage from outlier weights, which are common and catastrophic under a global scale (Dettmers et al., 2022, LLM.int8(), made the same argument for the 8-bit case with per-column scaling and an fp16 outlier path).
NF4 (Dettmers et al., 2023, QLoRA) chooses the codebook as the quantiles of a standard normal, so each of the 16 levels receives roughly equal probability mass under the empirical weight distribution — information-theoretically optimal for exactly-normal inputs, and close to it for real weights. Double quantisation then quantises the per-block scales themselves (blocks of 256, to 8 bits), reducing the scale overhead from 32/64 = 0.5 bits per parameter to about 0.127.
Effective bits per parameter under NF4 with double quantisation is therefore near 4.127, against 16 for fp16 — the ~3.9× ratio that shows up in practice only when linear layers dominate. Embeddings, the LM head, layer norms and biases stay 16-bit, which is why the 135M-parameter demo achieved 2.4×.
Where the accuracy goes
Inference is not integer arithmetic here: Linear4bit dequantises weights to compute_dtype on the fly and runs a standard 16-bit GEMM. Quantisation error is therefore a fixed perturbation of the weights, not an accumulating numerical drift. Perplexity degradation for 4-bit NF4 on multi-billion-parameter models is typically small and shrinks with scale — larger models are more redundant. Small models are the fragile case, which cuts against the intuition that small models are safe to compress.
The alternatives trade calibration effort for quality. GPTQ (Frantar et al., 2023) minimises layer-wise reconstruction error against a calibration set using approximate second-order information; AWQ (Lin et al., 2023) scales salient channels identified by activation magnitude before quantising. Both are post-training methods needing calibration data and produce portable artifacts. Bitsandbytes NF4 needs no calibration at all, which is precisely why it is the format used for QLoRA fine-tuning.
The throughput picture
Memory reduction does not imply speedup. Per-token decoding is memory-bandwidth-bound, so fewer weight bytes read per token can help; but the dequantisation step adds work, and bitsandbytes kernels are not the fastest available. At batch size 1 on a small model, 4-bit is frequently slower than fp16. Quantisation wins on latency in one regime: when it changes what is possible at all. a model that fits entirely in GPU memory instead of spilling to CPU, or a batch size that fits instead of overflowing. See latency and throughput for the general shape of the argument.
What to learn next
- QLoRA: fine-tuning a large model on one GPU — training on top of a 4-bit base.
- Quantization in practice — the same idea aimed at phones and edge devices.
- Model compression — the other levers: pruning and distillation.