QLoRA: fine-tuning a large model on one GPU
QLoRA freezes a 4-bit copy of the model and trains a few million extra numbers beside it, so fine-tuning fits on one card and the result saves as a few megabytes.
- 12 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
QLoRA freezes a squashed copy of a big model and trains a small set of new numbers alongside it. One graphics card is then enough.
Think of a rented flat you are not allowed to renovate. You cannot knock down walls, so you add your own furniture, curtains and shelves. The flat is unchanged; your things sit on top and make it yours. When you leave, you carry away a few boxes, not a building.
The frozen model is the flat. The adapter — the small set of trainable numbers — is your furniture. Training touches only your boxes.
Why it exists
Ordinary fine-tuning updates every number in the model. That needs room for the model, a copy of its corrections, and the optimizer's own bookkeeping. Often twelve to sixteen times the model's own size. A 7-billion-number model then wants around 100 GB. No consumer card comes close.
Two ideas together fix it. Store the frozen model in 4 bits, which shrinks it to a quarter. And only train a small add-on, which removes almost all of the bookkeeping.
The result is a fine-tune that fits on one ordinary card and saves as a file you can email.
How it works
full fine-tuning: [======== whole model, all of it trainable ========] huge memory
QLoRA: [======== frozen 4-bit model, untouched ========]
+ [tiny adapter] ← the only trainable part
↓
save just the adapter: megabytes, not gigabytesBecause the base never changes, one downloaded model can wear many adapters. A support-tone adapter, a summarising adapter, a Marathi adapter — swap them like SIM cards.
A real example you have seen
Phone cases. Millions of people own the same phone model and make it theirs with a case, a skin, a screen protector. Nobody manufactures a new phone to get a different colour.
Remember this
- The big model is frozen and stored in 4 bits; only a small adapter trains.
- Memory drops enough that one ordinary graphics card is sufficient.
- The saved result is megabytes, and one base model can carry many adapters.
What to learn next
- LoRA — the beginner-level story of the low-rank idea.
- Fine-tuning — when to fine-tune at all, versus prompting or retrieval.
- RLHF — the step after supervised fine-tuning.
Developer — Code and libraries.
Setup
pip install transformers torch datasets accelerate peft bitsandbytesTested with transformers 5.6, peft 0.20, accelerate 1.13 and bitsandbytes 0.50.
Two scripts follow. The first shows what an adapter is and runs anywhere, including a CPU-only laptop. The second is real QLoRA and needs an NVIDIA GPU, because 4-bit loading does.
What an adapter actually adds — CPU friendly
from transformers import AutoModelForSequenceClassification
from peft import LoraConfig, get_peft_model, TaskType
base = AutoModelForSequenceClassification.from_pretrained(
"google/bert_uncased_L-2_H-128_A-2", num_labels=2) # BERT-tiny, ~18 MB
model = get_peft_model(base, LoraConfig(
task_type=TaskType.SEQ_CLS,
r=8, # rank: how wide the two small matrices are
lora_alpha=16,
lora_dropout=0.05,
target_modules=["query", "value"], # which layers get an adapter
))
model.print_trainable_parameters()
for name, p in model.named_parameters():
if p.requires_grad:
print(f" {name.split('base_model.model.')[-1]:52} {tuple(p.shape)}")trainable params: 8,450 || all params: 4,394,628 || trainable%: 0.1923 bert.encoder.layer.0.attention.self.query.lora_A.default.weight (8, 128) bert.encoder.layer.0.attention.self.query.lora_B.default.weight (128, 8) bert.encoder.layer.0.attention.self.value.lora_A.default.weight (8, 128) bert.encoder.layer.0.attention.self.value.lora_B.default.weight (128, 8) bert.encoder.layer.1.attention.self.query.lora_A.default.weight (8, 128) bert.encoder.layer.1.attention.self.query.lora_B.default.weight (128, 8) bert.encoder.layer.1.attention.self.value.lora_A.default.weight (8, 128) bert.encoder.layer.1.attention.self.value.lora_B.default.weight (128, 8) classifier.modules_to_save.default.weight (2, 128) classifier.modules_to_save.default.bias (2,)
Eight thousand trainable numbers out of four million. Each targeted layer gained a pair: lora_A shaped (8, 128) and lora_B shaped (128, 8). Everything else has requires_grad=False.
Real QLoRA — needs a CUDA GPU
import os, torch
from datasets import Dataset
from transformers import (AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig,
DataCollatorForLanguageModeling, Trainer, TrainingArguments, set_seed)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
set_seed(0)
NAME = "HuggingFaceTB/SmolLM2-135M" # ~272 MB of files
tok = AutoTokenizer.from_pretrained(NAME)
tok.pad_token = tok.eos_token # this model ships without a pad token
qcfg = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained(NAME, quantization_config=qcfg, device_map="cuda:0")
model = prepare_model_for_kbit_training(model)
model = get_peft_model(model, LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"]))
model.print_trainable_parameters()
print("footprint:", round(model.get_memory_footprint() / 1e6), "MB")
pairs = [("Maharashtra", "Mumbai"), ("Karnataka", "Bengaluru"), ("Telangana", "Hyderabad"),
("Kerala", "Thiruvananthapuram"), ("Punjab", "Chandigarh"), ("Gujarat", "Gandhinagar")]
lines = [f"Q: What is the capital of {s}? A: {c}." for s, c in pairs] * 20
ds = Dataset.from_dict({"text": lines}).map(
lambda b: tok(b["text"], truncation=True, max_length=32), batched=True, remove_columns=["text"])
args = TrainingArguments(output_dir="qlora-run", num_train_epochs=3,
per_device_train_batch_size=8, learning_rate=2e-4,
logging_strategy="epoch", save_strategy="no",
report_to=[], seed=0, disable_tqdm=True, bf16=True)
Trainer(model=model, args=args, train_dataset=ds,
data_collator=DataCollatorForLanguageModeling(tok, mlm=False)).train()
model.save_pretrained("smol-capitals-adapter")
files = sorted(os.listdir("smol-capitals-adapter"))
size = sum(os.path.getsize(os.path.join("smol-capitals-adapter", f)) for f in files)
print("adapter files:", files)
print("adapter size :", round(size / 1e6, 1), "MB (base model on disk: ~272 MB)")trainable params: 1,843,200 || all params: 136,358,208 || trainable%: 1.3517
footprint: 174 MB
{'loss': '2.783', 'grad_norm': '2.423', 'learning_rate': '0.0001378', 'epoch': '1'}
{'loss': '1.341', 'grad_norm': '0.8432', 'learning_rate': '7.111e-05', 'epoch': '2'}
{'loss': '0.832', 'grad_norm': '0.671', 'learning_rate': '4.444e-06', 'epoch': '3'}
{'train_runtime': '21.86', 'train_samples_per_second': '16.47', 'train_steps_per_second': '2.058', 'train_loss': '1.652', 'epoch': '3'}
adapter files: ['README.md', 'adapter_config.json', 'adapter_model.safetensors']
adapter size : 7.4 MB (base model on disk: ~272 MB)Timings are from an RTX A6000; losses shift slightly with your bitsandbytes and CUDA build.
Did it learn anything?
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
NAME = "HuggingFaceTB/SmolLM2-135M"
tok = AutoTokenizer.from_pretrained(NAME)
qcfg = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16)
base = AutoModelForCausalLM.from_pretrained(NAME, quantization_config=qcfg, device_map="cuda:0")
model = PeftModel.from_pretrained(base, "smol-capitals-adapter") # injects into `base` in place
ids = tok("Q: What is the capital of Karnataka? A:", return_tensors="pt").to(model.device)
kw = dict(max_new_tokens=10, do_sample=False, pad_token_id=tok.eos_token_id)
with model.disable_adapter():
print("adapter off:", repr(tok.decode(model.generate(**ids, **kw)[0], skip_special_tokens=True)))
print("adapter on :", repr(tok.decode(model.generate(**ids, **kw)[0], skip_special_tokens=True)))adapter off: 'Q: What is the capital of Karnataka? A: It is the capital of Karnataka.\n\nQ' adapter on : 'Q: What is the capital of Karnataka? A: Bangalore. B: Bangalore. C:'
Be honest about what changed. Without the adapter the model dodges the question. With it, the model answers with a city and stops. It also says "Bangalore" where our 120 training lines all said "Bengaluru" — 7.4 MB of adapter learned the format in three epochs, and did not overwrite the base model's strong prior about that city's name. That gap is the honest picture of what a small adapter on tiny data can and cannot do.
The walkthrough
Two numbers explain the memory saving. trainable params: 1,843,200 out of 136 million means gradients and optimizer state exist for 1.4% of the model. Adam keeps roughly 8 bytes per trainable parameter; over the full model that would be about 1 GB, and here it is about 15 MB.
prepare_model_for_kbit_training is not optional. It casts layer-norm and head layers to a stable dtype, enables gradient checkpointing support, and makes the input embeddings pass gradients through the frozen quantised stack. Skip it and you often get a loss that never moves, with no error.
target_modules is the main quality dial after rank. Attention projections only (q_proj, v_proj) is the original recipe and the cheapest. Adding k_proj and o_proj costs little. Including the feed-forward layers (gate_proj, up_proj, down_proj) is measurably better for teaching new behaviour, at more memory. Names differ by architecture — query/value on BERT, q_proj/v_proj on Llama-family models.
lora_alpha is a scale, not a size. The adapter's contribution is multiplied by lora_alpha / r. The common convention is alpha = 2r, which is what r=16, lora_alpha=32 gives. Changing r without changing alpha silently changes the effective learning rate for the adapter.
Adapters want a bigger learning rate than full fine-tuning. 2e-4 here, against 2e-5 for updating every weight. The adapter starts at zero and has to travel; the base model is already where it wants to be.
PeftModel.from_pretrained modifies the base object in place. Keeping a separate reference to compare "before and after" does not work — both names point at the same modules. model.disable_adapter() is the correct comparison, which is why compare.py uses it.
Common mistakes
Expecting a merged model back. save_pretrained on a PEFT model writes the adapter, not the model. To ship one standalone file, load the base in 16-bit (not 4-bit), attach the adapter, and call merge_and_unload(). Merging into quantised weights is lossy and generally refused.
Training on text without a target. The collator with mlm=False scores every token, including your prompt. That works for teaching a style. For instruction tuning you usually want the loss masked to the answer only — TRL's SFTTrainer handles this, and doing it by hand means setting prompt positions to -100 as in data collators.
Believing loss going down means the task is learned. Our loss fell from 2.78 to 0.83 on 120 lines. That is memorisation of a small pattern. Hold out examples the adapter never saw before claiming anything — overfitting and underfitting is the whole story here.
Fine-tuning when retrieval was the answer. If the goal is "know about my documents", RAG is usually cheaper, more current and easier to audit. Fine-tuning teaches format, tone and task shape. It is a poor way to install facts.
Setting rank far too high. r=64 or r=128 on a small dataset burns memory and overfits faster. Start at 8 or 16 and raise it only when a held-out score says to.
Try it yourself
Rerun qlora.py with r=4 and again with r=32, and compare final loss against adapter size. Then add "gate_proj" and "down_proj" to target_modules and see what happens to both. Finally, change one training line's answer to something false and check whether 20 repetitions can overwrite the base model's prior — the boundary you find is the practical limit of small-adapter fine-tuning.
What to learn next
- LoRA — the beginner-level story of the low-rank idea.
- Fine-tuning — when to fine-tune at all, versus prompting or retrieval.
- RLHF — the step after supervised fine-tuning.
Researcher — Mathematics and papers.
LoRA, stated once
For a frozen weight matrix W₀ ∈ R^{d×k}, LoRA (Hu et al., 2021, LoRA: low-rank adaptation of large language models) parameterises the update as ΔW = BA with B ∈ R^{d×r}, A ∈ R^{r×k} and r ≪ min(d, k). The forward pass becomes h = W₀x + (α/r)·BAx, where α is lora_alpha. A is initialised from a random normal and B from zeros, so ΔW = 0 at step 0 and training starts exactly at the base model.
Trainable parameters per adapted matrix drop from dk to r(d+k) — for d = k = 4096 and r = 16, from 16.8M to 131k, a factor of 128. The α/r scaling exists so that changing r does not require retuning the learning rate; it makes the update's magnitude approximately rank-invariant.
Merging is exact: W = W₀ + (α/r)BA is computable once, after which inference has zero added latency. This is LoRA's structural advantage over adapter layers inserted in series (Houlsby et al., 2019), which add depth and therefore latency permanently.
What QLoRA adds
QLoRA (Dettmers et al., 2023) is LoRA over a 4-bit NF4 base, with three specific contributions:
- NF4 storage for the frozen base, described in 4-bit loading. Gradients flow through the dequantised weights to the adapters; the quantised weights themselves receive no update.
- Double quantisation of the block scales, ~0.37 bits per parameter saved.
- Paged optimizers, using NVIDIA unified memory so optimizer states page to CPU on a gradient-checkpointing spike rather than raising out-of-memory.
The memory budget of full fine-tuning is roughly 2 bytes (fp16 weights) + 2 (gradients) + 12 (Adam fp32 master weights and moments) = 16 bytes per parameter, plus activations. QLoRA's is ~0.5 bytes per frozen parameter + 16 bytes per adapter parameter + activations. At r = 16 on attention projections, adapters are near 0.1–1% of the model, so the frozen base dominates — which is why 4-bit storage is the load-bearing part, and why the paper could fine-tune a 65B model on a single 48 GB card.
The paper's central empirical claim is that NF4 + LoRA matches 16-bit full fine-tuning on their benchmark suite. Treat it as strong evidence for instruction tuning on mid-size models, not as a universal law — quantisation error is relatively more damaging on small models, and adapter capacity does bind on tasks far from the pretraining distribution.
The variants worth knowing
DoRA (Liu et al., 2024) decomposes weights into magnitude and direction, adapting each separately, and closes part of the LoRA-to-full-fine-tuning gap at low rank. rsLoRA replaces α/r with α/√r, which stabilises training at high rank. LoRA+ uses different learning rates for A and B on the argument that they play asymmetric roles. QA-LoRA and LoftQ address the merge problem directly — quantisation-aware initialisation so the adapter can be folded back into a quantised model without the accuracy loss that plain merging incurs.
Serving is where the adapter design pays a second time. Because ΔW is separable, a server can hold one base model in memory and thousands of adapters, batching requests for different adapters together (S-LoRA, Sheng et al., 2023). Full fine-tuning gives one model per task and no such sharing — see vLLM for the serving side of this argument.
What to learn next
- LoRA — the beginner-level story of the low-rank idea.
- Fine-tuning — when to fine-tune at all, versus prompting or retrieval.
- RLHF — the step after supervised fine-tuning.