LoRA
LoRA fine-tunes a large model by freezing it and training a small add-on beside it, which cuts the memory cost enormously and lets you swap behaviours like plug-in packs.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
LoRA is a way to fine-tune a big model without changing the big model.
Think about a library book you are not allowed to write in. You still want it marked up for your exam, so you cover it in sticky notes. Your notes sit on top of the page, you read both together, and the book underneath is untouched.
Peel the notes off and the book is exactly as you found it. Hand it to a friend with a different set of notes and it becomes their book.
LoRA does that to a model. The original weights are frozen and never edited. A small set of new numbers is trained beside them, and the two are read together.
LoRA stands for low-rank adaptation — "low-rank" meaning the add-on is deliberately kept small and simple.
Why it exists
Ordinary fine-tuning edits every number in the model. For a model with billions of numbers, three problems follow.
Memory. Training holds the model, a correction for every number, and the optimizer's own bookkeeping for every number. That is many times the memory of the model alone. It is why a model that runs on your laptop still cannot be trained on it.
Storage. A separate full copy per task. Ten tasks, ten copies, each many gigabytes.
Swapping. Serving ten fine-tuned models means loading ten models. Serving one model with ten sticky-note packs means loading one.
LoRA fixes all three at once, which is why it went from a research paper to the standard method very quickly.
How it works
The add-on sits beside the frozen layer, not inside it. Both run, and their answers are added together.
┌────────────────────────┐
input ──► │ frozen original layer │ ──► big answer ──┐
│ └────────────────────────┘ │
│ ├──► final answer
│ ┌────────────────────────┐ │
└───► │ small trained add-on │ ──► small nudge ─┘
└────────────────────────┘
(the sticky note)Only the add-on learns. The original layer never moves.
The add-on is built as two small pieces rather than one big one. Think of squeezing the information through a narrow gate: everything must pass through a handful of channels before spreading back out. That narrow gate is what keeps the add-on tiny, and it is the "low-rank" part of the name.
Training a typical add-on updates well under one percent of the numbers in the model. That is not a rounding error, it is the entire point.
Two things this buys you
It starts as a no-op. The add-on is set up so that on the first step it contributes exactly nothing. The model behaves identically to before training began. A LoRA run therefore starts from a known-good model, and drifts away from it only as far as training takes it.
It can be folded back in. Once trained, the add-on can be merged into the frozen weights to produce a single ordinary model. After merging, answering a question costs no more than it did before. You pay nothing at all at use time.
Where you have already seen it
- Image generators. The style packs people download for Stable Diffusion — a particular art style, a character, a look — are almost all LoRA files. They are a few megabytes against a base model of several gigabytes.
- Assistants tuned per customer. One base model in memory, one small adapter per client, swapped on the fly.
- Small teams fine-tuning at all. Most fine-tuning done outside large labs is LoRA, because the alternative needs hardware they do not have.
When NOT to use LoRA
Everything in the fine-tuning lesson still applies, and it applies harder because LoRA is cheap enough to reach for thoughtlessly.
It does not add facts. Being cheaper does not change what fine-tuning is bad at. If the model is missing information, you want RAG.
It does not rescue a vague goal. If you cannot write twenty examples of the output you want, you are not ready to train anything.
It is slightly weaker than full fine-tuning on hard target tasks, because the add-on is deliberately limited. In exchange it damages the model's other abilities far less. For most people that trade is worth taking, but it is a trade.
Remember this
- LoRA freezes the big model and trains a small add-on beside it.
- It cuts training memory and storage enormously, and adapters can be swapped or merged.
- It is still fine-tuning. Every reason not to fine-tune is still a reason not to use LoRA.
What to learn next
- Fine-tuning — the decision that comes before choosing a method.
- Hugging Face — where base models and adapters are published.
- What is RAG? — the alternative when the problem is missing knowledge.
Developer — Code and libraries.
Setup
pip install torchLoRA sounds complicated and is about fifteen lines of code. Building it once from scratch is worth more than reading any library's documentation, so that is what this does. No downloads, CPU only, runs in a few seconds.
LoRA from scratch
import torch
import torch.nn as nn
torch.manual_seed(0)
DIM, RANK, ALPHA = 256, 4, 8
class LoRALinear(nn.Module):
"""A frozen Linear layer with a small trainable side path bolted on."""
def __init__(self, base: nn.Linear, rank: int, alpha: int):
super().__init__()
self.base = base
for p in self.base.parameters():
p.requires_grad = False # the big layer never moves again
self.A = nn.Parameter(torch.randn(rank, base.in_features) * 0.01)
self.B = nn.Parameter(torch.zeros(base.out_features, rank)) # zero, so training starts as a no-op
self.scale = alpha / rank
def forward(self, x):
return self.base(x) + (x @ self.A.T @ self.B.T) * self.scale
base = nn.Linear(DIM, DIM)
layer = LoRALinear(base, RANK, ALPHA)
total = sum(p.numel() for p in layer.parameters())
trainable = sum(p.numel() for p in layer.parameters() if p.requires_grad)
print(f"weights in the frozen layer : {total - trainable}")
print(f"weights we actually train : {trainable} ({100 * trainable / total:.1f}%)")
# The job: recover a change to the base layer that happens to be low rank.
with torch.no_grad():
u = torch.randn(DIM, 2)
v = torch.randn(2, DIM)
W_true = base.weight + (u @ v) * 0.05 # a rank-2 edit to the frozen weights
X = torch.randn(64, DIM)
Y = X @ W_true.T + base.bias
before = base.weight.clone()
opt = torch.optim.Adam([p for p in layer.parameters() if p.requires_grad], lr=0.01)
lossfn = nn.MSELoss()
print()
print("adapter output identical to base at start?", torch.equal(layer(X), base(X)))
print(f"loss before any training : {lossfn(layer(X), Y).item():.6f}")
for step in range(1, 301):
opt.zero_grad()
loss = lossfn(layer(X), Y)
loss.backward()
opt.step()
if step in (50, 150, 300):
print(f"step {step:3d} loss : {loss.item():.6f}")
print()
print("frozen weights changed? ", not torch.equal(before, base.weight))
# Merging: fold the side path into the frozen weights, so inference costs nothing extra.
with torch.no_grad():
merged = nn.Linear(DIM, DIM)
merged.weight.copy_(base.weight + (layer.B @ layer.A) * layer.scale)
merged.bias.copy_(base.bias)
gap = (merged(X) - layer(X)).abs().max().item()
print(f"largest gap after merging: {gap:.2e}")weights in the frozen layer : 65792 weights we actually train : 2048 (3.0%) adapter output identical to base at start? True loss before any training : 1.166372 step 50 loss : 0.003430 step 150 loss : 0.000004 step 300 loss : 0.000001 frozen weights changed? False largest gap after merging: 3.81e-06
Every claim in the lesson, verified by that output
3.0% trainable. One 256 x 256 layer holds 65,792 weights. A rank-4 adapter holds 2,048. The saving grows with layer size, so on a real model the fraction is far smaller still.
Identical at the start. B is initialised to zeros, so the side path contributes exactly zero on step one. torch.equal confirms it bit for bit. This is why a LoRA run never starts from a damaged model.
A is random, B is zero — never both zero. If both were zero the gradient of each would also be zero and nothing would ever move. Exactly one of the two must be non-zero at the start.
The frozen weights never changed. requires_grad = False means autograd never fills their .grad, and the optimizer was handed only the trainable parameters. Both guards matter: passing model.parameters() to the optimizer while relying on requires_grad alone works, but it is easy to break later.
Merging is exact. The gap is 3.81e-06, which is float32 rounding, not a real difference. After merging you have a plain nn.Linear with no extra cost at inference.
The task was chosen honestly. The change to be learned was built as a rank-2 edit, and a rank-4 adapter recovers it. That is LoRA's core assumption made visible: it works when the change you need is simple in shape. Build W_true with a full-rank random edit instead and the loss will stall far above zero. Try it.
Rank and alpha, in practice
rank is the width of the narrow gate — how much the adapter can express. alpha scales the adapter's contribution; the effective strength is alpha / rank.
| Situation | Rank | Alpha |
|---|---|---|
| Style, tone, output format | 4 – 8 | 8 – 16 |
| A specific task, moderate data | 16 – 32 | 16 – 64 |
| New domain, large dataset | 64 – 128 | 64 – 128 |
Two habits keep you out of trouble. Change one of the two at a time, because raising both raises the effective strength twice over. And treat a high rank as a warning, not a fix: if rank 128 is needed, the honest question is whether LoRA is the right tool for this problem at all.
Using the real library
In practice you use Hugging Face peft, which wraps every matching layer for you.
from peft import LoraConfig, get_peft_model
config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"], # which layers get an adapter
task_type="CAUSAL_LM",
)
model = get_peft_model(base_model, config)
model.print_trainable_parameters()No output block here on purpose. The printed counts depend entirely on which base model you loaded, and an invented number would teach you to expect something that will not appear.
target_modules is the decision worth thinking about. The original paper adapted only the attention query and value projections. Later practice found that adapting all linear layers, including the feed-forward ones, works better for most tasks — that is what QLoRA recommends and what most recipes now do.
QLoRA, in one paragraph
QLoRA loads the frozen base model in 4-bit precision and trains the LoRA adapter in 16-bit on top of it. The frozen weights are never updated, so their reduced precision costs much less accuracy than it would during full training. This is what puts fine-tuning a large model onto a single consumer graphics card. It is slower per step than plain LoRA because weights are unpacked on the fly, and it is the reason most people can fine-tune at all.
Common mistakes
Forgetting to freeze the base. Without requires_grad = False, you are doing full fine-tuning with an extra side path attached, and you get none of the memory saving. Print your trainable count every run; that one line catches it immediately.
Initialising both A and B to zero. The adapter is then stuck at zero forever, and the loss sits flat while everything looks correctly configured.
Merging adapters trained on different base weights. An adapter is a change relative to one specific set of frozen weights. Applying it to a different base, or to a base loaded at a different quantisation, gives a result nobody can predict.
Expecting a memory saving on activations. LoRA removes optimizer state and gradients for the frozen weights. It does not remove the stored activations of the forward pass, which still scale with batch size and sequence length. Long sequences will still exhaust memory.
Raising rank when the loss is not falling. Check the learning rate first. LoRA adapters generally want a much higher learning rate than full fine-tuning — often around 1e-4 to 3e-4 against 1e-5 for full fine-tuning.
Try it yourself
In lora_from_scratch.py, replace the rank-2 construction of W_true with a full-rank edit:
W_true = base.weight + torch.randn(DIM, DIM) * 0.05Rerun. The loss will fall a little and then stop, far above zero. Then raise RANK to 64 and watch it fall further. You have measured the exact assumption LoRA rests on, and found where it breaks.
What to learn next
- Fine-tuning — the decision that comes before choosing a method.
- Hugging Face — where base models and adapters are published.
- What is RAG? — the alternative when the problem is missing knowledge.
Researcher — Mathematics and papers.
The parameterisation
For a frozen pretrained weight W₀ ∈ R^{d_out × d_in}, LoRA constrains the update to a low-rank factorisation:
h = W₀ x + ΔW x , ΔW = (α / r) · B A
A ∈ R^{r × d_in} initialised A ~ N(0, σ²)
B ∈ R^{d_out × r} initialised B = 0
r ≪ min(d_in, d_out)r is the rank, α a fixed scaling constant, and α/r the effective magnitude of the update. B = 0 gives ΔW = 0 at initialisation, so the adapted model is exactly the pretrained model at step zero. Exactly one factor may be zero-initialised, or the gradient of both vanishes identically.
Trainable parameters per adapted matrix fall from d_in · d_out to r(d_in + d_out). For d_in = d_out = 4096 and r = 16 that is 16.8M down to 131k, a factor of 128.
What is actually saved
The saving is in optimizer state and gradients, not in activations.
full fine-tune, Adam, mixed precision ≈ 16 bytes per parameter
LoRA ≈ 2 bytes per frozen parameter (weights only)
+ 16 bytes per adapter parameterActivation memory for the backward pass is unchanged: gradients must still flow through the frozen layers to reach adapters in earlier layers, so their inputs are still stored. Batch size and sequence length remain the binding constraints on long-context fine-tuning.
Hu et al. (2021) report 10,000× fewer trainable parameters and roughly 3× less GPU memory for GPT-3 175B compared with full fine-tuning under Adam, at equal or better quality on their benchmarks. Checkpoints shrink by the same factor as the parameter count, which is what makes per-task adapters practical to store and ship.
Why low rank is a reasonable prior
Aghajanyan, Zettlemoyer and Gupta (2020), Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, measured the smallest random subspace in which fine-tuning still reaches 90% of full performance. For many tasks it is in the hundreds or low thousands of dimensions, against hundreds of millions of parameters. Larger pretrained models had lower intrinsic dimension, which is the empirical result LoRA is built on.
The assumption is a prior, not a theorem. It holds well for adaptation to a task the base model can nearly already do, and degrades when the target genuinely requires new capability.
Variants that matter
QLoRA (Dettmers et al., 2023). Three contributions, all load-bearing: NF4, a 4-bit data type that is information-theoretically optimal for normally distributed weights; double quantisation, quantising the quantisation constants themselves for roughly 0.37 bits per parameter more; and paged optimizers, using unified memory to survive gradient-checkpointing memory spikes. Enabled 65B fine-tuning on a single 48GB card with 16-bit full-fine-tuning quality on their evaluations.
rsLoRA (Kalajdzievski, 2023). Scaling by α/r causes gradients to shrink as r grows, so high-rank adapters underperform their capacity. Scaling by α/√r fixes the collapse and makes rank behave monotonically.
DoRA (Liu et al., 2024). Decomposes the weight into magnitude and direction, applying LoRA only to the direction. Closes much of the remaining gap to full fine-tuning at low rank, for a small extra cost.
LoRA+ (Hayou et al., 2024). A and B have different roles under the zero-initialisation, and a single learning rate is suboptimal. Setting B's learning rate several times higher than A's improves convergence at no cost.
Serving many adapters
Because ΔW = BA is separable, a single base model in memory can serve many adapters concurrently by batching their side paths. S-LoRA (Sheng et al., 2023) and Punica demonstrate thousands of adapters on one GPU with unified paging for the adapter weights. Merging is the opposite choice: fold ΔW into W₀ for zero inference overhead, at the cost of losing the ability to switch.
Merging is exact only in the precision you merge at. Merging a bf16 adapter into a 4-bit base requires dequantising, merging, and requantising, which is lossy and a real source of quality regressions between training and deployment.
Known limitations
Biderman et al. (2024) trained matched LoRA and full fine-tuning runs on code and mathematics. LoRA underperformed on target-domain gains, particularly on continued pretraining with large token budgets, and simultaneously forgot less of the base model's general ability. They also found LoRA acts as a stronger regulariser than weight decay or dropout, and that full fine-tuning learns perturbations of far higher rank than typical LoRA settings — 10 to 100 times higher.
Shuttleworth et al. (2024) identify intruder dimensions: singular vectors in the adapted weight matrix that are near-orthogonal to anything in the pretrained spectrum. Full fine-tuning does not produce them. Models with intruder dimensions match on the target task but generalise worse out of distribution and forget more of the pretraining distribution, and higher rank with rank stabilisation reduces the effect.
The practical reading: LoRA is the right default for task adaptation on a fixed budget, and the wrong tool for teaching a model a genuinely new domain at scale.
Papers
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, 2021 — arxiv.org/abs/2106.09685
- Aghajanyan et al., Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, 2020 — arxiv.org/abs/2012.13255
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, 2023 — arxiv.org/abs/2305.14314
- Kalajdzievski, A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA, 2023 — arxiv.org/abs/2312.03732
- Liu et al., DoRA: Weight-Decomposed Low-Rank Adaptation, 2024 — arxiv.org/abs/2402.09353
- Hayou et al., LoRA+: Efficient Low Rank Adaptation of Large Models, 2024 — arxiv.org/abs/2402.12354
- Biderman et al., LoRA Learns Less and Forgets Less, 2024 — arxiv.org/abs/2405.09673
- Sheng et al., S-LoRA: Serving Thousands of Concurrent LoRA Adapters, 2023 — arxiv.org/abs/2311.03285
What to learn next
- Fine-tuning — the decision that comes before choosing a method.
- Hugging Face — where base models and adapters are published.
- What is RAG? — the alternative when the problem is missing knowledge.