Adapters beyond LoRA
LoRA is one of a family of tricks that train a tiny number of extra parameters instead of the whole model, and the others trade size against expressiveness differently.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Instead of retraining a whole model, you freeze it and train a very small add-on.
The analogy you have already lived
You buy a ready-made shirt and it fits everywhere except the waist. You do not have a new shirt stitched. You take it to a tailor, who takes in the sides.
Two seams. Ten minutes. The shirt is now yours, and the original fabric is untouched.
That is what these methods do to a model. The big frozen weights are the shirt. The adapter is the alteration.
Why it exists
Fine-tuning a whole model means three painful things.
Memory. Training a 7-billion-parameter model needs room for the weights, their gradients, and the optimiser's bookkeeping. That is many times the size of the model itself.
Storage. Every fine-tuned copy is another full-size model on disk. Ten customers, ten copies.
Risk. Changing every weight can damage abilities you were not trying to change.
Adapters fix all three. You train a few million numbers instead of billions. Each customer's version is a small file. The original model is untouched, so nothing you did not intend gets broken.
How it works
input
|
v
[ frozen weights ] + [ tiny trainable add-on ]
| |
+--------------> sum <---------+
|
v
outputThe add-on starts at zero, so on day one the model behaves exactly as before. Training moves only the add-on.
The family, in plain terms
Low-rank matrices. Two thin rectangles that multiply into a full-size correction. This is LoRA, and it is the default.
Scaling vectors. One number per channel, multiplied into the activations. Absurdly small. Sometimes enough.
Shared random matrices. Everyone uses the same frozen random rectangles, and each layer learns only two short vectors to reweight them. Smaller still.
Rotations. Instead of adding a correction, rotate the existing weights. Preserves more of the original behaviour.
Extra tokens. Do not touch the weights at all. Learn a few invisible tokens that get prepended to every input.
What is honestly hard here
There is a real trade-off and nobody has removed it.
Adapters are excellent when you are teaching a style, a format or a task the model already half knows. They are weaker when you are trying to install a genuinely new skill or a large body of new knowledge.
Careful measurements found two things at once. Adapters learn less than full fine-tuning on hard new material. They also forget less of what the model knew. Which half matters depends on your job.
Where you have already seen this
- A tailor altering a ready-made shirt.
- A phone case, which changes how the phone behaves without opening it.
- Photo filters, which sit on top of the original image file.
Remember this
- Freeze the big model, train a tiny add-on, keep the original intact.
- The add-on starts as a no-op, so nothing changes until you train it.
- Adapters learn less than full fine-tuning, and forget less.
What to learn next
- Catastrophic forgetting — the damage adapters are partly protecting you from.
- LoRA — the method this lesson builds on, in full.
- Merging model weights — combining several of these into one model.
Developer — Code and libraries.
Setup
pip install torch peftRuns on a CPU in a few seconds. Written against peft 0.20.0 and PyTorch 2.5.1.
Nine methods on the same model
import torch
import torch.nn as nn
from peft import (LoraConfig, IA3Config, VeraConfig, LoHaConfig, LoKrConfig,
OFTConfig, get_peft_model)
from collections import OrderedDict
def base():
torch.manual_seed(0)
layers = OrderedDict()
for i in range(4):
layers[f"lin{i}"] = nn.Linear(256, 256)
layers[f"act{i}"] = nn.ReLU()
return nn.Sequential(layers)
full = sum(p.numel() for p in base().parameters())
TARGETS = [f"lin{i}" for i in range(4)]
METHODS = {
"LoRA r=8": LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS),
"LoRA r=1": LoraConfig(r=1, lora_alpha=2, target_modules=TARGETS),
"DoRA r=8": LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS, use_dora=True),
"rsLoRA r=8": LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS,
use_rslora=True),
"VeRA r=8": VeraConfig(r=8, target_modules=TARGETS),
"IA3": IA3Config(target_modules=TARGETS, feedforward_modules=[]),
"LoHa r=8": LoHaConfig(r=8, target_modules=TARGETS),
"LoKr r=8": LoKrConfig(r=8, target_modules=TARGETS),
"OFT b=32": OFTConfig(oft_block_size=32, target_modules=TARGETS),
}
print(f"base model: {full:,} parameters, 4 Linear layers of 256x256\n")
print(f"{'method':<14} {'trainable':>10} {'% of base':>10} what it adds")
NOTES = {
"LoRA r=8": "two thin matrices per layer, B@A",
"LoRA r=1": "the same, at the smallest possible rank",
"DoRA r=8": "LoRA plus a learned per-column magnitude",
"rsLoRA r=8": "LoRA with alpha/sqrt(r) scaling",
"VeRA r=8": "frozen shared random A,B; learns two vectors",
"IA3": "one scaling vector per targeted activation",
"LoHa r=8": "Hadamard product of two low-rank pairs",
"LoKr r=8": "Kronecker product factorisation",
"OFT b=32": "a block-diagonal orthogonal rotation",
}
for name, cfg in METHODS.items():
m = get_peft_model(base(), cfg)
t = sum(p.numel() for p in m.parameters() if p.requires_grad)
print(f"{name:<14} {t:>10,} {100 * t / full:>9.3f}% {NOTES[name]}")
# What LoRA actually inserts, and the fact that it starts as a no-op
torch.manual_seed(0)
m = get_peft_model(base(), LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS))
layer = m.base_model.model.lin0
print(f"\ninside one adapted layer:")
print(f" frozen weight {tuple(layer.base_layer.weight.shape)}"
f" requires_grad={layer.base_layer.weight.requires_grad}")
print(f" lora_A {tuple(layer.lora_A['default'].weight.shape)}"
f" requires_grad={layer.lora_A['default'].weight.requires_grad}")
print(f" lora_B {tuple(layer.lora_B['default'].weight.shape)}"
f" requires_grad={layer.lora_B['default'].weight.requires_grad}")
print(f" B is initialised to all zeros: {bool((layer.lora_B['default'].weight == 0).all())}")
x = torch.randn(2, 256)
torch.manual_seed(0)
plain = base()
print(f" adapted output equals base output at step 0: "
f"{torch.allclose(m(x), plain(x), atol=1e-6)}")base model: 263,168 parameters, 4 Linear layers of 256x256 method trainable % of base what it adds LoRA r=8 16,384 6.226% two thin matrices per layer, B@A LoRA r=1 2,048 0.778% the same, at the smallest possible rank DoRA r=8 17,408 6.615% LoRA plus a learned per-column magnitude rsLoRA r=8 16,384 6.226% LoRA with alpha/sqrt(r) scaling VeRA r=8 1,056 0.401% frozen shared random A,B; learns two vectors IA3 1,024 0.389% one scaling vector per targeted activation LoHa r=8 32,768 12.451% Hadamard product of two low-rank pairs LoKr r=8 2,048 0.778% Kronecker product factorisation OFT b=32 15,872 6.031% a block-diagonal orthogonal rotation inside one adapted layer: frozen weight (256, 256) requires_grad=False lora_A (8, 256) requires_grad=True lora_B (256, 8) requires_grad=True B is initialised to all zeros: True adapted output equals base output at step 0: True
Reading the table
The percentages look large because the base model is tiny. 256×256 layers make LoRA r=8 look like 6%. On a real 4096-wide model the same r=8 is well under 0.1%, because LoRA's cost grows as 2·r·d while the frozen weight grows as d². Always compute this ratio for your dimensions before choosing a rank.
B starts at exactly zero, so the adapter starts as a no-op. The last line confirms it numerically: the adapted model's output is bit-identical to the base model's before training. That property is what makes adapters safe to attach to a production model — attaching one changes nothing until you train it.
IA3 is the smallest thing that still does something. 1,024 numbers, one scaling factor per output channel per layer. That is 0.4% of a tiny model and effectively nothing on a large one.
VeRA is smaller than LoRA r=1. It shares one frozen random A and B across all layers and learns only two per-layer vectors. The rank is still 8; the learned content is two vectors.
LoHa costs twice LoRA at the same r. It multiplies two low-rank pairs elementwise, which raises the effective rank beyond r at the price of two pairs of factors. "Rank 8" does not mean the same thing across these methods, so compare by trainable-parameter count, never by r.
DoRA is LoRA plus 1,024 numbers. One magnitude scalar per output column. That is the entire difference, and it is reported to close much of the gap to full fine-tuning at low rank.
Choosing, in practice
| Situation | Reasonable first choice |
|---|---|
| Standard task or style adaptation | LoRA, r=8–32, on all linear layers |
| Very low rank required | DoRA (recovers quality LoRA loses at r=4) |
| Serving hundreds of variants | VeRA or IA3 — adapter files of a few kilobytes |
| GPU memory is the binding limit | QLoRA: 4-bit frozen base plus LoRA |
| Preserving base behaviour matters most | OFT or BOFT — a rotation cannot change norms |
| Prompt-style steering only | prompt tuning or prefix tuning |
Two settings that matter more than the method:
Target all the linear layers, not only Q and V. The original LoRA paper adapted attention projections; later practice adapts the MLP layers as well, and it consistently helps.
Use a higher learning rate than full fine-tuning. Roughly 1e-4 for adapters against 1e-5 to 2e-5 for full fine-tuning, because only new parameters are being learned.
Common mistakes
Comparing methods by r. As the table shows, r means different things. Compare by trainable parameters and by measured quality.
Merging a QLoRA adapter into a 4-bit base. merge_and_unload() on a quantised base dequantises and reintroduces quantisation error. Merge into the full-precision base instead, then requantise.
Forgetting to save the adapter config. An adapter file without its adapter_config.json cannot be reattached. Both are needed.
Setting lora_alpha and then changing r. Effective scaling is alpha/r by default, so changing r changes the update magnitude. use_rslora=True uses alpha/sqrt(r), which makes higher ranks behave more predictably.
Expecting an adapter to teach a new language or domain from scratch. That is the case where full fine-tuning or continued pretraining wins. Adapters are for shaping, not for installing.
Try it yourself
Add LoraConfig(r=64, lora_alpha=128, target_modules=TARGETS) to METHODS. Compute where its parameter count crosses the base model's, then work out the width at which r=64 would still be under 1% of a real model.
What to learn next
- Catastrophic forgetting — the damage adapters are partly protecting you from.
- LoRA — the method this lesson builds on, in full.
- Merging model weights — combining several of these into one model.
Researcher — Mathematics and papers.
The low-rank hypothesis and its descendants
LoRA (Hu et al., 2022) reparameterises the weight update as
$$ W = W_0 + \Delta W = W_0 + \frac{\alpha}{r} BA, \qquad B \in \mathbb{R}^{d\times r},\ A \in \mathbb{R}^{r\times k} $$
with $W_0$ frozen, $B$ initialised to zero and $A$ to a random Gaussian. The claim is that the update has low intrinsic rank even when $W_0$ does not — supported by Aghajanyan et al., 2021 (Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning), who showed a few hundred to a few thousand parameters in a random subspace can reach 90% of full fine-tuning performance on GLUE.
rsLoRA (Kalajdzievski, 2023) points out the $\alpha/r$ scaling makes gradients vanish as $r$ grows, so higher ranks stop helping for the wrong reason. Replacing it with $\alpha/\sqrt r$ fixes the scaling and restores monotonic improvement with rank.
DoRA (Liu et al., 2024) decomposes the weight into magnitude and direction:
$$ W = m \cdot \frac{W_0 + BA}{\lVert W_0 + BA \rVert_c} $$
where $m \in \mathbb{R}^{1\times k}$ is a learned per-column magnitude and $\lVert\cdot\rVert_c$ is the column-wise norm. Their analysis shows full fine-tuning changes magnitude and direction largely independently, while LoRA couples them; decoupling closes most of the low-rank gap. The cost is $k$ extra parameters per layer — the 1,024 in the table.
LoRA+ (Hayou et al., 2024) shows the optimal learning rates for $A$ and $B$ differ by a factor that grows with width, and that setting $\eta_B = \lambda \eta_A$ with $\lambda \gg 1$ improves both speed and final quality. This is a one-line change most practitioners never make.
Smaller than low-rank
IA3 (Liu et al., 2022, Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning) learns three vectors per block, rescaling keys, values and the feed-forward intermediate:
$$ \text{attn} = \mathrm{softmax}!\left(\frac{Q(l_k \odot K)^\top}{\sqrt{d}}\right)(l_v \odot V), \qquad \text{FFN} = (l_{ff}\odot \gamma(W_1 x))W_2 $$
It uses roughly 0.01% of parameters and, in their few-shot setting, beat in-context learning at a fraction of the inference cost.
VeRA (Kopiczko et al., 2024) freezes a single pair of random matrices $A, B$ shared across all layers and learns only two scaling vectors $b, d$ per layer:
$$ \Delta W = \Lambda_b B \Lambda_d A $$
Parameter count becomes $O(d + r)$ per layer instead of $O(r(d+k))$, reported at roughly 10× smaller than LoRA at comparable quality on GLUE and E2E. The relevant use case is serving thousands of adapters.
BitFit (Ben Zaken et al., 2022) trains only bias terms — about 0.09% of parameters — and is competitive on small-data GLUE tasks. It is the strongest evidence that a large part of task adaptation is a shift rather than a rotation.
Structure-preserving methods
OFT and BOFT (Qiu et al., 2023; Liu et al., 2024) multiply rather than add:
$$ W = R\,W_0, \qquad R\ \text{orthogonal, block-diagonal} $$
Because $R$ is orthogonal, the singular values of $W_0$ are preserved exactly — the transformation is a rotation, hence norm-preserving and information-preserving in a precise sense. This is why OFT was first adopted in text-to-image work, where additive adapters visibly damage the base model's generative quality. BOFT adds a butterfly factorisation to raise expressiveness at fixed parameter count.
Prompt-space methods
Prefix tuning (Li and Liang, 2021) prepends trainable key–value vectors to every attention layer. Prompt tuning (Lester et al., 2021) prepends trainable embeddings to the input only, and shows the gap to full fine-tuning closes as model size grows, vanishing around 10B parameters. P-tuning v2 restores the per-layer prefixes for smaller models.
These consume context length, which is their practical drawback: a 100-token soft prompt costs 100 tokens of window on every request forever.
What adapters cost in quality
Biderman et al., 2024 (LoRA Learns Less and Forgets Less) is the most careful public comparison, on continued pretraining and instruction tuning in code and maths. Their findings:
- On continued pretraining with 10B–20B tokens of new domain data, LoRA falls substantially short of full fine-tuning. The gap is largest for code.
- On instruction tuning with smaller datasets, LoRA is close to full fine-tuning.
- LoRA forgets less — it better preserves performance outside the target domain — and acts as a stronger regulariser than classic weight decay or dropout.
- Full fine-tuning finds updates whose rank is 10–100× higher than typical LoRA configurations, which explains the first two points directly.
The practical rule follows from the ranking: adapters for shaping behaviour on modest data, full fine-tuning or continued pretraining for installing large new knowledge.
Serving many adapters
The property that made LoRA an infrastructure decision rather than a training trick: $B A$ can be computed as a separate small matmul on a shared frozen base, so many adapters can be served in one batch. S-LoRA (Sheng et al., 2024) and Punica implement this with custom kernels, serving thousands of adapters on a single GPU with the base weights loaded once. Merging is the alternative and gives up multi-tenancy.
Papers
- Houlsby et al., Parameter-Efficient Transfer Learning for NLP, ICML 2019 — arxiv.org/abs/1902.00751
- Li and Liang, Prefix-Tuning, ACL 2021 — arxiv.org/abs/2101.00190
- Lester et al., The Power of Scale for Parameter-Efficient Prompt Tuning, EMNLP 2021 — arxiv.org/abs/2104.08691
- Aghajanyan et al., Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, ACL 2021 — arxiv.org/abs/2012.13255
- Ben Zaken et al., BitFit, ACL 2022 — arxiv.org/abs/2106.10199
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, ICLR 2022 — arxiv.org/abs/2106.09685
- Liu et al., Few-Shot PEFT is Better and Cheaper than In-Context Learning (IA3), NeurIPS 2022 — arxiv.org/abs/2205.05638
- Dettmers et al., QLoRA, NeurIPS 2023 — arxiv.org/abs/2305.14314
- Qiu et al., Controlling Text-to-Image Diffusion by Orthogonal Finetuning (OFT), NeurIPS 2023 — arxiv.org/abs/2306.07280
- Kalajdzievski, A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA, 2023 — arxiv.org/abs/2312.03732
- Kopiczko et al., VeRA: Vector-based Random Matrix Adaptation, ICLR 2024 — arxiv.org/abs/2310.11454
- Liu et al., DoRA: Weight-Decomposed Low-Rank Adaptation, ICML 2024 — arxiv.org/abs/2402.09353
- Hayou et al., LoRA+, ICML 2024 — arxiv.org/abs/2402.12354
- Biderman et al., LoRA Learns Less and Forgets Less, TMLR 2024 — arxiv.org/abs/2405.09673
- Sheng et al., S-LoRA: Serving Thousands of Concurrent LoRA Adapters, MLSys 2024 — arxiv.org/abs/2311.03285
What to learn next
- Catastrophic forgetting — the damage adapters are partly protecting you from.
- LoRA — the method this lesson builds on, in full.
- Merging model weights — combining several of these into one model.