Post-training and Alignment

Adapters beyond LoRA

LoRA is one of a family of tricks that train a tiny number of extra parameters instead of the whole model, and the others trade size against expressiveness differently.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The family, in plain terms
  6. What is honestly hard here
  7. Where you have already seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Instead of retraining a whole model, you freeze it and train a very small add-on.

The analogy you have already lived

You buy a ready-made shirt and it fits everywhere except the waist. You do not have a new shirt stitched. You take it to a tailor, who takes in the sides.

Two seams. Ten minutes. The shirt is now yours, and the original fabric is untouched.

That is what these methods do to a model. The big frozen weights are the shirt. The adapter is the alteration.

Why it exists

Fine-tuning a whole model means three painful things.

Memory. Training a 7-billion-parameter model needs room for the weights, their gradients, and the optimiser's bookkeeping. That is many times the size of the model itself.

Storage. Every fine-tuned copy is another full-size model on disk. Ten customers, ten copies.

Risk. Changing every weight can damage abilities you were not trying to change.

Adapters fix all three. You train a few million numbers instead of billions. Each customer's version is a small file. The original model is untouched, so nothing you did not intend gets broken.

How it works

   input
     |
     v
   [ frozen weights ]  +  [ tiny trainable add-on ]
     |                              |
     +--------------> sum <---------+
                       |
                       v
                    output

The add-on starts at zero, so on day one the model behaves exactly as before. Training moves only the add-on.

The family, in plain terms

Low-rank matrices. Two thin rectangles that multiply into a full-size correction. This is LoRA, and it is the default.

Scaling vectors. One number per channel, multiplied into the activations. Absurdly small. Sometimes enough.

Shared random matrices. Everyone uses the same frozen random rectangles, and each layer learns only two short vectors to reweight them. Smaller still.

Rotations. Instead of adding a correction, rotate the existing weights. Preserves more of the original behaviour.

Extra tokens. Do not touch the weights at all. Learn a few invisible tokens that get prepended to every input.

What is honestly hard here

There is a real trade-off and nobody has removed it.

Adapters are excellent when you are teaching a style, a format or a task the model already half knows. They are weaker when you are trying to install a genuinely new skill or a large body of new knowledge.

Careful measurements found two things at once. Adapters learn less than full fine-tuning on hard new material. They also forget less of what the model knew. Which half matters depends on your job.

Where you have already seen this

  • A tailor altering a ready-made shirt.
  • A phone case, which changes how the phone behaves without opening it.
  • Photo filters, which sit on top of the original image file.

Remember this

  • Freeze the big model, train a tiny add-on, keep the original intact.
  • The add-on starts as a no-op, so nothing changes until you train it.
  • Adapters learn less than full fine-tuning, and forget less.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install torch peft

Runs on a CPU in a few seconds. Written against peft 0.20.0 and PyTorch 2.5.1.

Nine methods on the same model

peft_zoo.py
import torch
import torch.nn as nn
from peft import (LoraConfig, IA3Config, VeraConfig, LoHaConfig, LoKrConfig,
                  OFTConfig, get_peft_model)

from collections import OrderedDict


def base():
    torch.manual_seed(0)
    layers = OrderedDict()
    for i in range(4):
        layers[f"lin{i}"] = nn.Linear(256, 256)
        layers[f"act{i}"] = nn.ReLU()
    return nn.Sequential(layers)


full = sum(p.numel() for p in base().parameters())
TARGETS = [f"lin{i}" for i in range(4)]

METHODS = {
    "LoRA r=8":      LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS),
    "LoRA r=1":      LoraConfig(r=1, lora_alpha=2, target_modules=TARGETS),
    "DoRA r=8":      LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS, use_dora=True),
    "rsLoRA r=8":    LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS,
                                use_rslora=True),
    "VeRA r=8":      VeraConfig(r=8, target_modules=TARGETS),
    "IA3":           IA3Config(target_modules=TARGETS, feedforward_modules=[]),
    "LoHa r=8":      LoHaConfig(r=8, target_modules=TARGETS),
    "LoKr r=8":      LoKrConfig(r=8, target_modules=TARGETS),
    "OFT b=32":      OFTConfig(oft_block_size=32, target_modules=TARGETS),
}

print(f"base model: {full:,} parameters, 4 Linear layers of 256x256\n")
print(f"{'method':<14} {'trainable':>10} {'% of base':>10}  what it adds")
NOTES = {
    "LoRA r=8": "two thin matrices per layer, B@A",
    "LoRA r=1": "the same, at the smallest possible rank",
    "DoRA r=8": "LoRA plus a learned per-column magnitude",
    "rsLoRA r=8": "LoRA with alpha/sqrt(r) scaling",
    "VeRA r=8": "frozen shared random A,B; learns two vectors",
    "IA3": "one scaling vector per targeted activation",
    "LoHa r=8": "Hadamard product of two low-rank pairs",
    "LoKr r=8": "Kronecker product factorisation",
    "OFT b=32": "a block-diagonal orthogonal rotation",
}
for name, cfg in METHODS.items():
    m = get_peft_model(base(), cfg)
    t = sum(p.numel() for p in m.parameters() if p.requires_grad)
    print(f"{name:<14} {t:>10,} {100 * t / full:>9.3f}%  {NOTES[name]}")

# What LoRA actually inserts, and the fact that it starts as a no-op
torch.manual_seed(0)
m = get_peft_model(base(), LoraConfig(r=8, lora_alpha=16, target_modules=TARGETS))
layer = m.base_model.model.lin0
print(f"\ninside one adapted layer:")
print(f"  frozen weight      {tuple(layer.base_layer.weight.shape)}"
      f"  requires_grad={layer.base_layer.weight.requires_grad}")
print(f"  lora_A             {tuple(layer.lora_A['default'].weight.shape)}"
      f"  requires_grad={layer.lora_A['default'].weight.requires_grad}")
print(f"  lora_B             {tuple(layer.lora_B['default'].weight.shape)}"
      f"  requires_grad={layer.lora_B['default'].weight.requires_grad}")
print(f"  B is initialised to all zeros: {bool((layer.lora_B['default'].weight == 0).all())}")

x = torch.randn(2, 256)
torch.manual_seed(0)
plain = base()
print(f"  adapted output equals base output at step 0: "
      f"{torch.allclose(m(x), plain(x), atol=1e-6)}")
Output
base model: 263,168 parameters, 4 Linear layers of 256x256

method          trainable  % of base  what it adds
LoRA r=8           16,384     6.226%  two thin matrices per layer, B@A
LoRA r=1            2,048     0.778%  the same, at the smallest possible rank
DoRA r=8           17,408     6.615%  LoRA plus a learned per-column magnitude
rsLoRA r=8         16,384     6.226%  LoRA with alpha/sqrt(r) scaling
VeRA r=8            1,056     0.401%  frozen shared random A,B; learns two vectors
IA3                 1,024     0.389%  one scaling vector per targeted activation
LoHa r=8           32,768    12.451%  Hadamard product of two low-rank pairs
LoKr r=8            2,048     0.778%  Kronecker product factorisation
OFT b=32           15,872     6.031%  a block-diagonal orthogonal rotation

inside one adapted layer:
  frozen weight      (256, 256)  requires_grad=False
  lora_A             (8, 256)  requires_grad=True
  lora_B             (256, 8)  requires_grad=True
  B is initialised to all zeros: True
  adapted output equals base output at step 0: True

Reading the table

The percentages look large because the base model is tiny. 256×256 layers make LoRA r=8 look like 6%. On a real 4096-wide model the same r=8 is well under 0.1%, because LoRA's cost grows as 2·r·d while the frozen weight grows as d². Always compute this ratio for your dimensions before choosing a rank.

B starts at exactly zero, so the adapter starts as a no-op. The last line confirms it numerically: the adapted model's output is bit-identical to the base model's before training. That property is what makes adapters safe to attach to a production model — attaching one changes nothing until you train it.

IA3 is the smallest thing that still does something. 1,024 numbers, one scaling factor per output channel per layer. That is 0.4% of a tiny model and effectively nothing on a large one.

VeRA is smaller than LoRA r=1. It shares one frozen random A and B across all layers and learns only two per-layer vectors. The rank is still 8; the learned content is two vectors.

LoHa costs twice LoRA at the same r. It multiplies two low-rank pairs elementwise, which raises the effective rank beyond r at the price of two pairs of factors. "Rank 8" does not mean the same thing across these methods, so compare by trainable-parameter count, never by r.

DoRA is LoRA plus 1,024 numbers. One magnitude scalar per output column. That is the entire difference, and it is reported to close much of the gap to full fine-tuning at low rank.

Choosing, in practice

SituationReasonable first choice
Standard task or style adaptationLoRA, r=8–32, on all linear layers
Very low rank requiredDoRA (recovers quality LoRA loses at r=4)
Serving hundreds of variantsVeRA or IA3 — adapter files of a few kilobytes
GPU memory is the binding limitQLoRA: 4-bit frozen base plus LoRA
Preserving base behaviour matters mostOFT or BOFT — a rotation cannot change norms
Prompt-style steering onlyprompt tuning or prefix tuning

Two settings that matter more than the method:

Target all the linear layers, not only Q and V. The original LoRA paper adapted attention projections; later practice adapts the MLP layers as well, and it consistently helps.

Use a higher learning rate than full fine-tuning. Roughly 1e-4 for adapters against 1e-5 to 2e-5 for full fine-tuning, because only new parameters are being learned.

Common mistakes

Comparing methods by r. As the table shows, r means different things. Compare by trainable parameters and by measured quality.

Merging a QLoRA adapter into a 4-bit base. merge_and_unload() on a quantised base dequantises and reintroduces quantisation error. Merge into the full-precision base instead, then requantise.

Forgetting to save the adapter config. An adapter file without its adapter_config.json cannot be reattached. Both are needed.

Setting lora_alpha and then changing r. Effective scaling is alpha/r by default, so changing r changes the update magnitude. use_rslora=True uses alpha/sqrt(r), which makes higher ranks behave more predictably.

Expecting an adapter to teach a new language or domain from scratch. That is the case where full fine-tuning or continued pretraining wins. Adapters are for shaping, not for installing.

Try it yourself

Add LoraConfig(r=64, lora_alpha=128, target_modules=TARGETS) to METHODS. Compute where its parameter count crosses the base model's, then work out the width at which r=64 would still be under 1% of a real model.

What to learn next

Researcher — Mathematics and papers.

The low-rank hypothesis and its descendants

LoRA (Hu et al., 2022) reparameterises the weight update as

$$ W = W_0 + \Delta W = W_0 + \frac{\alpha}{r} BA, \qquad B \in \mathbb{R}^{d\times r},\ A \in \mathbb{R}^{r\times k} $$

with $W_0$ frozen, $B$ initialised to zero and $A$ to a random Gaussian. The claim is that the update has low intrinsic rank even when $W_0$ does not — supported by Aghajanyan et al., 2021 (Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning), who showed a few hundred to a few thousand parameters in a random subspace can reach 90% of full fine-tuning performance on GLUE.

rsLoRA (Kalajdzievski, 2023) points out the $\alpha/r$ scaling makes gradients vanish as $r$ grows, so higher ranks stop helping for the wrong reason. Replacing it with $\alpha/\sqrt r$ fixes the scaling and restores monotonic improvement with rank.

DoRA (Liu et al., 2024) decomposes the weight into magnitude and direction:

$$ W = m \cdot \frac{W_0 + BA}{\lVert W_0 + BA \rVert_c} $$

where $m \in \mathbb{R}^{1\times k}$ is a learned per-column magnitude and $\lVert\cdot\rVert_c$ is the column-wise norm. Their analysis shows full fine-tuning changes magnitude and direction largely independently, while LoRA couples them; decoupling closes most of the low-rank gap. The cost is $k$ extra parameters per layer — the 1,024 in the table.

LoRA+ (Hayou et al., 2024) shows the optimal learning rates for $A$ and $B$ differ by a factor that grows with width, and that setting $\eta_B = \lambda \eta_A$ with $\lambda \gg 1$ improves both speed and final quality. This is a one-line change most practitioners never make.

Smaller than low-rank

IA3 (Liu et al., 2022, Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning) learns three vectors per block, rescaling keys, values and the feed-forward intermediate:

$$ \text{attn} = \mathrm{softmax}!\left(\frac{Q(l_k \odot K)^\top}{\sqrt{d}}\right)(l_v \odot V), \qquad \text{FFN} = (l_{ff}\odot \gamma(W_1 x))W_2 $$

It uses roughly 0.01% of parameters and, in their few-shot setting, beat in-context learning at a fraction of the inference cost.

VeRA (Kopiczko et al., 2024) freezes a single pair of random matrices $A, B$ shared across all layers and learns only two scaling vectors $b, d$ per layer:

$$ \Delta W = \Lambda_b B \Lambda_d A $$

Parameter count becomes $O(d + r)$ per layer instead of $O(r(d+k))$, reported at roughly 10× smaller than LoRA at comparable quality on GLUE and E2E. The relevant use case is serving thousands of adapters.

BitFit (Ben Zaken et al., 2022) trains only bias terms — about 0.09% of parameters — and is competitive on small-data GLUE tasks. It is the strongest evidence that a large part of task adaptation is a shift rather than a rotation.

Structure-preserving methods

OFT and BOFT (Qiu et al., 2023; Liu et al., 2024) multiply rather than add:

$$ W = R\,W_0, \qquad R\ \text{orthogonal, block-diagonal} $$

Because $R$ is orthogonal, the singular values of $W_0$ are preserved exactly — the transformation is a rotation, hence norm-preserving and information-preserving in a precise sense. This is why OFT was first adopted in text-to-image work, where additive adapters visibly damage the base model's generative quality. BOFT adds a butterfly factorisation to raise expressiveness at fixed parameter count.

Prompt-space methods

Prefix tuning (Li and Liang, 2021) prepends trainable key–value vectors to every attention layer. Prompt tuning (Lester et al., 2021) prepends trainable embeddings to the input only, and shows the gap to full fine-tuning closes as model size grows, vanishing around 10B parameters. P-tuning v2 restores the per-layer prefixes for smaller models.

These consume context length, which is their practical drawback: a 100-token soft prompt costs 100 tokens of window on every request forever.

What adapters cost in quality

Biderman et al., 2024 (LoRA Learns Less and Forgets Less) is the most careful public comparison, on continued pretraining and instruction tuning in code and maths. Their findings:

  • On continued pretraining with 10B–20B tokens of new domain data, LoRA falls substantially short of full fine-tuning. The gap is largest for code.
  • On instruction tuning with smaller datasets, LoRA is close to full fine-tuning.
  • LoRA forgets less — it better preserves performance outside the target domain — and acts as a stronger regulariser than classic weight decay or dropout.
  • Full fine-tuning finds updates whose rank is 10–100× higher than typical LoRA configurations, which explains the first two points directly.

The practical rule follows from the ranking: adapters for shaping behaviour on modest data, full fine-tuning or continued pretraining for installing large new knowledge.

Serving many adapters

The property that made LoRA an infrastructure decision rather than a training trick: $B A$ can be computed as a separate small matmul on a shared frozen base, so many adapters can be served in one batch. S-LoRA (Sheng et al., 2024) and Punica implement this with custom kernels, serving thousands of adapters on a single GPU with the base weights loaded once. Merging is the alternative and gives up multi-tenancy.

Papers

What to learn next