LoRA (low-rank adaptation)
In one sentence LoRA fine-tunes a large model by training a small pair of extra matrices while the original weights stay frozen.
Updated
LoRA, short for low-rank adaptation, fine-tunes a large model by training a small set of extra weights and leaving every original weight untouched.
You have a printed textbook and you want it to cover your syllabus. Reprinting the whole book for each change would be absurd. Instead you write on sticky notes and attach them to the relevant pages. The book is unchanged, the notes are tiny, and anyone reading gets your version. Swap the notes and the same book serves a different course.
Technically, the sticky note is a pair of small matrices placed beside a big weight matrix. The big matrix is frozen. During training only the small pair changes, and at inference their product is added to the original weights. "Low-rank" is the reason the pair can be small: the useful adjustment for one task turns out to be far simpler than the full weight matrix, so it can be expressed compactly.
Why this matters in practice
Full fine-tune of a 7B model
weights 14 GB + gradients 14 GB + Adam optimizer state 56 GB → needs a big multi-GPU box
LoRA on the same model
frozen weights 14 GB + a few million trainable weights → fits on one 24 GB card
saved adapter file → tens of megabytesThe optimizer state is the part that disappears, and it is the biggest part. Adam keeps two extra values per trainable weight, so freezing 99% of the model removes almost all of that.
Two more things follow. Adapters are small enough to store dozens of them and load one per customer or task on top of a single shared base model. And QLoRA goes further by keeping the frozen base in 4-bit quantized form, which brings fine-tuning of a 7B model within reach of a single consumer GPU. The library people use for this is Hugging Face PEFT, and the original paper is Hu et al., 2021.
Where to go next
- Full lesson: LoRA
- Related terms: fine-tuning, quantization, parameter, llm