PEFT (parameter-efficient fine-tuning)
In one sentence PEFT is the family of methods that adapt a large model by training only a small fraction of its weights, with LoRA as the star member.
Updated
Parameter-efficient fine-tuning, or PEFT, adapts a large pretrained model by training a small added or selected set of weights while the original billions stay frozen.
Renovating a rented flat is the analogy. You cannot rebuild the walls — not yours to change, and far too costly — but curtains, shelves and lighting transform how the place works for you. PEFT treats the pretrained model as the building: frozen structure, small trainable additions, most of the value of a renovation at a sliver of the cost.
Why the family exists: full fine-tuning of a 70B-parameter model must store gradients and optimizer state for every weight — hundreds of gigabytes of GPU memory — and produces a full-size copy per task. PEFT methods train well under 1% of the parameters, run on a single modest GPU, and each task's adaptation is a file measured in megabytes.
The members you will actually meet:
LoRA train small low-rank matrices alongside frozen weights — the default
QLoRA LoRA on top of a 4-bit quantized base — fine-tune 70B on one GPU
prompt/prefix tuning learn a few virtual tokens prepended to the input
adapters small bottleneck layers inserted between frozen onesThe name is also a library: Hugging Face's peft package implements these behind one interface. Rule of thumb: LoRA or QLoRA covers most needs; research comparing PEFT to full fine-tuning finds near-parity for most task adaptation, with full fine-tuning still ahead when behaviour must change deeply.
Where to go next
- Full lesson: LoRA
- Related terms: lora, fine-tuning, quantization, transfer-learning