AI glossary

PEFT (parameter-efficient fine-tuning)

In one sentence PEFT is the family of methods that adapt a large model by training only a small fraction of its weights, with LoRA as the star member.

By Updated

Parameter-efficient fine-tuning, or PEFT, adapts a large pretrained model by training a small added or selected set of weights while the original billions stay frozen.

Renovating a rented flat is the analogy. You cannot rebuild the walls — not yours to change, and far too costly — but curtains, shelves and lighting transform how the place works for you. PEFT treats the pretrained model as the building: frozen structure, small trainable additions, most of the value of a renovation at a sliver of the cost.

Why the family exists: full fine-tuning of a 70B-parameter model must store gradients and optimizer state for every weight — hundreds of gigabytes of GPU memory — and produces a full-size copy per task. PEFT methods train well under 1% of the parameters, run on a single modest GPU, and each task's adaptation is a file measured in megabytes.

The members you will actually meet:

LoRA        train small low-rank matrices alongside frozen weights — the default
QLoRA       LoRA on top of a 4-bit quantized base — fine-tune 70B on one GPU
prompt/prefix tuning   learn a few virtual tokens prepended to the input
adapters    small bottleneck layers inserted between frozen ones

The name is also a library: Hugging Face's peft package implements these behind one interface. Rule of thumb: LoRA or QLoRA covers most needs; research comparing PEFT to full fine-tuning finds near-parity for most task adaptation, with full fine-tuning still ahead when behaviour must change deeply.

Where to go next