ValueError: Target modules not found in the base model (PEFT / LoRA)
Your LoRA config names layers that do not exist in this architecture — module names differ between model families. Print the real module names, or use target_modules="all-linear".
Updated
The error
ValueError: Target modules ['query_key_value'] not found in the base model. Please check the target modules and try again.
What it means
LoRA fine-tunes a model by attaching small trainable adapters to specific layers, and target_modules names those layers. PEFT searched your model for modules matching the names you gave and found none. The names are architecture-specific — the config you copied was written for a different model family than the one you loaded.
Why it happens
Attention layers carry different names in different families. Falcon and the Bloom generation call the fused attention projection query_key_value. Llama, Mistral and Qwen split it into q_proj, k_proj, v_proj, o_proj. GPT-2 uses c_attn. A LoRA recipe travels between tutorials, the model underneath changes, and the names stop matching.
How to fix it
1. Look at the real module names in your model.
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-Instruct-v0.3")
names = {n.split(".")[-1] for n, m in model.named_modules()
if m.__class__.__name__ == "Linear"}
print(names){'q_proj', 'k_proj', 'v_proj', 'o_proj', 'gate_proj', 'up_proj', 'down_proj', 'lm_head'}Those are the valid targets for this model.
2. Set target_modules from what you saw.
from peft import LoraConfig
config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_dropout=0.05,
task_type="CAUSAL_LM",
)Targeting the four attention projections is the common starting point; adding the MLP projections (gate_proj, up_proj, down_proj) increases capacity at more memory cost.
3. Or let PEFT pick every linear layer.
config = LoraConfig(r=16, lora_alpha=32,
target_modules="all-linear", task_type="CAUSAL_LM")"all-linear" targets all linear layers except the output head — architecture-independent, so it survives model swaps. It trains more parameters than a hand-picked list; on tight GPU budgets, prefer the explicit list.
4. Quick reference for common families.
| Family | Attention module names |
|---|---|
| Llama / Mistral / Qwen | q_proj, k_proj, v_proj, o_proj |
| Falcon / Bloom style | query_key_value |
| GPT-2 style | c_attn |
Verify with fix 1 rather than trusting tables — including this one — since new releases change layouts.
How to prevent it
Whenever the base model changes, re-run the module listing before training. Keep the target_modules choice next to the model name in your config file so the pair travels together. And when borrowing a fine-tuning script, the module names are the first thing to audit.
Related errors
- Missing key(s) / Unexpected key(s) in state_dict — the same "names must match" rule at load time
- Model type not recognized by Transformers
- CUDA out of memory — the usual next constraint during fine-tuning