Merging model weights
Two models fine-tuned from the same starting point can be combined by averaging their weights, with no training and no data.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Take the numbers inside two models and average them. The result often works.
The analogy you have already lived
Open two tins of paint from the same base white, one tinted blue and one tinted yellow. Pour some of each into a third tin and stir.
You get green. Not blue, not yellow, and not mud — a usable colour.
Now try mixing paint with cooking oil. You get a mess. The trick works because both tins started from the same base.
Model merging has exactly that condition attached, and it is the whole story.
Why it exists
Suppose two teams fine-tuned the same open model. One taught it Hindi. One taught it to write code.
You want both. The normal answer is to gather both datasets and train again. That is expensive and slow, and impossible without the data.
Merging skips all of it. Average the weights. No training, no data, no GPU. It takes about as long as copying the files.
How it works
The simplest version really is an average.
model A weights [0.4, -0.2, 0.9, ...]
model B weights [0.6, 0.1, 0.5, ...]
| | |
averaged [0.5, -0.05, 0.7, ...]A more careful version thinks in terms of changes. Subtract the original model from each fine-tuned one, and what is left is "what this fine-tune did". That difference is called a task vector.
what fine-tuning A changed = A - original
what fine-tuning B changed = B - original
merged = original + (A's changes) + (B's changes)Written that way, other ideas become natural. You can scale a change up or down. You can subtract one, to remove a behaviour. You can keep only the biggest changes and throw the rest away.
The condition that makes it work
Both models must be fine-tunes of the same starting model, and neither may have wandered too far from it.
Suppose they started from different places. Or one was trained so hard it ended up somewhere else entirely. Averaging then produces nonsense. The developer section below shows exactly that failure.
What is honestly hard here
Merging is not free multi-task learning, and it gets sold that way.
In the measurement below, the merged model beats either specialist and loses to a model trained properly on both datasets. That is the realistic result.
Merging is what you do when retraining is not available. If you have the data and the compute, train.
Where you have already seen this
- Mixing two tins of tinted paint from a common base.
- Blending two spice mixes that share the same core masala.
- Averaging two people's estimates of a distance.
Remember this
- Merging averages the weights of models fine-tuned from a common starting point.
- It needs no data, no training and no GPU.
- It beats either specialist, and loses to properly retraining on both datasets.
What to learn next
- Reward hacking and sycophancy — the last thing that can go wrong in post-training.
- Catastrophic forgetting — the problem merging partly repairs.
- Adapters beyond LoRA — the other way to combine specialisations.
Developer — Code and libraries.
Setup
pip install torchRuns on a CPU in about thirty seconds.
Four merge methods, and the condition they depend on
import copy
import torch
import torch.nn as nn
import torch.nn.functional as F
# Same two-task setup as the forgetting lesson: disjoint labels, a task marker.
D, C = 41, 10
def task(which, n, gen):
x = torch.randn(n, D, generator=gen)
x[:, -1] = 0.0 if which == "A" else 1.0
half = x[:, :20] if which == "A" else x[:, 20:40]
b = (half.sum(1) / 3 + 2.5).clamp(0, 5 - 1e-3).long()
return x, b + (0 if which == "A" else 5)
def net(seed=0):
torch.manual_seed(seed)
return nn.Sequential(nn.Linear(D, 128), nn.ReLU(),
nn.Linear(128, 128), nn.ReLU(), nn.Linear(128, C))
gen = torch.Generator().manual_seed(0)
ax, ay = task("A", 4000, gen)
bx, by = task("B", 4000, gen)
tax, tay = task("A", 2000, torch.Generator().manual_seed(5))
tbx, tby = task("B", 2000, torch.Generator().manual_seed(6))
def fit(model, x, y, steps, lr=3e-3):
opt = torch.optim.AdamW(model.parameters(), lr=lr)
for _ in range(steps):
loss = F.cross_entropy(model(x), y)
opt.zero_grad(); loss.backward(); opt.step()
return model
def acc(m, x, y):
with torch.no_grad():
return (m(x).argmax(-1) == y).float().mean().item()
def sd(m):
return {k: v.clone() for k, v in m.state_dict().items()}
def report(base_steps):
base = fit(net(), torch.cat([ax, bx]), torch.cat([ay, by]), base_steps)
theta0 = sd(base)
thetaA = sd(fit(copy.deepcopy(base), ax, ay, 600)) # two specialists, same start
thetaB = sd(fit(copy.deepcopy(base), bx, by, 600))
tauA = {k: thetaA[k] - theta0[k] for k in theta0} # task vectors
tauB = {k: thetaB[k] - theta0[k] for k in theta0}
def task_arithmetic(lam):
return {k: theta0[k] + lam * (tauA[k] + tauB[k]) for k in theta0}
def average():
return {k: 0.5 * thetaA[k] + 0.5 * thetaB[k] for k in theta0}
def ties(density=0.2, lam=1.0):
out = {}
for k in theta0:
flat = torch.stack([tauA[k], tauB[k]]).reshape(2, -1)
keep = max(1, int(density * flat.shape[1]))
trimmed = torch.zeros_like(flat)
for i in range(2): # TRIM: keep the largest
idx = flat[i].abs().topk(keep).indices
trimmed[i, idx] = flat[i, idx]
sign = trimmed.sum(0).sign() # ELECT: majority sign
agree = (trimmed.sign() == sign) & (trimmed != 0)
merged = (trimmed * agree).sum(0) / agree.sum(0).clamp(min=1)
out[k] = theta0[k] + lam * merged.reshape(theta0[k].shape)
return out
def dare(p=0.9, lam=1.0, seed=3):
g = torch.Generator().manual_seed(seed)
out = {}
for k in theta0:
total = torch.zeros_like(theta0[k])
for tau in (tauA, tauB): # DROP p, RESCALE by 1/(1-p)
mask = (torch.rand(tau[k].shape, generator=g) > p).float()
total = total + mask * tau[k] / (1 - p)
out[k] = theta0[k] + lam * total
return out
joint = sd(fit(copy.deepcopy(base), torch.cat([ax, bx]), torch.cat([ay, by]), 600))
rows = [("shared base, before specialising", theta0),
("specialist A", thetaA), ("specialist B", thetaB),
("average of A and B", average()),
("task arithmetic, lambda=0.5", task_arithmetic(0.5)),
("task arithmetic, lambda=1.0", task_arithmetic(1.0)),
("TIES, density 0.2", ties()),
("DARE, drop 0.9", dare()),
("joint training on A+B data", joint)]
print(f"\n=== base trained for {base_steps} steps ===")
print(f"{'model':<34} {'task A':>8} {'task B':>8} {'mean':>8}")
for name, state in rows:
m = copy.deepcopy(base)
m.load_state_dict(state)
a, b = acc(m, tax, tay), acc(m, tbx, tby)
print(f"{name:<34} {a:>8.3f} {b:>8.3f} {(a + b) / 2:>8.3f}")
report(40) # a reasonably strong shared base
report(10) # a weak shared base: the specialists drift into different basins=== base trained for 40 steps === model task A task B mean shared base, before specialising 0.822 0.855 0.839 specialist A 0.891 0.000 0.445 specialist B 0.000 0.904 0.452 average of A and B 0.633 0.581 0.607 task arithmetic, lambda=0.5 0.633 0.581 0.607 task arithmetic, lambda=1.0 0.525 0.397 0.461 TIES, density 0.2 0.573 0.387 0.480 DARE, drop 0.9 0.273 0.145 0.209 joint training on A+B data 0.882 0.896 0.889 === base trained for 10 steps === model task A task B mean shared base, before specialising 0.382 0.335 0.359 specialist A 0.876 0.000 0.438 specialist B 0.000 0.892 0.446 average of A and B 0.398 0.269 0.334 task arithmetic, lambda=0.5 0.398 0.269 0.334 task arithmetic, lambda=1.0 0.368 0.214 0.291 TIES, density 0.2 0.272 0.311 0.292 DARE, drop 0.9 0.282 0.020 0.151 joint training on A+B data 0.890 0.895 0.892
Written against PyTorch 2.5.1, CPU, all seeds fixed — reproducible on this build.
Six honest readings
Merging works: 0.607 mean, against 0.445 and 0.452 for the specialists. Two models that had each lost one task became one model that does both partly, using no data and no gradient steps. That is a real result.
Averaging and task arithmetic at lambda=0.5 gave identical numbers, to three decimals. They are the same operation. For two models, θ₀ + 0.5(τ_A + τ_B) = 0.5·θ_A + 0.5·θ_B exactly. Task arithmetic's value is that it generalises — you can scale, subtract, and combine more than two.
lambda=1.0 was worse than lambda=0.5. Adding both task vectors at full strength overshoots. Every merge method has a scaling coefficient and it is the most important knob; sweep it.
TIES did not beat plain averaging, and DARE was much worse. With only two task vectors there is very little sign conflict to resolve, so trimming to 20% mostly throws information away, and dropping 90% of a two-vector merge is severe. These methods are designed for many models with conflicting updates. Applying them to two models is applying the medicine without the disease.
Joint training won, by a lot: 0.889 against 0.607. If you have both datasets, use both datasets. Merging is the answer when you cannot.
The second table is the condition, demonstrated. With a weak shared base, merging fell below the base itself — 0.334 against 0.359 — and every method degraded. The specialists drifted into different regions of weight space, and averaging two unrelated points gives a third unrelated point. A strong shared base and modest fine-tuning are not nice-to-haves; they are the prerequisite.
Doing it on real models
mergekit (Arcee AI, Apache 2.0) is the standard tool, driven by a YAML file:
models:
- model: org/model-a
parameters: {weight: 0.5, density: 0.5}
- model: org/model-b
parameters: {weight: 0.5, density: 0.5}
merge_method: ties
base_model: org/base-model
dtype: bfloat16No output block — this needs the real checkpoints, tens of gigabytes of download, and mergekit installed.
Methods you will see in that merge_method field:
| Method | What it does | Models |
|---|---|---|
linear | weighted average | any number |
slerp | spherical interpolation along the arc | exactly 2 |
task_arithmetic | scaled sum of task vectors | any number |
ties | trim, elect sign, then average the survivors | many |
dare_ties | random drop and rescale, then TIES | many |
model_stock, della, … | later refinements | many |
Common mistakes
Merging models with different architectures or tokenizers. The weights must line up key by key. Different vocabulary sizes make the embedding matrices incompatible.
Merging models from different bases. Two Llama fine-tunes merge. A Llama fine-tune and a Mistral fine-tune do not.
Not sweeping the scaling coefficient. As the table shows, lambda moves results substantially. Evaluate at 0.3, 0.5, 0.7, 1.0 before believing any merge.
Reaching for TIES or DARE by default. They help when many task vectors conflict. On two models they can lose to a plain average, as measured above.
Skipping evaluation because "merging is free". The merge is free. Being wrong about it is not.
Merging away safety behaviour. A merge with an unaligned model can dilute alignment. Re-test refusals on any merged model you ship.
Try it yourself
Add a subtract row: theta0 + tauA - tauB. Evaluate it on both tasks. Subtracting a task vector is how "unlearning" by merging works, and seeing the number is the fastest way to understand what a task vector is.
What to learn next
- Reward hacking and sycophancy — the last thing that can go wrong in post-training.
- Catastrophic forgetting — the problem merging partly repairs.
- Adapters beyond LoRA — the other way to combine specialisations.
Researcher — Mathematics and papers.
Why averaging weights works at all
Averaging the parameters of two independently trained networks normally produces a model at chance, because of permutation symmetry: hidden units can be reordered without changing the function, so two solutions are almost never aligned.
The escape is linear mode connectivity. Frankle et al., 2020 showed that networks which share a short period of training before diverging remain connected by a low-loss linear path. Neyshabur et al., 2020 extended this to fine-tuning: models fine-tuned from a common pretrained checkpoint stay in one basin, and the linear path between them has no loss barrier.
That is the theorem behind the condition the code demonstrates. Merging is not a general property of neural networks; it is a property of fine-tunes of a shared pretrained model.
Ainsworth et al., 2023 (Git Re-Basin) attack the general case by solving for the permutation that aligns two independently trained models before merging, making the barrier disappear. It works, and it is far more expensive than averaging.
Model soups
Wortsman et al., 2022 (Model Soups) average many fine-tunes of the same base that differ only in hyperparameters. Two recipes:
- Uniform soup: average everything.
- Greedy soup: sort by validation accuracy and add a model only if it improves the running average.
Greedy soup on fine-tuned ViT-G reached a then-state-of-the-art 90.94% ImageNet top-1, beating the best single member, and improved out-of-distribution robustness. Crucially, inference cost is that of one model, unlike an ensemble.
Task arithmetic
Ilharco et al., 2023 define the task vector $\tau = \theta_{\text{ft}} - \theta_{\text{pre}}$ and show three operations behave as the name suggests:
- Negation ($\theta_{\text{pre}} - \lambda\tau$) reduces the target behaviour with little collateral damage. Their headline case: removing toxic generation.
- Addition ($\theta_{\text{pre}} + \lambda\sum_i \tau_i$) builds multi-task models.
- Analogy ($\tau_C \approx \tau_B - \tau_A + \tau_{A'}$) transfers a task relationship to a new domain.
Resolving interference
Naive summation of many task vectors degrades as the count grows, from two causes identified by Yadav et al., 2023 (TIES-Merging): redundant small-magnitude changes, and sign disagreement where two task vectors want a parameter to move in opposite directions. TIES:
- Trim each task vector to its top-$k$% by magnitude (typically 20%).
- Elect a sign per parameter by summed magnitude.
- Merge by averaging only the entries agreeing with the elected sign.
Yu et al., 2024 (DARE — Language Models are Super Mario) show delta parameters are extremely redundant: randomly dropping 90% (and up to 99% for small deltas) and rescaling the survivors by $1/(1-p)$ preserves performance. DARE is a preprocessing step and composes with TIES.
The measurement in the code is a fair illustration of the boundary condition: both methods are designed for high-interference regimes, and at two models with disjoint tasks they have nothing useful to do. Reported gains in the papers come from merging 3–11 task vectors.
Where merging is genuinely load-bearing
- Frontier post-training. Llama 3's post-training averages models from several RLHF and DPO runs at each round. The models are all fine-tunes of one base and the averaging reduces variance.
- Alignment–capability trade-offs. Interpolating a fine-tuned model back toward its base (Wortsman et al., 2022, WiSE-FT) recovers robustness lost to fine-tuning, and mitigates catastrophic forgetting.
- Replacing the learning-rate decay phase. Hägele et al., 2024 and the WSM line report that averaging checkpoints from a constant-learning-rate phase recovers most of a cooldown's benefit.
- The open-weights leaderboard ecosystem, where a large fraction of top entries are merges. This is also where merging's weakness is clearest: merged models can score well on benchmarks whose test data influenced the selection of merge coefficients, which is benchmark overfitting by another route.
Open problems
- No theory of coefficients. $\lambda$ and density are chosen by search against a validation set. Nothing predicts them.
- Scaling in the number of models. Yadav et al., 2024 (What Matters for Model Merging at Scale?) find merging improves with base-model size and with the number of tasks, but that held-out generalisation and interference behave differently from in-task performance.
- Merging across architectures remains unsolved outside the Git Re-Basin line.
Papers and tools
- Frankle et al., Linear Mode Connectivity and the Lottery Ticket Hypothesis, ICML 2020 — arxiv.org/abs/1912.05671
- Neyshabur et al., What is being transferred in transfer learning?, NeurIPS 2020 — arxiv.org/abs/2008.11687
- Wortsman et al., Model Soups, ICML 2022 — arxiv.org/abs/2203.05482
- Wortsman et al., Robust fine-tuning of zero-shot models (WiSE-FT), CVPR 2022 — arxiv.org/abs/2109.01903
- Ilharco et al., Editing Models with Task Arithmetic, ICLR 2023 — arxiv.org/abs/2212.04089
- Ainsworth et al., Git Re-Basin, ICLR 2023 — arxiv.org/abs/2209.04836
- Yadav et al., TIES-Merging, NeurIPS 2023 — arxiv.org/abs/2306.01708
- Yu et al., Language Models are Super Mario (DARE), ICML 2024 — arxiv.org/abs/2311.03099
- Goddard et al., Arcee's MergeKit, 2024 — arxiv.org/abs/2403.13257
- Yadav et al., What Matters for Model Merging at Scale?, 2024 — arxiv.org/abs/2410.03617
What to learn next
- Reward hacking and sycophancy — the last thing that can go wrong in post-training.
- Catastrophic forgetting — the problem merging partly repairs.
- Adapters beyond LoRA — the other way to combine specialisations.