Image matting and background removal
Hair and soft edges are part foreground and part background in the same pixel, so matting asks how much rather than which, and a hard mask can never answer it.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Matting asks how much of each pixel is foreground, instead of forcing a yes-or-no answer.
The analogy
Photograph a friend with loose hair against a bright window. Look closely at a single dot on the screen at the edge of a strand.
That dot is not hair and it is not window. It is part of both, mixed together, because the strand is thinner than the dot.
Now cut them out with scissors. Every cut has to go one side of the strand or the other, and the result looks pasted on.
A soft answer says "this dot is one third hair". That is what matting produces. It is why hair cut-outs look real in films and fake in most apps.
Why a hard mask fails
Take the hard cut-out and place it on a dark background. Every edge pixel still carries the colour of the old background it was photographed against.
You get a pale outline around the whole person. Everybody has seen that glow, and now you know exactly what causes it.
The old background was never removed from those pixels. It was carried across, because a yes-or-no mask has no way to say "partly".
How the soft answer is used
for each pixel, take some of the foreground colour
and take the rest from the new background
a pixel that is all hair -> take all of the hair colour
a pixel that is all window -> take all of the new background
a pixel that is a third hair -> take a third hair, two thirds new backgroundFully inside and fully outside are easy. Everything interesting lives in the thin band between them.
The reason it is genuinely hard
For each mixed pixel you see one colour. From it you must recover three things. The true foreground colour, the true background colour, and how much of each.
One observation, three unknowns. Endless combinations produce exactly the colour you see. Nothing in the picture alone tells you which is right.
This is not a temporary limitation. It is the shape of the problem. Every method adds an extra assumption to pin it down.
The two ways people pin it down
Tell it where to look. A trimap is a rough three-way sketch. Definitely foreground, definitely background, and a band of "not sure" between them. The method only has to solve the hard problem inside the band. It borrows colours from the certain regions on either side.
Control the background. Paint the wall a single bright green. Now the background colour is known everywhere, and two unknowns become one. That is a green screen, and it is why studios still use them.
Modern tools skip the trimap for portraits. They learned from thousands of examples what a person's soft edge looks like. That works well on people and much less well on a bicycle wheel or a glass.
Where you have already seen this
- Video call backgrounds, and the halo around your shoulders when the light is poor.
- Weather forecasts, filmed against green.
- Product photos for online shops, cut out and placed on white.
- "Remove background" buttons in phone galleries and design tools.
Remember this
- Edge pixels are mixtures, not one thing or the other.
- A hard mask carries the old background into every edge pixel.
- One colour, three unknowns: matting needs an extra assumption, such as a trimap or a green screen.
What to learn next
- Cleaning up predicted masks — the hard-mask toolbox this lesson argues past.
- How images are stored — channels, alpha and why PNG behaves differently from JPEG.
- Diffusion models — the other route to replacing a background, by generating one.
Developer — Code and libraries.
Setup
pip install numpyRuns instantly on CPU. Working with a synthetic scene means the true alpha is known, so every method can be scored rather than admired.
The whole problem, in one script
import numpy as np
# ---------- 1. why matting is under-determined ----------
I = 0.50 # one observed grey pixel
print("observed pixel value 0.50. Every row below explains it perfectly:")
print(f"{'foreground':>12}{'background':>12}{'alpha':>8}")
for F, B in [(0.00, 1.00), (0.20, 0.80), (0.40, 0.60), (0.45, 0.95), (0.10, 0.62)]:
a = (I - B) / (F - B)
print(f"{F:>12.2f}{B:>12.2f}{a:>8.2f} rebuilds to {a*F + (1-a)*B:.2f}")
print("one equation, three unknowns per channel. A trimap is how you cut that down.\n")
# ---------- 2. a strand of hair with a known true alpha ----------
H, W = 14, 40
rng = np.random.default_rng(3)
alpha = np.zeros((H, W))
alpha[:, :14] = 1.0 # solid head
for r in range(H): # strands fading out to the right
reach = 14 + int(8 + 10 * rng.random())
alpha[r, 14:reach] = np.linspace(1.0, 0.0, reach - 14)
alpha = (alpha + 0.05 * rng.random((H, W)) * (alpha > 0) * (alpha < 1)).clip(0, 1)
F_col = 0.15 # dark hair
wall = 0.85 - 0.25 * (np.arange(W) / W) # the studio wall is not one flat colour
B_old = np.tile(wall, (H, 1))
B_new = 0.35 # the new background
image = alpha * F_col + (1 - alpha) * B_old # what the camera recorded
BAR = " .:-=+*#@"
def bar(row):
return "".join(BAR[min(8, int(v * 8.999))] for v in row)
print("true alpha (@ = only hair, space = only background):")
for r in range(0, H, 2):
print(" " + bar(alpha[r]))
# ---------- 3. four ways to move it onto the new background ----------
ideal = alpha * F_col + (1 - alpha) * B_new
results = {
"hard mask, cut and paste": np.where(alpha > 0.5, image, B_new),
"alpha from one average wall colour": None,
"alpha from the true wall colour": None,
"true alpha": alpha * F_col + (1 - alpha) * B_new,
}
a_avg = np.clip((image - B_old.mean()) / (F_col - B_old.mean()), 0, 1)
results["alpha from one average wall colour"] = a_avg * F_col + (1 - a_avg) * B_new
a_true = np.clip((image - B_old) / (F_col - B_old), 0, 1)
results["alpha from the true wall colour"] = a_true * F_col + (1 - a_true) * B_new
print()
for name, out in results.items():
err = np.abs(out - ideal)
print(f"{name:36s} mean error {err.mean():.4f} worst pixel {err.max():.4f}")
print("\nrow 6 on the new background (space = dark, @ = light):")
for name in ["true alpha", "hard mask, cut and paste", "alpha from one average wall colour"]:
print(f" {name:36s}" + bar(results[name][6]))
print(" the hard mask carries the OLD wall colour into every edge pixel")
# ---------- 4. the whole argument lives in a thin band ----------
band = (alpha > 0.02) & (alpha < 0.98)
hard_err = np.abs(results["hard mask, cut and paste"] - ideal)
print(f"\npixels in the soft band: {band.sum()} of {alpha.size} ({band.mean():.0%})")
print(f"hard-mask error inside the band : {hard_err[band].mean():.4f}")
print(f"hard-mask error outside the band: {hard_err[~band].mean():.4f}")observed pixel value 0.50. Every row below explains it perfectly:
foreground background alpha
0.00 1.00 0.50 rebuilds to 0.50
0.20 0.80 0.50 rebuilds to 0.50
0.40 0.60 0.50 rebuilds to 0.50
0.45 0.95 0.90 rebuilds to 0.50
0.10 0.62 0.23 rebuilds to 0.50
one equation, three unknowns per channel. A trimap is how you cut that down.
true alpha (@ = only hair, space = only background):
@@@@@@@@@@@@@@@@*+=:.
@@@@@@@@@@@@@@@@@#**+===-::.
@@@@@@@@@@@@@@@@*+=:.
@@@@@@@@@@@@@@@@#**+=-:..
@@@@@@@@@@@@@@@@@#*++==-:...
@@@@@@@@@@@@@@@@#*+=-:..
@@@@@@@@@@@@@@@@#**==-:..
hard mask, cut and paste mean error 0.0192 worst pixel 0.1855
alpha from one average wall colour mean error 0.0108 worst pixel 0.0422
alpha from the true wall colour mean error 0.0000 worst pixel 0.0000
true alpha mean error 0.0000 worst pixel 0.0000
row 6 on the new background (space = dark, @ = light):
true alpha ...................::::::---------------
hard mask, cut and paste ................::----------------------
alpha from one average wall colour ...................::::::---::::::::::::
the hard mask carries the OLD wall colour into every edge pixel
pixels in the soft band: 133 of 560 (24%)
hard-mask error inside the band : 0.0807
hard-mask error outside the band: 0.0000Reading the output
The first table is the impossibility result, made concrete. Five completely different explanations of one observed pixel, each exact. Two of them share alpha 0.50 with different colours; two others give alpha 0.90 and 0.23. Nothing in a single pixel distinguishes them.
Every practical method breaks the tie with an assumption. Trimaps assume nearby known pixels supply the colours. Green screens assume the background is known. Learned matting assumes objects look like the training set.
The hard mask is exactly right everywhere except the band, and wrong inside it. Error 0.0000 outside, 0.0807 inside. All the difficulty in matting lives in the 24 per cent of pixels that are mixtures. There, a hard mask is a perfect answer to the wrong question.
Row 6 shows the halo, in text. The correct result fades through :::::: before reaching the new background. The hard mask jumps almost straight from foreground to background, with a short :: stub. That abrupt step is what your eye reads as "pasted on".
The wall gradient shows why real matting is hard. With the true per-pixel wall colour, recovery is exact: error 0.0000. That is the green-screen case, where the background is known.
Use a single average wall colour instead and errors appear far from the hair, printed as :::::::::::: where it should read ---------------. The method assigns a small non-zero alpha to pure background pixels, because those pixels are darker than the average it assumed. Background tinting is the classic artefact of naive matting.
Real tools, honestly described
Portrait background removal today is generally a trimap-free network trained on composited data. Two families you will meet on the Hub:
- RMBG-1.4 and RMBG-2.0 from BRIA, at 176 MB and 885 MB.
- BiRefNet, at 444 MB, aimed at high-resolution dichotomous segmentation.
The popular rembg package wraps models of this kind behind one function call.
No output block for those, because they need a network download and an input photo, and the result depends on both. Two things are worth knowing before you use one.
Most of them return a hard-ish mask, not a true alpha matte. The output has soft values at the boundary, but it is trained against binary masks, so the softness is uncertainty rather than genuine coverage. On hair, this shows.
Only the alpha is predicted, not the foreground colour. Correct compositing needs both. Reusing the original pixel colour in the band leaves colour spill from the old background. It is obvious when a green-screen shot moves onto a red backdrop.
Saving the result
# alpha as the fourth channel of a PNG
from PIL import Image
import numpy as np
rgb = (np.random.rand(64, 64, 3) * 255).astype(np.uint8) # your foreground colours
a = (np.random.rand(64, 64) * 255).astype(np.uint8) # your alpha, 0 to 255
Image.fromarray(np.dstack([rgb, a]), mode="RGBA").save("cutout.png")PNG stores alpha; JPEG does not. Saving a cut-out as JPEG discards the transparency and fills it with black, which is a support ticket waiting to happen.
Common mistakes
Scoring matting with IoU. IoU needs a binary mask, so it throws away the soft values that are the whole point. Use mean absolute difference on alpha, and gradient or connectivity errors, which are the standard matting metrics.
Training on composited data only. Almost every matting dataset is made by compositing foregrounds onto random backgrounds. Networks learn to detect compositing artefacts rather than object boundaries, and they transfer poorly to real photographs. This is a known and repeatedly reported issue.
Feeding a badly registered trimap. If the unknown band does not contain the true boundary, no method can find it. Generate the band by dilating and eroding a segmentation mask, and make it wide enough to be certain.
Compositing in sRGB without thinking. Blending is physically correct in linear light. Blending gamma-encoded values darkens soft edges slightly. For most product work this is acceptable; for film work it is not.
Assuming a segmentation model gives you matting. Cleaning up predicted masks makes a mask tidier, not softer. A tidy hard edge is still a hard edge.
Try it yourself
Change B_new from 0.35 to 0.05, so the new background is much darker than the old wall. Re-run. The hard-mask error inside the band grows, because the halo is now a bright fringe against a dark backdrop. The stronger the contrast between old and new background, the more visible the artefact.
What to learn next
- Cleaning up predicted masks — the hard-mask toolbox this lesson argues past.
- How images are stored — channels, alpha and why PNG behaves differently from JPEG.
- Diffusion models — the other route to replacing a background, by generating one.
Researcher — Mathematics and papers.
The compositing equation
Porter and Duff (1984) formalised compositing; matting inverts it. For each pixel $i$ and colour channel:
$$ I_i = \alpha_i F_i + (1 - \alpha_i) B_i, \qquad \alpha_i \in [0, 1] $$
In RGB this is 3 equations in 7 unknowns: $F \in \mathbb{R}^3$, $B \in \mathbb{R}^3$, $\alpha \in \mathbb{R}$. The problem is under-determined by 4 degrees of freedom per pixel, which is the formal version of the first table in the developer section.
Every method supplies the missing constraints:
- Trimap: $\alpha = 1$ and $\alpha = 0$ are given on most pixels; $F$ and $B$ in the unknown band are estimated from nearby known pixels.
- Known background: $B$ is measured, reducing the problem to 3 equations in 4 unknowns.
- Colour-line prior: foreground and background colours lie on a line locally.
- Learned prior: a network trained on ground-truth mattes.
Closed-form matting
Levin, Lischinski, Weiss (2008) derive the most influential classical result. Assume $F$ and $B$ are locally approximately linear in the image colours within a small window. That is the colour-line model. Then $F$ and $B$ can be eliminated analytically, leaving a quadratic cost in $\alpha$ alone:
$$ J(\alpha) = \alpha^{\top} L \alpha + \lambda (\alpha - \hat{\alpha})^{\top} D (\alpha - \hat{\alpha}) $$
where $L$ is the matting Laplacian, a sparse symmetric positive semi-definite matrix whose entries derive from local colour covariance:
$$ L_{ij} = \sum_{k | (i,j) \in w_k} \left( \delta_{ij} - \frac{1}{|w_k|}\left(1 + (I_i - \mu_k)^{\top}\left(\Sigma_k + \frac{\varepsilon}{|w_k|}\mathbb{I}\right)^{-1}(I_j - \mu_k)\right)\right) $$
with $\mu_k$ and $\Sigma_k$ the mean and covariance in window $w_k$. Minimising gives a sparse linear system, so matting reduces to solving $(L + \lambda D)\alpha = \lambda D \hat{\alpha}$.
Two reasons this matters beyond history. The colour-line assumption is stated and falsifiable, which is rare. And the matting Laplacian reappears as a regulariser inside modern deep methods and in unrelated tasks such as image editing and depth refinement.
Learned matting
Xu et al. (2017), Deep Image Matting, established the modern setup. An encoder-decoder takes image and trimap, followed by a refinement stage. The loss combines alpha prediction with a compositional term, penalising $\alpha F + (1-\alpha) B$ against the observed image. They released Composition-1k, built by compositing 493 foregrounds onto COCO backgrounds.
Later directions:
- IndexNet (Lu et al., 2019) learns the unpooling indices, treating upsampling itself as the thing to learn.
- Trimap-free portrait matting (MODNet, and later work) exploits a narrow domain prior to remove the trimap, at the cost of generality.
- Background matting (Sengupta et al., 2020; Lin et al., 2021) asks for a photograph of the empty background instead of a trimap. This is often much easier to obtain — a video call, a fixed security camera — and reduces the unknowns substantially.
- Matting with a mask prompt. A segmentation mask, including one from Segment Anything, can be dilated and eroded into a trimap automatically. This is the standard bridge from segmentation to matting today.
The evaluation problem
The alphamatting.com benchmark (Rhemann et al., 2009) defines the four standard errors: sum of absolute differences, mean squared error, gradient error, and connectivity error. Gradient and connectivity exist because pixel-wise error tolerates a matte that is numerically close and visually wrong.
The deeper problem is data. Ground-truth alpha requires either careful studio capture or manual creation, so nearly all training data is composited. Networks then key on compositing cues — resampling artefacts, mismatched noise, inconsistent lighting — rather than on the boundary physics. Reported numbers on composited test sets overstate real-photograph performance, and this discrepancy is documented across the matting literature. Evaluate on real photographs from your own domain before believing any benchmark number.
Practical notes
- Spill suppression is a separate step. Even with a perfect alpha, $F$ recovered from a green-screen shot carries a green cast. Production pipelines run explicit despill after matting.
- Estimate $F$, not only $\alpha$. Compositing with the original $I$ inside the band reintroduces the old background. Methods that predict $F$ alongside $\alpha$ composite correctly.
- Resolution matters more than in segmentation. The band is a few pixels wide, so downsampling to 512 pixels destroys the signal. Guided or patch-based inference at native resolution is standard.
- Alpha is not a probability. Alpha 0.5 means "half covered", not "50 per cent confident". Thresholding it is a category error.
Papers
- Porter, Duff, Compositing Digital Images, SIGGRAPH 1984
- Chuang, Curless, Salesin, Szeliski, A Bayesian Approach to Digital Matting, CVPR 2001
- Levin, Lischinski, Weiss, A Closed-Form Solution to Natural Image Matting, TPAMI 2008
- Rhemann et al., A Perceptually Motivated Online Benchmark for Image Matting, CVPR 2009
- Xu, Price, Cohen, Huang, Deep Image Matting, CVPR 2017 — arxiv.org/abs/1703.03872
- Lin et al., Real-Time High-Resolution Background Matting, CVPR 2021 — arxiv.org/abs/2012.07810
What to learn next
- Cleaning up predicted masks — the hard-mask toolbox this lesson argues past.
- How images are stored — channels, alpha and why PNG behaves differently from JPEG.
- Diffusion models — the other route to replacing a background, by generating one.