Cleaning up predicted masks
A raw predicted mask is speckled, holed and ragged, and a few classical operations fix most of it — while quietly deleting real small objects.
- 15 min read
- 3 reading levels
- Updated
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A raw mask comes out speckled and holed, and a few old, simple operations clean it up.
The analogy
Think of a bad photocopy. Black specks scattered across the white margin, and tiny white pinholes inside the black text.
You would fix it in two moves. Wipe away the specks, which are small and isolated. Fill the pinholes, which are small and surrounded.
Those are exactly the two operations used on masks, and they are over fifty years old. One removes small bits of foreground. The other seals small bits of background.
Why this is needed
A segmentation model decides each pixel largely on its own. Nothing in the model enforces "objects are usually one connected piece" or "objects do not have random holes".
So the raw output has isolated bright pixels, where noise looked like an object. It has pinholes inside real objects, where the model wavered. Its edges wobble in and out by a pixel.
None of that is a modelling failure exactly. It is the absence of a rule the model was never given.
The toolbox
erode shave a layer off the edge of every shape
dilate add a layer to the edge of every shape
open is erode, then dilate -> small specks vanish, big shapes survive
close is dilate, then erode -> small holes seal, big gaps survive
fill holes -> close off any background not connected to the border
drop small pieces -> delete connected blobs below a size
keep largest -> keep one blob, delete the restOpening works because a speck disappears entirely during the shave, so there is nothing left to grow back. A big shape shrinks and then regrows to almost its original size.
Closing is the same trick reversed, and it seals a pinhole for the same reason.
The knob that matters most
Before any of that, there is the cut-off. The model gives each pixel a score, and something has to decide how high counts as inside.
The usual answer is the middle of the range, and the usual answer is not always right. Nudge it and both the noise level and the object size change together.
Tuning that one number is often worth more than every cleaning step combined, and it is free.
The cost nobody mentions
Every rule that deletes small blobs deletes real small blobs.
If your pictures contain one large object, deleting everything else is a fine rule. If they contain a big lesion and a small one, that rule throws away the small lesion. Your metrics look better and your product gets worse.
This is why cleaning steps must be chosen against the actual job, and measured, not copied from a tutorial.
Where you have already seen this
- Document scanners on phones, removing speckle before saving a page.
- Number plate readers, cleaning up the character shapes before reading them.
- Green screen tools, where the removed background is tidied before compositing.
Remember this
- Opening removes small specks. Closing seals small holes.
- Tune the cut-off before reaching for anything fancier.
- Any small-blob rule deletes real small objects too. Decide on purpose.
What to learn next
- Image matting and background removal — when a hard yes-or-no mask is the wrong output entirely.
- OpenCV — the wider toolbox these operations come from.
- Model evaluation — choosing what to measure before you tune anything.
Developer — Code and libraries.
Setup
pip install numpy scipy opencv-pythonRuns in a second on CPU. The scripts use a synthetic noisy prediction with a known answer, so every step can be scored rather than eyeballed.
Each operation, scored
import numpy as np
from scipy import ndimage as ndi
rng = np.random.default_rng(0)
S = 48
yy, xx = np.mgrid[0:S, 0:S]
# ---- ground truth: one round object plus a genuinely small second one ----
gt = ((yy - 22)**2 + (xx - 20)**2 < 12**2)
gt |= ((yy - 8)**2 + (xx - 40)**2 < 3**2) # small, but real
# ---- what a model actually returns: soft scores, not a clean mask ----
prob = np.clip(np.where(gt, 0.80, 0.12) + rng.normal(0, 0.22, (S, S)), 0, 1)
raw = prob > 0.5
def iou(a, b):
return (a & b).sum() / (a | b).sum()
def holes(m):
return ndi.label(ndi.binary_fill_holes(m) & ~m)[1]
def report(name, m):
print(f"{name:32s} IoU {iou(m, gt):.3f} pieces {ndi.label(m)[1]:>3} holes {holes(m):>3}")
print("threshold sweep on the raw scores:")
for t in (0.3, 0.4, 0.5, 0.6, 0.7):
print(f" threshold {t} IoU {iou(prob > t, gt):.3f} pieces {ndi.label(prob > t)[1]:>3}")
print()
report("raw mask at 0.5", raw)
opened = ndi.binary_opening(raw, np.ones((3, 3))) # erode then dilate: removes specks
report("opening 3x3", opened)
closed = ndi.binary_closing(opened, np.ones((3, 3))) # dilate then erode: seals pinholes
report("then closing 3x3", closed)
filled = ndi.binary_fill_holes(closed)
report("then fill holes", filled)
lab, n = ndi.label(filled)
sizes = ndi.sum_labels(np.ones_like(lab), lab, range(1, n + 1)).astype(int)
print("\nsizes of the surviving pieces:", sorted(sizes.tolist(), reverse=True))
for floor in (10, 30, 60):
keep = np.isin(lab, [i + 1 for i, sz in enumerate(sizes) if sz >= floor])
print(f" drop pieces under {floor:>3} px -> IoU {iou(keep, gt):.3f} pieces kept {ndi.label(keep)[1]}")
biggest = lab == (1 + int(np.argmax(sizes)))
print(f" keep only the largest -> IoU {iou(biggest, gt):.3f} pieces kept 1")
print(f"a floor of 30 deletes a real object of {sizes.min()} pixels. Area floors are not free.")
despeck = np.isin(lab, [i + 1 for i, sz in enumerate(sizes) if sz >= 10])
print("\nraw (left) and cleaned (right), rows 4-16")
for r in range(4, 17):
a = "".join("#" if v else "." for v in raw[r])
b = "".join("#" if v else "." for v in despeck[r])
print(f" {a} {b}")threshold sweep on the raw scores: threshold 0.3 IoU 0.537 pieces 222 threshold 0.4 IoU 0.669 pieces 150 threshold 0.5 IoU 0.771 pieces 79 threshold 0.6 IoU 0.734 pieces 24 threshold 0.7 IoU 0.620 pieces 22 raw mask at 0.5 IoU 0.771 pieces 79 holes 24 opening 3x3 IoU 0.782 pieces 2 holes 5 then closing 3x3 IoU 0.834 pieces 2 holes 1 then fill holes IoU 0.853 pieces 2 holes 0 sizes of the surviving pieces: [373, 23] drop pieces under 10 px -> IoU 0.853 pieces kept 2 drop pieces under 30 px -> IoU 0.803 pieces kept 1 drop pieces under 60 px -> IoU 0.803 pieces kept 1 keep only the largest -> IoU 0.803 pieces kept 1 a floor of 30 deletes a real object of 23 pixels. Area floors are not free. raw (left) and cleaned (right), rows 4-16 ...................#.......#.................... ................................................ ....#..#...........#..........#................. ................................................ ......................................#####..... ......................................#####..... ...............#.........#............#####..... ......................................#####..... ...................#..................#####...#. ......................................#####..... ...........#...........................####..... .......................................####..... .#..................................##.####...#. .......................................####..... .....#..........#########....................... ................#####........................... ...............######..##.##................#... ...............######........................... ............#########.######..........#......... ............#########...###..................... ............#######.#########................... ............#########...###..................... ........#..#####.###.##.###.##....#..#......#... ............#########...###..................... ..........###.#################................. ..........####################..................
Reading the numbers
The threshold sweep is the largest single effect on the page. IoU runs 0.537, 0.669, 0.771, 0.734, 0.620 as the cut-off rises. That is a swing of 0.234 from one number, larger than everything the morphology achieves.
Notice how the piece count collapses from 222 to 22 as the threshold rises. Raising the threshold is itself a de-speckling operation, and a free one.
Opening removed 77 of the 79 pieces and moved IoU by 0.011. That gap is the point. Opening is a structural fix, not an accuracy fix. Two hundred stray pixels barely dent an IoU dominated by a 400-pixel object. They wreck any code that counts objects, measures areas, or exports polygons.
Score the property you care about. If you count objects, count pieces.
Closing and filling gave the real accuracy gain: 0.782 to 0.834 to 0.853. Those pinholes sit inside the object, and each one is a false negative. Sealing them is pure profit here, because the true object genuinely has no holes.
That last clause is a domain assumption. On a mask of a doughnut, or a leaf with insect damage, filling holes destroys the answer.
The size floor demonstrates the cost. At 10 pixels, both pieces survive and IoU is 0.853. At 30, the real 23-pixel object is deleted and IoU falls to 0.803. The rule looks tidy and lost a true object.
If your data contains objects near the floor, no floor is safe. Set it from the smallest object you must detect, not from what makes the output look clean.
Turning a clean mask into polygons
Masks are heavy to store and awkward to edit. Most annotation formats and most map exports want outlines.
import numpy as np, cv2
from scipy import ndimage as ndi
rng = np.random.default_rng(0)
S = 48
yy, xx = np.mgrid[0:S, 0:S]
gt = ((yy - 22)**2 + (xx - 20)**2 < 12**2)
gt |= ((yy - 8)**2 + (xx - 40)**2 < 3**2)
prob = np.clip(np.where(gt, 0.80, 0.12) + rng.normal(0, 0.22, (S, S)), 0, 1)
clean = ndi.binary_fill_holes(
ndi.binary_closing(ndi.binary_opening(prob > 0.5, np.ones((3, 3))), np.ones((3, 3))))
mask = (clean * 255).astype(np.uint8) # OpenCV wants uint8
contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
print("objects found:", len(contours))
for i, c in enumerate(sorted(contours, key=cv2.contourArea, reverse=True)):
approx = cv2.approxPolyDP(c, epsilon=0.01 * cv2.arcLength(c, True), closed=True)
print(f" object {i}: area {cv2.contourArea(c):6.1f} perimeter {cv2.arcLength(c, True):6.1f}"
f" points {len(c):>3} -> {len(approx):>2} after simplification")objects found: 2 object 0: area 330.5 perimeter 90.0 points 43 -> 23 after simplification object 1: area 14.5 perimeter 15.4 points 6 -> 6 after simplification
RETR_EXTERNAL returns outer boundaries only and ignores holes. Use RETR_CCOMP when holes matter, and read the hierarchy array to tell an outer boundary from an inner one.
approxPolyDP is the Ramer-Douglas-Peucker algorithm. The epsilon is a distance tolerance, expressed here as a fraction of the perimeter, so it scales with object size. Larger epsilon means fewer points and more corner-cutting.
The contour area of 330.5 is smaller than the 373 pixels the connected-component count reported. That is expected: contour area measures the polygon through pixel centres, while a pixel count measures whole pixels. Do not mix the two in one report.
Where a conditional random field still fits
A conditional random field rescores every pixel using the model's confidence and the colours of its neighbours. Boundaries then snap to real edges in the photo. Applied after DeepLab v1 and v2, it was worth several points.
It is largely gone from modern pipelines, for reasons worth knowing. Strong networks already produce boundaries that a CRF barely improves. It is slow, adds several hyperparameters, and the standard Python binding is awkward to install.
It still earns its place in two situations. One is a weak or weakly supervised network. The other is work where boundaries must follow colour edges exactly. See image matting and background removal.
Common mistakes
Cleaning before tuning the threshold. The threshold does more, costs nothing, and changes what the cleaning has to do.
Using a large structuring element. A 7x7 opening rounds corners and eats thin structures. Start at 3x3, and increase only with a measurement in hand.
Filling holes on classes that have holes. Doughnuts, windows, tyres, pipes, cell nuclei with vacuoles. Make this decision per class.
Running morphology on a multi-class label map. Erosion and dilation on integer class ids are meaningless, since 3 is not between 2 and 4 in any real sense. Split into per-class binary masks, clean each, then recombine and resolve conflicts.
Reporting only IoU. Piece count, hole count and boundary error each capture something IoU hides. The output above shows a change that improved structure enormously and IoU by 0.011.
Try it yourself
Set the noise standard deviation from 0.22 to 0.35 and re-run. The best threshold shifts, and opening starts to matter far more. Then change the 3x3 structuring element to np.ones((5, 5)) and watch the small real object disappear before any size floor is applied.
What to learn next
- Image matting and background removal — when a hard yes-or-no mask is the wrong output entirely.
- OpenCV — the wider toolbox these operations come from.
- Model evaluation — choosing what to measure before you tune anything.
Researcher — Mathematics and papers.
Morphology, stated properly
For a binary set $A \subseteq \mathbb{Z}^2$ and structuring element $B$:
$$ A \ominus B = {z : B_z \subseteq A} \quad \text{(erosion)}, \qquad A \oplus B = \bigcup_{b \in B} A_b \quad \text{(dilation)} $$
$$ A \circ B = (A \ominus B) \oplus B \quad \text{(opening)}, \qquad A \bullet B = (A \oplus B) \ominus B \quad \text{(closing)} $$
Properties that justify the operations rather than only describing them:
- Idempotence. $(A \circ B) \circ B = A \circ B$. Applying opening twice changes nothing, so iterating is pointless.
- Anti-extensivity and extensivity. $A \circ B \subseteq A \subseteq A \bullet B$. Opening only removes, closing only adds.
- Duality. $(A \circ B)^c = A^c \bullet \check{B}$. Opening the foreground is closing the background with the reflected element.
- Granulometry. Opening with elements of increasing size gives a size distribution of the shapes present. This is the principled way to choose a size floor, rather than by eye: plot area removed against element size and look for the knee.
Erosion and dilation are separable for rectangular elements, so a $k \times k$ operation costs $O(2k)$ per pixel rather than $O(k^2)$. scipy.ndimage and OpenCV both exploit this.
Thresholding is a decision problem
The 0.5 default assumes equal costs for false positives and false negatives and a calibrated model. Both assumptions usually fail.
With cost $c_{FP}$ and $c_{FN}$, the Bayes-optimal threshold on a calibrated probability is $t^* = \frac{c_{FP}}{c_{FP} + c_{FN}}$. Dice-trained networks are typically poorly calibrated (Mehrtash et al., 2020), which breaks that formula. It is a further argument for keeping a cross-entropy term, as discussed in Dice and other segmentation losses.
In practice, sweep the threshold on validation data against the metric you actually report, and fix it per class. Reporting a single dataset-level threshold when class frequencies differ by orders of magnitude leaves easy accuracy on the table.
Conditional random fields
Krähenbühl and Koltun (2011) define a fully connected pairwise CRF over all pixel pairs with Gaussian edge potentials:
$$ E(x) = \sum_i \psi_u(x_i) + \sum_{i<j} \psi_p(x_i, x_j) $$
$$ \psi_p(x_i, x_j) = \mu(x_i, x_j)\left[ w_1 \exp\left(-\frac{|p_i - p_j|^2}{2\theta_\alpha^2} - \frac{|I_i - I_j|^2}{2\theta_\beta^2}\right) + w_2 \exp\left(-\frac{|p_i - p_j|^2}{2\theta_\gamma^2}\right)\right] $$
where $p$ is position and $I$ is colour. The contribution is the inference method: mean-field updates whose message passing is a high-dimensional Gaussian filter, computable in linear time with the permutohedral lattice. That turned a hopeless $O(N^2)$ model into a practical one.
Zheng et al. (2015) unrolled the mean-field iterations as recurrent layers, making the CRF trainable end to end. Both are worth reading as examples of turning an inference procedure into an architecture, even though neither is common in current pipelines.
Alternatives that usually beat post-processing
Post-processing patches a symptom. Three approaches attack the cause.
- Test-time augmentation. Average predictions over flips and scales. Consistently worth a point or so of IoU, at a linear cost in inference passes. It reduces speckle for the same reason self-training does: inconsistent errors cancel.
- Higher-resolution prediction. Output stride 8 rather than 32, or PointRend-style adaptive refinement at boundaries. More expensive, and it fixes the boundary rather than smoothing it.
- Boundary-aware losses. Penalise distance from the true boundary during training, as in Kervadec et al. (2019).
A good rule: if morphology improves your metric substantially, the model is underfitting the boundary and you should fix the model. If morphology improves structure but barely moves the metric, keep it, because structure is what downstream code consumes.
Evaluation for the properties morphology fixes
- Boundary IoU (Cheng et al., 2021) restricts IoU to a band of width $d$ around the boundary, so large objects stop masking boundary error.
- Normalised surface distance / surface Dice with a tolerance, standard in medical challenges because tolerance can be stated in millimetres.
- Counting metrics. Segment count error, and the split-and-merge rates from the connectomics literature, which measure exactly what speckle and pinholes damage.
Papers
- Krähenbühl, Koltun, Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials, NeurIPS 2011 — arxiv.org/abs/1210.5644
- Zheng et al., Conditional Random Fields as Recurrent Neural Networks, ICCV 2015 — arxiv.org/abs/1502.03240
- Cheng, Girshick, Dollár, Berg, Kirillov, Boundary IoU: Improving Object-Centric Image Segmentation Evaluation, CVPR 2021 — arxiv.org/abs/2103.16562
- Mehrtash et al., Confidence Calibration and Predictive Uncertainty Estimation for Deep Medical Image Segmentation, IEEE TMI 2020 — arxiv.org/abs/1911.13273
What to learn next
- Image matting and background removal — when a hard yes-or-no mask is the wrong output entirely.
- OpenCV — the wider toolbox these operations come from.
- Model evaluation — choosing what to measure before you tune anything.