Segmentation in Depth

Cleaning up predicted masks

A raw predicted mask is speckled, holed and ragged, and a few classical operations fix most of it — while quietly deleting real small objects.

Read these first

On this page 9
  1. The short answer
  2. The analogy
  3. Why this is needed
  4. The toolbox
  5. The knob that matters most
  6. The cost nobody mentions
  7. Where you have already seen this
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A raw mask comes out speckled and holed, and a few old, simple operations clean it up.

The analogy

Think of a bad photocopy. Black specks scattered across the white margin, and tiny white pinholes inside the black text.

You would fix it in two moves. Wipe away the specks, which are small and isolated. Fill the pinholes, which are small and surrounded.

Those are exactly the two operations used on masks, and they are over fifty years old. One removes small bits of foreground. The other seals small bits of background.

Why this is needed

A segmentation model decides each pixel largely on its own. Nothing in the model enforces "objects are usually one connected piece" or "objects do not have random holes".

So the raw output has isolated bright pixels, where noise looked like an object. It has pinholes inside real objects, where the model wavered. Its edges wobble in and out by a pixel.

None of that is a modelling failure exactly. It is the absence of a rule the model was never given.

The toolbox

   erode    shave a layer off the edge of every shape
   dilate   add a layer to the edge of every shape

   open   is erode, then dilate  -> small specks vanish, big shapes survive
   close  is dilate, then erode  -> small holes seal, big gaps survive

   fill holes        -> close off any background not connected to the border
   drop small pieces -> delete connected blobs below a size
   keep largest      -> keep one blob, delete the rest

Opening works because a speck disappears entirely during the shave, so there is nothing left to grow back. A big shape shrinks and then regrows to almost its original size.

Closing is the same trick reversed, and it seals a pinhole for the same reason.

The knob that matters most

Before any of that, there is the cut-off. The model gives each pixel a score, and something has to decide how high counts as inside.

The usual answer is the middle of the range, and the usual answer is not always right. Nudge it and both the noise level and the object size change together.

Tuning that one number is often worth more than every cleaning step combined, and it is free.

The cost nobody mentions

Every rule that deletes small blobs deletes real small blobs.

If your pictures contain one large object, deleting everything else is a fine rule. If they contain a big lesion and a small one, that rule throws away the small lesion. Your metrics look better and your product gets worse.

This is why cleaning steps must be chosen against the actual job, and measured, not copied from a tutorial.

Where you have already seen this

  • Document scanners on phones, removing speckle before saving a page.
  • Number plate readers, cleaning up the character shapes before reading them.
  • Green screen tools, where the removed background is tidied before compositing.

Remember this

  • Opening removes small specks. Closing seals small holes.
  • Tune the cut-off before reaching for anything fancier.
  • Any small-blob rule deletes real small objects too. Decide on purpose.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scipy opencv-python

Runs in a second on CPU. The scripts use a synthetic noisy prediction with a known answer, so every step can be scored rather than eyeballed.

Each operation, scored

clean_masks.py
import numpy as np
from scipy import ndimage as ndi

rng = np.random.default_rng(0)
S = 48
yy, xx = np.mgrid[0:S, 0:S]

# ---- ground truth: one round object plus a genuinely small second one ----
gt = ((yy - 22)**2 + (xx - 20)**2 < 12**2)
gt |= ((yy - 8)**2 + (xx - 40)**2 < 3**2)          # small, but real

# ---- what a model actually returns: soft scores, not a clean mask ----
prob = np.clip(np.where(gt, 0.80, 0.12) + rng.normal(0, 0.22, (S, S)), 0, 1)
raw = prob > 0.5

def iou(a, b):
    return (a & b).sum() / (a | b).sum()

def holes(m):
    return ndi.label(ndi.binary_fill_holes(m) & ~m)[1]

def report(name, m):
    print(f"{name:32s} IoU {iou(m, gt):.3f}   pieces {ndi.label(m)[1]:>3}   holes {holes(m):>3}")

print("threshold sweep on the raw scores:")
for t in (0.3, 0.4, 0.5, 0.6, 0.7):
    print(f"   threshold {t}   IoU {iou(prob > t, gt):.3f}   pieces {ndi.label(prob > t)[1]:>3}")

print()
report("raw mask at 0.5", raw)
opened = ndi.binary_opening(raw, np.ones((3, 3)))      # erode then dilate: removes specks
report("opening 3x3", opened)
closed = ndi.binary_closing(opened, np.ones((3, 3)))   # dilate then erode: seals pinholes
report("then closing 3x3", closed)
filled = ndi.binary_fill_holes(closed)
report("then fill holes", filled)

lab, n = ndi.label(filled)
sizes = ndi.sum_labels(np.ones_like(lab), lab, range(1, n + 1)).astype(int)
print("\nsizes of the surviving pieces:", sorted(sizes.tolist(), reverse=True))
for floor in (10, 30, 60):
    keep = np.isin(lab, [i + 1 for i, sz in enumerate(sizes) if sz >= floor])
    print(f"   drop pieces under {floor:>3} px  ->  IoU {iou(keep, gt):.3f}   pieces kept {ndi.label(keep)[1]}")
biggest = lab == (1 + int(np.argmax(sizes)))
print(f"   keep only the largest    ->  IoU {iou(biggest, gt):.3f}   pieces kept 1")
print(f"a floor of 30 deletes a real object of {sizes.min()} pixels. Area floors are not free.")

despeck = np.isin(lab, [i + 1 for i, sz in enumerate(sizes) if sz >= 10])
print("\nraw (left) and cleaned (right), rows 4-16")
for r in range(4, 17):
    a = "".join("#" if v else "." for v in raw[r])
    b = "".join("#" if v else "." for v in despeck[r])
    print(f"   {a}  {b}")
Output
threshold sweep on the raw scores:
   threshold 0.3   IoU 0.537   pieces 222
   threshold 0.4   IoU 0.669   pieces 150
   threshold 0.5   IoU 0.771   pieces  79
   threshold 0.6   IoU 0.734   pieces  24
   threshold 0.7   IoU 0.620   pieces  22

raw mask at 0.5                  IoU 0.771   pieces  79   holes  24
opening 3x3                      IoU 0.782   pieces   2   holes   5
then closing 3x3                 IoU 0.834   pieces   2   holes   1
then fill holes                  IoU 0.853   pieces   2   holes   0

sizes of the surviving pieces: [373, 23]
   drop pieces under  10 px  ->  IoU 0.853   pieces kept 2
   drop pieces under  30 px  ->  IoU 0.803   pieces kept 1
   drop pieces under  60 px  ->  IoU 0.803   pieces kept 1
   keep only the largest    ->  IoU 0.803   pieces kept 1
a floor of 30 deletes a real object of 23 pixels. Area floors are not free.

raw (left) and cleaned (right), rows 4-16
   ...................#.......#....................  ................................................
   ....#..#...........#..........#.................  ................................................
   ......................................#####.....  ......................................#####.....
   ...............#.........#............#####.....  ......................................#####.....
   ...................#..................#####...#.  ......................................#####.....
   ...........#...........................####.....  .......................................####.....
   .#..................................##.####...#.  .......................................####.....
   .....#..........#########.......................  ................#####...........................
   ...............######..##.##................#...  ...............######...........................
   ............#########.######..........#.........  ............#########...###.....................
   ............#######.#########...................  ............#########...###.....................
   ........#..#####.###.##.###.##....#..#......#...  ............#########...###.....................
   ..........###.#################.................  ..........####################..................

Reading the numbers

The threshold sweep is the largest single effect on the page. IoU runs 0.537, 0.669, 0.771, 0.734, 0.620 as the cut-off rises. That is a swing of 0.234 from one number, larger than everything the morphology achieves.

Notice how the piece count collapses from 222 to 22 as the threshold rises. Raising the threshold is itself a de-speckling operation, and a free one.

Opening removed 77 of the 79 pieces and moved IoU by 0.011. That gap is the point. Opening is a structural fix, not an accuracy fix. Two hundred stray pixels barely dent an IoU dominated by a 400-pixel object. They wreck any code that counts objects, measures areas, or exports polygons.

Score the property you care about. If you count objects, count pieces.

Closing and filling gave the real accuracy gain: 0.782 to 0.834 to 0.853. Those pinholes sit inside the object, and each one is a false negative. Sealing them is pure profit here, because the true object genuinely has no holes.

That last clause is a domain assumption. On a mask of a doughnut, or a leaf with insect damage, filling holes destroys the answer.

The size floor demonstrates the cost. At 10 pixels, both pieces survive and IoU is 0.853. At 30, the real 23-pixel object is deleted and IoU falls to 0.803. The rule looks tidy and lost a true object.

If your data contains objects near the floor, no floor is safe. Set it from the smallest object you must detect, not from what makes the output look clean.

Turning a clean mask into polygons

Masks are heavy to store and awkward to edit. Most annotation formats and most map exports want outlines.

mask_to_polygon.py
import numpy as np, cv2
from scipy import ndimage as ndi

rng = np.random.default_rng(0)
S = 48
yy, xx = np.mgrid[0:S, 0:S]
gt = ((yy - 22)**2 + (xx - 20)**2 < 12**2)
gt |= ((yy - 8)**2 + (xx - 40)**2 < 3**2)
prob = np.clip(np.where(gt, 0.80, 0.12) + rng.normal(0, 0.22, (S, S)), 0, 1)
clean = ndi.binary_fill_holes(
    ndi.binary_closing(ndi.binary_opening(prob > 0.5, np.ones((3, 3))), np.ones((3, 3))))

mask = (clean * 255).astype(np.uint8)                      # OpenCV wants uint8
contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
print("objects found:", len(contours))
for i, c in enumerate(sorted(contours, key=cv2.contourArea, reverse=True)):
    approx = cv2.approxPolyDP(c, epsilon=0.01 * cv2.arcLength(c, True), closed=True)
    print(f"   object {i}: area {cv2.contourArea(c):6.1f}  perimeter {cv2.arcLength(c, True):6.1f}"
          f"  points {len(c):>3} -> {len(approx):>2} after simplification")
Output
objects found: 2
   object 0: area  330.5  perimeter   90.0  points  43 -> 23 after simplification
   object 1: area   14.5  perimeter   15.4  points   6 ->  6 after simplification

RETR_EXTERNAL returns outer boundaries only and ignores holes. Use RETR_CCOMP when holes matter, and read the hierarchy array to tell an outer boundary from an inner one.

approxPolyDP is the Ramer-Douglas-Peucker algorithm. The epsilon is a distance tolerance, expressed here as a fraction of the perimeter, so it scales with object size. Larger epsilon means fewer points and more corner-cutting.

The contour area of 330.5 is smaller than the 373 pixels the connected-component count reported. That is expected: contour area measures the polygon through pixel centres, while a pixel count measures whole pixels. Do not mix the two in one report.

Where a conditional random field still fits

A conditional random field rescores every pixel using the model's confidence and the colours of its neighbours. Boundaries then snap to real edges in the photo. Applied after DeepLab v1 and v2, it was worth several points.

It is largely gone from modern pipelines, for reasons worth knowing. Strong networks already produce boundaries that a CRF barely improves. It is slow, adds several hyperparameters, and the standard Python binding is awkward to install.

It still earns its place in two situations. One is a weak or weakly supervised network. The other is work where boundaries must follow colour edges exactly. See image matting and background removal.

Common mistakes

Cleaning before tuning the threshold. The threshold does more, costs nothing, and changes what the cleaning has to do.

Using a large structuring element. A 7x7 opening rounds corners and eats thin structures. Start at 3x3, and increase only with a measurement in hand.

Filling holes on classes that have holes. Doughnuts, windows, tyres, pipes, cell nuclei with vacuoles. Make this decision per class.

Running morphology on a multi-class label map. Erosion and dilation on integer class ids are meaningless, since 3 is not between 2 and 4 in any real sense. Split into per-class binary masks, clean each, then recombine and resolve conflicts.

Reporting only IoU. Piece count, hole count and boundary error each capture something IoU hides. The output above shows a change that improved structure enormously and IoU by 0.011.

Try it yourself

Set the noise standard deviation from 0.22 to 0.35 and re-run. The best threshold shifts, and opening starts to matter far more. Then change the 3x3 structuring element to np.ones((5, 5)) and watch the small real object disappear before any size floor is applied.

What to learn next

Researcher — Mathematics and papers.

Morphology, stated properly

For a binary set $A \subseteq \mathbb{Z}^2$ and structuring element $B$:

$$ A \ominus B = {z : B_z \subseteq A} \quad \text{(erosion)}, \qquad A \oplus B = \bigcup_{b \in B} A_b \quad \text{(dilation)} $$

$$ A \circ B = (A \ominus B) \oplus B \quad \text{(opening)}, \qquad A \bullet B = (A \oplus B) \ominus B \quad \text{(closing)} $$

Properties that justify the operations rather than only describing them:

  • Idempotence. $(A \circ B) \circ B = A \circ B$. Applying opening twice changes nothing, so iterating is pointless.
  • Anti-extensivity and extensivity. $A \circ B \subseteq A \subseteq A \bullet B$. Opening only removes, closing only adds.
  • Duality. $(A \circ B)^c = A^c \bullet \check{B}$. Opening the foreground is closing the background with the reflected element.
  • Granulometry. Opening with elements of increasing size gives a size distribution of the shapes present. This is the principled way to choose a size floor, rather than by eye: plot area removed against element size and look for the knee.

Erosion and dilation are separable for rectangular elements, so a $k \times k$ operation costs $O(2k)$ per pixel rather than $O(k^2)$. scipy.ndimage and OpenCV both exploit this.

Thresholding is a decision problem

The 0.5 default assumes equal costs for false positives and false negatives and a calibrated model. Both assumptions usually fail.

With cost $c_{FP}$ and $c_{FN}$, the Bayes-optimal threshold on a calibrated probability is $t^* = \frac{c_{FP}}{c_{FP} + c_{FN}}$. Dice-trained networks are typically poorly calibrated (Mehrtash et al., 2020), which breaks that formula. It is a further argument for keeping a cross-entropy term, as discussed in Dice and other segmentation losses.

In practice, sweep the threshold on validation data against the metric you actually report, and fix it per class. Reporting a single dataset-level threshold when class frequencies differ by orders of magnitude leaves easy accuracy on the table.

Conditional random fields

Krähenbühl and Koltun (2011) define a fully connected pairwise CRF over all pixel pairs with Gaussian edge potentials:

$$ E(x) = \sum_i \psi_u(x_i) + \sum_{i<j} \psi_p(x_i, x_j) $$

$$ \psi_p(x_i, x_j) = \mu(x_i, x_j)\left[ w_1 \exp\left(-\frac{|p_i - p_j|^2}{2\theta_\alpha^2} - \frac{|I_i - I_j|^2}{2\theta_\beta^2}\right) + w_2 \exp\left(-\frac{|p_i - p_j|^2}{2\theta_\gamma^2}\right)\right] $$

where $p$ is position and $I$ is colour. The contribution is the inference method: mean-field updates whose message passing is a high-dimensional Gaussian filter, computable in linear time with the permutohedral lattice. That turned a hopeless $O(N^2)$ model into a practical one.

Zheng et al. (2015) unrolled the mean-field iterations as recurrent layers, making the CRF trainable end to end. Both are worth reading as examples of turning an inference procedure into an architecture, even though neither is common in current pipelines.

Alternatives that usually beat post-processing

Post-processing patches a symptom. Three approaches attack the cause.

  • Test-time augmentation. Average predictions over flips and scales. Consistently worth a point or so of IoU, at a linear cost in inference passes. It reduces speckle for the same reason self-training does: inconsistent errors cancel.
  • Higher-resolution prediction. Output stride 8 rather than 32, or PointRend-style adaptive refinement at boundaries. More expensive, and it fixes the boundary rather than smoothing it.
  • Boundary-aware losses. Penalise distance from the true boundary during training, as in Kervadec et al. (2019).

A good rule: if morphology improves your metric substantially, the model is underfitting the boundary and you should fix the model. If morphology improves structure but barely moves the metric, keep it, because structure is what downstream code consumes.

Evaluation for the properties morphology fixes

  • Boundary IoU (Cheng et al., 2021) restricts IoU to a band of width $d$ around the boundary, so large objects stop masking boundary error.
  • Normalised surface distance / surface Dice with a tolerance, standard in medical challenges because tolerance can be stated in millimetres.
  • Counting metrics. Segment count error, and the split-and-merge rates from the connectomics literature, which measure exactly what speckle and pinholes damage.

Papers

  • Krähenbühl, Koltun, Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials, NeurIPS 2011 — arxiv.org/abs/1210.5644
  • Zheng et al., Conditional Random Fields as Recurrent Neural Networks, ICCV 2015 — arxiv.org/abs/1502.03240
  • Cheng, Girshick, Dollár, Berg, Kirillov, Boundary IoU: Improving Object-Centric Image Segmentation Evaluation, CVPR 2021 — arxiv.org/abs/2103.16562
  • Mehrtash et al., Confidence Calibration and Predictive Uncertainty Estimation for Deep Medical Image Segmentation, IEEE TMI 2020 — arxiv.org/abs/1911.13273

What to learn next