Semantic, instance and panoptic segmentation
Three different jobs hide behind the word segmentation, and picking the wrong one is the most expensive mistake in a vision project.
- 13 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Three different jobs are called segmentation, and each answers a different question.
The analogy
Take a printed group photo and a set of marker pens. Your first job is to shade every region by what it is. Sky in blue, road in grey, people in green.
You never lift the pen to ask how many people there are. Now the second job. Circle each person separately, and write a number beside each circle.
You ignore the sky and the road completely, because sky cannot be counted. The third job is both at once, under two strict rules. Shade every part of the paper, and shade no part twice.
Those three jobs have names. Semantic segmentation shades by what a thing is. Instance segmentation circles and numbers each object. Panoptic segmentation does both, with no gaps and no double-shading.
Why the split exists
Some parts of a picture can be counted. People, cars, dogs, plates. Vision researchers call these things, meaning countable objects with a clear edge.
Other parts cannot be counted. Sky, road, grass, wall. These are called stuff, meaning material that spreads out with no natural unit.
Asking how many skies are in a photo is not a sensible question. Asking how much sky is in it makes perfect sense.
That one difference is why the field split into three tasks instead of one.
How each one answers
photo
|
+--> semantic : every pixel gets a label
| "sky, sky, road, car, car, road"
| tells you WHAT. Cannot tell you HOW MANY.
|
+--> instance : each countable object gets its own outline
| "car A here, car B there, person A there"
| tells you HOW MANY. Says nothing about sky.
|
+--> panoptic : a label AND an object identity per pixel
"sky, car A, car B, person A"
tells you both. Every pixel used exactly once.Where you have already seen this
Portrait mode on your phone is semantic segmentation in disguise. It shades "person" against "everything else", then blurs the second part.
Self-driving demos show panoptic segmentation. Road and sky get shaded, and every car and pedestrian gets its own colour and identity.
Photo editors that let you tap one person and move them use instance segmentation. Tapping needs a separate outline per person.
The trap that costs projects months
A team wants to count damaged tiles on a roof from a drone photo. They train semantic segmentation, because that is the tutorial everybody starts with.
The model works beautifully. It shades every damaged patch. Then somebody asks how many damaged tiles there are, and the answer is not in the output.
Two damaged tiles touching each other became one blob. Nobody can tell two from one without rebuilding the pipeline.
Decide first: does anyone need to count? If the answer is yes, semantic segmentation will never get you there.
What is honestly hard here
Which classes count as things and which count as stuff is a decision somebody made. It is not a fact about the world.
Are trees countable? In a park, maybe. In a forest, nobody wants ten thousand tree outlines.
Different datasets answer this differently, and a model trained on one answer behaves oddly under the other. Read that twice if it feels slippery — the feeling is correct, because the boundary really is a judgement call.
Remember this
- Semantic shades by category, and cannot count.
- Instance outlines and counts objects, and ignores background material.
- Panoptic does both, giving every pixel exactly one owner.
What to learn next
- U-Net — the architecture that made semantic segmentation practical.
- Image segmentation — the shorter overview, if this lesson moved fast.
- Object detection — boxes first, masks second, and where instance segmentation grew from.
Developer — Code and libraries.
Setup
pip install numpyThe three tasks differ in what their output can express, not in how hard they look. Building all three by hand for one tiny scene makes the difference impossible to miss.
One scene, three outputs
import numpy as np
H, W = 8, 12
CLASSES = {0: "sky", 1: "road", 2: "car", 3: "person"}
IS_THING = {0: False, 1: False, 2: True, 3: True}
semantic = np.zeros((H, W), dtype=int)
semantic[3:, :] = 1 # road below the horizon
instance = np.zeros((H, W), dtype=int) # 0 means "not part of any countable object"
def put(cls, inst_id, rows, cols):
semantic[rows, cols] = cls
instance[rows, cols] = inst_id
put(2, 1, slice(3, 6), slice(1, 5)) # car 1
put(2, 2, slice(3, 6), slice(6, 9)) # car 2
put(3, 3, slice(4, 7), slice(10, 12)) # person
GLYPH = {0: ".", 1: "-", 2: "C", 3: "P"}
def show(grid, mapping, title):
print(title)
for row in grid:
print(" " + " ".join(mapping[v] for v in row))
show(semantic, GLYPH, "semantic map (. sky, - road, C car, P person)")
print()
show(instance, {i: str(i) for i in range(4)}, "instance map (0 = stuff, no instance id)")
print("\n--- the same three questions, asked of each output ---")
print("how many cars? semantic map says:", "cannot tell")
print("how many cars? instance map says:",
len({i for i in np.unique(instance) if i != 0 and semantic[instance == i][0] == 2}))
print("how much of the frame is road? semantic map says:",
f"{(semantic == 1).mean():.1%}")
print("how much of the frame is road? instance map says:", "nothing, road has no instances")
# a panoptic label is a (class, instance) pair, one per pixel
panoptic = np.stack([semantic, instance], axis=-1)
segments = {}
for cls, inst in panoptic.reshape(-1, 2):
segments[(cls, inst)] = segments.get((cls, inst), 0) + 1
print("\npanoptic segments, one row per segment:")
print(f"{'class':8s} {'inst id':8s} {'pixels':7s} {'kind'}")
for (cls, inst), n in sorted(segments.items()):
kind = "thing" if IS_THING[cls] else "stuff"
print(f"{CLASSES[cls]:8s} {inst:<8d} {n:<7d} {kind}")
print("\nevery pixel belongs to exactly one segment:",
sum(segments.values()) == H * W)
# the rule instance segmentation is allowed to break
mask_a = np.zeros((H, W), bool); mask_a[3:6, 1:5] = True
mask_b = np.zeros((H, W), bool); mask_b[3:6, 4:9] = True # a wrong, overlapping guess
overlap = (mask_a & mask_b).sum()
print(f"\ntwo instance masks may overlap: {overlap} pixels claimed twice")
print("a panoptic map cannot store that, so a merge step must pick a winner")semantic map (. sky, - road, C car, P person) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . - C C C C - C C C - - - - C C C C - C C C - P P - C C C C - C C C - P P - - - - - - - - - - P P - - - - - - - - - - - - instance map (0 = stuff, no instance id) 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 0 2 2 2 0 0 0 0 1 1 1 1 0 2 2 2 0 3 3 0 1 1 1 1 0 2 2 2 0 3 3 0 0 0 0 0 0 0 0 0 0 3 3 0 0 0 0 0 0 0 0 0 0 0 0 --- the same three questions, asked of each output --- how many cars? semantic map says: cannot tell how many cars? instance map says: 2 how much of the frame is road? semantic map says: 34.4% how much of the frame is road? instance map says: nothing, road has no instances panoptic segments, one row per segment: class inst id pixels kind sky 0 36 stuff road 0 33 stuff car 1 12 thing car 2 9 thing person 3 6 thing every pixel belongs to exactly one segment: True two instance masks may overlap: 3 pixels claimed twice a panoptic map cannot store that, so a merge step must pick a winner
Reading the output
The semantic map genuinely cannot count. The two cars sit in separate columns here, so your eye separates them. Move them one column closer and they become one blob of class car. No post-processing recovers the split reliably, because the information was never stored.
The instance map genuinely cannot measure area. It holds no id for road, so the area question returns nothing. Many instance datasets do not label stuff at all.
The panoptic table is the shape you usually want. Each row is a segment: a class, an instance id for things, and a pixel count. Stuff classes carry id 0, meaning one region rather than a counted object.
The last two lines are why panoptic merging is a real algorithm. Instance heads produce masks independently, so two masks can claim one pixel. Panoptic output forbids that, so something must decide the winner. Panoptic segmentation covers that decision in detail.
Output formats you will meet
| Task | Stored as | Typical file |
|---|---|---|
| Semantic | one integer per pixel | PNG, single channel or palette |
| Instance | a list of masks, each with a class and a score | COCO JSON with RLE masks |
| Panoptic | a PNG of segment ids plus a JSON segment list | COCO panoptic format |
RLE means run-length encoding. Instead of every pixel you store "300 zeros, 12 ones, 44 zeros". Masks compress enormously that way, which is why COCO uses it.
Common mistakes
Saving a label map as JPEG. JPEG is lossy, so class id 7 becomes 6 or 8 along every edge. Label maps must be PNG. This bug is silent and it poisons training.
Reading a palette PNG as RGB. Many semantic datasets store class ids in a paletted PNG. Opening it with .convert("RGB") hands you colours, not ids. Open it with no conversion, then call np.array.
Treating instance ids as classes. Instance id 3 is not class 3. Keep them in separate channels, files or fields.
Assuming instance masks are disjoint. They are not. Code that flattens them by overwriting loses whichever mask is written first.
Counting blobs in a semantic map and calling it instance segmentation. Connected-component labelling works only while objects never touch. In crowds, shelves and car parks, they always touch.
Try it yourself
Change car 2 to slice(5, 9) so the two cars touch. Re-run. The semantic map still looks correct, and a blob count would now report one car. The panoptic table still reports two, because the ids were never merged.
What to learn next
- U-Net — the architecture that made semantic segmentation practical.
- Image segmentation — the shorter overview, if this lesson moved fast.
- Object detection — boxes first, masks second, and where instance segmentation grew from.
Researcher — Mathematics and papers.
The formal definitions
Let $\mathcal{L} = {0, \dots, K-1}$ be the label set, partitioned into stuff classes $\mathcal{L}^{st}$ and thing classes $\mathcal{L}^{th}$, with $\mathcal{L}^{st} \cap \mathcal{L}^{th} = \varnothing$. Let $\Omega \subset \mathbb{Z}^2$ be the pixel grid.
Semantic segmentation learns $f: \Omega \to \mathcal{L}$. The output is a function on pixels, so two objects of one class are indistinguishable by construction.
Instance segmentation learns a set ${(m_i, c_i, s_i)}_{i=1}^{N}$, where $m_i \subset \Omega$ is a mask, $c_i \in \mathcal{L}^{th}$ its class and $s_i \in [0,1]$ its confidence. Neither $\bigcup_i m_i = \Omega$ nor $m_i \cap m_j = \varnothing$ is required, and $N$ is predicted rather than fixed.
Panoptic segmentation (Kirillov et al., 2019) learns $f: \Omega \to \mathcal{L} \times \mathbb{N}$, mapping each pixel to a class and an instance id, with the id ignored when $c \in \mathcal{L}^{st}$. The output partitions $\Omega$: segments are non-overlapping and cover everything except an explicit void label.
That partition constraint is the entire source of difficulty. It turns set prediction into set prediction plus a global consistency requirement.
Why the metrics are not comparable
- Semantic: $\text{mIoU} = \frac{1}{K}\sum_c \frac{TP_c}{TP_c + FP_c + FN_c}$, accumulated over pixel counts across the whole dataset. Class-balanced, resolution-sensitive, blind to instances.
- Instance: mask average precision, averaged over IoU thresholds $0.5{:}0.05{:}0.95$ on COCO. Rank-sensitive, so calibration of $s_i$ matters as much as mask quality.
- Panoptic: $\text{PQ} = \text{SQ} \times \text{RQ}$, derived in panoptic segmentation.
A model can raise mIoU while its AP falls. Splitting one merged blob into two correct instances leaves mIoU untouched and raises AP substantially. One number hides that.
The thing/stuff boundary is a dataset convention
COCO panoptic uses 80 thing classes and 53 stuff classes. Cityscapes evaluates 19 classes of which 8 are things. ADE20K's 150 classes mix both without a formal split, so its semantic and panoptic protocols handle labels differently.
Three consequences bite in practice:
- Amodal versus modal masks. Standard datasets label only visible pixels, called modal. A car behind a pole becomes two disconnected components sharing one instance id, so code that assumes connectivity breaks. Amodal datasets (Zhu et al., 2017, Semantic Amodal Segmentation) label the occluded extent instead, and the two conventions are not interchangeable.
- Crowd regions. COCO marks dense groups with
iscrowd=1and one region rather than per-instance masks. Evaluation ignores detections inside them. Training code that treats crowds as ordinary instances learns to merge people. - Part hierarchies. A person contains a face contains an eye. Flat label sets cannot express that, which motivated part-aware panoptic segmentation (de Geus et al., 2021).
Where the architectures converged
The three tasks were built by three separate communities.
| Task | Lineage |
|---|---|
| Semantic | FCN (Long et al., 2015) → U-Net → DeepLab → SegFormer |
| Instance | R-CNN → Faster R-CNN → Mask R-CNN → YOLACT, SOLO |
| Panoptic | Panoptic FPN (Kirillov et al., 2019) → UPSNet → Panoptic-DeepLab |
Then they collapsed into one. MaskFormer (Cheng et al., 2021) reframed all three as mask classification. Predict a fixed set of binary masks, give each one class label, then post-process differently per task. Mask2Former refined it and matched or beat the task-specific architectures on all three.
The lesson generalises past segmentation. Per-pixel classification was an architectural commitment that quietly ruled out counting. Set prediction makes no such commitment, so one architecture covers every task.
Papers
- Long, Shelhamer, Darrell, Fully Convolutional Networks for Semantic Segmentation, CVPR 2015 — arxiv.org/abs/1411.4038
- He, Gkioxari, Dollár, Girshick, Mask R-CNN, ICCV 2017 — arxiv.org/abs/1703.06870
- Kirillov, He, Girshick, Rother, Dollár, Panoptic Segmentation, CVPR 2019 — arxiv.org/abs/1801.00868
- Cheng, Schwing, Kirillov, Per-Pixel Classification is Not All You Need for Semantic Segmentation, NeurIPS 2021 — arxiv.org/abs/2107.06278
- Cheng, Misra, Schwing, Kirillov, Girdhar, Masked-attention Mask Transformer for Universal Image Segmentation, CVPR 2022 — arxiv.org/abs/2112.01527
What to learn next
- U-Net — the architecture that made semantic segmentation practical.
- Image segmentation — the shorter overview, if this lesson moved fast.
- Object detection — boxes first, masks second, and where instance segmentation grew from.