Panoptic segmentation
Panoptic segmentation gives every pixel exactly one owner, which means overlapping predictions must be merged and a new score is needed to grade the result.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Panoptic segmentation covers the whole picture with pieces that never overlap and never leave a gap.
The analogy
Finish a jigsaw puzzle on a table. When it is done, three things are true at once.
Every part of the picture is covered. No piece sits on top of another piece. No bare table shows through anywhere.
That is exactly what a panoptic output must be. Every pixel belongs to one region, and to one region only.
Now imagine two people building the same puzzle from opposite ends, each holding a piece for the same gap. Somebody has to decide whose piece goes in. That decision is most of the work in panoptic segmentation.
Why it exists
Two separate systems already existed, and each was half an answer.
One shaded the picture by category. It knew where the sky and the road were, and could not count cars. The other outlined countable objects, and had nothing to say about sky.
Running both and stapling the outputs together sounds easy. It is not, because they disagree.
The category system says a pixel is road. The object system says the same pixel is inside car number two. The overlapping object outlines disagree with each other as well.
Panoptic segmentation is the task of producing one answer that cannot contradict itself.
How the merge works
overlapping object outlines a shaded category map
(each with a confidence score) (sky, road, wall, grass)
\ /
\ /
v v
first place the most confident object
then place the next, but only on pixels still free
then if too little of an object is left free, drop it
then fill everything still empty from the category map
last delete any leftover speck that is too small
|
v
one map. Every pixel owned once.The second step is the important one. Pixels are handed out on a first-come basis, and confidence decides who comes first.
The third step prevents a nonsense result. Suppose a more confident car has taken nearly all of another car outline's pixels. That outline was a duplicate. Keeping a sliver of it would invent a car made of scraps.
How you grade it
Two things can go wrong, and a good score has to separate them.
You can miss a region entirely, or invent one that is not there. That is a recognition mistake.
You can find the right region but draw it badly. That is a shape mistake.
The panoptic score multiplies a recognition mark by a shape mark. One model finds everything and draws it roughly. Another draws beautifully and misses half the scene. Their scores should differ. Multiplying keeps both visible.
Where you have already seen this
- Self-driving demos where the road, footpath, sky and every individual car are all coloured in.
- Robot vacuums building a map of floor, wall and each separate obstacle.
- Sports broadcast graphics that shade the pitch and outline each player.
What is honestly hard here
Merging by confidence is a rule of thumb, not a truth. The most confident outline is not always the correct one.
A confident but slightly wrong car outline steals pixels from a less confident, better outline that comes second. Nothing later can undo that.
This is why the field moved away from stapling two systems together. If one model produces all the regions at once, it can learn to make them fit. That is the next lesson.
Remember this
- Panoptic output has one owner per pixel: no gaps, no overlaps.
- Overlaps are resolved by confidence order, and low-survival regions get dropped.
- The score separates finding regions from shaping them, then multiplies the two.
What to learn next
- Mask2Former and universal segmentation — one model that removes the merge heuristic.
- Model evaluation — where precision, recall and F-scores come from.
- Mask R-CNN — the instance masks this lesson was merging.
Developer — Code and libraries.
Setup
pip install numpyThe merge and the metric are both short algorithms. Writing them once removes the mystery from every panoptic evaluation you will ever read.
Merge, then score
import numpy as np
H, W = 10, 16
SKY, ROAD, TREE, CAR, PERSON = 0, 1, 2, 3, 4
NAME = {SKY: "sky", ROAD: "road", TREE: "tree", CAR: "car", PERSON: "person"}
STUFF = (SKY, ROAD, TREE)
# ---------------- ground truth panoptic map ----------------
gt_seg = np.zeros((H, W), int) # one segment id per pixel
gt_cls = {1: SKY, 2: ROAD, 3: CAR, 4: CAR, 5: PERSON}
gt_seg[:4] = 1
gt_seg[4:] = 2
gt_seg[4:8, 1:6] = 3
gt_seg[4:8, 8:12] = 4
gt_seg[5:9, 13:16] = 5
# ---------------- what the two heads predicted ----------------
sem = np.zeros((5, H, W)) # semantic head score per class per pixel
sem[SKY, :4] = 3.0
sem[ROAD, 4:] = 3.0
sem[ROAD, 3, :] = 3.5 # the horizon is one row too low
sem[TREE, 9, 0:3] = 3.5 # three stray tree pixels
def box(r0, r1, c0, c1):
m = np.zeros((H, W), bool); m[r0:r1, c0:c1] = True; return m
instances = [ # (score, class, mask) from the instance head
(0.95, CAR, box(4, 8, 1, 6)), # car 1, exactly right
(0.88, CAR, box(4, 8, 8, 13)), # car 2, one column too wide
(0.71, CAR, box(4, 8, 4, 10)), # a third car nobody asked for
(0.55, PERSON, box(5, 9, 12, 16)), # person, one column too wide
]
# ---------------- merge them into one map ----------------
def merge(instances, sem, overlap_keep=0.5, min_stuff_area=8):
pan = np.zeros((H, W), int) # 0 means still unclaimed
cls_of, next_id = {}, 1
for score, cls, mask in sorted(instances, key=lambda t: -t[0]):
free = mask & (pan == 0) # pixels a higher-scoring mask has not taken
kept = free.sum() / mask.sum()
if kept < overlap_keep:
print(f" dropped {NAME[cls]:6s} score {score:.2f}: only {kept:.0%} of it was still free")
continue
pan[free] = next_id; cls_of[next_id] = cls; next_id += 1
print(f" kept {NAME[cls]:6s} score {score:.2f}: {free.sum():>3} pixels ({kept:.0%} of its mask)")
for c in STUFF: # all pixels of one stuff class = ONE segment
region = (sem.argmax(0) == c) & (pan == 0)
if region.sum() >= min_stuff_area:
pan[region] = next_id; cls_of[next_id] = c; next_id += 1
elif region.sum():
print(f" dropped {NAME[c]:6s} : {region.sum()} pixels is under the area floor")
return pan, cls_of
print("merge order is score, highest first:")
pred_seg, pred_cls = merge(instances, sem)
GLYPH = {SKY: ".", ROAD: "-", TREE: "T", CAR: "C", PERSON: "P"}
def show(seg, cls_of, title):
print("\n" + title)
for row in seg:
print(" " + " ".join(GLYPH[cls_of[v]] if v else "?" for v in row))
show(gt_seg, gt_cls, "ground truth")
show(pred_seg, pred_cls, "merged prediction (? = void, claimed by nothing)")
# ---------------- panoptic quality ----------------
def pq(gt_seg, gt_cls, pr_seg, pr_cls):
out = {}
for c in sorted(set(gt_cls.values()) | set(pr_cls.values())):
g = [i for i, k in gt_cls.items() if k == c]
p = [i for i, k in pr_cls.items() if k == c]
hit_g, hit_p, ious = set(), set(), []
for gi in g:
for pi in p:
inter = np.sum((gt_seg == gi) & (pr_seg == pi))
if not inter:
continue
iou = inter / np.sum((gt_seg == gi) | (pr_seg == pi))
if iou > 0.5: # above one half the match can only be unique
hit_g.add(gi); hit_p.add(pi); ious.append(iou)
tp, fp, fn = len(ious), len(p) - len(hit_p), len(g) - len(hit_g)
sq = float(np.mean(ious)) if tp else 0.0
rq = tp / (tp + 0.5 * fp + 0.5 * fn)
out[c] = (tp, fp, fn, sq, rq, sq * rq)
return out
print(f"\n{'class':8s} {'TP':>3} {'FP':>3} {'FN':>3} {'SQ':>6} {'RQ':>6} {'PQ':>6}")
rows = pq(gt_seg, gt_cls, pred_seg, pred_cls)
for c, (tp, fp, fn, sq, rq, p) in rows.items():
print(f"{NAME[c]:8s} {tp:>3} {fp:>3} {fn:>3} {sq:>6.3f} {rq:>6.3f} {p:>6.3f}")
print(f"{'PQ':8s} {'':>3} {'':>3} {'':>3} {'':>6} {'':>6} "
f"{np.mean([r[5] for r in rows.values()]):>6.3f}")
# what a missed car does, with everything else identical
second_car = pred_seg[4, 9] # the segment id sitting on the second car
worse_seg = np.where(pred_seg == second_car, 0, pred_seg)
worse_cls = {i: c for i, c in pred_cls.items() if i != second_car}
rows2 = pq(gt_seg, gt_cls, worse_seg, worse_cls)
print(f"\ndelete the second car and only recognition quality moves:")
print(f" car SQ {rows[CAR][3]:.3f} -> {rows2[CAR][3]:.3f} "
f"RQ {rows[CAR][4]:.3f} -> {rows2[CAR][4]:.3f} PQ {rows[CAR][5]:.3f} -> {rows2[CAR][5]:.3f}")merge order is score, highest first: kept car score 0.95: 20 pixels (100% of its mask) kept car score 0.88: 20 pixels (100% of its mask) dropped car score 0.71: only 33% of it was still free kept person score 0.55: 13 pixels (81% of its mask) dropped tree : 3 pixels is under the area floor ground truth . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . - C C C C C - - C C C C - - - - - C C C C C - - C C C C - P P P - C C C C C - - C C C C - P P P - C C C C C - - C C C C - P P P - - - - - - - - - - - - - P P P - - - - - - - - - - - - - - - - merged prediction (? = void, claimed by nothing) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . - - - - - - - - - - - - - - - - - C C C C C - - C C C C C - - - - C C C C C - - C C C C C P P P - C C C C C - - C C C C C P P P - C C C C C - - C C C C C P P P - - - - - - - - - - - - P P P P ? ? ? - - - - - - - - - - - - - class TP FP FN SQ RQ PQ sky 1 0 0 0.750 1.000 0.750 road 1 0 0 0.625 1.000 0.625 car 2 0 0 0.900 1.000 0.900 person 1 0 0 0.923 1.000 0.923 PQ 0.800 delete the second car and only recognition quality moves: car SQ 0.900 -> 1.000 RQ 1.000 -> 0.667 PQ 0.900 -> 0.667
Reading the merge log
The duplicate car was dropped by the 50 per cent rule, not by any special duplicate detector. It scored 0.71, arrived third, and found only 33 per cent of its pixels still free. That threshold is the entire mechanism, and COCO's official merger uses the same idea.
The person was kept, but truncated to 81 per cent of its mask. The second car had already taken the shared column. Panoptic output has no way to record that argument, so the person's segment is permanently one column smaller than the model believed.
The tree fell below the area floor. Three pixels of a stuff class is noise, so those pixels became void, printed as ?. Void pixels are excluded from evaluation rather than counted as errors — which is worth knowing, because it means an over-aggressive floor is not directly punished.
Stuff of one class forms a single segment, however scattered. All sky pixels share one id even if a chimney splits the region in two. That rule surprises people who expect connected components.
Reading the score
SQ is the average IoU over matched segments; RQ is a familiar F-score in disguise. Look at the last two lines to see them separate. Deleting a correct car raises SQ from 0.900 to 1.000, because the remaining match is the perfect one. RQ falls from 1.000 to 0.667, because there is now one true positive and one false negative.
A single averaged number would have hidden that trade completely. PQ falls from 0.900 to 0.667, which is the honest summary.
Road scores 0.625 while the picture looks nearly right. The horizon is off by one row and a car is one column too wide, and road is a large, thin-margin region. Stuff classes with long boundaries lose IoU quickly, which is why panoptic scores are usually reported split as PQ-things and PQ-stuff.
The > 0.5 in the matching loop is not a tunable knob. Above one half, a predicted segment can overlap at most one ground-truth segment of its class by more than half. The match is therefore unique, and greedy matching is provably optimal. Below one half that guarantee vanishes, and the metric would need Hungarian matching instead.
Common mistakes
Sorting by class score instead of mask quality. A confident class with a poor mask wins the pixel. Some implementations sort by the product of class score and mean mask probability, which usually helps.
Forgetting that ordering is not commutative. Swap two instances with near-identical scores and the output changes. Panoptic merging is stable only when scores are well separated.
Treating void as background. Void means "do not grade this". Feeding void pixels into a per-pixel loss teaches the model that void is a class.
Comparing PQ across datasets. COCO panoptic and Cityscapes panoptic have different class counts and different stuff conventions, so their PQ numbers are not on the same scale.
Building panoptic output from a semantic model with connected components. Touching objects merge, and the count is wrong exactly when it matters.
Try it yourself
Change the third car's score from 0.71 to 0.99. It now runs first, takes the pixels, and pushes the two correct cars below the survival threshold. Watch PQ collapse. That fragility is the strongest argument for the unified models in the next lesson.
What to learn next
- Mask2Former and universal segmentation — one model that removes the merge heuristic.
- Model evaluation — where precision, recall and F-scores come from.
- Mask R-CNN — the instance masks this lesson was merging.
Researcher — Mathematics and papers.
The metric
Kirillov, He, Girshick, Rother and Dollár (2019) define a matching for one class, between predicted segments $\mathcal{P}$ and ground-truth segments $\mathcal{G}$. A prediction $p$ matches $g$ if $\text{IoU}(p, g) > 0.5$. Write the resulting sets as $TP$, $FP$ and $FN$. Then
$$ \text{PQ} = \frac{\sum_{(p,g) \in TP} \text{IoU}(p,g)}{|TP| + \tfrac{1}{2}|FP| + \tfrac{1}{2}|FN|} $$
which factorises exactly as
$$ \text{PQ} = \underbrace{\frac{\sum_{(p,g) \in TP} \text{IoU}(p,g)}{|TP|}}{\text{SQ}} \times \underbrace{\frac{|TP|}{|TP| + \tfrac{1}{2}|FP| + \tfrac{1}{2}|FN|}}{\text{RQ}} $$
$\text{SQ}$ is the mean IoU of matched segments. $\text{RQ}$ is the $F_1$ score over segments. The dataset-level PQ is the unweighted mean over classes. A class appearing in one image counts as much as a class appearing in every image.
Uniqueness lemma. If $\text{IoU}(p, g) > 0.5$, then $p$ can match at most one $g$ and vice versa. Proof sketch: $\text{IoU} > 0.5$ implies $|p \cap g| > \frac{1}{2}|p|$; two disjoint ground-truth segments cannot each take more than half of $p$. This is why the matching needs no optimisation.
Known criticisms of PQ
- Small segments dominate. Every segment contributes equally to RQ regardless of area, so a 12-pixel sign and a 200,000-pixel building carry the same weight. Reported PQ can be driven by tiny-object recall.
- The 0.5 threshold is a cliff. A segment at IoU 0.49 contributes a false positive and a false negative; at 0.51 it contributes a true positive with SQ 0.51. Small changes near the threshold produce large score jumps.
- Void handling gives free wins. Predicting void is never penalised directly. A model that abstains on hard regions is punished only through the false negatives it accumulates.
Three alternatives address these. PQ† uses semantic IoU for stuff classes rather than segment matching. Parsing Covering (Yang et al., 2019) weights segments by area, and is region-based and threshold-free.
The merge as an approximation
The heuristic in the developer section is the COCO panoptic combination baseline. Formally, given instance proposals ${(m_i, c_i, s_i)}$ and a semantic map, it greedily solves
$$ \max_{\text{assignment}} \sum_i s_i \cdot |m_i \cap \text{free}_i| \quad \text{subject to a partition constraint} $$
by sorting on $s_i$ alone. It has no view of the global optimum. UPSNet (Xiong et al., 2019) replaced the heuristic with a parameter-free panoptic head. It computes logits for stuff classes and instances in one tensor, plus an explicit unknown channel. Conflicts are resolved by softmax rather than by ordering. Panoptic-DeepLab (Cheng et al., 2020) went bottom-up instead. It predicts a semantic map plus instance centres and centre offsets, then groups pixels to the nearest predicted centre. That is a merge with no ordering at all.
The trajectory ends at mask classification. If one model emits all segments jointly, the training loss can penalise inconsistent overlaps directly, and the merge becomes an argmax rather than a policy. See Mask2Former and universal segmentation.
Reference numbers
Read these as landmarks, not as a leaderboard. On COCO panoptic validation: Panoptic FPN with ResNet-50 scores around 39 PQ, and UPSNet around 42. Panoptic-DeepLab lands around 39 to 41 depending on backbone. Mask2Former with Swin-L is reported at 57.8 PQ in Cheng et al. (2022). Cityscapes numbers run roughly 20 points higher on the same methods, because it has 19 classes and highly consistent imagery.
Implementation notes
- The reference implementation is
panopticapifrom the COCO authors. Its PNG format packs a segment id into three channels as $id = R + 256G + 256^2B$. Reading it as an image and forgetting that packing is the most common integration bug. label_ids_to_fusein the Hugging Face post-processing merges all instances of chosen classes into one segment. Use it to make sky or road behave as stuff when your model predicts them as things.- Evaluate PQ, PQ-things and PQ-stuff separately from the start. An aggregate PQ regression is nearly uninterpretable without the split.
Papers
- Kirillov, He, Girshick, Rother, Dollár, Panoptic Segmentation, CVPR 2019 — arxiv.org/abs/1801.00868
- Kirillov, Girshick, He, Dollár, Panoptic Feature Pyramid Networks, CVPR 2019 — arxiv.org/abs/1901.02446
- Xiong et al., UPSNet: A Unified Panoptic Segmentation Network, CVPR 2019 — arxiv.org/abs/1901.03784
- Cheng et al., Panoptic-DeepLab, CVPR 2020 — arxiv.org/abs/1911.10194
What to learn next
- Mask2Former and universal segmentation — one model that removes the merge heuristic.
- Model evaluation — where precision, recall and F-scores come from.
- Mask R-CNN — the instance masks this lesson was merging.