Heatmaps for keypoints
Instead of predicting a joint's coordinates directly, a pose model predicts a picture of where the joint probably is — and the way you read a coordinate back out of that picture costs more accuracy than the network does.
- 17 min read
- 3 reading levels
- Updated
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A heatmap is a picture the model draws, glowing brightest where it thinks a joint is.
Think of asking somebody where they left their keys. One answer is a precise claim: "on the second shelf, thirty centimetres from the left edge." Another is a gesture: "somewhere around here," with a hand sweeping over a region.
The second answer is less committal and more honest. It shows how sure they are. A wide sweep means uncertain; a jab of the finger means confident.
A heatmap is that gesture, drawn as a picture.
Why a picture beats a number
You could ask a network to output two numbers per joint, an across and a down. This works badly, and the reasons are worth understanding.
The task fights the architecture. The layers of a vision network hold a grid, matching the picture. Squashing that grid into two numbers throws away the spatial arrangement the network spent every layer building.
It cannot express doubt. If an elbow is hidden behind a bag, two numbers must still name a spot. A heatmap can glow faintly over a wide area, and that faintness is usable information.
It cannot say "two". Two people in frame means two left elbows. Two numbers name one. A heatmap can have two bright spots.
The learning signal is thin. With a picture as the target, every position on the grid gets told whether it should be bright or dark. Thousands of small corrections per joint per image, rather than two.
How it works
During training, the target is drawn, not written down. Take the true joint position and paint a soft blob of light centred on it. Bright in the middle, fading outwards. The network is asked to reproduce that picture.
image crop
│
▼
network
│
▼
one heatmap per joint
left wrist right knee
. . . . . . . . . . . . . . . .
. . . o o . . . . . . . . . . .
. . o O O o . . . . . . o o . .
. . o O O o . . . . . o O O o .
. . . o o . . . . . . . o o . .
. . . . . . . . . . . . . . . .
│
▼
find the brightest spot in each
│
▼
scale back up to the original pictureThe problem hiding in that last step
Here is the part that surprises people, and it is the real content of this lesson.
The heatmap is smaller than the picture. To keep the model affordable, it is usually a quarter of the size in each direction. So one heatmap square covers a block of four picture squares.
Take the brightest square, scale it back up, and your answer lands on a multiple of four. It can never land between them. Every joint you report is snapped to a coarse grid. That snapping error adds to whatever the network got wrong.
Measured on a well-trained model, this reading-back error is comparable to the network's own error. Half the mistakes in the finished system come from the last three lines of code, not from the model.
The fix is to look at the brightest square's neighbours. Suppose the square to the right is brighter than the one on the left. The true position then sits right of centre. That lets you place the answer between grid points, and the error drops sharply.
Where you have already seen this
- A fitness app counting squats and telling you your knees have caved in.
- Motion capture for a game character, driven by a phone camera.
- Sports analysis drawing a skeleton over a bowler's action.
- Face filters that keep a moustache on your lip as you talk.
Remember this
- A heatmap is a picture of where a joint probably is, with brightness meaning confidence.
- It keeps the grid the network built, expresses doubt, and handles more than one person.
- Reading a coordinate back out of a small heatmap is a large source of error. Use the neighbours.
What to learn next
- CNN — the grid-preserving architecture that makes heatmaps natural.
- Image segmentation — the other task that outputs a picture rather than a label.
- Object detection — where keypoint models get their person boxes from.
Developer — Code and libraries.
Setup
pip install numpyRun against numpy 1.26. This measures the three standard decoders on a fixed grid of sub-pixel targets, so the numbers are exactly reproducible.
Encoding, and three ways to decode
import numpy as np
H = W = 64 # heatmap size
STRIDE = 4 # so the input image is 256x256
SIGMA = 2.0
def encode(x, y, h=H, w=W, sigma=SIGMA):
"""Ground truth: a Gaussian blob centred on the true keypoint."""
yy, xx = np.mgrid[0:h, 0:w]
return np.exp(-((xx - x) ** 2 + (yy - y) ** 2) / (2 * sigma ** 2))
def decode_argmax(hm):
j = int(np.argmax(hm))
return float(j % hm.shape[1]), float(j // hm.shape[1])
def decode_quarter(hm):
"""The standard trick: nudge a quarter pixel toward the second-highest neighbour."""
x, y = decode_argmax(hm)
xi, yi = int(x), int(y)
if 0 < xi < hm.shape[1] - 1:
x += 0.25 * np.sign(hm[yi, xi + 1] - hm[yi, xi - 1])
if 0 < yi < hm.shape[0] - 1:
y += 0.25 * np.sign(hm[yi + 1, xi] - hm[yi - 1, xi])
return x, y
def decode_dark(hm, sigma=SIGMA):
"""DARK: a Gaussian in log space is a parabola, so solve for its true peak."""
x, y = decode_argmax(hm)
xi, yi = int(x), int(y)
if not (0 < xi < hm.shape[1] - 1 and 0 < yi < hm.shape[0] - 1):
return x, y
L = np.log(np.maximum(hm, 1e-10))
dx = 0.5 * (L[yi, xi + 1] - L[yi, xi - 1])
dy = 0.5 * (L[yi + 1, xi] - L[yi - 1, xi])
dxx = L[yi, xi + 1] - 2 * L[yi, xi] + L[yi, xi - 1]
dyy = L[yi + 1, xi] - 2 * L[yi, xi] + L[yi - 1, xi]
if dxx < 0: x -= dx / dxx
if dyy < 0: y -= dy / dyy
return x, y
# 400 keypoints at sub-pixel positions, on a fixed grid so this is fully reproducible.
offsets = np.linspace(0.0, 0.95, 20)
truth = [(20 + a, 30 + b) for a in offsets for b in offsets]
def evaluate(decoder, **kw):
errs = []
for tx, ty in truth:
px, py = decoder(encode(tx, ty), **kw)
errs.append(np.hypot(px - tx, py - ty))
return np.array(errs)
print(f"heatmap {H}x{W}, network stride {STRIDE}, so 1 heatmap pixel = {STRIDE} image pixels\n")
print(f"{'decoder':>16} {'mean err (hm px)':>18} {'mean err (image px)':>21} {'worst':>8}")
for name, fn in (("argmax", decode_argmax), ("quarter offset", decode_quarter),
("DARK", decode_dark)):
e = evaluate(fn)
print(f"{name:>16} {e.mean():>18.4f} {e.mean()*STRIDE:>21.4f} {e.max():>8.4f}")
# Why argmax alone cannot win: the answer is forced onto whole heatmap pixels.
hm = encode(20.5, 30.5)
print("\na keypoint truly at (20.50, 30.50):")
print(" argmax says ", decode_argmax(hm))
print(" quarter offset says", tuple(round(v, 4) for v in decode_quarter(hm)))
print(" DARK says ", tuple(round(v, 4) for v in decode_dark(hm)))
# And the cost of the resolution choice itself.
print("\nheatmap resolution vs argmax error, for the same 256x256 input:")
for h, stride in ((32, 8), (64, 4), (128, 2)):
e = np.array([np.hypot(*(np.array(decode_argmax(encode(20 + a, 30 + b, h, h))) - (20 + a, 30 + b)))
for a in offsets for b in offsets])
print(f" {h:>3}x{h:<3} stride {stride}: {e.mean()*stride:.4f} image pixels of error")
# Honesty check: a real network's heatmap is not a perfect Gaussian.
rng = np.random.default_rng(0)
print("\nsame test, with 3% noise added to every heatmap (closer to a real network):")
for name, fn in (("argmax", decode_argmax), ("quarter offset", decode_quarter), ("DARK", decode_dark)):
e = []
for tx, ty in truth:
hm = np.clip(encode(tx, ty) + rng.normal(0, 0.03, (H, W)), 1e-10, None)
px, py = fn(hm)
e.append(np.hypot(px - tx, py - ty))
print(f" {name:>14}: {np.mean(e)*STRIDE:.4f} image pixels of error")heatmap 64x64, network stride 4, so 1 heatmap pixel = 4 image pixels
decoder mean err (hm px) mean err (image px) worst
argmax 0.3833 1.5332 0.7071
quarter offset 0.1760 0.7041 0.3536
DARK 0.0000 0.0000 0.0000
a keypoint truly at (20.50, 30.50):
argmax says (20.0, 30.0)
quarter offset says (20.25, 30.25)
DARK says (20.5, 30.5)
heatmap resolution vs argmax error, for the same 256x256 input:
32x32 stride 8: 3.0664 image pixels of error
64x64 stride 4: 1.5332 image pixels of error
128x128 stride 2: 0.7666 image pixels of error
same test, with 3% noise added to every heatmap (closer to a real network):
argmax: 1.7168 image pixels of error
quarter offset: 1.0151 image pixels of error
DARK: 0.6496 image pixels of errorReading it
Argmax costs 1.53 image pixels on a perfect heatmap. The network made no error at all here; the heatmap is the exact ground truth. Every one of those 1.53 pixels comes from rounding to the grid. On a COCO-scale evaluation that is enough to move AP by several points.
The quarter-offset trick more than halves it, to 0.70. Two lines of code. This is why the trick appears in nearly every pose repository, usually without explanation.
DARK reaches exactly zero — and read the caveat before believing it. The target here is a Gaussian, and DARK's model is a Gaussian, so the fit is exact. That is a property of the experiment, not a promise about your model.
The noise block is the honest version. With 3% noise the ordering holds and the magnitudes are realistic: 1.72, 1.02, 0.65. DARK still wins by a clear margin, and it does not reach zero. Published gains on COCO are of the order of one to two AP, which is consistent with this and is a large gain for a decoding change.
The resolution table doubles the error every time you halve the heatmap. Stride 8 gives 3.07 pixels of pure rounding error. This is the argument for high-resolution branches: HRNet keeps a stride-4 representation alive through the whole network rather than downsampling and upsampling back.
Line by line, for the parts that trip people up
decode_argmax uses j % width for x and j // width for y. Swapping these is the most common bug in hand-written decoders, and it produces a transposed skeleton that looks almost plausible.
decode_dark works in log space because the logarithm of a Gaussian is a parabola. Fitting a parabola through three log-values and solving for its vertex is the whole method. dx / dxx is a one-step Newton update; dxx < 0 confirms the point is a maximum rather than a saddle.
np.maximum(hm, 1e-10) guards the logarithm. A real network emits zeros and small negatives after ReLU and interpolation, and log(0) poisons the whole computation.
sigma is a training hyperparameter, not a free choice at inference. DARK's Gaussian fit assumes the model learned to produce blobs of the sigma you trained with. Change it in the loss and change it here too.
Common mistakes
Applying the quarter offset and DARK together. They are alternatives. Doing both double-counts the shift.
Forgetting the flip-test alignment. Test-time augmentation flips the image, runs the model, and flips the heatmap back. A horizontal flip of an array is off by one pixel relative to a flip of the coordinate system. Failing to shift by one before averaging the two heatmaps is a persistent, silent accuracy loss.
Ignoring the affine transform back to the original image. The crop was resized and possibly rotated to reach the model. The inverse of that transform must be applied to the decoded coordinates, not to the heatmap.
Choosing sigma without thinking. Too small and the blob covers a couple of pixels, giving almost no gradient signal. Too large and neighbouring joints merge. Two to three pixels at stride 4 is the usual range.
Treating peak height as a calibrated confidence. It correlates with correctness and is not a probability. Threshold it if you must, and do not report it as one.
Try it yourself
Set SIGMA = 1.0 and rerun. DARK's advantage narrows, because a narrow blob gives the parabola fit fewer usable neighbours. Then set SIGMA = 4.0 and watch it widen again. You have found the interaction between the training hyperparameter and the decoder, which is exactly why the two have to be chosen together.
What to learn next
- CNN — the grid-preserving architecture that makes heatmaps natural.
- Image segmentation — the other task that outputs a picture rather than a label.
- Object detection — where keypoint models get their person boxes from.
Researcher — Mathematics and papers.
The representation
For keypoint $k$ at true location $(x, y)$ in image coordinates, and network stride $s$, the target heatmap on the $H \times W$ grid is
$$ G_k(u, v) = \exp\left(-\frac{(u - x/s)^2 + (v - y/s)^2}{2\sigma^2}\right) $$
$u, v$ index heatmap cells, $\sigma$ is the blob width in heatmap units. Training minimises mean squared error between predicted and target heatmaps, typically masked to visible joints.
The formulation is from Tompson et al. (NeurIPS 2014) and the stacked hourglass architecture of Newell et al. (ECCV 2016), and it displaced direct coordinate regression (DeepPose, Toshev and Szegedy, CVPR 2014) almost completely.
The quantisation error, quantified
Standard decoding takes $\hat{p} = \arg\max_{u,v} \hat{G}(u,v)$ and maps back as $s \cdot \hat{p}$. The error from rounding alone, for a target uniformly distributed within a cell, has expectation
$$ \mathbb{E}\big[\lVert \epsilon \rVert\big] = s \cdot \mathbb{E}\big[\sqrt{\delta_x^2 + \delta_y^2}\big], \qquad \delta \sim \mathcal{U}[-\tfrac12, \tfrac12]^2 $$
which evaluates to roughly $0.3826\,s$. The measured $0.3833$ heatmap pixels above matches this to three decimal places, confirming the experiment is measuring the quantisation term and nothing else.
The empirical quarter offset shifts the estimate by $0.25$ toward the larger of the two neighbours:
$$ \hat{p} \leftarrow \hat{p} + 0.25 \cdot \operatorname{sign}\big(\nabla \hat{G}(\hat{p})\big) $$
It halves the error and is unprincipled — the shift magnitude is a constant, unrelated to the actual blob shape.
DARK
Zhang et al. (CVPR 2020), Distribution-Aware Coordinate Representation for Human Pose Estimation (arxiv.org/abs/1910.06278), treat the heatmap as a sampled Gaussian and solve for its continuous peak.
Take $\mathcal{P} = \log \hat{G}$. For a Gaussian, $\mathcal{P}$ is a quadratic, so a second-order Taylor expansion at the discrete maximum $m$ is exact:
$$ \mathcal{P}(p) = \mathcal{P}(m) + \mathcal{D}(m)^\top (p - m) + \tfrac{1}{2}(p - m)^\top \mathcal{D}''(m)(p - m) $$
Setting the derivative to zero gives
$$ \hat{p} = m - \big(\mathcal{D}''(m)\big)^{-1} \mathcal{D}(m) $$
$\mathcal{D}(m)$ is the first derivative and $\mathcal{D}''(m)$ the Hessian of the log-heatmap at $m$, both estimated by finite differences. The code above uses the diagonal approximation; the full paper uses the complete $2 \times 2$ Hessian.
DARK also corrects the encoding side: the standard practice of centring the Gaussian on a rounded integer coordinate injects a bias before training even starts. Generating unbiased targets at the true sub-pixel location removes it. The paper reports both halves as necessary, and it is model-agnostic — reported gains of roughly 1 to 2 AP on COCO across several backbones.
UDP (Huang et al., CVPR 2020), The Devil is in the Details, addresses a related bias: the conventional size = (w - 1) versus w ambiguity in the image-to-heatmap affine transform introduces a systematic flip-test misalignment. Both papers are about arithmetic in the data pipeline, not architecture, and both produce gains comparable to a substantial model change.
Resolution and architecture
Quantisation error scales linearly with stride, which is a direct architectural constraint.
- Hourglass (Newell et al., 2016) — symmetric downsample/upsample with skip connections, repeated.
- Simple Baseline (Xiao et al., ECCV 2018) — ResNet plus three transposed convolutions; a strong and deliberately plain baseline.
- HRNet (Sun et al., CVPR 2019, arxiv.org/abs/1902.09212) — maintains a high-resolution branch throughout, fusing repeatedly across resolutions rather than recovering resolution at the end. HRNet-W48 reaches 76.3 AP on COCO val at 384x288 input.
- ViTPose (Xu et al., NeurIPS 2022, arxiv.org/abs/2204.12484) — a plain ViT backbone with a lightweight decoder, reaching 80.9 AP on COCO test-dev at the largest scale.
Alternatives to argmax decoding
Integral / soft-argmax (Sun et al., ECCV 2018) replaces the discrete maximum with a spatial expectation over the softmax-normalised heatmap:
$$ \hat{p} = \sum_{u,v} (u, v) \cdot \frac{\exp(\hat{G}(u,v)/T)}{\sum_{u',v'} \exp(\hat{G}(u',v')/T)} $$
This is differentiable end to end, removing the need for an explicit heatmap loss. The drawback is a bias induced by background activation across the whole map, and a documented sensitivity to the temperature $T$.
SimCC (Li et al., ECCV 2022) reformulates the problem as two 1D classifications, over $x$ and over $y$ separately, at sub-pixel bin resolution. This removes 2D quantisation entirely and is what RTMPose (Jiang et al., 2023, arxiv.org/abs/2303.07399) builds on to reach 75.8 AP on COCO at over 90 FPS on a desktop CPU.
Regression, revisited. RLE (Li et al., ICCV 2021), Residual Log-likelihood Estimation, makes direct coordinate regression competitive by learning the output distribution with a normalising flow instead of assuming a fixed Laplacian. Regression is cheaper at inference and is the right choice for tight latency budgets.
Evaluation
COCO uses OKS — object keypoint similarity:
$$ \text{OKS} = \frac{\sum_i \exp\left(-d_i^2 / (2 s^2 \kappa_i^2)\right) \cdot \mathbb{1}[v_i > 0]}{\sum_i \mathbb{1}[v_i > 0]} $$
$d_i$ is the distance between predicted and true keypoint $i$, $s$ the square root of the person's segment area, $\kappa_i$ a per-keypoint constant reflecting annotator disagreement, and $v_i$ the visibility flag. AP is then averaged over OKS thresholds from 0.50 to 0.95.
The $\kappa_i$ values matter and are frequently ignored: eyes have small $\kappa$ and hips large, so a hip error of the same magnitude costs far less OKS than an eye error. MPII instead uses PCKh, thresholding at a fraction of head segment length.
Papers
- Toshev and Szegedy, DeepPose, CVPR 2014 — arxiv.org/abs/1312.4659
- Tompson et al., Joint Training of a CNN and a Graphical Model for Human Pose Estimation, NeurIPS 2014
- Newell et al., Stacked Hourglass Networks, ECCV 2016 — arxiv.org/abs/1603.06937
- Xiao et al., Simple Baselines for Human Pose Estimation and Tracking, ECCV 2018 — arxiv.org/abs/1804.06208
- Sun et al., Deep High-Resolution Representation Learning (HRNet), CVPR 2019 — arxiv.org/abs/1902.09212
- Zhang et al., Distribution-Aware Coordinate Representation (DARK), CVPR 2020 — arxiv.org/abs/1910.06278
- Huang et al., The Devil is in the Details (UDP), CVPR 2020 — arxiv.org/abs/1911.07524
- Li et al., Human Pose Regression with Residual Log-likelihood Estimation, ICCV 2021 — arxiv.org/abs/2107.11291
- Li et al., SimCC: a Simple Coordinate Classification Perspective, ECCV 2022 — arxiv.org/abs/2107.03332
What to learn next
- CNN — the grid-preserving architecture that makes heatmaps natural.
- Image segmentation — the other task that outputs a picture rather than a label.
- Object detection — where keypoint models get their person boxes from.