Evaluating Vision Models

Keypoint metrics: OKS and PCK

Pose models are scored by how far each predicted joint sits from the true one, with a tolerance that changes per joint and shrinks as the person gets smaller.

On this page 10
  1. The short answer
  2. The analogy you have lived
  3. What a pose model produces
  4. The first idea: PCK
  5. The better idea: OKS
  6. How it gets used
  7. Where you have seen this
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

These metrics score a pose model by how far each joint sits from the true one. Each joint gets its own tolerance.

The analogy you have lived

A tailor measures you for a shirt. He marks points on your body: the shoulder, the elbow, the wrist, the collar.

If his shoulder mark is a centimetre off, the shirt still fits. Nobody notices. If his collar mark is a centimetre off, the shirt is wrong and you will feel it all day.

Same error, in centimetres. Completely different consequence. A good scoring rule has to know that.

What a pose model produces

A pose model does not draw a box. It puts a dot on each joint of a person: nose, eyes, shoulders, elbows, wrists, hips, knees, ankles. The COCO standard uses seventeen of them.

            o  nose
          /   \
    shoulder--shoulder
       |    |    |
    elbow   |   elbow
       |    |    |
    wrist  hip  wrist
          /   \
       knee   knee
         |      |
      ankle   ankle

So the question is not "did the boxes overlap". It is "how far off is each dot, and does that distance matter here".

The first idea: PCK

PCK stands for percentage of correct keypoints. It is the straightforward version.

Pick a tolerance, say a fifth of the person's torso length. Any dot landing inside that distance counts as correct. Count the fraction that pass.

PCK is easy to explain, and that is its main virtue. It has two weaknesses. Every joint gets the same tolerance, which the tailor already told us is wrong. And the answer depends on what you divide by. Torso length, head size and box width all differ.

The better idea: OKS

OKS stands for object keypoint similarity. It fixes both problems.

First, it gives each joint its own tolerance, measured from how much real human annotators disagree about that joint. Two people asked to mark a nose will agree closely. Two people asked to mark a hip will not. A hip is inside the body, with no visible landmark.

So the nose gets a tight tolerance and the hip gets a loose one. The numbers are not invented; they were measured from repeat annotations of the same photographs.

Second, OKS scales the tolerance with the size of the person. Being ten pixels off on someone filling the frame is nothing. Being ten pixels off on someone standing far away is the whole person.

Instead of a hard pass or fail, each joint gets a smooth score between zero and one. Perfect placement scores one. Far away scores near zero. The person's overall score is the average across their joints.

How it gets used

OKS plays the same role for pose that overlap plays for boxes. It decides whether a predicted person counts as a match for a real one.

Then the whole ranked-list machinery runs on top, exactly as it does for detection. You get average precision at a strict tolerance, at a loose one, and averaged over a range of them.

   predicted person  →  OKS against each real person
                     →  best match above the cut-off? yes → correct
                     →  no → invented
   real person with no match at all → missed

Where you have seen this

  • A fitness app counting your squats from the phone camera.
  • A dance or yoga app telling you your arm is too low.
  • Motion capture for animation, without the suit covered in markers.
  • A factory camera checking whether a worker is reaching into a machine.

Every one of them ships or fails on the wrists and ankles. Those are the joints that move.

The honest part

OKS averages over joints. Get fourteen out of seventeen joints perfect and three badly wrong, and the average stays high.

For a fitness app, three wrong joints may be the three that matter. For a safety system, one wrong wrist is the entire product. The average is a summary, and summaries hide exactly the thing you were worried about.

Look at per-joint scores. Always. The average is for the leaderboard.

Remember this

  • Pose is scored by distance per joint, not by overlap.
  • PCK uses one tolerance for every joint. OKS uses a measured tolerance per joint, scaled by person size.
  • The average across joints hides the few joints that usually matter most.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

The tolerance constants below are copied from pycocotools/cocoeval.py. They are not adjustable knobs; changing them makes your numbers incomparable to every published result.

OKS and PCK, written out

keypoint_metrics.py
import numpy as np

NAMES = ["nose", "left_eye", "right_eye", "left_ear", "right_ear",
         "left_shoulder", "right_shoulder", "left_elbow", "right_elbow",
         "left_wrist", "right_wrist", "left_hip", "right_hip",
         "left_knee", "right_knee", "left_ankle", "right_ankle"]

# The per-keypoint tolerances COCO ships with, in pycocotools/cocoeval.py.
SIGMAS = np.array([.26, .25, .25, .35, .35, .79, .79, .72, .72,
                   .62, .62, 1.07, 1.07, .87, .87, .89, .89]) / 10.0

# One standing person, roughly 60 wide and 180 tall.
gt = np.array([
    [100,  20], [ 96,  16], [104,  16], [ 90,  18], [110,  18],
    [ 80,  45], [120,  45], [ 70,  80], [130,  80], [ 65, 115], [135, 115],
    [ 88, 110], [112, 110], [ 86, 150], [114, 150], [ 84, 190], [116, 190],
], dtype=float)
vis = np.ones(17)                      # every joint is labelled and visible
area = 60.0 * 180.0                    # the person's bounding-box area

pred = gt.copy()
pred[9] += [12, 0]                     # left wrist off by 12 pixels
pred[0] += [12, 0]                     # nose off by the same 12 pixels
pred[15] += [12, 0]                    # left ankle off by the same 12 pixels


def oks(gt, pred, vis, area, sigmas=SIGMAS):
    d2 = ((gt - pred) ** 2).sum(axis=1)
    var = (sigmas * 2) ** 2            # cocoeval uses vars = (sigmas*2)**2
    e = d2 / var / (area + np.spacing(1)) / 2
    e = e[vis > 0]                     # unlabelled joints are skipped entirely
    return float(np.exp(-e).mean())


print(f"{'joint':<15} {'sigma':>6} {'error px':>9} {'per-joint score':>16}")
for i in (0, 9, 15):
    d2 = ((gt[i] - pred[i]) ** 2).sum()
    e = d2 / ((SIGMAS[i] * 2) ** 2) / area / 2
    print(f"{NAMES[i]:<15} {SIGMAS[i]:>6.3f} {np.sqrt(d2):>9.1f} {np.exp(-e):>16.3f}")

print(f"\nOKS for the whole person: {oks(gt, pred, vis, area):.4f}")

print("\nOKS as the person gets smaller (same 12-pixel errors):")
for scale in (1.0, 0.5, 0.25):
    g = gt * scale
    p = g.copy()
    p[[0, 9, 15]] += [12, 0]
    print(f"  person {int(60*scale)}x{int(180*scale)} px -> OKS {oks(g, p, vis, area*scale**2):.4f}")

# PCK: a hit if the error is under a fraction of a reference length.
torso = np.linalg.norm(gt[5] - gt[12])        # left shoulder to right hip
err = np.linalg.norm(gt - pred, axis=1)
print(f"\ntorso diagonal = {torso:.1f} px")
for a in (0.05, 0.1, 0.2):
    hits = (err <= a * torso).sum()
    print(f"  PCK@{a:<4} -> {hits}/17 = {hits/17:.3f}")

# The same predictions scored at a sweep of OKS thresholds, COCO style.
print("\nAP-style pass/fail against OKS thresholds:")
score = oks(gt, pred, vis, area)
for t in np.arange(0.5, 1.0, 0.05):
    print(f"  OKS >= {t:.2f}: {'pass' if score >= t else 'fail'}")
Output
joint            sigma  error px  per-joint score
nose             0.026      12.0            0.085
left_wrist       0.062      12.0            0.648
left_ankle       0.089      12.0            0.810

OKS for the whole person: 0.9143

OKS as the person gets smaller (same 12-pixel errors):
  person 60x180 px -> OKS 0.9143
  person 30x90 px -> OKS 0.8593
  person 15x45 px -> OKS 0.8256

torso diagonal = 72.4 px
  PCK@0.05 -> 14/17 = 0.824
  PCK@0.1  -> 14/17 = 0.824
  PCK@0.2  -> 17/17 = 1.000

AP-style pass/fail against OKS thresholds:
  OKS >= 0.50: pass
  OKS >= 0.55: pass
  OKS >= 0.60: pass
  OKS >= 0.65: pass
  OKS >= 0.70: pass
  OKS >= 0.75: pass
  OKS >= 0.80: pass
  OKS >= 0.85: pass
  OKS >= 0.90: pass
  OKS >= 0.95: fail

Reading the output carefully

The same 12-pixel error scores 0.085, 0.648 and 0.810. Nose, wrist, ankle. That is the tailor's point, made numerically. The nose has the tightest sigma at 0.026, so twelve pixels destroys it. The ankle has 0.089 and shrugs the same error off. Any metric using one tolerance for all joints would have called these three errors identical.

The whole-person OKS is 0.9143 with three joints wrong. Fourteen perfect joints each contribute a score of exactly 1.0, and the mean drags the three failures upward. This is the averaging weakness, visible in a single number.

Shrinking the person only pushes OKS from 0.914 to 0.826. You might expect a collapse. It cannot happen, because fourteen of seventeen joints are still exact, giving a floor of 14/17 = 0.824. The three bad joints have already bottomed out. Read this as a warning: OKS compresses hard at the bad end, and the difference between "somewhat wrong" and "catastrophically wrong" is nearly invisible in the aggregate.

PCK@0.05 and PCK@0.1 give the same 0.824. Five percent of the torso is 3.6 pixels and ten percent is 7.2; both are under the twelve-pixel error, so the same fourteen joints pass. PCK is a step function, and between two thresholds it can be completely blind.

PCK@0.2 reports a perfect 1.000. Twenty percent of the torso is 14.5 pixels, which is more than the error, so everything passes. The most commonly quoted PCK threshold declares this model flawless while OKS is still discriminating. Be very suspicious of a PCK@0.2 headline.

The threshold sweep is where the metric earns its keep. The prediction passes everything up to 0.90 and fails at 0.95. Averaging pass rates over that sweep is what COCO's keypoint AP does, and it is why a single-threshold number tells you so little.

The details that bite

vars = (sigmas * 2) ** 2 — the factor of two is in the reference implementation and is easy to drop when reimplementing. Without it every score is far too harsh.

area is the ground-truth segmentation area in COCO's own evaluation, not the bounding-box area. Using box area instead makes your scores mildly optimistic and incomparable to published results.

e = e[vis > 0] drops joints the annotator never labelled. COCO's visibility flag takes three values: 0 unlabelled, 1 labelled but occluded, 2 labelled and visible. Occluded joints are scored. Treating flag 1 as unlabelled is a common bug that inflates scores.

The maxDets for keypoints is [20], not [1, 10, 100] as for boxes, and the area ranges have no "small" bucket. Keypoint evaluation is a different configuration of the same code path, and copying detection settings across gives silently wrong numbers.

Use the reference implementation

bash
pip install pycocotools
python
from pycocotools.coco import COCO
from pycocotools.cocoeval import COCOeval

gt = COCO("person_keypoints_val2017.json")
dt = gt.loadRes("keypoint_results.json")
e = COCOeval(gt, dt, iouType="keypoints")
e.evaluate(); e.accumulate(); e.summarize()

No output block: it depends entirely on your annotations and model, and a fabricated table would teach you to expect numbers that will not appear. Note iouType="keypoints" — passing "bbox" here runs a completely different metric without complaining.

Your keypoint_results.json entries need "keypoints" as a flat list of 51 numbers, ordered x1, y1, v1, x2, y2, v2, …. The visibility values in a prediction are ignored by the scorer, but the list still has to be the right length.

Common mistakes

Reporting PCK without saying what it was normalised by. PCKh@0.5 (head-segment length, from MPII) and PCK@0.2 (torso) are different metrics with similar names. Quote the reference length every time.

Evaluating on ground-truth boxes and calling it end-to-end. Top-down pose models need a person detector first. Feeding them perfect boxes measures the pose head alone, and the number will be several points higher than the deployed system.

Averaging OKS over people instead of running the AP machinery. Mean OKS ignores false positives entirely. A model that outputs a skeleton for every shadow can have an excellent mean OKS.

Left-right flips. Swapping left and right wrist gives two large errors and is invisible in an aggregate number. Check it explicitly; it is one of the most common pose failures and it comes from horizontal-flip augmentation without swapping the joint labels.

Try it yourself

Move the two ear keypoints by 12 pixels instead of the nose, wrist and ankle. Predict whether OKS goes up or down before running it, using the sigma table. Then set the visibility of the three broken joints to zero and watch OKS jump to 1.0 — which is exactly why the visibility flag has to come from the annotation, never from your model.

What to learn next

Researcher — Mathematics and papers.

Object keypoint similarity

For a ground-truth person with keypoints $g_i$, visibility flags $v_i$ and area $s$, and predicted keypoints $p_i$:

$$ \mathrm{OKS} = \frac{\sum_i \exp!\left(-\dfrac{d_i^2}{2 s^2 \kappa_i^2}\right)\delta(v_i > 0)}{\sum_i \delta(v_i > 0)} $$

with $d_i = \lVert p_i - g_i \rVert_2$ the Euclidean error for joint $i$, $s$ the object area in pixels squared, and $\kappa_i = 2\sigma_i$ the per-joint falloff. $\delta(\cdot)$ is one when the condition holds and zero otherwise.

The $\sigma_i$ were estimated by COCO from the standard deviation of redundant human annotations of the same images, normalised by object scale. Reported values, in the canonical joint order:

$$ \sigma = 0.1 \cdot [\,0.26, 0.25, 0.25, 0.35, 0.35, 0.79, 0.79, 0.72, 0.72, 0.62, 0.62, 1.07, 1.07, 0.87, 0.87, 0.89, 0.89\,] $$

The ordering is informative on its own. Facial landmarks are three to four times tighter than hips. Any model or loss weighting joints uniformly is fighting the metric.

Substituting $d_i = 2\sigma_i s^{1/2}$ gives a per-joint score of $e^{-1/2} \approx 0.607$, which is the sense in which $2\sigma_i$ is "the tolerance".

PCK and its variants

$$ \mathrm{PCK}@\alpha = \frac{1}{K}\sum_{i=1}^{K} \delta!\left(\lVert p_i - g_i\rVert_2 \le \alpha \cdot L\right) $$

$L$ is a normalising length. The literature uses at least four:

Name$L$Dataset
PCK@0.2torso diameterFLIC, LSP
PCKh@0.5head-segment lengthMPII
PCPlimb length, both endpointsearly Pascal work
NMEinter-ocular or inter-pupil distanceface alignment

PCK is a thresholded step function, so it is not differentiable, discards all information about how wrong a miss is, and saturates. Yang and Ramanan (2013) introduced PCP; MPII's switch to PCKh was motivated by torso length being unstable under foreshortening.

Face alignment's NME has its own trap: inter-ocular normalisation collapses for profile faces, which is why Sagonas et al. (2016) and later work report the area under the cumulative error curve rather than a single NME.

Keypoint AP

COCO applies the detection AP machinery with OKS replacing IoU, sweeping $\mathrm{OKS} \in {0.50, 0.55, \dots, 0.95}$ and reporting $\mathrm{AP}$, $\mathrm{AP}^{50}$, $\mathrm{AP}^{75}$, $\mathrm{AP}^{M}$ and $\mathrm{AP}^{L}$. maxDets is $[20]$ and there is no small-object bucket, since keypoints are not annotated on small people.

A subtlety with real consequences: OKS depends on the ground-truth area, so a detection cannot be scored without being matched first, and the matching uses the score being computed. cocoeval.computeOks resolves this by evaluating all pairs and letting the greedy assignment pick, which means an over-confident prediction can consume a ground truth that a better prediction would have matched.

Top-down versus bottom-up, and what the metric rewards

Top-down methods (Mask R-CNN keypoints, HRNet, ViTPose) detect people, crop, and regress joints in the crop. Accuracy is high because the crop normalises scale, which is exactly the quantity OKS divides by. Cost grows linearly with the number of people.

Bottom-up methods (OpenPose, Associative Embedding, HigherHRNet) detect all joints then group them. Constant cost in the number of people, weaker on small people, and grouping failures produce the left-right and person-mixing errors that OKS penalises heavily.

Sun et al. (2019), HRNet, is the architectural turning point: maintaining high-resolution representations throughout, rather than recovering resolution by upsampling, directly attacks the localisation precision that the strict end of the OKS sweep measures.

Known problems with the metric

Sigma values are dataset-specific and rarely re-estimated. They were measured on COCO images with COCO annotators. Applying them to infants, to animals, or to overhead camera views is an unexamined assumption. Cao et al. (2019) and the AP-10K animal-pose work both had to derive their own.

Area normalisation conflates scale with difficulty. A large but heavily occluded person gets a generous tolerance because they are large. Occlusion, not size, is what drives error.

OKS has no notion of anatomical plausibility. A prediction with the elbow beyond the wrist can score well if the individual distances are small. Bone-length and joint-angle consistency checks are a separate, and worthwhile, evaluation.

Aggregate scores conceal per-joint collapse. Reporting per-joint AP is cheap and is what actually tells you whether the model has learned wrists. Most papers include the table; most readers skip it.

Papers

  • Lin et al., Microsoft COCO, ECCV 2014 — arxiv.org/abs/1405.0312
  • Andriluka et al., 2D Human Pose Estimation: New Benchmark and State of the Art Analysis, CVPR 2014 (MPII, PCKh)
  • Yang and Ramanan, Articulated Human Detection with Flexible Mixtures of Parts, TPAMI 2013
  • Cao et al., OpenPose, TPAMI 2019 — arxiv.org/abs/1812.08008
  • Sun et al., Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019 — arxiv.org/abs/1902.09212
  • Ruggero Ronchi and Perona, Benchmarking and Error Diagnosis in Multi-Instance Pose Estimation, ICCV 2017 — arxiv.org/abs/1707.05388

What to learn next