Faces, People and Pose

MediaPipe face, hand and body landmarks

MediaPipe ships ready-made models that return 478 face points, 21 hand points and 33 body points on a CPU in real time, and the useful work starts after you have the numbers.

Read these first

On this page 8
  1. The short answer
  2. What you get
  3. The two kinds of numbers
  4. Where the real work is
  5. What it does not do
  6. Where you have already seen this
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

MediaPipe is a free Google toolkit that finds body, hand and face points fast enough for a phone.

Think of buying vegetables already chopped. Somebody else did the slow, fiddly part. You still have to decide what to cook, and that decision is now the whole job.

MediaPipe is the pre-chopped ingredients. It hands you numbers describing where the body parts are. What those numbers mean for your problem is entirely yours to work out.

What you get

Three separate models, each doing one thing.

Body. Thirty-three points: shoulders, elbows, wrists, hips, knees, ankles, plus some on the face and hands. Enough for posture, exercise counting and gesture.

Hands. Twenty-one points per hand: the wrist, then four points along each finger from knuckle to tip. Enough for sign language work, pinch gestures and finger counting.

Face. Four hundred and seventy-eight points, forming a mesh over the whole face including the eyes. Enough for filters, gaze direction and expression.

The two kinds of numbers

The body model returns each point twice, and the difference matters.

Screen coordinates. Where the point is in the picture, from one edge to the other. Useful for drawing on top of the image.

World coordinates. Where the point is in real space, measured in metres, with the origin between the hips. Useful for measuring an angle at a joint.

Reach for the second whenever you are measuring the body rather than drawing on the picture. Screen coordinates change when the person walks toward the camera. World coordinates do not.

Where the real work is

Getting the points is the easy part now. Everything interesting happens afterwards.

   camera frame
        │
        ▼
   MediaPipe  ─────►  a list of points
        │
        ▼
   YOUR code:
        compute angles at the joints
        smooth the numbers across frames
        decide what counts as "one squat"
        decide what to do when a limb is hidden
        │
        ▼
   something useful

That list of decisions is where projects succeed or fail. A raw point stream is noisy. It jitters when the person is still. It jumps when an arm passes behind the body. Nobody hands you a solution to that; you build it.

What it does not do

It does not tell you who somebody is. The face mesh describes shape, not identity. A separate recognition model would be needed. That carries all the consent and legal duties of biometric identification.

It handles one or a few people, not a crowd. These models are built for a phone user in front of a camera. Point one at a stadium and it will not do what you want.

It is not a medical instrument. The joint positions are estimates from a single camera. Clinical gait analysis uses marker systems and multiple cameras for good reasons.

Where you have already seen this

  • Video-call background blur and hand-raise detection.
  • Face filters in camera apps.
  • Fitness apps counting repetitions and flagging bad form.
  • Sign language and gesture demos on the web.

Remember this

  • Three models: thirty-three body points, twenty-one per hand, four hundred and seventy-eight on the face.
  • Body points come in screen coordinates and in real-world metres. Use metres for angles.
  • Getting the points is the easy part. Smoothing, angles and rules are your job.

What to learn next

  • OpenCV — reading frames, colour conversion and drawing the skeleton back on.
  • What is edge AI — why these models are built the way they are.
  • TensorFlow Lite — the runtime inside every .task bundle.

Developer — Code and libraries.

Setup

bash
pip install mediapipe numpy

Written against mediapipe 0.10.21 on Python 3.10. Note two things before you start.

The models are separate downloads. The pip package does not include them. Each task needs a .task bundle:

TaskFileURL
Faceface_landmarker.taskhttps://storage.googleapis.com/mediapipe-models/face_landmarker/face_landmarker/float16/latest/face_landmarker.task
Handhand_landmarker.taskhttps://storage.googleapis.com/mediapipe-models/hand_landmarker/hand_landmarker/float16/latest/hand_landmarker.task
Posepose_landmarker_full.taskhttps://storage.googleapis.com/mediapipe-models/pose_landmarker/pose_landmarker_full/float16/latest/pose_landmarker_full.task

Pose also offers pose_landmarker_lite.task and pose_landmarker_heavy.task at the matching paths.

MediaPipe pins an old protobuf. Version 0.10.x requires protobuf below 5. With a newer protobuf installed the import still succeeds and prints AttributeError: 'MessageFactory' object has no attribute 'GetPrototype' to stderr, along with absl and oneDNN startup lines. Install MediaPipe into its own virtual environment.

Part 1: the landmark layout, straight from the installed package

This runs with no model file and no image. It is the reference you will keep going back to.

landmark_layout.py
import numpy as np
from mediapipe.python.solutions.pose import PoseLandmark
from mediapipe.python.solutions.hands import HandLandmark
from mediapipe.python.solutions.face_mesh import FACEMESH_TESSELATION, FACEMESH_IRISES

print("pose landmarks:", len(PoseLandmark))
print("  ", [PoseLandmark(i).name for i in (0, 11, 12, 23, 24, 27, 28)])
print("hand landmarks:", len(HandLandmark))
print("  ", [HandLandmark(i).name for i in (0, 4, 8, 12, 16, 20)])

mesh = {i for e in FACEMESH_TESSELATION for i in e}
iris = {i for e in FACEMESH_IRISES for i in e}
print("face mesh vertices:", len(mesh), f"(indices 0-{max(mesh)})")
print("iris ring vertices:", sorted(iris))
print("iris centres 468 and 473 have no edges, so they are absent above")
print("total landmarks the Face Landmarker returns:", len(mesh) + len(iris) + 2)

# ---- Geometry on landmarks. This part is plain numpy; it never touches MediaPipe.
# World landmarks are metres, with the midpoint of the hips as the origin.
world = {                       # one frame of a person with a bent right elbow
    "RIGHT_SHOULDER": (-0.18, -0.52, -0.02),
    "RIGHT_ELBOW":    (-0.22, -0.25,  0.01),
    "RIGHT_WRIST":    (-0.05, -0.12,  0.06),
    "LEFT_SHOULDER":  ( 0.18, -0.52, -0.02),
    "LEFT_ELBOW":     ( 0.20, -0.24, -0.01),
    "LEFT_WRIST":     ( 0.21,  0.03, -0.02),
}

def angle(a, b, c):
    """Interior angle at b, in degrees."""
    u = np.array(a) - np.array(b)
    v = np.array(c) - np.array(b)
    cos = u @ v / (np.linalg.norm(u) * np.linalg.norm(v))
    return float(np.degrees(np.arccos(np.clip(cos, -1, 1))))

for side in ("RIGHT", "LEFT"):
    a = angle(world[f"{side}_SHOULDER"], world[f"{side}_ELBOW"], world[f"{side}_WRIST"])
    print(f"{side:>5} elbow angle: {a:6.2f} degrees")

upper = np.linalg.norm(np.subtract(world["RIGHT_SHOULDER"], world["RIGHT_ELBOW"]))
fore  = np.linalg.norm(np.subtract(world["RIGHT_ELBOW"], world["RIGHT_WRIST"]))
print(f"right upper arm: {upper*100:.1f} cm, forearm: {fore*100:.1f} cm")
Output
pose landmarks: 33
   ['NOSE', 'LEFT_SHOULDER', 'RIGHT_SHOULDER', 'LEFT_HIP', 'RIGHT_HIP', 'LEFT_ANKLE', 'RIGHT_ANKLE']
hand landmarks: 21
   ['WRIST', 'THUMB_TIP', 'INDEX_FINGER_TIP', 'MIDDLE_FINGER_TIP', 'RING_FINGER_TIP', 'PINKY_TIP']
face mesh vertices: 468 (indices 0-467)
iris ring vertices: [469, 470, 471, 472, 474, 475, 476, 477]
iris centres 468 and 473 have no edges, so they are absent above
total landmarks the Face Landmarker returns: 478
RIGHT elbow angle: 119.59 degrees
 LEFT elbow angle: 175.40 degrees
right upper arm: 27.5 cm, forearm: 22.0 cm

MediaPipe writes several startup lines to stderr before this — an absl logging warning, a oneDNN notice, and on a mismatched protobuf the MessageFactory line described above. They are noise, and the numbers above are what lands on stdout.

Reading it

468 mesh vertices plus 10 iris vertices makes 478. The tesselation covers indices 0 to 467. The iris additions are 468 to 477, and the two centres (468 and 473) appear in no edge, which is why the printed ring lists only eight. Getting this arithmetic right saves an afternoon when your indices are off by two.

The bent elbow reads 119.59 degrees, the straight one 175.40. These are real numbers computed from real coordinates. Notice that the "straight" arm is not 180 — a straight arm never measures 180 in practice, because estimated landmarks carry error. Any rule of the form "the arm is straight when the angle is 180" will never fire. Threshold at 160 or 165.

The limb lengths, 27.5 cm and 22.0 cm, are plausible and they are the honest use of world coordinates. They are also a free sanity check: if a frame reports a 60 cm forearm, that frame's landmarks are wrong and you should discard it rather than act on it.

Part 2: the actual MediaPipe call

run_pose.py
import mediapipe as mp

BaseOptions       = mp.tasks.BaseOptions
PoseLandmarker    = mp.tasks.vision.PoseLandmarker
PoseLandmarkerOptions = mp.tasks.vision.PoseLandmarkerOptions
VisionRunningMode = mp.tasks.vision.RunningMode

options = PoseLandmarkerOptions(
    base_options=BaseOptions(model_asset_path="pose_landmarker_full.task"),
    running_mode=VisionRunningMode.IMAGE)

with PoseLandmarker.create_from_options(options) as landmarker:
    image = mp.Image.create_from_file("person.jpg")
    result = landmarker.detect(image)

    for person in result.pose_world_landmarks:          # metres, hip-midpoint origin
        for i, lm in enumerate(person):
            print(i, round(lm.x, 3), round(lm.y, 3), round(lm.z, 3), round(lm.visibility, 3))

No output block for this one, deliberately. The numbers depend entirely on your photograph and on the model version you downloaded. Printing invented coordinates would teach you to expect values that will not appear.

Two API notes worth carrying: result.pose_landmarks gives normalised screen coordinates in the range zero to one, while result.pose_world_landmarks gives metres. VisionRunningMode.VIDEO requires detect_for_video(image, timestamp_ms) with monotonically increasing timestamps, and LIVE_STREAM requires detect_async plus a result_callback.

Common mistakes

The protobuf conflict. Symptom: AttributeError: 'MessageFactory' object has no attribute 'GetPrototype' on import. Fix: a fresh virtual environment, or pin protobuf<5. It often does not stop the code from working, which makes it worse, because it looks like a harmless warning until something downstream fails.

Passing a NumPy BGR array straight from OpenCV. mp.Image expects RGB. Convert with cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) and wrap it as mp.Image(image_format=mp.ImageFormat.SRGB, data=rgb). Skipping the conversion degrades accuracy silently.

Mixing up normalised and world landmarks. Computing a joint angle from normalised screen coordinates gives an answer that changes when the person moves toward the camera, because the aspect ratio is not preserved. Use world landmarks for anything geometric.

Non-monotonic timestamps in VIDEO mode. The tracker uses timestamps to decide whether to re-detect. Feed them out of order and it throws.

Ignoring visibility and presence. Each landmark carries confidence scores. An occluded wrist still gets coordinates, and they are a guess. Gate your logic on visibility.

Acting on a single frame. Landmarks jitter. Apply a one-euro filter or an exponential moving average before any threshold-based rule, or your squat counter will register forty repetitions while the user stands still.

Try it yourself

Take the world-coordinate dictionary above and write a classify_arm(angle) function returning "straight", "bent" or "folded". Then add plus or minus three degrees of noise to every angle and rerun it a hundred times. Count how often a borderline pose flips category. That number tells you how much hysteresis your real rule needs.

What to learn next

  • OpenCV — reading frames, colour conversion and drawing the skeleton back on.
  • What is edge AI — why these models are built the way they are.
  • TensorFlow Lite — the runtime inside every .task bundle.

Researcher — Mathematics and papers.

Architecture

MediaPipe Tasks packages a pipeline, not a single model. The .task bundle is a zip containing several TFLite models plus metadata.

Pose Landmarker bundles a detector (BlazePose-family) and a landmark model. The published design (Bazarevsky et al., 2020, BlazePose, arxiv.org/abs/2006.10204) uses a detector–tracker split: the detector runs on the first frame and whenever tracking is lost; thereafter the landmark model's own output crops the next frame. This is what makes real-time CPU operation possible, since the detector is the expensive component and it rarely runs.

BlazePose's landmark head combines heatmap supervision during training with direct regression at inference. The heatmap branch is discarded after training, keeping the training signal density of a heatmap and the latency of regression. This is a clean answer to the trade-off analysed in heatmaps for keypoints.

Face Landmarker bundles three models: BlazeFace short-range detection, a 478-point mesh model, and a 52-coefficient blendshape predictor. It optionally returns a facial transformation matrix for effect rendering. The mesh derives from Kartynnik et al. (2019), Real-time Facial Surface Geometry from Monocular Video on Mobile GPUs (arxiv.org/abs/1907.06724); the iris refinement adding indices 468–477 comes from Ablavatski et al. (2020), Real-time Pupil Tracking from Monocular Video (arxiv.org/abs/2006.11341).

Hand Landmarker bundles a palm detector and a 21-point hand model (Zhang et al., 2020, MediaPipe Hands, arxiv.org/abs/2006.10214). Detecting palms rather than hands is a deliberate choice: palms are near-rigid and square, so a square anchor scheme works, whereas articulated fingers make hand bounding boxes highly variable.

The world-landmark coordinate frame

Pose world landmarks are metric, origin at the hip midpoint, and are not obtained by triangulation. They come from a regression head trained against 3D ground truth captured with a synchronised rig and augmented with synthetic data.

Two consequences follow:

  • The depth axis is the weakest. Monocular depth for a body is fundamentally ambiguous up to a scale-and-pose family, and the network resolves it with a learned prior over human proportions. Expect larger error in $z$ than in $x$ or $y$.
  • The output is metric but not calibrated to this subject. Limb lengths are predicted from a population prior, so an unusually tall or short person will have systematically biased lengths.

For measurement applications, treat world landmarks as a strong prior rather than a measurement, and validate against your own population.

Accuracy, honestly

MediaPipe does not publish COCO AP alongside academic pose methods, and the comparison would be misleading in either direction. BlazePose targets single-person, near-frontal, fitness-style imagery at 30+ FPS on a mobile CPU. HRNet and ViTPose target multi-person COCO AP with a GPU budget one to two orders of magnitude larger.

Independent comparisons of pose stacks consistently report that MediaPipe leads on latency-per-watt and trails HRNet-class top-down methods on accuracy, particularly for unusual poses, heavy occlusion and multi-person scenes. Choose it for on-device latency, not for benchmark accuracy.

The 33-point topology also differs from COCO's 17. Converting between them loses information in one direction and requires interpolation in the other, and any cross-comparison must state the mapping used.

API stability

The Tasks API (mediapipe.tasks.python.vision) superseded the legacy mediapipe.solutions API. The legacy solutions modules still ship in 0.10.x and are the source of the landmark enums used in the code above, but they are the older interface — build new work on Tasks.

MediaPipe's release cadence and its pinned dependencies (notably protobuf) change between minor versions. Pin the version in your project and re-test on upgrade. Model bundles served from the latest path can also change beneath you; for reproducible results, download the bundle once, hash it, and vendor it alongside your code rather than fetching latest at runtime.

Filtering

The one-euro filter (Casiez, Roussel and Vogel, CHI 2012) is the standard choice for landmark smoothing and is what MediaPipe uses internally for its own tracking. It adapts its cut-off frequency to the observed speed: heavy smoothing when the signal is slow, light smoothing when it moves fast, trading jitter against lag in the way a human perceiver finds acceptable.

A fixed low-pass filter cannot do both. If your application feels either jittery at rest or laggy in motion, a one-euro filter is the first thing to try.

Papers

What to learn next

  • OpenCV — reading frames, colour conversion and drawing the skeleton back on.
  • What is edge AI — why these models are built the way they are.
  • TensorFlow Lite — the runtime inside every .task bundle.