3D human pose and SMPL
Predicting a body in 3D from one camera is ambiguous, so modern methods stop predicting points and instead predict the few dozen numbers that drive a statistical body model.
- 17 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
SMPL is a model of the human body that a few dozen numbers can reshape and pose.
Think of a tailor's dummy with dials on the back. One set of dials changes the build — taller, broader, narrower shoulders. Another set bends the joints. Turn the dials and the whole dummy reshapes smoothly.
Now the problem becomes easy to state. Instead of guessing where every point on a body is, look at a photo and work out the dial settings.
Why guessing points does not work
Take a photograph of somebody with an arm pointing toward the camera. Is the forearm reaching toward you, or away from you?
Both produce almost the same picture. This is not a hard problem the models have failed to solve. The information is not present in a single image. A camera throws depth away, and no amount of cleverness recovers what was discarded.
There is a second, related trap. A person could be small and near, or large and far. Those two also look the same. Size and distance come as a pair from one camera, and you cannot separate them.
So a system that predicts three numbers per joint has to invent depth, with nothing to constrain the invention. The results are what you would expect. Limbs of impossible lengths. Elbows bending backwards. A skeleton no body could adopt.
What the model fixes
The dummy with dials fixes all of it, because the dials cannot produce a body that is not a body.
There is no dial for "make the forearm twice as long". There is no dial for "bend the knee the wrong way". Every setting produces a plausible human. The model absorbed the shape of real people from thousands of body scans. That knowledge limits what it can express.
The prediction problem shrinks too. Instead of hundreds of free numbers, the system predicts around ten for build and around seventy for posture. Fewer things to guess, and every guess is anchored.
photo
│
▼
network predicts the dial settings
│
▼
dials feed the body model
│
▼
a full body in three dimensions: every surface point, every jointHow the dummy works inside
Two ideas, both simple.
Start from an average body and add adjustments. There is one mean shape, computed from many scans. Each build dial says "and now move each surface point a little, in this direction". Adding the adjustments gives your body, at rest.
Bend the joints, and drag the skin along. Each point on the surface is attached to nearby bones with different strengths. A point in the middle of the upper arm follows the upper arm entirely. A point on the elbow follows both bones about half each, so the skin creases rather than tearing.
There is a third idea, and it exists to patch a flaw in the second. Dragging skin by weighted averages pinches the surface shut at a bent joint, like a twisted sweet wrapper. So the model adds a correction that depends on the pose: when this joint bends, push these points outward. Those corrections were learned from scans of real people bending.
The catch nobody mentions
The SMPL model files are not free to use for anything you sell.
You register on the research website, agree to a licence, and download them. The licence permits non-commercial research, education and artistic work. Commercial use requires a separate agreement.
This surprises teams late, after a prototype works. If you are building a product, sort the licence out first. Some open alternatives exist, and their quality and community are not comparable.
Where you have already seen this
- A game character copying your movements from a phone camera.
- Virtual clothes try-on showing a garment on a body like yours.
- Film and game studios turning video into animation data.
- Sports analysis showing a full body rather than a stick figure.
Remember this
- One camera cannot see depth, so predicting loose points in space produces impossible bodies.
- SMPL turns the problem into predicting a few dozen dial settings that can only make real bodies.
- The model files require registration and are licensed for non-commercial use.
What to learn next
- Vision transformers — the backbone behind the current best mesh-recovery models.
- Autoencoders — learning a compact code for a complicated object, which is what the shape parameters are.
- Linear algebra — rotations, blend weights and the principal components behind the shape space.
Developer — Code and libraries.
Setup
pip install numpyRun against numpy 1.26. The real SMPL model file requires registration at smpl.is.tue.mpg.de and is licensed for non-commercial use, so this lesson builds a miniature model with the same four components. Six vertices, two joints, one shape parameter, in two dimensions — small enough to read every number, structured exactly like the real thing.
SMPL in miniature
import numpy as np
# A toy body model with the same four pieces SMPL has, at a thousandth of the size.
# SMPL: 6890 vertices, 24 joints, 10 shape parameters. Here: 6 vertices, 2 joints, 1 shape.
T = np.array([[0.0, 0.0], [0.0, 1.0], [0.4, 1.0], # template ("mean") body, in 2D
[0.8, 1.0], [0.8, 0.0], [0.4, 0.0]])
BS = np.array([[-0.3, 0.0], [-0.3, 0.0], [0.0, 0.0], # shape blend shape: beta widens the body
[ 0.3, 0.0], [ 0.3, 0.0], [0.0, 0.0]])
BP = np.array([[0.0, 0.0], [0.0, 0.0], [0.0, 0.10], # pose blend shape: fixes the bent joint
[0.0, 0.0], [0.0, 0.0], [0.0, -0.10]])
W = np.array([[1.0, 0.0], [1.0, 0.0], [0.5, 0.5], # skinning weights, each row sums to 1
[0.0, 1.0], [0.0, 1.0], [0.5, 0.5]])
def joints(beta):
"""J(beta): joint centres are predicted FROM the shaped body, not fixed in advance."""
V = T + beta * BS
return np.array([V[[0, 1, 2, 5]].mean(0), V[[2, 3, 4, 5]].mean(0)])
def rot(deg):
r = np.deg2rad(deg)
return np.array([[np.cos(r), -np.sin(r)], [np.sin(r), np.cos(r)]])
def model(beta, theta, use_pose_blend=True):
V = T + beta * BS + (abs(theta) / 90.0) * BP * use_pose_blend
J = joints(beta)
R = [np.eye(2), rot(theta)] # joint 0 is the root; joint 1 bends
out = np.zeros_like(V)
for i, v in enumerate(V): # linear blend skinning
for j in (0, 1):
out[i] += W[i, j] * (R[j] @ (v - J[j]) + J[j])
return out, J
print("M(beta, theta) = LBS( T + B_S(beta) + B_P(theta), J(beta), theta, W )\n")
for beta, theta in ((0.0, 0.0), (1.0, 0.0), (0.0, 60.0), (1.0, 60.0)):
V, J = model(beta, theta)
print(f"beta={beta:+.1f} theta={theta:+5.1f} joints {np.round(J,3).tolist()}")
print(f" vertices {np.round(V,3).tolist()}")
print("\nwidth across the bending joint (vertex 2 to vertex 5):")
for theta in (0, 30, 60, 90):
a, _ = model(0.0, theta, use_pose_blend=False)
b, _ = model(0.0, theta, use_pose_blend=True)
print(f" theta={theta:>3} deg skinning alone {np.linalg.norm(a[2]-a[5]):.4f}"
f" with pose blend shape {np.linalg.norm(b[2]-b[5]):.4f}")
print("\nSkinning alone pinches the joint shut. That collapse is what B_P(theta) is for.")M(beta, theta) = LBS( T + B_S(beta) + B_P(theta), J(beta), theta, W )
beta=+0.0 theta= +0.0 joints [[0.2, 0.5], [0.6, 0.5]]
vertices [[0.0, 0.0], [0.0, 1.0], [0.4, 1.0], [0.8, 1.0], [0.8, 0.0], [0.4, 0.0]]
beta=+1.0 theta= +0.0 joints [[0.05, 0.5], [0.75, 0.5]]
vertices [[-0.3, 0.0], [-0.3, 1.0], [0.4, 1.0], [1.1, 1.0], [1.1, 0.0], [0.4, 0.0]]
beta=+0.0 theta=+60.0 joints [[0.2, 0.5], [0.6, 0.5]]
vertices [[0.0, 0.0], [0.0, 1.0], [0.205, 0.838], [0.267, 0.923], [1.133, 0.423], [0.695, -0.012]]
beta=+1.0 theta=+60.0 joints [[0.05, 0.5], [0.75, 0.5]]
vertices [[-0.3, 0.0], [-0.3, 1.0], [0.242, 0.773], [0.492, 1.053], [1.358, 0.553], [0.733, -0.077]]
width across the bending joint (vertex 2 to vertex 5):
theta= 0 deg skinning alone 1.0000 with pose blend shape 1.0000
theta= 30 deg skinning alone 0.9659 with pose blend shape 1.0303
theta= 60 deg skinning alone 0.8660 with pose blend shape 0.9815
theta= 90 deg skinning alone 0.7071 with pose blend shape 0.8485
Skinning alone pinches the joint shut. That collapse is what B_P(theta) is for.Reading it
Changing beta moved the joints. At beta=0 the joints sit at 0.20 and 0.60; at beta=1 they move to 0.05 and 0.75. This is the piece people miss. SMPL does not have a fixed skeleton with a skin stretched over it — the joint locations are a learned function of the body shape. A broader person's shoulder joint is further out, and the model knows that.
The pinching is real, and it is arithmetic. Skinning alone takes the joint width from 1.0000 to 0.7071 at 90 degrees. That is exactly $\cos(45°)$, because averaging two rotations by half each gives a chord, not an arc. This is the classic "candy wrapper" artefact of linear blend skinning, and it is why a naive skinned character's elbow looks deflated when bent.
The pose blend shape recovers it to 0.8485. Not to 1.0000, because our correction is a crude two-vertex nudge. In the real model these corrections are learned from thousands of registered scans of people in many poses, and they close the gap far more convincingly.
Notice what the correction depends on. The pose blend shape is a function of theta, not of beta. Bending drives it. In real SMPL the pose blend shape is linear in the elements of the joint rotation matrices, which is what makes the whole model differentiable and trainable end to end.
What the real model looks like
| SMPL | SMPL-X | |
|---|---|---|
| Vertices | 6,890 | 10,475 |
| Joints | 24 (23 plus root) | 54 |
| Shape parameters | 10 | 10, jointly trained for body, face and hands |
| Extras | — | Articulated hands (MANO) and expressive face (FLAME) |
Pose is parameterised as an axis-angle vector per joint, giving 72 numbers for SMPL (24 joints times 3), plus 3 for global translation. Around 85 numbers describe a complete posed human body.
Working with it in practice
# pip install smplx torch
# Model files: register at smpl.is.tue.mpg.de or smpl-x.is.tue.mpg.de and download.
# Licence: non-commercial research, education and artistic use.
import torch, smplx
body = smplx.create("models", model_type="smpl", gender="neutral", batch_size=1)
out = body(betas=torch.zeros(1, 10), body_pose=torch.zeros(1, 69),
global_orient=torch.zeros(1, 3))
print(out.vertices.shape, out.joints.shape)No output block here, deliberately. This code cannot run without the licensed model files, and printing shapes I have not executed against those files would be inventing a result. Once the files are in place, vertices is a (1, 6890, 3) tensor and joints is (1, 45, 3) for the SMPL configuration — check the values yourself rather than trusting a tutorial, since joint counts differ between model variants and library versions.
Common mistakes
Treating axis-angle as a vector space. Rotations do not add, and averaging two axis-angle vectors gives something meaningless. Convert to rotation matrices or 6D representations before interpolating. Networks that regress axis-angle directly train worse than those regressing 6D rotations, for exactly this reason.
Confusing the model's joints with a keypoint dataset's joints. SMPL's 24 joints are not COCO's 17 and are not MediaPipe's 33. Every pipeline needs an explicit regressor or mapping, and silently mismatching them produces a systematic offset that looks like poor accuracy.
Fitting shape and camera at the same time from one image. Body size and camera distance are entangled. Standard practice fixes a weak-perspective camera with an assumed focal length, and accepts that absolute scale is unrecoverable.
Assuming the mesh is measurement-grade. A monocular reconstruction is a plausible body consistent with the image. Taking tailoring measurements from it requires validation against the population you care about.
Ignoring the licence until launch. The most expensive mistake on this page. Check it in week one.
Try it yourself
Add a second shape parameter BS2 that makes the body taller by moving the top two vertices upward. Recompute joints(beta) with both parameters. Confirm the joint centres move with the shape. Then bend the joint and check that the pose blend shape, which was tuned for the default build, over- or under-corrects. That mismatch is why the real model's pose corrections have to be learned across many body shapes rather than one.
What to learn next
- Vision transformers — the backbone behind the current best mesh-recovery models.
- Autoencoders — learning a compact code for a complicated object, which is what the shape parameters are.
- Linear algebra — rotations, blend weights and the principal components behind the shape space.
Researcher — Mathematics and papers.
The SMPL function
Loper, Mahmood, Romero, Pons-Moll and Black (SIGGRAPH Asia 2015), SMPL: A Skinned Multi-Person Linear Model (files.is.tue.mpg.de/black/papers/SMPL2015.pdf), define
$$ M(\beta, \theta) = W\big(T_P(\beta, \theta),\; J(\beta),\; \theta,\; \mathcal{W}\big) $$
$$ T_P(\beta, \theta) = \bar{T} + B_S(\beta) + B_P(\theta) $$
- $\bar{T} \in \mathbb{R}^{3N}$ — the mean template, $N = 6890$ vertices.
- $\beta \in \mathbb{R}^{10}$ — shape coefficients; $B_S(\beta) = \sum_n \beta_n S_n$ with $S_n$ the principal components of registered body scans in the rest pose.
- $\theta \in \mathbb{R}^{3K+3}$ — axis-angle rotations for $K = 23$ joints plus a global orientation.
- $B_P(\theta) = \sum_n \big(R_n(\theta) - R_n(\theta^)\big) P_n$ — pose corrective blend shapes, linear in the elements of the joint rotation matrices relative to the rest pose $\theta^$. This linearity is the paper's central design choice: it keeps the model differentiable and cheap while capturing pose-dependent deformation.
- $J(\beta) = \mathcal{J}\,(\bar{T} + B_S(\beta))$ — a learned sparse regressor mapping the shaped surface to joint centres. Joints move with shape.
- $\mathcal{W} \in \mathbb{R}^{N \times K}$ — skinning weights, rows summing to one.
Linear blend skinning is then
$$ \bar{t}'i = \sum{k=1}^{K} w_{k,i}\, G'_k(\theta, J)\; \bar{t}_i $$
$G'_k$ is the world transform of joint $k$ with the rest-pose transform removed. The "candy wrapper" collapse is the failure mode of this weighted average of rigid transforms, and $B_P$ is the learned correction for it.
SMPL-X (Pavlakos et al., CVPR 2019, arxiv.org/abs/1904.05866) extends this to $N = 10475$ vertices and $K = 54$ joints, adding articulated hands from MANO and an expressive face from FLAME, with shape parameters trained jointly across body, face and hands.
Why monocular 3D pose is ill-posed
A perspective camera maps $\mathbb{R}^3 \to \mathbb{R}^2$, so depth is unrecoverable without priors. Two specific ambiguities dominate:
Depth reflection. For a limb whose 2D projection is known, two 3D configurations — toward and away from the camera — project identically. With $J$ joints this generates up to $2^{J}$ consistent skeletons, pruned only by kinematic and anthropometric constraints.
Scale–depth entanglement. Under weak perspective, a body of size $s$ at depth $d$ is indistinguishable from one of size $\alpha s$ at depth $\alpha d$. Absolute metric scale is not identifiable from a single image without a known focal length or an object of known size.
A parametric body model resolves the first by construction, since $\beta$ and $\theta$ cannot express a non-anatomical body, and constrains the second by supplying a population prior over body size.
Method lineage
Optimisation. SMPLify (Bogo et al., ECCV 2016, arxiv.org/abs/1607.08128) detects 2D joints, then minimises reprojection error over $(\beta, \theta)$ with a pose prior fitted as a Gaussian mixture over MoCap data and an interpenetration penalty. Accurate, slow, and dependent on initialisation.
Direct regression. HMR (Kanazawa et al., CVPR 2018, arxiv.org/abs/1712.06584) regresses $(\beta, \theta, \text{camera})$ from image features with an iterative error-feedback loop, supervised by 2D reprojection loss plus an adversarial prior discriminating real from implausible poses. The adversarial prior removed the need for paired 3D annotations, which is what made in-the-wild training viable.
Hybrid. SPIN (Kolotouros et al., ICCV 2019) runs SMPLify inside the training loop, using the optimiser's output as pseudo-ground-truth for the regressor, which in turn initialises the optimiser. Each component fixes the other's weakness.
Transformers. HMR 2.0 / 4D-Humans (Goel, Pavlakos, Rajasegaran, Kanazawa and Malik, ICCV 2023, arxiv.org/abs/2305.20091) replaces the convolutional backbone and iterative head with a plain ViT encoder and transformer decoder, reporting substantially improved robustness on unusual poses, and feeds the per-frame reconstructions into a 3D tracker that maintains identity through occlusion.
Probabilistic output. Because the problem is genuinely ambiguous, several methods output a distribution rather than a point estimate — ProHMR (Kolotouros et al., ICCV 2021) with normalising flows, and diffusion-based samplers more recently. This is the intellectually correct framing, and it is under-used in applications that want one answer.
Evaluation
- MPJPE — mean per-joint position error in millimetres, after aligning the root joint.
- PA-MPJPE — the same after a full Procrustes alignment, removing global rotation, translation and scale. It measures pose quality only and is the fairer comparison across methods with different camera handling.
- PVE / MPVPE — per-vertex error, which additionally measures shape quality.
- Datasets — Human3.6M (indoor, MoCap), 3DPW (in the wild, IMU-derived ground truth), EMDB and AGORA (synthetic with exact ground truth).
Two cautions. Human3.6M is indoor, small-population and largely saturated; report 3DPW. And PA-MPJPE flatters methods with poor global orientation, so quote both.
Licensing
SMPL and SMPL-X model files are distributed by the Max Planck Institute under a non-commercial licence, requiring registration and granting use for non-commercial scientific research, non-commercial education and non-commercial artistic projects. Commercial use requires a separate agreement, and this applies to the model parameters themselves, not only to the code.
The practical implication for a product team: the model file is a dependency with a licence, exactly like a proprietary library. Resolve it before it is embedded in a codebase. Open alternatives exist and differ in quality and ecosystem support; evaluate them early rather than late.
Papers
- Loper et al., SMPL: A Skinned Multi-Person Linear Model, SIGGRAPH Asia 2015
- Bogo et al., Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image, ECCV 2016 — arxiv.org/abs/1607.08128
- Kanazawa et al., End-to-end Recovery of Human Shape and Pose, CVPR 2018 — arxiv.org/abs/1712.06584
- Pavlakos et al., Expressive Body Capture: 3D Hands, Face, and Body from a Single Image (SMPL-X), CVPR 2019 — arxiv.org/abs/1904.05866
- Kolotouros et al., Learning to Reconstruct 3D Human Pose and Shape via Model-fitting in the Loop (SPIN), ICCV 2019 — arxiv.org/abs/1909.12828
- Kolotouros et al., Probabilistic Modeling for Human Mesh Recovery, ICCV 2021 — arxiv.org/abs/2108.11944
- Goel et al., Humans in 4D: Reconstructing and Tracking Humans with Transformers, ICCV 2023 — arxiv.org/abs/2305.20091
What to learn next
- Vision transformers — the backbone behind the current best mesh-recovery models.
- Autoencoders — learning a compact code for a complicated object, which is what the shape parameters are.
- Linear algebra — rotations, blend weights and the principal components behind the shape space.