3D Vision and Depth

Structure from motion

Move one camera around an object and you can recover both the 3D shape and every camera position — but never the real-world size, unless you measure one thing yourself.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The catch you cannot avoid
  6. The second catch
  7. Where you have already seen it
  8. The honest part
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Structure from motion builds a 3D model from many ordinary photos, taken as you walk around.

The analogy you have already lived

Walk around a statue in a park and you understand its shape. Not from one glance — from the way its outline changes as you move.

You were also, without noticing, keeping track of where you had walked. Shape and your own path came together.

That is exactly the trick. The name says it: structure (the shape) from motion (the camera moving).

Why it exists

Stereo needs two cameras bolted a known distance apart. Most people have one camera, in a phone.

Structure from motion removes the rig. Move one camera, and consecutive photos act as a stereo pair. You get the same information without any special hardware.

How it works

   many photos
       |
   find the same physical points across them   (feature matching)
       |
   work out how the camera moved between two photos
       |
   with the motion known, work out where each point sits in 3D
       |
   add the next photo: its position comes from points you already have
       |
   repeat, then polish everything together

The last step is important. Errors build up as you add photos. At the end, the shape and all the camera positions are adjusted at once.

The catch you cannot avoid

You get the shape, and you never get the size.

Move the camera 60 centimetres and photograph a real chair. Or move it 6 metres and photograph a giant model of that chair. The photos are identical.

The run below shows this plainly. The recovered camera motion always has a length of exactly 1, no matter what the truth was. Not approximately one. Exactly.

To get metres, you must supply one real measurement from outside the pictures: a ruler in the scene, the actual step you took, GPS, or a wheel sensor.

The second catch

Everything depends on matching the same points across photos. That means the object needs texture.

A statue with carved detail works beautifully. A smooth white sculpture, a glass vase or a plain wall gives nothing to match. The reconstruction then has holes exactly where the surface is featureless.

Where you have already seen it

  • Photogrammetry apps building a 3D model from a phone video.
  • Google Earth's 3D buildings, built from aerial photos.
  • Archaeological and heritage scanning of monuments.
  • Drone surveys measuring a construction site.
  • The first stage of every gaussian splatting capture.

The honest part

The numbers in the code below look excellent. They come from a clean synthetic scene with plenty of matched points.

Real captures fail in ways this demo cannot show. Repeated architecture confuses matching. Moving people break the assumption that the world is still. A camera that walks in a straight line without turning gives almost no triangulation angle, so depth becomes very uncertain.

Good capture technique matters more than the software.

Remember this

  • Many photos of one scene give both the shape and the camera path.
  • The size is never recoverable; one real measurement supplies it.
  • No texture means no matches, and no matches means no reconstruction.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python==4.10.0.84 numpy==1.26.4

We generate a point cloud, photograph it from three known viewpoints with realistic pixel noise, then recover the geometry and score it against the truth.

Two views, then a third

sfm.py
import cv2
import numpy as np

rng = np.random.default_rng(3)
K = np.array([[600.0, 0, 320.0], [0, 600.0, 240.0], [0, 0, 1.0]])

# A cloud of 3-D points, roughly 4 to 8 metres in front of the first camera.
N = 140
points = np.column_stack([rng.uniform(-2.0, 2.0, N),
                          rng.uniform(-1.5, 1.5, N),
                          rng.uniform(4.0, 8.0, N)])

# Camera 2: steps 0.6 m to the right and turns slightly to keep the cloud in view.
R_true, _ = cv2.Rodrigues(np.array([0.03, -0.10, 0.01]))
t_true = np.array([[-0.60], [0.02], [0.05]])
print(f"true camera motion: baseline {np.linalg.norm(t_true):.3f} m")


def project(P, R=np.eye(3), t=np.zeros((3, 1))):
    cam = (R @ P.T + t).T
    px = (K @ (cam / cam[:, 2:3]).T).T[:, :2]
    return px + rng.normal(0, 0.4, px.shape)        # half a pixel of matching noise


p1 = project(points)
p2 = project(points, R_true, t_true)
print(f"{N} matched points, with 0.4 px of noise added to every one\n")

E, mask = cv2.findEssentialMat(p1, p2, K, method=cv2.USAC_MAGSAC, prob=0.9999,
                               threshold=1.0)
n_in, R_est, t_est, mask_pose = cv2.recoverPose(E, p1, p2, K, mask=mask)
print(f"essential matrix inliers: {int(mask.sum())} of {N}")
print(f"points passing the in-front-of-both-cameras test: {n_in}\n")

angle = np.degrees(np.arccos((np.trace(R_est @ R_true.T) - 1) / 2))
print("rotation recovered:")
print(f"   angle between the true and estimated rotation: {angle:.4f} degrees")
print("translation recovered:")
print(f"   true direction      {np.round((t_true / np.linalg.norm(t_true)).ravel(), 4)}")
print(f"   estimated direction {np.round(t_est.ravel(), 4)}")
print(f"   estimated length    {np.linalg.norm(t_est):.4f}  <- always exactly 1")
print("   the length is not recoverable. Two views fix the shape, never the size.\n")

P1 = K @ np.hstack([np.eye(3), np.zeros((3, 1))])
P2 = K @ np.hstack([R_est, t_est])
X = cv2.triangulatePoints(P1, P2, p1.T, p2.T)
X = (X[:3] / X[3]).T

print("triangulated cloud, before any scale is applied:")
print(f"   depth range  {X[:, 2].min():.2f} to {X[:, 2].max():.2f}")
print(f"   true depths  {points[:, 2].min():.2f} to {points[:, 2].max():.2f}")
scale = np.median(points[:, 2] / X[:, 2])
b_true = float(np.linalg.norm(t_true))
print(f"   one number fixes it: multiply everything by {scale:.4f}")
print(f"   with a perfect pose that number would be the true baseline, {b_true:.4f}")
print(f"   the {100 * abs(scale - b_true) / b_true:.1f}% gap is what 0.4 px of noise costs\n")

X_scaled = X * scale
err = np.linalg.norm(X_scaled - points, axis=1)
print("after applying the scale, how good is the 3-D structure?")
print(f"   median point error {np.median(err) * 100:.2f} cm")
print(f"   90th percentile    {np.percentile(err, 90) * 100:.2f} cm")
print(f"   worst point        {err.max() * 100:.2f} cm\n")

reproj = (K @ (X_scaled / X_scaled[:, 2:3]).T).T[:, :2]
print(f"reprojection error in view 1: {np.median(np.linalg.norm(reproj - p1, axis=1)):.3f} px")
print("   that is about the noise we injected, which is what a good fit looks like\n")

print("adding a third view: solve its pose against the points we already have")
R3_true, _ = cv2.Rodrigues(np.array([-0.02, 0.18, 0.0]))
t3_true = np.array([[0.55], [-0.03], [0.10]])
p3 = project(points, R3_true, t3_true)
ok, rvec, tvec, inl = cv2.solvePnPRansac(X_scaled.astype(np.float64), p3, K, None,
                                         reprojectionError=2.0)
R3, _ = cv2.Rodrigues(rvec)
ang3 = np.degrees(np.arccos(np.clip((np.trace(R3 @ R3_true.T) - 1) / 2, -1, 1)))
print(f"   solvePnPRansac used {len(inl)} of {N} points")
print(f"   rotation error {ang3:.4f} degrees")
print(f"   translation error {np.linalg.norm(tvec - t3_true) * 100:.2f} cm")
print("   note the third camera comes out in METRES, because the cloud already had a scale")
print("   this is exactly how incremental structure from motion grows a reconstruction")
Output
true camera motion: baseline 0.602 m
140 matched points, with 0.4 px of noise added to every one

essential matrix inliers: 138 of 140
points passing the in-front-of-both-cameras test: 138

rotation recovered:
   angle between the true and estimated rotation: 0.2836 degrees
translation recovered:
   true direction      [-0.996   0.0332  0.083 ]
   estimated direction [-0.9951  0.0329  0.0932]
   estimated length    1.0000  <- always exactly 1
   the length is not recoverable. Two views fix the shape, never the size.

triangulated cloud, before any scale is applied:
   depth range  6.84 to 14.01
   true depths  4.01 to 7.97
   one number fixes it: multiply everything by 0.5735
   with a perfect pose that number would be the true baseline, 0.6024
   the 4.8% gap is what 0.4 px of noise costs

after applying the scale, how good is the 3-D structure?
   median point error 5.33 cm
   90th percentile    16.00 cm
   worst point        33.00 cm

reprojection error in view 1: 0.204 px
   that is about the noise we injected, which is what a good fit looks like

adding a third view: solve its pose against the points we already have
   solvePnPRansac used 127 of 140 points
   rotation error 0.1618 degrees
   translation error 1.66 cm
   note the third camera comes out in METRES, because the cloud already had a scale
   this is exactly how incremental structure from motion grows a reconstruction

Reading the output carefully

estimated length 1.0000. Not close to one. Exactly one, because recoverPose normalises the translation. The essential matrix is defined only up to scale, so the length carries no information and the library returns a unit vector as an honest admission of that.

The rotation came back to 0.2836 degrees, the direction to about a degree. Rotation is far better determined than translation from a two-view solve, and that asymmetry persists in real pipelines. It is why visual odometry drifts in position long before it drifts in heading.

0.5735 against an ideal 0.6024 — a 4.8% scale error. This one deserves attention. The pose is slightly wrong, so the reconstruction is not an exact scaled copy of the truth; it is a slightly distorted one, and no single scale factor fits it perfectly. Bundle adjustment does not remove this: the solution above already sits at the minimum of reprojection error, and 0.4 pixels of matching noise on a 0.6 m baseline does not carry enough information to do better.

Median point error 5.33 cm, worst point 33 cm. The distribution has a long tail. The worst points are the ones with the smallest triangulation angle — near the direction of camera motion, where the two rays are almost parallel. Every SfM pipeline filters points by triangulation angle for this reason, typically discarding anything under one or two degrees.

Reprojection error 0.204 px against injected noise of 0.4 px per coordinate. Being below the noise level is expected, not suspicious: the triangulated point is fitted to both observations, so it absorbs part of the noise. A reprojection error far below the noise means overfitting; far above means a wrong model.

The third camera came back to 1.66 cm and 0.16 degrees, in metres. This is the incremental loop. Once a scaled cloud exists, solvePnP places each new camera in that same frame, and new points triangulated from it inherit the scale. Register, triangulate, adjust, repeat.

Doing it on real photos

Do not write this pipeline. Use COLMAP:

bash
colmap automatic_reconstructor --workspace_path ./work --image_path ./images

COLMAP handles feature extraction, exhaustive or vocabulary-tree matching, geometric verification, incremental registration, repeated bundle adjustment, and dense multi-view stereo. Its output — camera poses and a sparse cloud — is the standard input to NeRF and gaussian splatting pipelines.

Capture advice that matters more than any setting:

  • Move sideways, not only forwards. Forward motion gives near-zero triangulation angle at the image centre.
  • Overlap by 60 to 80% between consecutive shots.
  • Lock focus and exposure. Autofocus changes the intrinsics between frames.
  • Avoid motion blur. A blurred frame contributes bad matches, not missing ones, which is worse.
  • Include a ruler or an object of known size if you need metres.

Common mistakes

Passing unnormalised points to findEssentialMat. Pass pixel coordinates and K, or normalised coordinates and no K. Mixing the two silently produces nonsense.

Skipping the cheirality check. An essential matrix decomposes into four pose candidates. recoverPose picks the one placing points in front of both cameras. Doing the decomposition by hand and forgetting this gives a mirrored reconstruction.

Triangulating from a nearly-pure rotation. With no translation there is no parallax and no depth. Check the median triangulation angle before trusting anything.

Forgetting the essential matrix is degenerate on a plane. A planar scene has infinitely many valid essential matrices. Detect it, and use a homography-based initialisation instead, as DEGENSAC does.

Assuming the scene is static. Every moving object is a set of outliers with a consistent, wrong geometry — the hardest kind for RANSAC to reject.

Try it yourself

Change t_true to [[-0.05], [0.02], [0.05]], a 7 cm baseline instead of 60 cm, and rerun. The rotation stays accurate and the 3D errors explode. Small baselines are the single most common cause of bad reconstructions from a phone video walked in a straight line.

What to learn next

Researcher — Mathematics and papers.

The two-view solve

For calibrated cameras, corresponding normalised points $\hat{\mathbf{x}}_1, \hat{\mathbf{x}}_2$ satisfy the epipolar constraint

$$ \hat{\mathbf{x}}_2^{\top} E \, \hat{\mathbf{x}}1 = 0, \qquad E = [\mathbf{t}]\times R $$

Where $[\mathbf{t}]_\times$ is the skew-symmetric matrix of $\mathbf{t}$. $E$ is $3\times3$, defined up to scale, and satisfies $\det E = 0$ plus the two equal-singular-value condition, leaving 5 degrees of freedom — 3 for rotation, 2 for translation direction. The magnitude of $\mathbf{t}$ is absent from the equation, which is the algebraic statement of the scale ambiguity.

Nistér (TPAMI 2004), An efficient solution to the five-point relative pose problem, gives the minimal solver: 5 correspondences, up to 10 solutions from a tenth-degree polynomial. Five points is why RANSAC on $E$ converges quickly, following the iteration-count argument in RANSAC for robust fitting.

The uncalibrated analogue is the fundamental matrix $F = K_2^{-\top} E K_1^{-1}$, with 7 degrees of freedom and the classic 8-point algorithm (Longuet-Higgins, 1981; Hartley, 1997 for the normalisation that makes it usable).

Decomposing $E = U \operatorname{diag}(1,1,0) V^\top$ yields four $(R, \mathbf{t})$ candidates. Only one places points in front of both cameras — the cheirality constraint (Hartley, 1998).

Triangulation

Given $P_1, P_2$ and a correspondence, each observation gives two linear constraints on the homogeneous point $\mathbf{X}$ from $\mathbf{x} \times P\mathbf{X} = \mathbf{0}$. Stacking four rows and taking the smallest right singular vector is the DLT triangulation in cv2.triangulatePoints.

This minimises an algebraic error. The statistically correct objective is

$$ \min_{\mathbf{X}} \; d\bigl(\mathbf{x}_1, P_1\mathbf{X}\bigr)^2 + d\bigl(\mathbf{x}_2, P_2\mathbf{X}\bigr)^2 $$

which Hartley and Sturm (1997) solve exactly for two views by reducing it to a degree-6 polynomial. The difference matters at small triangulation angles, which is precisely where the errors in the output above come from.

Uncertainty scales as $\sigma_Z \propto \frac{Z^2 \sigma_x}{f b}$ with $b$ the baseline — the same relation as stereo, and the reason a small baseline is fatal.

Bundle adjustment

The global refinement minimises reprojection error over all cameras and points jointly:

$$ \min_{{R_j, \mathbf{t}_j},\, {\mathbf{X}i}} \sum{i}\sum_{j} w_{ij} \, \rho\Bigl(\bigl\lVert \mathbf{x}_{ij} - \pi(K_j, R_j, \mathbf{t}_j, \mathbf{X}_i) \bigr\rVert^2\Bigr) $$

with $w_{ij} = 1$ when point $i$ is seen in view $j$, and $\rho$ a robust kernel (Huber or Cauchy) to limit the influence of surviving mismatches.

Two structural facts make this tractable at scale.

Sparsity. Each residual touches one camera and one point, so the Jacobian is block-sparse and $J^\top J$ has an arrowhead structure. The Schur complement eliminates the point block, which is block-diagonal and cheap to invert, leaving a much smaller system in the camera parameters. This is the difference between a solvable and an unsolvable problem for thousands of images.

Gauge freedom. The objective is invariant under any global similarity transformation: 7 degrees of freedom, being 3 rotation, 3 translation and 1 scale. $J^\top J$ is therefore rank-deficient by 7. Implementations either fix one camera and one distance, or add a gauge-fixing prior. Ignoring it gives a singular normal-equation system.

Triggs et al. (2000), Bundle adjustment — a modern synthesis, remains the definitive treatment. Ceres Solver is the standard implementation and is what COLMAP uses.

Incremental against global

Incremental SfM (COLMAP; Schönberger and Frahm, CVPR 2016) starts from a well-conditioned pair, then repeatedly registers the image seeing the most existing points, triangulates, and runs local plus periodic global bundle adjustment. It is robust and its cost is roughly quadratic in image count because of repeated re-adjustment.

Global SfM estimates all rotations at once from pairwise relative rotations (rotation averaging), then all translations (translation averaging), then triangulates and adjusts once. Far faster, historically less robust — translation averaging is sensitive to the fact that pairwise translations are known only up to scale.

GLOMAP (Pan et al., ECCV 2024) narrowed that gap substantially by replacing separate translation averaging with a joint global positioning step over cameras and points, reaching incremental-level accuracy at global-level speed. It is a drop-in alternative to COLMAP's mapper.

The feed-forward turn

DUSt3R (Wang et al., CVPR 2024) reframed two-view reconstruction as direct regression of pointmaps: for an image pair, predict both images' 3D points in a common frame, with no correspondence, no pose and no intrinsics as input. MASt3R (Leroy et al., ECCV 2024) added a matching head for accuracy.

VGGT (Wang et al., CVPR 2025, best paper award) extended this to hundreds of views in a single forward pass, predicting intrinsics, extrinsics, depth maps, pointmaps and tracks jointly. Follow-on work through 2025 and 2026 has targeted its limits — higher resolution, longer streams, larger scenes.

Two honest observations about this shift.

  • It is a genuine capability change for hard cases: few images, wide baselines, textureless surfaces, no overlap that classical matching can find.
  • It is not yet a replacement for metric survey work. Accuracy on well-textured, well-covered captures still favours COLMAP plus bundle adjustment, and a common pattern is to use a feed-forward model for initialisation followed by classical refinement. The scale ambiguity is unchanged by any of it — a learned model predicts a scale from priors, exactly as monocular depth does, with the same failure modes.

What to learn next