3D Vision and Depth

Stereo depth from two cameras

Two cameras a known distance apart recover depth from how far each point shifts between the images, and the accuracy falls off with the square of the distance.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The catch that decides everything
  6. The second catch
  7. Where you have already seen it
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Two cameras side by side can measure distance, because near things shift more between the pictures than far things.

The analogy you have already lived

Hold a finger up in front of your face. Close your left eye, then your right eye, alternately.

Your finger jumps sideways. Now hold it at arm's length and do it again — it jumps less. Look at a building through a window and it barely jumps at all.

That jump is the whole method. Big jump means near. Small jump means far. Your brain has been doing this every waking second of your life.

Why it exists

A single photo cannot tell distance. We saw exactly why in camera models and intrinsics: dividing by depth throws depth away.

Two cameras a known distance apart put the information back. Each one throws away depth differently, and comparing them recovers it.

How it works

Set the two cameras level and pointing the same way, a known distance apart. That distance is called the baseline.

Now a point in the world lands on both pictures, but at different sideways positions. The difference is called the disparity.

   left image:    the corner of a box sits at column 200
   right image:   the same corner sits at column 175
                       |
                  disparity = 25 pixels
                       |
   distance = (how zoomed in) x (baseline) / disparity

That last line is the entire method. One division.

Notice its shape. Disparity in the bottom means a big shift gives a small distance. That matches the finger experiment exactly.

The catch that decides everything

Look at the division again. Disparity is on the bottom, so when disparity is small, small errors in it become big errors in distance.

The run below measures this. The near box at 2 metres comes out perfect. The box at 3 metres is off by 5 centimetres. The wall at 6 metres is off by 20 centimetres. Same camera, same code.

The error grows with the square of the distance. Double the distance, four times the error.

This is why self-driving cars carry radar and lidar as well as stereo cameras. Stereo is excellent close up and useless far away.

The second catch

The computer has to work out which pixel in the left picture is which pixel in the right picture. To do that, the pixels have to be distinguishable.

Point a stereo camera at a plain white wall and it fails completely. Every patch looks like every other patch. In the run below, two blank grey images produce depth for exactly 0% of pixels.

This is why depth-sensing cameras like the older Kinect project a pattern of infrared dots onto the room. They are giving themselves texture to match on.

Where you have already seen it

  • Your own two eyes.
  • Portrait mode on a dual-camera phone.
  • Reversing sensors and parking cameras on newer cars.
  • Warehouse robots avoiding obstacles.
  • A VR headset tracking your hands.

Remember this

  • Distance comes from how far a point shifts between the two images.
  • Accuracy falls off with the square of the distance.
  • A texture-free surface gives no depth at all.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install opencv-python==4.10.0.84 numpy==1.26.4

We build a stereo pair from a depth map we chose, so every recovered number can be scored. The right image is rendered by sampling the texture at fractional positions, which means the true disparities are not whole numbers — otherwise the matcher would look better than it is.

A stereo pair with a known answer

stereo.py
import cv2
import numpy as np

H, W = 240, 320
FOCAL = 500.0          # pixels
BASELINE = 0.10        # metres between the two camera centres
PAD = 64
rng = np.random.default_rng(1)

# A depth map for the left camera: a back wall and two boxes in front of it.
depth_left = np.full((H, W), 6.0, np.float32)
depth_left[70:170, 60:150] = 3.0
depth_left[100:200, 180:280] = 2.0

# Speckle texture, wider than the image so the right camera has something to see.
texture = cv2.GaussianBlur(rng.integers(0, 256, (H, W + PAD), np.uint8), (3, 3), 0)
left = texture[:, PAD:PAD + W].copy()               # left pixel x samples texture x+PAD

# Where does each surface land in the right image? Paint far first, so a near
# surface correctly covers what is behind it.
disp_left = FOCAL * BASELINE / depth_left
depth_right = np.full((H, W), 6.0, np.float32)      # holes default to the back wall
order = np.argsort(-depth_left, axis=None)
ys, xs = np.unravel_index(order, depth_left.shape)
tx = np.round(xs - disp_left[ys, xs]).astype(int)
ok = (tx >= 0) & (tx < W)
depth_right[ys[ok], tx[ok]] = depth_left[ys[ok], xs[ok]]

# Now sample the texture at a fractional position, so disparities are not integers.
grid_x, grid_y = np.meshgrid(np.arange(W, dtype=np.float32),
                             np.arange(H, dtype=np.float32))
disp_right = FOCAL * BASELINE / depth_right         # the true disparity of each right pixel
map_x = grid_x + disp_right + PAD
right = cv2.remap(texture, map_x, grid_y, cv2.INTER_LINEAR)

print("disparity = focal * baseline / depth")
for z in [6.0, 3.0, 2.0]:
    print(f"   {z:.1f} m  ->  {FOCAL * BASELINE / z:6.2f} pixels of shift")
print("\ntwo of those are not whole numbers, so a matcher has to work below a pixel\n")

matcher = cv2.StereoSGBM_create(
    minDisparity=0,
    numDisparities=48,          # must be a multiple of 16
    blockSize=7,
    P1=8 * 7 * 7,               # small-change penalty, keeps surfaces smooth
    P2=32 * 7 * 7,              # large-change penalty, still allows real edges
    uniquenessRatio=8,
    speckleWindowSize=64,
    speckleRange=2,
)
disp = matcher.compute(left, right).astype(np.float32) / 16.0   # SGBM uses 4 fixed bits

valid = disp > 0
est_depth = np.where(valid, FOCAL * BASELINE / np.maximum(disp, 1e-6), np.nan)
print(f"pixels given a disparity: {100 * valid.mean():.1f}%")
print("(a strip on the left has no match, because that scene is off the right image)\n")

print("estimated depth in the middle of each surface:")
for name, (y, x), truth in [("back wall", (30, 250), 6.0),
                            ("middle box", (120, 105), 3.0),
                            ("near box", (150, 230), 2.0)]:
    patch = est_depth[y - 8:y + 8, x - 8:x + 8]
    good = patch[~np.isnan(patch)]
    print(f"   {name:<11} true {truth:.2f} m   measured {np.median(good):.3f} m"
          f"   error {abs(np.median(good) - truth) * 100:5.1f} cm")

inner = valid.copy()
inner[:, :52] = False                            # ignore the unmatchable left strip
err = np.abs(est_depth - depth_right)[inner]
print("\nover every matched pixel away from that strip:")
print(f"   median depth error {np.median(err) * 100:.1f} cm")
print(f"   90th percentile    {np.percentile(err, 90) * 100:.1f} cm")
print(f"   median disparity error {np.median(np.abs(disp - disp_right)[inner]):.3f} px")

print("\nprecision falls off with distance, and the maths says by how much:")
print("   one pixel of disparity error costs  depth^2 / (focal * baseline)  metres")
for z in [1, 2, 5, 10, 20]:
    print(f"   at {z:>2} m: {z * z / (FOCAL * BASELINE) * 100:7.1f} cm per pixel of error")

print("\nwhat happens without texture to match on:")
plain = np.full((H, W), 128, np.uint8)
plain_disp = matcher.compute(plain, plain).astype(np.float32) / 16.0
print(f"   two blank grey images: {100 * (plain_disp > 0).mean():.1f}% of pixels matched")
print("   a blank wall gives a stereo camera nothing to work with, at any distance")
Output
disparity = focal * baseline / depth
   6.0 m  ->    8.33 pixels of shift
   3.0 m  ->   16.67 pixels of shift
   2.0 m  ->   25.00 pixels of shift

two of those are not whole numbers, so a matcher has to work below a pixel

pixels given a disparity: 82.2%
(a strip on the left has no match, because that scene is off the right image)

estimated depth in the middle of each surface:
   back wall   true 6.00 m   measured 6.202 m   error  20.2 cm
   middle box  true 3.00 m   measured 2.952 m   error   4.8 cm
   near box    true 2.00 m   measured 2.000 m   error   0.0 cm

over every matched pixel away from that strip:
   median depth error 15.4 cm
   90th percentile    20.2 cm
   median disparity error 0.208 px

precision falls off with distance, and the maths says by how much:
   one pixel of disparity error costs  depth^2 / (focal * baseline)  metres
   at  1 m:     2.0 cm per pixel of error
   at  2 m:     8.0 cm per pixel of error
   at  5 m:    50.0 cm per pixel of error
   at 10 m:   200.0 cm per pixel of error
   at 20 m:   800.0 cm per pixel of error

what happens without texture to match on:
   two blank grey images: 0.0% of pixels matched
   a blank wall gives a stereo camera nothing to work with, at any distance

Reading the output carefully

The three surface errors are the lesson: 0.0, 4.8, 20.2 cm at 2, 3 and 6 metres. One matcher, one image pair. The only thing that changed is distance, and the error grew roughly with its square. From 2 m to 6 m is a factor of 3 in distance and roughly 9 in error — which is exactly what the table below it predicts.

median disparity error 0.208 px. SGBM interpolates a parabola through the three matching costs around the winner, giving sub-pixel precision. That 0.2-pixel figure is typical for a good matcher on well-textured input, and it is the number that everything else follows from.

82.2% of pixels got a disparity. The missing 18% are mostly a strip down the left edge. A point visible at the left of the left image shifts out of frame in the right image, so there is nothing to match it against. The width of that dead strip is exactly the maximum disparity — 48 pixels here. Every stereo rig loses a band on one side, and cropping it is standard.

0.0% on two blank images. Not a low number, zero. uniquenessRatio=8 rejects any match whose best cost fails to beat the runner-up by a margin, and on uniform grey every candidate ties. This is the correct behaviour: refusing to answer is better than inventing a surface.

numDisparities=48 had to be at least 25. The nearest surface has a disparity of 25 pixels, and a search range that does not reach it would leave the near box unmatched. Setting this from your scene's minimum depth is the first thing to get right, and it costs linearly in compute.

Deriving the depth formula

Two cameras with the same focal length $f$, separated by baseline $B$ along the x-axis, both looking straight ahead. A point at depth $Z$ and lateral offset $X$ projects to

$$ u_L = f\frac{X}{Z} + c_x, \qquad u_R = f\frac{X - B}{Z} + c_x $$

Subtracting gives the disparity $d = u_L - u_R = fB/Z$, so

$$ Z = \frac{fB}{d} $$

Differentiating, $\left|\frac{\partial Z}{\partial d}\right| = \frac{fB}{d^2} = \frac{Z^2}{fB}$. That is the row of numbers in the output, and it is why stereo range is fundamentally limited.

Real cameras need rectification first

The formula above assumes the two cameras are perfectly aligned, so a point sits on the same image row in both. Real rigs are never that precise.

Rectification warps both images so that corresponding points share a row, which turns matching from a 2D search into a 1D one:

python
R1, R2, P1, P2, Q, _, _ = cv2.stereoRectify(K1, d1, K2, d2, size, R, T)
map1x, map1y = cv2.initUndistortRectifyMap(K1, d1, R1, P1, size, cv2.CV_32FC1)
left_rect = cv2.remap(left_raw, map1x, map1y, cv2.INTER_LINEAR)

R and T come from cv2.stereoCalibrate, which is chessboard calibration run on both cameras at once. The Q matrix then converts a disparity map straight to 3D points with cv2.reprojectImageTo3D.

Common mistakes

Skipping rectification. Without it, the true match is a few rows away and SGBM never looks there. The symptom is a noisy, mostly-empty disparity map on real hardware that works fine in simulation.

Forgetting to divide by 16. StereoSGBM.compute returns CV_16S with 4 fractional bits. Treating it as integer pixels makes every distance 16 times too small.

Choosing the baseline badly. A wide baseline improves far-range accuracy and increases occlusion and the search range. A narrow one does the reverse. Pick it from the depth range you actually need.

Using StereoBM and expecting SGBM quality. StereoBM is block matching only, with no smoothness term. It is faster and considerably worse on low-texture regions. P1 and P2 are what make SGBM good, and tuning them matters.

Treating the disparity map as if every pixel is reliable. Invalidate low-confidence pixels, and check disp > minDisparity rather than disp > 0 when minDisparity is not zero.

Assuming both cameras have the same exposure. SGBM's Birchfield-Tomasi cost has some tolerance, but a stereo pair with automatic exposure running independently on each camera will match badly. Lock exposure and gain together.

Try it yourself

Change BASELINE from 0.10 to 0.30 and rerun. The disparities triple, so numDisparities must rise to at least 96, the dead strip on the left triples in width, and the error at 6 m falls to roughly a third. That single edit is the whole baseline trade-off.

What to learn next

Researcher — Mathematics and papers.

The correspondence problem

Given rectified images $I_L, I_R$, find for each pixel the disparity $d$ minimising a matching cost, subject to a smoothness prior. Formally, minimise an energy over the disparity field $D$:

$$ E(D) = \sum_{\mathbf{p}} C(\mathbf{p}, D_{\mathbf{p}}) + \sum_{\mathbf{p}}\sum_{\mathbf{q} \in N_{\mathbf{p}}} \Bigl( P_1 \, \mathbb{1}\bigl[|D_{\mathbf{p}} - D_{\mathbf{q}}| = 1\bigr] + P_2 \, \mathbb{1}\bigl[|D_{\mathbf{p}} - D_{\mathbf{q}}| > 1\bigr] \Bigr) $$

Where $C$ is the per-pixel matching cost and $N_{\mathbf{p}}$ the 4- or 8-neighbourhood. The two-level penalty is deliberate: $P_1$ is small so slanted surfaces are not over-penalised, while $P_2$ is large and — crucially — constant regardless of the jump size, so a true depth discontinuity costs no more than a two-pixel one. That is what preserves object boundaries.

Exact minimisation of this energy in 2D is NP-hard.

Semi-global matching

Hirschmüller (CVPR 2005; TPAMI 2008), Stereo processing by semiglobal matching and mutual information, is the practical answer, and it is what StereoSGBM implements.

Instead of a 2D optimisation, aggregate costs along 8 or 16 1D paths through each pixel, each solved exactly by dynamic programming:

$$ L_{\mathbf{r}}(\mathbf{p}, d) = C(\mathbf{p}, d) + \min\bigl(L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, d),\; L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, d\pm1) + P_1,\; \min_i L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, i) + P_2\bigr) - \min_i L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, i) $$

The trailing subtraction only prevents numeric overflow and does not change the argmin. Summing $L_{\mathbf{r}}$ over all directions $\mathbf{r}$ approximates the 2D energy well, at $O(WHD)$ cost. OpenCV's default mode uses 5 paths; StereoSGBM.MODE_HH uses 8 at higher memory cost.

The single-path predecessor, scanline dynamic programming, produces the characteristic horizontal streaking artefact, because rows are optimised independently. Aggregating over multiple directions is what removes it.

Matching costs

  • SAD / SSD over a window. Fast, and assumes identical brightness.
  • Census transform. Encode each pixel as the bit pattern of comparisons with its neighbours, then use Hamming distance. Invariant to any monotonic intensity change, which makes it robust to exposure and vignetting differences. It is the standard cost in embedded stereo.
  • Mutual information. Hirschmüller's original, robust to arbitrary intensity mappings, expensive and iterative.
  • Birchfield-Tomasi. Samples the linearly interpolated intensity, making the cost insensitive to image sampling. OpenCV's SGBM uses a BT-based cost over a small window.

Sub-pixel and the quadratic fit

After the winner $d^$ is found, fit a parabola through the costs at $d^-1, d^, d^+1$:

$$ \delta = \frac{C(d^-1) - C(d^+1)}{2\bigl(C(d^-1) - 2C(d^) + C(d^*+1)\bigr)} $$

This is where the 0.208-pixel median error in the output comes from. It introduces a known bias, pixel-locking, where estimates cluster toward integer disparities because the cost function is not locally parabolic. Correcting it matters for high-precision metrology and rarely otherwise.

Fundamental limits

Depth uncertainty propagates as

$$ \sigma_Z = \frac{Z^2}{fB}\,\sigma_d $$

With $\sigma_d \approx 0.2$ px, $f = 500$, $B = 0.1$ m, the standard deviation is 4 mm at 1 m and 40 cm at 10 m. No amount of algorithmic cleverness changes the $Z^2$ term; only $f$ (resolution), $B$ (rig geometry) or $\sigma_d$ (matching quality) can move.

Three failure modes are structural rather than algorithmic:

  • Textureless regions. No information exists in the images. Learned methods fill these by prior, not by measurement, which is a different claim about the world.
  • Occlusion. Points visible in one view only have no true disparity. Left-right consistency checking detects them; nothing recovers them.
  • Repetitive structure. A brick wall gives periodic cost minima. uniquenessRatio rejects these, at the cost of holes.

Learned stereo

End-to-end networks now dominate the KITTI and Middlebury leaderboards. The lineage: DispNet (Mayer et al., CVPR 2016) with a correlation layer; GC-Net (Kendall et al., ICCV 2017) introducing an explicit 3D cost volume with 3D convolutions and a differentiable soft-argmin; PSMNet adding spatial pyramid pooling; RAFT-Stereo (Lipson et al., 3DV 2021) applying iterative recurrent refinement to a correlation pyramid, which remains a strong and widely used baseline.

Two honest caveats before choosing one over SGBM.

  • Cross-domain generalisation is the weak point. Networks trained on synthetic data or KITTI can degrade on a different sensor, baseline or scene type. SGBM has no training distribution to leave.
  • Learned methods interpolate textureless regions. They produce a smooth, plausible surface where SGBM produces a hole. For a rendering task that is better. For an obstacle detector it is a hallucinated surface, and the hole was the honest answer.

Active stereo sidesteps the texture problem entirely by projecting a pattern — the approach of the Intel RealSense D400 series and the original Kinect — and remains the most reliable way to get dense indoor depth.

What to learn next

What to learn next

These follow on from what you just read.

  • 3D Vision and Depth

    Monocular depth estimation

    A network can guess depth from a single photo by recognising what things usually look like, but it returns relative depth — the two numbers that turn it into metres must come from somewhere else.

  • 3D Vision and Depth

    Structure from motion

    Move one camera around an object and you can recover both the 3D shape and every camera position — but never the real-world size, unless you measure one thing yourself.

  • 3D Vision and Depth

    Visual SLAM

    Visual SLAM builds a map while working out where the camera is inside it, and the thing that keeps it honest is recognising a place you have already been.