Stereo depth from two cameras
Two cameras a known distance apart recover depth from how far each point shifts between the images, and the accuracy falls off with the square of the distance.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Two cameras side by side can measure distance, because near things shift more between the pictures than far things.
The analogy you have already lived
Hold a finger up in front of your face. Close your left eye, then your right eye, alternately.
Your finger jumps sideways. Now hold it at arm's length and do it again — it jumps less. Look at a building through a window and it barely jumps at all.
That jump is the whole method. Big jump means near. Small jump means far. Your brain has been doing this every waking second of your life.
Why it exists
A single photo cannot tell distance. We saw exactly why in camera models and intrinsics: dividing by depth throws depth away.
Two cameras a known distance apart put the information back. Each one throws away depth differently, and comparing them recovers it.
How it works
Set the two cameras level and pointing the same way, a known distance apart. That distance is called the baseline.
Now a point in the world lands on both pictures, but at different sideways positions. The difference is called the disparity.
left image: the corner of a box sits at column 200
right image: the same corner sits at column 175
|
disparity = 25 pixels
|
distance = (how zoomed in) x (baseline) / disparityThat last line is the entire method. One division.
Notice its shape. Disparity in the bottom means a big shift gives a small distance. That matches the finger experiment exactly.
The catch that decides everything
Look at the division again. Disparity is on the bottom, so when disparity is small, small errors in it become big errors in distance.
The run below measures this. The near box at 2 metres comes out perfect. The box at 3 metres is off by 5 centimetres. The wall at 6 metres is off by 20 centimetres. Same camera, same code.
The error grows with the square of the distance. Double the distance, four times the error.
This is why self-driving cars carry radar and lidar as well as stereo cameras. Stereo is excellent close up and useless far away.
The second catch
The computer has to work out which pixel in the left picture is which pixel in the right picture. To do that, the pixels have to be distinguishable.
Point a stereo camera at a plain white wall and it fails completely. Every patch looks like every other patch. In the run below, two blank grey images produce depth for exactly 0% of pixels.
This is why depth-sensing cameras like the older Kinect project a pattern of infrared dots onto the room. They are giving themselves texture to match on.
Where you have already seen it
- Your own two eyes.
- Portrait mode on a dual-camera phone.
- Reversing sensors and parking cameras on newer cars.
- Warehouse robots avoiding obstacles.
- A VR headset tracking your hands.
Remember this
- Distance comes from how far a point shifts between the two images.
- Accuracy falls off with the square of the distance.
- A texture-free surface gives no depth at all.
What to learn next
- Monocular depth estimation — guessing depth from one image, and what that costs.
- Structure from motion — one moving camera doing the job of two.
- Camera calibration with a chessboard — the step a real stereo rig needs first.
Developer — Code and libraries.
Setup
pip install opencv-python==4.10.0.84 numpy==1.26.4We build a stereo pair from a depth map we chose, so every recovered number can be scored. The right image is rendered by sampling the texture at fractional positions, which means the true disparities are not whole numbers — otherwise the matcher would look better than it is.
A stereo pair with a known answer
import cv2
import numpy as np
H, W = 240, 320
FOCAL = 500.0 # pixels
BASELINE = 0.10 # metres between the two camera centres
PAD = 64
rng = np.random.default_rng(1)
# A depth map for the left camera: a back wall and two boxes in front of it.
depth_left = np.full((H, W), 6.0, np.float32)
depth_left[70:170, 60:150] = 3.0
depth_left[100:200, 180:280] = 2.0
# Speckle texture, wider than the image so the right camera has something to see.
texture = cv2.GaussianBlur(rng.integers(0, 256, (H, W + PAD), np.uint8), (3, 3), 0)
left = texture[:, PAD:PAD + W].copy() # left pixel x samples texture x+PAD
# Where does each surface land in the right image? Paint far first, so a near
# surface correctly covers what is behind it.
disp_left = FOCAL * BASELINE / depth_left
depth_right = np.full((H, W), 6.0, np.float32) # holes default to the back wall
order = np.argsort(-depth_left, axis=None)
ys, xs = np.unravel_index(order, depth_left.shape)
tx = np.round(xs - disp_left[ys, xs]).astype(int)
ok = (tx >= 0) & (tx < W)
depth_right[ys[ok], tx[ok]] = depth_left[ys[ok], xs[ok]]
# Now sample the texture at a fractional position, so disparities are not integers.
grid_x, grid_y = np.meshgrid(np.arange(W, dtype=np.float32),
np.arange(H, dtype=np.float32))
disp_right = FOCAL * BASELINE / depth_right # the true disparity of each right pixel
map_x = grid_x + disp_right + PAD
right = cv2.remap(texture, map_x, grid_y, cv2.INTER_LINEAR)
print("disparity = focal * baseline / depth")
for z in [6.0, 3.0, 2.0]:
print(f" {z:.1f} m -> {FOCAL * BASELINE / z:6.2f} pixels of shift")
print("\ntwo of those are not whole numbers, so a matcher has to work below a pixel\n")
matcher = cv2.StereoSGBM_create(
minDisparity=0,
numDisparities=48, # must be a multiple of 16
blockSize=7,
P1=8 * 7 * 7, # small-change penalty, keeps surfaces smooth
P2=32 * 7 * 7, # large-change penalty, still allows real edges
uniquenessRatio=8,
speckleWindowSize=64,
speckleRange=2,
)
disp = matcher.compute(left, right).astype(np.float32) / 16.0 # SGBM uses 4 fixed bits
valid = disp > 0
est_depth = np.where(valid, FOCAL * BASELINE / np.maximum(disp, 1e-6), np.nan)
print(f"pixels given a disparity: {100 * valid.mean():.1f}%")
print("(a strip on the left has no match, because that scene is off the right image)\n")
print("estimated depth in the middle of each surface:")
for name, (y, x), truth in [("back wall", (30, 250), 6.0),
("middle box", (120, 105), 3.0),
("near box", (150, 230), 2.0)]:
patch = est_depth[y - 8:y + 8, x - 8:x + 8]
good = patch[~np.isnan(patch)]
print(f" {name:<11} true {truth:.2f} m measured {np.median(good):.3f} m"
f" error {abs(np.median(good) - truth) * 100:5.1f} cm")
inner = valid.copy()
inner[:, :52] = False # ignore the unmatchable left strip
err = np.abs(est_depth - depth_right)[inner]
print("\nover every matched pixel away from that strip:")
print(f" median depth error {np.median(err) * 100:.1f} cm")
print(f" 90th percentile {np.percentile(err, 90) * 100:.1f} cm")
print(f" median disparity error {np.median(np.abs(disp - disp_right)[inner]):.3f} px")
print("\nprecision falls off with distance, and the maths says by how much:")
print(" one pixel of disparity error costs depth^2 / (focal * baseline) metres")
for z in [1, 2, 5, 10, 20]:
print(f" at {z:>2} m: {z * z / (FOCAL * BASELINE) * 100:7.1f} cm per pixel of error")
print("\nwhat happens without texture to match on:")
plain = np.full((H, W), 128, np.uint8)
plain_disp = matcher.compute(plain, plain).astype(np.float32) / 16.0
print(f" two blank grey images: {100 * (plain_disp > 0).mean():.1f}% of pixels matched")
print(" a blank wall gives a stereo camera nothing to work with, at any distance")disparity = focal * baseline / depth 6.0 m -> 8.33 pixels of shift 3.0 m -> 16.67 pixels of shift 2.0 m -> 25.00 pixels of shift two of those are not whole numbers, so a matcher has to work below a pixel pixels given a disparity: 82.2% (a strip on the left has no match, because that scene is off the right image) estimated depth in the middle of each surface: back wall true 6.00 m measured 6.202 m error 20.2 cm middle box true 3.00 m measured 2.952 m error 4.8 cm near box true 2.00 m measured 2.000 m error 0.0 cm over every matched pixel away from that strip: median depth error 15.4 cm 90th percentile 20.2 cm median disparity error 0.208 px precision falls off with distance, and the maths says by how much: one pixel of disparity error costs depth^2 / (focal * baseline) metres at 1 m: 2.0 cm per pixel of error at 2 m: 8.0 cm per pixel of error at 5 m: 50.0 cm per pixel of error at 10 m: 200.0 cm per pixel of error at 20 m: 800.0 cm per pixel of error what happens without texture to match on: two blank grey images: 0.0% of pixels matched a blank wall gives a stereo camera nothing to work with, at any distance
Reading the output carefully
The three surface errors are the lesson: 0.0, 4.8, 20.2 cm at 2, 3 and 6 metres. One matcher, one image pair. The only thing that changed is distance, and the error grew roughly with its square. From 2 m to 6 m is a factor of 3 in distance and roughly 9 in error — which is exactly what the table below it predicts.
median disparity error 0.208 px. SGBM interpolates a parabola through the three matching costs around the winner, giving sub-pixel precision. That 0.2-pixel figure is typical for a good matcher on well-textured input, and it is the number that everything else follows from.
82.2% of pixels got a disparity. The missing 18% are mostly a strip down the left edge. A point visible at the left of the left image shifts out of frame in the right image, so there is nothing to match it against. The width of that dead strip is exactly the maximum disparity — 48 pixels here. Every stereo rig loses a band on one side, and cropping it is standard.
0.0% on two blank images. Not a low number, zero. uniquenessRatio=8 rejects any match whose best cost fails to beat the runner-up by a margin, and on uniform grey every candidate ties. This is the correct behaviour: refusing to answer is better than inventing a surface.
numDisparities=48 had to be at least 25. The nearest surface has a disparity of 25 pixels, and a search range that does not reach it would leave the near box unmatched. Setting this from your scene's minimum depth is the first thing to get right, and it costs linearly in compute.
Deriving the depth formula
Two cameras with the same focal length $f$, separated by baseline $B$ along the x-axis, both looking straight ahead. A point at depth $Z$ and lateral offset $X$ projects to
$$ u_L = f\frac{X}{Z} + c_x, \qquad u_R = f\frac{X - B}{Z} + c_x $$
Subtracting gives the disparity $d = u_L - u_R = fB/Z$, so
$$ Z = \frac{fB}{d} $$
Differentiating, $\left|\frac{\partial Z}{\partial d}\right| = \frac{fB}{d^2} = \frac{Z^2}{fB}$. That is the row of numbers in the output, and it is why stereo range is fundamentally limited.
Real cameras need rectification first
The formula above assumes the two cameras are perfectly aligned, so a point sits on the same image row in both. Real rigs are never that precise.
Rectification warps both images so that corresponding points share a row, which turns matching from a 2D search into a 1D one:
R1, R2, P1, P2, Q, _, _ = cv2.stereoRectify(K1, d1, K2, d2, size, R, T)
map1x, map1y = cv2.initUndistortRectifyMap(K1, d1, R1, P1, size, cv2.CV_32FC1)
left_rect = cv2.remap(left_raw, map1x, map1y, cv2.INTER_LINEAR)R and T come from cv2.stereoCalibrate, which is chessboard calibration run on both cameras at once. The Q matrix then converts a disparity map straight to 3D points with cv2.reprojectImageTo3D.
Common mistakes
Skipping rectification. Without it, the true match is a few rows away and SGBM never looks there. The symptom is a noisy, mostly-empty disparity map on real hardware that works fine in simulation.
Forgetting to divide by 16. StereoSGBM.compute returns CV_16S with 4 fractional bits. Treating it as integer pixels makes every distance 16 times too small.
Choosing the baseline badly. A wide baseline improves far-range accuracy and increases occlusion and the search range. A narrow one does the reverse. Pick it from the depth range you actually need.
Using StereoBM and expecting SGBM quality. StereoBM is block matching only, with no smoothness term. It is faster and considerably worse on low-texture regions. P1 and P2 are what make SGBM good, and tuning them matters.
Treating the disparity map as if every pixel is reliable. Invalidate low-confidence pixels, and check disp > minDisparity rather than disp > 0 when minDisparity is not zero.
Assuming both cameras have the same exposure. SGBM's Birchfield-Tomasi cost has some tolerance, but a stereo pair with automatic exposure running independently on each camera will match badly. Lock exposure and gain together.
Try it yourself
Change BASELINE from 0.10 to 0.30 and rerun. The disparities triple, so numDisparities must rise to at least 96, the dead strip on the left triples in width, and the error at 6 m falls to roughly a third. That single edit is the whole baseline trade-off.
What to learn next
- Monocular depth estimation — guessing depth from one image, and what that costs.
- Structure from motion — one moving camera doing the job of two.
- Camera calibration with a chessboard — the step a real stereo rig needs first.
Researcher — Mathematics and papers.
The correspondence problem
Given rectified images $I_L, I_R$, find for each pixel the disparity $d$ minimising a matching cost, subject to a smoothness prior. Formally, minimise an energy over the disparity field $D$:
$$ E(D) = \sum_{\mathbf{p}} C(\mathbf{p}, D_{\mathbf{p}}) + \sum_{\mathbf{p}}\sum_{\mathbf{q} \in N_{\mathbf{p}}} \Bigl( P_1 \, \mathbb{1}\bigl[|D_{\mathbf{p}} - D_{\mathbf{q}}| = 1\bigr] + P_2 \, \mathbb{1}\bigl[|D_{\mathbf{p}} - D_{\mathbf{q}}| > 1\bigr] \Bigr) $$
Where $C$ is the per-pixel matching cost and $N_{\mathbf{p}}$ the 4- or 8-neighbourhood. The two-level penalty is deliberate: $P_1$ is small so slanted surfaces are not over-penalised, while $P_2$ is large and — crucially — constant regardless of the jump size, so a true depth discontinuity costs no more than a two-pixel one. That is what preserves object boundaries.
Exact minimisation of this energy in 2D is NP-hard.
Semi-global matching
Hirschmüller (CVPR 2005; TPAMI 2008), Stereo processing by semiglobal matching and mutual information, is the practical answer, and it is what StereoSGBM implements.
Instead of a 2D optimisation, aggregate costs along 8 or 16 1D paths through each pixel, each solved exactly by dynamic programming:
$$ L_{\mathbf{r}}(\mathbf{p}, d) = C(\mathbf{p}, d) + \min\bigl(L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, d),\; L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, d\pm1) + P_1,\; \min_i L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, i) + P_2\bigr) - \min_i L_{\mathbf{r}}(\mathbf{p}-\mathbf{r}, i) $$
The trailing subtraction only prevents numeric overflow and does not change the argmin. Summing $L_{\mathbf{r}}$ over all directions $\mathbf{r}$ approximates the 2D energy well, at $O(WHD)$ cost. OpenCV's default mode uses 5 paths; StereoSGBM.MODE_HH uses 8 at higher memory cost.
The single-path predecessor, scanline dynamic programming, produces the characteristic horizontal streaking artefact, because rows are optimised independently. Aggregating over multiple directions is what removes it.
Matching costs
- SAD / SSD over a window. Fast, and assumes identical brightness.
- Census transform. Encode each pixel as the bit pattern of comparisons with its neighbours, then use Hamming distance. Invariant to any monotonic intensity change, which makes it robust to exposure and vignetting differences. It is the standard cost in embedded stereo.
- Mutual information. Hirschmüller's original, robust to arbitrary intensity mappings, expensive and iterative.
- Birchfield-Tomasi. Samples the linearly interpolated intensity, making the cost insensitive to image sampling. OpenCV's SGBM uses a BT-based cost over a small window.
Sub-pixel and the quadratic fit
After the winner $d^$ is found, fit a parabola through the costs at $d^-1, d^, d^+1$:
$$ \delta = \frac{C(d^-1) - C(d^+1)}{2\bigl(C(d^-1) - 2C(d^) + C(d^*+1)\bigr)} $$
This is where the 0.208-pixel median error in the output comes from. It introduces a known bias, pixel-locking, where estimates cluster toward integer disparities because the cost function is not locally parabolic. Correcting it matters for high-precision metrology and rarely otherwise.
Fundamental limits
Depth uncertainty propagates as
$$ \sigma_Z = \frac{Z^2}{fB}\,\sigma_d $$
With $\sigma_d \approx 0.2$ px, $f = 500$, $B = 0.1$ m, the standard deviation is 4 mm at 1 m and 40 cm at 10 m. No amount of algorithmic cleverness changes the $Z^2$ term; only $f$ (resolution), $B$ (rig geometry) or $\sigma_d$ (matching quality) can move.
Three failure modes are structural rather than algorithmic:
- Textureless regions. No information exists in the images. Learned methods fill these by prior, not by measurement, which is a different claim about the world.
- Occlusion. Points visible in one view only have no true disparity. Left-right consistency checking detects them; nothing recovers them.
- Repetitive structure. A brick wall gives periodic cost minima.
uniquenessRatiorejects these, at the cost of holes.
Learned stereo
End-to-end networks now dominate the KITTI and Middlebury leaderboards. The lineage: DispNet (Mayer et al., CVPR 2016) with a correlation layer; GC-Net (Kendall et al., ICCV 2017) introducing an explicit 3D cost volume with 3D convolutions and a differentiable soft-argmin; PSMNet adding spatial pyramid pooling; RAFT-Stereo (Lipson et al., 3DV 2021) applying iterative recurrent refinement to a correlation pyramid, which remains a strong and widely used baseline.
Two honest caveats before choosing one over SGBM.
- Cross-domain generalisation is the weak point. Networks trained on synthetic data or KITTI can degrade on a different sensor, baseline or scene type. SGBM has no training distribution to leave.
- Learned methods interpolate textureless regions. They produce a smooth, plausible surface where SGBM produces a hole. For a rendering task that is better. For an obstacle detector it is a hallucinated surface, and the hole was the honest answer.
Active stereo sidesteps the texture problem entirely by projecting a pattern — the approach of the Intel RealSense D400 series and the original Kinect — and remains the most reliable way to get dense indoor depth.
What to learn next
- Monocular depth estimation — guessing depth from one image, and what that costs.
- Structure from motion — one moving camera doing the job of two.
- Camera calibration with a chessboard — the step a real stereo rig needs first.