Monocular depth estimation
A network can guess depth from a single photo by recognising what things usually look like, but it returns relative depth — the two numbers that turn it into metres must come from somewhere else.
- 16 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A model can guess which parts of one photo are near and which are far. It cannot tell you how many metres.
The analogy you have already lived
Look at a photograph of a street in a city you have never visited. You know instantly that the car is closer than the building. You know the person is standing on the pavement, not floating.
You have never measured that street. You are using a lifetime of knowing how big cars and people are, and how roads recede.
That is exactly what these models do. Learned expectation, not measurement.
Why it works at all
We showed in camera models and intrinsics that a single photo genuinely cannot contain distance. A small near box and a big far box produce the same pixels, exactly.
So how can any model do it?
Because a photo of the real world is not an arbitrary arrangement. Doors are about two metres tall. Roads narrow as they recede. Textures get finer with distance. Things nearer the bottom of the frame are usually nearer to you.
The model has seen millions of scenes. It bets on the usual answer.
What you get, and what you do not
You get the ordering. Which parts are nearer than which. This is genuinely useful for blurring a background, inserting an object into a scene, or reasoning about layout.
You do not get metres. The output has no units. It tells you a pixel is "nearer" without saying how much nearer in any real measure.
Two numbers are missing: how much to scale the answer, and where the zero is. Getting them requires something outside the picture. A second camera, a known object size, a wheel sensor, or a distance you already know.
What that means in practice
one photo -> [ model ] -> a relative depth picture
(near is bright, far is dark)
|
plus ONE known distance -> metresSay you know the person in the photo is 1.7 metres tall. Or that the camera sits 1.5 metres above the floor. Either fact pins the whole map down. Without one, you cannot.
The run below measures this exactly. Fitting those two missing numbers brings the model's answer to within 4% of the truth.
The honest part
Where the guess is wrong, it is confidently wrong, and it looks smooth and plausible.
A mirror is read as a room continuing behind it. A large poster of a corridor is read as a corridor. A puddle reflecting the sky is read as a hole. These are not edge cases; they are consequences of the method being recognition, not measurement.
For a vehicle, a robot arm or a drone, a wrong distance is dangerous. There, monocular depth is a helper, not the sensor of record.
Where you have already seen it
- Portrait mode on a single-camera phone.
- Background blur in a video call.
- Photo apps that add a 3D parallax wobble to a still image.
- AR apps placing an object roughly on a floor.
Remember this
- One photo cannot contain distance; the model supplies it from experience.
- The output is relative. Metres need one extra fact from outside the image.
- Where it guesses wrong, it guesses smoothly and confidently.
What to learn next
- Structure from motion — recovering real geometry from a moving camera.
- Stereo depth from two cameras — measurement rather than inference.
- Vision transformers — the backbone these models are built on.
Developer — Code and libraries.
Setup
pip install transformers==5.6.2 torch==2.5.1 pillow==11.0.0 numpy==1.26.4Depth-Anything-V2-Small-hf is about 100 MB and runs on CPU in a few seconds per image.
We render a corridor by ray-casting, so we know the true depth of every pixel exactly. Then we can score the model instead of admiring a pretty colour map.
Rendering a scene with known depth, then measuring the model
import numpy as np
import torch
from PIL import Image
from transformers import pipeline
print("part 1 - why one photo can never give you metres\n")
K = np.array([[400.0, 0, 224.0], [0, 400.0, 168.0], [0, 0, 1.0]])
small_near = np.array([[-0.5, -0.5, 4.0], [0.5, -0.5, 4.0], [0.5, 0.5, 4.0]])
big_far = small_near * 3.0 # three times bigger, three times further
def project(p):
return (K @ (p / p[:, 2:3]).T).T[:, :2]
print(" a 1 m box at 4 m ->", np.round(project(small_near)[0], 2))
print(" a 3 m box at 12 m ->", np.round(project(big_far)[0], 2))
print(" identical pixels. No algorithm can separate them from the image alone.")
print(" so a single-image depth model predicts RELATIVE depth, not metres.\n")
print("part 2 - render a corridor whose true depth we know exactly\n")
H, W = 336, 448
u, v = np.meshgrid(np.arange(W, dtype=np.float64), np.arange(H, dtype=np.float64))
dx, dy = (u - K[0, 2]) / K[0, 0], (v - K[1, 2]) / K[1, 1] # ray direction, dz = 1
FLOOR_Y, WALL_X, BACK_Z = 1.5, 2.0, 30.0
candidates = np.full((4, H, W), np.inf)
with np.errstate(divide="ignore", invalid="ignore"):
candidates[0] = np.where(dy > 1e-6, FLOOR_Y / np.maximum(dy, 1e-9), np.inf)
candidates[1] = np.where(dy < -1e-6, -FLOOR_Y / np.minimum(dy, -1e-9), np.inf)
candidates[2] = np.where(dx < -1e-6, -WALL_X / np.minimum(dx, -1e-9), np.inf)
candidates[3] = np.where(dx > 1e-6, WALL_X / np.maximum(dx, 1e-9), np.inf)
surface = candidates.argmin(axis=0)
depth = np.minimum(candidates.min(axis=0), BACK_Z)
# A checkerboard painted in world units, so the corridor has usable texture.
world_a = np.where(surface < 2, u * 0 + dx * depth, depth) # across the surface
world_b = depth
tile = ((np.floor(world_a / 0.5) + np.floor(world_b / 0.8)) % 2)
shade = np.array([0.55, 0.85, 0.70, 0.62])[surface] # each face lit differently
img = np.clip((90 + 110 * tile) * shade, 0, 255).astype(np.uint8)
img[depth >= BACK_Z] = 40 # the far wall
rgb = np.dstack([img] * 3)
print(f" true depth range in the render: {depth.min():.2f} m to {depth.max():.2f} m")
print(" the corridor, coarsely sampled:")
for y in range(0, H, 18):
print(" " + "".join("#" if img[y, x] > 140 else ("." if img[y, x] > 60 else " ")
for x in range(0, W, 7)))
print("\npart 3 - what Depth Anything V2 predicts on it\n")
pipe = pipeline("depth-estimation", model="depth-anything/Depth-Anything-V2-Small-hf",
device=-1)
out = pipe(Image.fromarray(rgb))
pred = out["predicted_depth"].squeeze().numpy()
pred = np.array(Image.fromarray(pred).resize((W, H), Image.BILINEAR))
print(f" raw output range: {pred.min():.2f} to {pred.max():.2f} (no units at all)")
print(" larger values mean NEARER: the model predicts inverse depth\n")
for name, (y, x) in [("far, down the corridor", (170, 224)),
("mid floor", (250, 224)),
("near floor", (325, 224)),
("left wall, mid height", (168, 40))]:
print(f" {name:<24} true {depth[y, x]:6.2f} m model says {pred[y, x]:7.2f}")
order_true = np.argsort([depth[170, 224], depth[250, 224], depth[325, 224]])
order_pred = np.argsort([-pred[170, 224], -pred[250, 224], -pred[325, 224]])
print(f"\n near-to-far ordering, truth {order_true.tolist()}, "
f"model {order_pred.tolist()} match: {order_true.tolist() == order_pred.tolist()}")
print("\npart 4 - fitting the two missing numbers\n")
inv_true = 1.0 / depth
m = depth < BACK_Z
a, b = np.polyfit(pred[m].ravel(), inv_true[m].ravel(), 1)
print(f" least-squares fit of 1/depth = a * prediction + b")
print(f" a = {a:.6f} b = {b:.6f}")
fitted = 1.0 / np.maximum(a * pred + b, 1e-6)
err = np.abs(fitted - depth)[m]
rel = (err / depth[m])
print(f" after fitting: median absolute error {np.median(err):.2f} m")
print(f" median relative error {100 * np.median(rel):.1f}%")
print(f" correlation between predicted and true inverse depth: "
f"{np.corrcoef(pred[m].ravel(), inv_true[m].ravel())[0, 1]:.4f}")
print("\n those two numbers, a and b, are what a single image cannot supply.")
print(" in a real system they come from a known object size, a second camera,")
print(" a wheel odometer, or an IMU.")part 1 - why one photo can never give you metres
a 1 m box at 4 m -> [174. 118.]
a 3 m box at 12 m -> [174. 118.]
identical pixels. No algorithm can separate them from the image alone.
so a single-image depth model predicts RELATIVE depth, not metres.
part 2 - render a corridor whose true depth we know exactly
true depth range in the render: 3.57 m to 30.00 m
the corridor, coarsely sampled:
........#########........#######........########........########
....#######.......#######.......########.......#######..........
.......#######......######......#######......######.............
................######.....#####......#####......##### . ......
..............#####....#####....#####.....####...... . ......
..................###....####...####....####... ... . ......
........................###...##...###...### . ... . ......
.........................#...##.##..##... . ... . ......
............................#.#.##.#... . . ... . ......
............................. .. . . ... . ......
............................. .. . . ... . ......
............................ . .. .. . . . ... . ......
.......................... .. ... .. . ... . ......
....................... ... .... ... . ... . ......
................. ..... ... .... .... ... . ......
.................. ..... ..... ..... . ......
.......... ...... ..... ...... ...... ......
............. ....... ....... ...... .....
... ........ ....... ....... .........
part 3 - what Depth Anything V2 predicts on it
raw output range: -0.00 to 4.55 (no units at all)
larger values mean NEARER: the model predicts inverse depth
far, down the corridor true 30.00 m model says 0.00
mid floor true 7.32 m model says 1.55
near floor true 3.82 m model says 3.80
left wall, mid height true 4.35 m model says 3.59
near-to-far ordering, truth [2, 1, 0], model [2, 1, 0] match: True
part 4 - fitting the two missing numbers
least-squares fit of 1/depth = a * prediction + b
a = 0.056769 b = 0.039914
after fitting: median absolute error 0.20 m
median relative error 3.7%
correlation between predicted and true inverse depth: 0.9886
those two numbers, a and b, are what a single image cannot supply.
in a real system they come from a known object size, a second camera,
a wheel odometer, or an IMU.Reading the output carefully
Part 1 is the whole justification. [174. 118.] twice. A 1 m box at 4 m and a 3 m box at 12 m are pixel-for-pixel identical. Whatever the model is doing, it is not extracting information that is present in the image, because the information is not there.
raw output range: -0.00 to 4.55. Not metres, not centimetres, not anything. Depth Anything predicts inverse depth up to an unknown scale and shift. A common bug is to treat the output as distance directly, which inverts your entire scene — near becomes far.
The far end reads 0.00 and the near floor reads 3.80. Larger means nearer, consistently. That is the sign convention to check first with any depth model, because it varies between families.
The ordering matched exactly. Nothing in the image determines the true depths, and the model still got the ordering right on all three test points. That is the useful thing monocular depth actually provides.
a = 0.056769, b = 0.039914 and then 3.7% relative error. These two numbers convert the model's opinion into metres. b is not zero, which matters: a scale-only correction would not have worked. Models in this family are trained with a loss that is invariant to both scale and shift, so both are genuinely undetermined.
Correlation 0.9886 on inverse depth. The model captured the geometry of a scene it has never seen, from a single 448-pixel-wide render. This is a genuine capability and it is worth being impressed by, alongside the caveat that it comes from prior knowledge rather than measurement.
Getting to metres in a real system
Four ways, in descending order of reliability.
- A metric model. Depth Anything 3 ships a metric variant, and Metric3D and UniDepth target the same problem. They predict absolute depth by also conditioning on camera intrinsics. Accuracy degrades when the true intrinsics differ from what the model assumes.
- One known distance. Camera height above a flat floor is the classic one for ground robots and cars. Fit
aandbfrom the floor region, apply everywhere. - A second sensor. Sparse lidar points, wheel odometry, or a stereo pair, used to fit the same two numbers per frame.
- A known object. A detected face, a standard doorway, a number plate of known width.
Common mistakes
Using the pipeline's depth output for measurement. pipe(image)["depth"] is a PIL image, normalised to 0-255 for display. ["predicted_depth"] is the tensor you want.
Forgetting the model output is smaller than the input. These models predict at a reduced resolution and the pipeline resizes for display. Resize deliberately, and use bilinear rather than nearest so edges do not stair-step.
Comparing depth maps across frames. Scale and shift are estimated per image. Two consecutive video frames can come back with different scales, so an object that looks like it moved may only have been rescaled. Temporal consistency needs explicit handling, or a video depth model.
Trusting depth at object boundaries. The strongest errors are at edges, where the model interpolates between two surfaces. Erode your depth mask before using it for compositing.
Feeding a resized image without keeping the aspect ratio. These models are sensitive to the apparent geometry of the scene. Squashing an image changes the perspective cues it relies on.
Try it yourself
Flip the rendered corridor upside down with rgb[::-1] and rerun. The true depths are unchanged in structure, but the model's usual "the bottom of the frame is near" prior now works against it. The size of the accuracy drop tells you how much of the result was geometry and how much was prior.
What to learn next
- Structure from motion — recovering real geometry from a moving camera.
- Stereo depth from two cameras — measurement rather than inference.
- Vision transformers — the backbone these models are built on.
Researcher — Mathematics and papers.
Why it is ill-posed, precisely
The projection $\pi(\mathbf{X}) = K\mathbf{X}/Z$ satisfies $\pi(s\mathbf{X}) = \pi(\mathbf{X})$ for any $s > 0$. Scene and scale are therefore unidentifiable from a single view: the likelihood is flat along the scale direction.
Monocular depth estimation is consequently not inference but prior-dominated regression. The model learns $p(Z \mid I)$ under the empirical distribution of scenes it was trained on, and its output is the conditional mean of that prior, not a measurement. Everything about its failure modes follows from that sentence.
The scale-and-shift-invariant loss
Ranftl et al. (TPAMI 2020), Towards robust monocular depth estimation: mixing datasets for zero-shot cross-dataset transfer (MiDaS), made large-scale training possible by removing the incompatibility between datasets. Different sources give depth in metres, in disparity, or up to an unknown scale, so a common target is needed.
Work in disparity space $d = 1/Z$ and align prediction to ground truth per image before computing the loss:
$$ \mathcal{L} = \frac{1}{M}\sum_{i} \rho\bigl(s\,\hat{d}_i + t - d_i^{}\bigr), \qquad (s, t) = \arg\min_{s,t} \sum_i \bigl(s\hat{d}_i + t - d^{}_i\bigr)^2 $$
The alignment $(s, t)$ has a closed form, so the loss remains cheap. Because the loss cannot see scale or shift, the network is never trained to produce them — which is exactly why the code above had to fit a and b.
Disparity space rather than depth space is deliberate: it compresses the far field, matching both the $1/Z$ error behaviour of stereo supervision and the fact that relative error matters more than absolute error at distance.
Architectures
MiDaS (2019-2022) established the recipe: a strong pretrained encoder, a dense prediction decoder, and mixed multi-dataset training with the affine-invariant loss.
DPT (Ranftl et al., ICCV 2021), Vision transformers for dense prediction, replaced the convolutional encoder with a ViT and reassembled tokens into multi-scale feature maps. The constant global receptive field at every stage is what improved long-range consistency, which convolutional decoders struggle with.
Depth Anything (Yang et al., CVPR 2024) scaled the data instead of the architecture: 1.5M labelled images plus 62M unlabelled ones, pseudo-labelled by a teacher, with strong perturbations applied to the student and an auxiliary loss keeping features aligned to DINOv2 semantics. V2 (NeurIPS 2024) switched the labelled portion to synthetic data for cleaner boundaries, then distilled through a large teacher into pseudo-labelled real images.
Depth Anything 3 (ByteDance Seed, arXiv 2511.10647, released November 2025) generalises the problem to any number of input views, with or without known poses, predicting spatially consistent geometry plus camera parameters. It ships any-view models from 0.08B to 1.15B parameters, plus metric and monocular specialisations. Licensing differs by size — Apache 2.0 for the smaller models, CC BY-NC 4.0 for the larger ones — which matters for commercial use.
Metric depth
To output metres, a model must break the scale ambiguity with extra information. Two routes:
- Condition on intrinsics. Metric3D (Yin et al., ICCV 2023) and Metric3Dv2 canonicalise the input to a reference camera using the known focal length, so apparent size becomes informative. Accuracy then depends on the intrinsics being right.
- Fine-tune per domain. Train the affine-invariant backbone, then fine-tune a metric head on a single dataset such as NYUv2 or KITTI. Accurate in-domain, and the scale prior does not transfer.
Depth Anything 3's metric variant documents a conversion of the form metric_depth = focal * output / c for a model-specific constant, again making the focal length the carrier of scale.
Evaluation, and how it misleads
Standard metrics on NYUv2 and KITTI: AbsRel $\frac{1}{N}\sum |Z_i - Z^_i|/Z^_i$, RMSE, and threshold accuracy $\delta_k$, the fraction of pixels with $\max(Z/Z^, Z^/Z) < 1.25^k$.
Three cautions that change how published numbers should be read.
- Median scaling. Most zero-shot evaluations align the prediction to ground truth by median ratio, or by full least-squares affine fit, before scoring. Those numbers therefore describe relative accuracy and say nothing about metric accuracy. The alignment method used is frequently reported only in the appendix, and results are not comparable across different choices.
- Sparse ground truth. KITTI's lidar covers the lower part of the frame and misses sky, glass and thin structures. Errors in the uncovered regions are unmeasured.
- Benchmark saturation and leakage. NYUv2 and KITTI are small and old, and modern models train on tens of millions of images from web sources. Independent evaluation on your own domain is not optional.
Where the failures come from
The prior-dominated framing predicts them all.
- Mirrors and glass are read as the reflected or transmitted geometry, because that is what the pixel statistics look like.
- Printed images of scenes — posters, screens, painted murals — get depth appropriate to the depicted scene.
- Unusual scale breaks it: a doll's house, a scale model, an extreme close-up.
- Absolute scale drifts across frames, since each image is scaled independently.
None of these are fixed by more data. They are consequences of estimating a scene from a prior over scenes, and the fix is a second view or a second sensor — which is what structure from motion and stereo provide.
What to learn next
- Structure from motion — recovering real geometry from a moving camera.
- Stereo depth from two cameras — measurement rather than inference.
- Vision transformers — the backbone these models are built on.