Camera models and intrinsics
A camera turns 3D points into pixels by dividing by depth, and the four numbers in its intrinsic matrix are everything you need to describe that squashing.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A camera turns a 3D world into a flat picture by dividing everything by how far away it is.
The analogy you have already lived
Hold your thumb up at arm's length and it covers a distant building. Bring it close to your eye and it covers the whole room.
Your thumb never changed size. Its distance did, and the image on your retina got bigger in exactly the proportion the distance shrank.
Every camera does that same arithmetic. Twice as far away means half as big in the picture. That single rule is the entire pinhole camera model.
Why we need to write it down
A photo is a grid of coloured dots and nothing else. Measure a room, place a virtual object, drive a robot: all of them need to know how the world became those dots.
That relationship needs four numbers.
How zoomed in the camera is, horizontally and vertically. These are the focal lengths, written fx and fy, measured in pixels.
Where the centre of the picture sits. This is the principal point, written cx and cy. It is usually near the middle, but never exactly.
Together they are called the intrinsics — the properties of the camera itself, as opposed to where it is standing.
How it works
a point 5 metres away, half a metre to the left
|
step 1: divide sideways distance by depth
-0.5 / 5 = -0.1
|
step 2: multiply by how zoomed in the camera is, and shift to the centre
-0.1 * 500 + 320 = 270
|
pixel column 270That is it. Divide by depth, scale, shift.
The thing you lose
Look at step 1 again. Once you divide by 5, the number 5 is gone.
A one-metre square at five metres and a two-metre square at ten metres land on exactly the same pixels. Not similar pixels. Identical ones.
So a single photo cannot tell you how far anything is. This is not a limitation of cheap cameras or weak algorithms. It is a fact about the arithmetic.
Everything in this section is a way of getting that lost number back. Two cameras. A moving camera. Learned guesswork about what usual objects look like.
Where you have already seen it
- A phone lens spec quoted as "26 mm equivalent" — that is a focal length.
- Wide-angle photos where faces near the edge look stretched.
- A measuring app that needs you to move the phone before it can work.
- Panorama mode failing when you rotate the phone off its centre.
Remember this
- A camera divides by depth. That is the whole model.
- Four numbers describe it: two focal lengths and the picture centre.
- Dividing by depth throws depth away, and one photo can never get it back.
What to learn next
- Calibrating a camera with a chessboard — measuring these four numbers for a real camera.
- Stereo depth from two cameras — the direct way to get the lost depth back.
- How images are stored — what a pixel coordinate actually indexes.
Developer — Code and libraries.
Setup
pip install numpy==1.26.4Nothing else. Projection is four lines of arithmetic, and writing it out once removes most of the mystery from everything that follows.
The pinhole model, and the ambiguity it creates
import numpy as np
# A 640x480 camera whose lens gives a focal length of 500 pixels.
W, H = 640, 480
K = np.array([[500.0, 0.0, 320.0], # fx, 0, cx
[0.0, 500.0, 240.0], # 0, fy, cy
[0.0, 0.0, 1.0]])
print("the camera matrix K:")
print(K)
print("fx =", K[0, 0], " fy =", K[1, 1], " cx =", K[0, 2], " cy =", K[1, 2])
fov_x = 2 * np.degrees(np.arctan(W / (2 * K[0, 0])))
fov_y = 2 * np.degrees(np.arctan(H / (2 * K[1, 1])))
print(f"horizontal field of view {fov_x:.1f} deg, vertical {fov_y:.1f} deg\n")
def project(points, K):
"""Camera-frame 3-D points -> pixels. Divide by depth, then apply K."""
xy = points[:, :2] / points[:, 2:3] # this division is the perspective
homogeneous = np.hstack([xy, np.ones((len(points), 1))])
return (K @ homogeneous.T).T[:, :2]
# Four points on a 1-metre-wide square, 5 metres in front of the camera.
square = np.array([[-0.5, -0.5, 5.0],
[0.5, -0.5, 5.0],
[0.5, 0.5, 5.0],
[-0.5, 0.5, 5.0]])
px = project(square, K)
print("a 1 m square at 5 m:")
for p3, p2 in zip(square, px):
print(f" world {p3} -> pixel [{p2[0]:7.2f} {p2[1]:7.2f}]")
side = px[1, 0] - px[0, 0]
print(f" it is {side:.0f} pixels wide on the sensor\n")
print("move the same square further away and it shrinks by exactly the ratio:")
for depth in [5, 10, 20, 40]:
far = square.copy()
far[:, 2] = depth
p = project(far, K)
print(f" at {depth:>2} m: {p[1, 0] - p[0, 0]:6.1f} pixels wide")
print(" double the distance, half the size. That is all perspective is.\n")
print("the ambiguity you cannot escape with one camera:")
big_far = np.array([[-1.0, -1.0, 10.0], [1.0, -1.0, 10.0],
[1.0, 1.0, 10.0], [-1.0, 1.0, 10.0]])
print(" 1 m square at 5 m ->", np.round(project(square, K)[0], 2))
print(" 2 m square at 10 m ->", np.round(project(big_far, K)[0], 2))
print(" identical pixels. One image can never tell them apart.\n")
print("what each entry of K actually does, by changing one at a time:")
for name, mod in [("fx doubled to 1000", np.array([[1000, 0, 320], [0, 500, 240], [0, 0, 1.]])),
("cx moved to 400", np.array([[500, 0, 400], [0, 500, 240], [0, 0, 1.]])),
("fy halved to 250", np.array([[500, 0, 320], [0, 250, 240], [0, 0, 1.]]))]:
p = project(square, mod)
print(f" {name:<22} corner 0 -> [{p[0, 0]:7.2f} {p[0, 1]:7.2f}]"
f" width {p[1, 0] - p[0, 0]:6.1f} height {p[3, 1] - p[0, 1]:6.1f}")
print("\nprojection throws depth away, and the maths shows exactly where:")
p = np.array([[2.0, 1.0, 8.0]])
print(" point in front of the camera:", p[0])
print(" after dividing by depth: ", np.round(p[0, :2] / p[0, 2], 4))
print(" the number 8.0 is now gone. Recovering it is the whole of 3-D vision.")the camera matrix K: [[500. 0. 320.] [ 0. 500. 240.] [ 0. 0. 1.]] fx = 500.0 fy = 500.0 cx = 320.0 cy = 240.0 horizontal field of view 65.2 deg, vertical 51.3 deg a 1 m square at 5 m: world [-0.5 -0.5 5. ] -> pixel [ 270.00 190.00] world [ 0.5 -0.5 5. ] -> pixel [ 370.00 190.00] world [0.5 0.5 5. ] -> pixel [ 370.00 290.00] world [-0.5 0.5 5. ] -> pixel [ 270.00 290.00] it is 100 pixels wide on the sensor move the same square further away and it shrinks by exactly the ratio: at 5 m: 100.0 pixels wide at 10 m: 50.0 pixels wide at 20 m: 25.0 pixels wide at 40 m: 12.5 pixels wide double the distance, half the size. That is all perspective is. the ambiguity you cannot escape with one camera: 1 m square at 5 m -> [270. 190.] 2 m square at 10 m -> [270. 190.] identical pixels. One image can never tell them apart. what each entry of K actually does, by changing one at a time: fx doubled to 1000 corner 0 -> [ 220.00 190.00] width 200.0 height 100.0 cx moved to 400 corner 0 -> [ 350.00 190.00] width 100.0 height 100.0 fy halved to 250 corner 0 -> [ 270.00 215.00] width 100.0 height 50.0 projection throws depth away, and the maths shows exactly where: point in front of the camera: [2. 1. 8.] after dividing by depth: [0.25 0.125] the number 8.0 is now gone. Recovering it is the whole of 3-D vision.
Reading the output carefully
100 pixels wide at 5 m, 50 at 10 m, 12.5 at 40 m. Exactly inversely proportional to depth, with no approximation anywhere. Every apparent-size calculation in computer vision is this line.
fx doubled widened the square from 100 to 200 pixels but left the height at 100. fx and fy scale independently. On almost every real camera they are within a fraction of a percent of each other, because pixels are square. When a calibration returns fx and fy differing by more than about 1%, something is wrong — usually the image was resized in one direction only.
cx moved to 400 shifted every point by 80 pixels and changed no size. The principal point is a translation. It matters for accuracy and never for shape.
[0.25, 0.125] is what remains of [2.0, 1.0, 8.0]. Three numbers went in, two came out. The lost dimension is not recoverable from this image, and every technique in this section is a way of supplying it from somewhere else.
65.2 and 51.3 degrees come from fov = 2 * atan(width / (2 * fx)). This is the single most useful sanity check on a calibration result. If your calibration reports a field of view of 20 degrees for a phone camera, the calibration is wrong, and you would not have noticed from the raw numbers.
Where the camera is standing
Intrinsics describe the camera. Extrinsics describe where it is: a rotation R and a translation t that take world coordinates into camera coordinates.
p_camera = R @ p_world + t
pixels = project(p_camera, K)The two are often written as one 3x4 matrix P = K [R | t], called the projection matrix. Keeping them mentally separate saves a great deal of confusion — intrinsics stay fixed while the camera moves, extrinsics change every frame.
Lens distortion
Real lenses bend straight lines, and the pinhole model has no term for it. OpenCV adds a polynomial correction applied to the normalised coordinates before K:
- Radial,
k1, k2, k3— barrel or pincushion bowing, growing with distance from the centre. - Tangential,
p1, p2— from a lens not perfectly parallel to the sensor, usually tiny.
For a normal phone lens k1 and k2 are enough. For a fisheye or an action camera, the polynomial model breaks down and you need cv2.fisheye, which uses an angular model instead.
Common mistakes
Reusing intrinsics after resizing an image. Halve the image and fx, fy, cx, cy all halve. Crop it and only cx, cy change. This is the most frequent silent error in 3D code.
Confusing focal length in millimetres with focal length in pixels. They relate by fx = f_mm * image_width_px / sensor_width_mm. A "26 mm" phone camera does not have fx = 26.
Assuming the principal point is the image centre. It is close, and using the exact centre costs accuracy in triangulation. Calibrate.
Forgetting that OpenCV pixel centres are at integer coordinates. The corner of the image is at (-0.5, -0.5). Off-by-half-pixel errors compound badly through a stereo pipeline.
Projecting points behind the camera. Any point with z <= 0 produces a meaningless pixel, and the division by a near-zero depth explodes. Filter on depth before projecting.
Try it yourself
Give the square a rotation before projecting, by multiplying the points by a rotation matrix about the y-axis. Watch the two vertical edges land at different widths. That difference is what tells you the square is turned, and it is the signal every pose-estimation method feeds on.
What to learn next
- Calibrating a camera with a chessboard — measuring these four numbers for a real camera.
- Stereo depth from two cameras — the direct way to get the lost depth back.
- How images are stored — what a pixel coordinate actually indexes.
Researcher — Mathematics and papers.
The model
A point $\mathbf{X}_w \in \mathbb{R}^3$ in world coordinates maps to a pixel $\mathbf{x}$ by
$$ s \begin{bmatrix} u \ v \ 1 \end{bmatrix} = K \begin{bmatrix} R & \mathbf{t} \end{bmatrix} \begin{bmatrix} X_w \ Y_w \ Z_w \ 1 \end{bmatrix} $$
Where $s$ is an arbitrary positive scalar (the depth in camera coordinates), $R \in SO(3)$ and $\mathbf{t} \in \mathbb{R}^3$ are the extrinsics, and
$$ K = \begin{bmatrix} f_x & \gamma & c_x \ 0 & f_y & c_y \ 0 & 0 & 1 \end{bmatrix} $$
$f_x, f_y$ are focal lengths in pixel units, $c_x, c_y$ the principal point, and $\gamma$ the skew, which accounts for non-perpendicular sensor axes and is zero for every camera manufactured in the last several decades. Most calibration routines fix it at zero, and OpenCV's does.
$K$ has 4 degrees of freedom, $[R|\mathbf{t}]$ has 6, and $P = K[R|\mathbf{t}]$ has 11 (a 3x4 matrix defined up to scale).
Why this is a projective map
The equality holds only up to scale, which is what makes projection a map from $\mathbb{P}^3$ to $\mathbb{P}^2$ rather than a linear map between vector spaces. Three consequences follow, and they explain most of what is strange about multi-view geometry.
- The fibre of a pixel is a ray. $P^{-1}$ does not exist; the preimage of a pixel is the one-parameter family ${\mathbf{C} + \lambda \, \mathbf{d}}$ where $\mathbf{C}$ is the camera centre and $\mathbf{d} = R^{\top} K^{-1} \mathbf{x}$.
- Straight lines stay straight, ratios do not. Projective transformations preserve incidence and cross-ratio; they do not preserve length, angle, or the midpoint of a segment.
- Points at infinity are ordinary. A direction $[\mathbf{d}^\top, 0]^\top$ projects to a finite pixel — the vanishing point. This is not a degenerate case but a useful one, and it is how horizon-based calibration works.
Normalised coordinates and distortion
Define normalised image coordinates $\tilde{\mathbf{x}} = (X/Z, Y/Z)$, so that $\mathbf{x} = K [\tilde{\mathbf{x}}^\top, 1]^\top$ in the distortion-free case. The Brown-Conrady model applies distortion in normalised coordinates, before $K$:
$$ \tilde{x}_d = \tilde{x}\,\frac{1 + k_1 r^2 + k_2 r^4 + k_3 r^6}{1 + k_4 r^2 + k_5 r^4 + k_6 r^6} + 2p_1 \tilde{x}\tilde{y} + p_2(r^2 + 2\tilde{x}^2) $$
with $r^2 = \tilde{x}^2 + \tilde{y}^2$, and symmetrically for $\tilde{y}_d$. The rational denominator terms $k_4, k_5, k_6$ are OpenCV's extension for wide-angle lenses and are off by default.
The order matters and is a frequent bug: distortion is applied to normalised coordinates, not to pixels. Undistorting requires an iterative inverse, since the polynomial is not analytically invertible — cv2.undistortPoints runs a fixed-point iteration.
For fields of view beyond roughly 120 degrees the polynomial fits poorly and the correct model is the equidistant or Kannala-Brandt fisheye model, $r_d = \theta(1 + k_1\theta^2 + k_2\theta^4 + \dots)$ with $\theta$ the incidence angle. Fitting a fisheye with the pinhole-plus-polynomial model produces a plausible reprojection error and badly wrong geometry at the periphery.
Degrees of freedom and what is recoverable
From a single image of an unknown scene, $K$ is not recoverable. From a single image with known structure — a plane, a calibration target, three orthogonal vanishing points — it is.
The image of the absolute conic $\omega = (K K^\top)^{-1}$ is the object that makes this precise. It is a fixed conic under any camera rotation, so anything that constrains $\omega$ constrains $K$. Two orthogonal vanishing points $\mathbf{v}_1, \mathbf{v}_2$ give one linear constraint $\mathbf{v}_1^\top \omega \, \mathbf{v}_2 = 0$; three mutually orthogonal directions give three, enough for $K$ with zero skew and unit aspect ratio. This is the theory behind chessboard calibration and behind single-image calibration from architectural scenes.
Hartley and Zisserman (2004), Multiple View Geometry in Computer Vision, 2nd edition, is the standard reference for all of this, and chapter 6 covers the camera model in full.
Calibrated versus uncalibrated
The distinction shapes the whole field.
- With $K$ known, image points can be converted to rays and the geometry becomes metric up to a global scale. The essential matrix applies, and pose decomposes into a rotation and a unit translation.
- With $K$ unknown, only projective reconstruction is possible. The fundamental matrix applies, and the reconstruction is correct only up to an arbitrary 3D projective transformation — parallel lines need not be parallel, angles are meaningless. Upgrading to metric requires additional constraints, a process called auto-calibration.
Modern feed-forward reconstruction models such as DUSt3R and VGGT blur this line by predicting point maps directly and recovering intrinsics as a by-product, learned from data rather than derived from constraints. The classical distinction still governs what information is present in the images, which is why it remains worth knowing.
What to learn next
- Calibrating a camera with a chessboard — measuring these four numbers for a real camera.
- Stereo depth from two cameras — the direct way to get the lost depth back.
- How images are stored — what a pixel coordinate actually indexes.