Computer Vision

OpenCV

OpenCV is the free toolbox that reads, reshapes and measures images for you, so you never write resize or blur or edge detection by hand.

On this page 8
  1. Why it exists
  2. What is actually inside it
  3. The tool that everything else stands on
  4. The trap that catches everyone once
  5. Where you have already seen the results
  6. What is honestly annoying about it
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

OpenCV is a free toolbox of ready-made operations for pictures and video.

When a tap leaks at home, you open the toolbox and take out a spanner. You do not go outside, find iron ore and forge one. Somebody already made the spanner, and it is better than anything you would make in an afternoon.

OpenCV is that toolbox for images. Reading a photo from disk, shrinking it, rotating it, blurring it, finding its edges. All of it is already built, tested and fast.

Why it exists

In the late 1990s every research lab was rewriting the same twenty things. Load a picture. Change its size. Convert colour to grey. Find the edges. Open a camera.

Each lab wrote its own version. Each version had its own bugs. Nobody could share code, so nobody could build on anybody else's work.

Intel released OpenCV in 2000 to end that duplication. It is open source, meaning free to read, use and change. Twenty-five years later it is still the default toolbox, and it is in phones, factories, cars and hospitals.

Open source is worth defining properly, because it is why this lesson exists at all. It means the code is public, anyone may use it in their own work, and anyone may fix it. You pay nothing, and you are not asking permission.

What is actually inside it

   OpenCV
     |
     +-- read and write        pictures and video files
     +-- resize, crop, rotate  change the shape of the grid
     +-- colour conversion     colour to grey, and between colour systems
     +-- blur and sharpen      smooth out noise, or bring out detail
     +-- find edges            where does brightness change sharply
     +-- find shapes           trace the outline of a blob and measure it
     +-- open a camera         pull frames off a webcam or phone stream
     +-- run a trained model   load a model somebody else trained, and use it

That last line surprises people, so it is worth being exact. OpenCV can run a trained model. It does not train one. Training is what PyTorch and similar libraries are for.

The tool that everything else stands on

Almost every computer vision project starts with the same three lines of thought. Get the picture in. Get it into the right shape and size. Then hand it to whatever does the clever part.

OpenCV is the first two of those three, every single time.

   file on disk  ->  OpenCV reads it  ->  grid of numbers
   grid          ->  OpenCV resizes   ->  smaller grid the model expects
   smaller grid  ->  model            ->  "this is a dog"

Get the first two wrong and the third one cannot save you. A model trained on 224-dot squares, fed a stretched rectangle, quietly performs worse and never says why.

The trap that catches everyone once

Here is a thing that has cost thousands of people an evening. There is no shame in falling for it.

Most of the world stores a colour pixel as red, then green, then blue. OpenCV stores it as blue, then green, then red. Backwards.

The reason is historical. Camera makers in 2000 used that order, so OpenCV matched them. By the time everyone else agreed on the other order, it was too late to change.

Load a photo with OpenCV, show it with any other tool, and it looks like an underwater shot. Skin goes blue. Sky goes orange. Nothing crashes, nothing warns you.

Fixing it is one line of code, shown in the Developer section. Knowing it exists is the hard part.

Where you have already seen the results

  • A shop's CCTV counting how many people walked in today.
  • A toll booth reading a number plate in rain and headlight glare.
  • A phone banking app finding the edges of your cheque before it reads the text.
  • A factory camera rejecting a bottle with a crooked cap.
  • A video call background blur. The blur itself is exactly this kind of operation.

What is honestly annoying about it

Two things, and both are worth knowing before you start.

The way you say things is inconsistent. Some functions take a point as column-then-row. The grid itself is indexed row-then-column. Both are correct, in their own worlds, and you will mix them up.

It fails quietly. Ask OpenCV to read a file that is not there and it does not shout. It hands you back an empty nothing and lets your next line crash with a confusing message. Always check that a picture actually loaded.

Neither of these is a reason to avoid it. They are reasons to expect them.

Remember this

  • OpenCV is a free, very fast toolbox for reading and reshaping images and video.
  • It handles everything around a model. It does not train models.
  • Its colour order is blue, green, red. Convert before showing or saving elsewhere.
  • It fails silently on a missing file. Check before you use the result.

What to learn next

Developer — Code and libraries.

Everything below runs on a CPU in under a second, draws its own test images, and downloads nothing at runtime.

Setup

bash
pip install opencv-python-headless numpy pillow

Use the headless build unless you genuinely need OpenCV to open desktop windows. It leaves out the GUI code, installs cleanly inside Docker and on servers, and avoids a common class of "cannot connect to display" errors.

Be aware of the size. The wheel is roughly 35 to 40 MB to download and around 130 MB once unpacked. That is a real cost on a metered connection, so install it once and keep the environment.

Pick exactly one of opencv-python, opencv-python-headless and opencv-contrib-python. Installing two puts two copies of cv2 on your path and the resulting import errors are genuinely hard to read.

Draw your own test image

You do not need a photo to learn OpenCV. Making the picture yourself means you know the right answer before you measure anything.

draw.py
import cv2
import numpy as np

print("OpenCV version:", cv2.__version__)

# A blank canvas: 120 rows, 200 columns, 3 channels. Zero is black.
canvas = np.zeros((120, 200, 3), dtype=np.uint8)

# OpenCV colours are (blue, green, red), not (red, green, blue).
RED = (0, 0, 255)
GREEN = (0, 255, 0)

# Drawing points are (x, y): column first. Array indexing is row first. Both are correct.
cv2.rectangle(canvas, (20, 20), (80, 100), RED, thickness=-1)   # -1 fills the shape
cv2.circle(canvas, (140, 60), 35, GREEN, thickness=-1)

print("canvas shape:", canvas.shape)
print("pixel at row 60, column  50:", canvas[60, 50])
print("pixel at row 60, column 140:", canvas[60, 140])
print()

# Shrink to something a terminal can print, then draw it with characters.
thumb = cv2.resize(canvas, (50, 15), interpolation=cv2.INTER_AREA)
for row in thumb:
    print("".join("R" if r > 100 else "G" if g > 100 else "." for b, g, r in row))

cv2.imwrite("shapes.png", canvas)
Output
OpenCV version: 4.11.0
canvas shape: (120, 200, 3)
pixel at row 60, column  50: [  0   0 255]
pixel at row 60, column 140: [  0 255   0]

..................................................
..................................................
.....RRRRRRRRRRRRRRR..............................
.....RRRRRRRRRRRRRRR...........GGGGGGGG...........
.....RRRRRRRRRRRRRRR.........GGGGGGGGGGGGG........
.....RRRRRRRRRRRRRRR.......GGGGGGGGGGGGGGGG.......
.....RRRRRRRRRRRRRRR.......GGGGGGGGGGGGGGGGG......
.....RRRRRRRRRRRRRRR......GGGGGGGGGGGGGGGGGG......
.....RRRRRRRRRRRRRRR......GGGGGGGGGGGGGGGGGG......
.....RRRRRRRRRRRRRRR.......GGGGGGGGGGGGGGGG.......
.....RRRRRRRRRRRRRRR........GGGGGGGGGGGGGG........
.....RRRRRRRRRRRRRRR...........GGGGGGGG...........
.....RRRRRRRRRRRRRRR..............................
..................................................
..................................................

Your version line will differ. Everything else will match, because none of this is random.

Look closely at the two coordinate systems living side by side. cv2.rectangle(canvas, (20, 20), (80, 100), ...) means columns 20 to 80 and rows 20 to 100. So the shape is taller than it is wide, which the drawing confirms. canvas[60, 50] means row 60, column 50. Drawing functions take (x, y). Array indexing takes [row, column]. Read every call twice until this stops feeling wrong.

Notice cv2.imshow is nowhere in this file. It needs a desktop window, and it hangs or crashes inside notebooks, containers and remote shells. Save with cv2.imwrite and open the file, or print numbers. It is the habit that keeps your code working everywhere.

Measure what you drew

analyse.py
import cv2
import numpy as np

canvas = np.zeros((120, 200, 3), dtype=np.uint8)
cv2.rectangle(canvas, (20, 20), (80, 100), (0, 0, 255), -1)
cv2.circle(canvas, (140, 60), 35, (0, 255, 0), -1)

gray = cv2.cvtColor(canvas, cv2.COLOR_BGR2GRAY)
print("colour shape:", canvas.shape, " grey shape:", gray.shape)

# Anything brighter than 30 becomes white, everything else black.
_, mask = cv2.threshold(gray, 30, 255, cv2.THRESH_BINARY)
print("white pixels in the mask:", int((mask == 255).sum()))

contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
print("shapes found:", len(contours))
for i, c in enumerate(sorted(contours, key=cv2.contourArea, reverse=True)):
    x, y, w, h = cv2.boundingRect(c)
    print(f"  shape {i}: area {int(cv2.contourArea(c)):5d} px   box x={x} y={y} w={w} h={h}")

small = cv2.resize(canvas, (50, 30), interpolation=cv2.INTER_AREA)
print("resized  :", small.shape)
edges = cv2.Canny(gray, 50, 150)
print("edge pixels:", int((edges > 0).sum()))
Output
colour shape: (120, 200, 3)  grey shape: (120, 200)
white pixels in the mask: 8794
shapes found: 2
  shape 0: area  4800 px   box x=20 y=20 w=61 h=81
  shape 1: area  3754 px   box x=105 y=25 w=71 h=71
resized  : (30, 50, 3)
edge pixels: 512

Line by line, the parts worth stopping on

cvtColor drops the third axis entirely. The shape goes from (120, 200, 3) to (120, 200). Code that assumes three axes will break on grayscale input, and this is the most common cause of "too many indices for array".

A threshold returns two things. The first is the threshold value actually used, which only matters with automatic methods such as Otsu. Assigning it to _ is the convention when you passed the number yourself.

Contour area is not pixel count. The rectangle covers 61 by 81 pixels, which is 4941 pixels, but contourArea reports 4800. contourArea measures the polygon traced through the pixel centres, so it comes up short by roughly half a pixel around the whole border. For blob comparison this rarely matters. For anything calibrated to real-world millimetres, it does.

The circle reports 3754 against a true area of about 3848 for radius 35. Same reason. A rasterised circle traced through pixel centres is slightly smaller than the ideal circle.

interpolation=cv2.INTER_AREA is not decoration. When shrinking an image, INTER_AREA averages the pixels being merged. The default, INTER_LINEAR, samples without averaging and throws away detail in a way that produces jagged, aliased results. Use INTER_AREA to shrink and INTER_LINEAR or INTER_CUBIC to enlarge.

Canny wants a blurred input. The original algorithm starts with a Gaussian smoothing step, and OpenCV's implementation expects you to have done that yourself. On this synthetic image there is no noise, so it does not matter. On a real photo, run cv2.GaussianBlur(gray, (5, 5), 0) first or you will detect every grain of sensor noise as an edge.

The colour order trap, demonstrated

channel_order.py
import cv2
import numpy as np
from PIL import Image

# One pixel that we intend to be RED, written the way most libraries expect: R, G, B.
rgb = np.array([[[255, 0, 0]]], dtype=np.uint8)

cv2.imwrite("wrong.png", rgb)                       # handed straight to OpenCV
cv2.imwrite("right.png", cv2.cvtColor(rgb, cv2.COLOR_RGB2BGR))

print("what we meant       :", rgb[0, 0], "= red")
print("wrong.png reads as  :", np.array(Image.open("wrong.png"))[0, 0], "= blue")
print("right.png reads as  :", np.array(Image.open("right.png"))[0, 0], "= red")
Output
what we meant       : [255   0   0] = red
wrong.png reads as  : [  0   0 255] = blue
right.png reads as  : [255   0   0] = red

Red went in. Blue came out. No error, no warning, no clue.

The rule that removes this problem for good: convert at the edges. Convert to RGB the moment a picture enters your code from OpenCV, and convert back to BGR the moment it leaves for OpenCV. Everything in between stays RGB, which is what PyTorch, Pillow and matplotlib expect.

Face detection without any deep learning

OpenCV ships a face detector from 2001, and it runs on a laptop in milliseconds. It is worth meeting, both because it still gets used and because it shows what the field looked like before neural networks.

This one needs an actual photograph. matplotlib ships a real photo inside the package, so nothing is downloaded.

faces.py
import cv2
import matplotlib.cbook as cbook

# A photograph that ships inside matplotlib, so there is nothing to download.
path = cbook.get_sample_data("grace_hopper.jpg").name

img = cv2.imread(path)
if img is None:                      # imread returns None on failure, it does not raise
    raise SystemExit("could not read the image")
print("loaded:", img.shape)

gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
detector = cv2.CascadeClassifier(
    cv2.data.haarcascades + "haarcascade_frontalface_default.xml")
print("detector loaded:", not detector.empty())

faces = detector.detectMultiScale(gray, scaleFactor=1.1, minNeighbors=5, minSize=(40, 40))
print("faces found:", len(faces))
for (x, y, w, h) in faces:
    print(f"  face at x={x} y={y} width={w} height={h}")
Output
loaded: (600, 512, 3)
detector loaded: True
faces found: 1
  face at x=155 y=105 width=222 height=222

The box may shift by a pixel or two on a different OpenCV build. One face, found in a few milliseconds, with no model download and no training.

Now the honest part, because this detector is often oversold. It is a Haar cascade, a hand-designed detector from 2001 that looks for fixed patterns of light and dark rectangles. It works on faces looking straight at the camera, in reasonable light. It falls apart on tilted heads, profiles, poor light and low-contrast faces, and its published error rates were never measured evenly across skin tones. Do not build anything that matters on it. Use it to understand where the field came from.

Reading a camera

This one cannot have an output block. It depends on hardware that varies between machines, and inventing a plausible printout would teach you to expect something that may not happen.

camera.py
import cv2

cap = cv2.VideoCapture(0)            # 0 is the first camera the system reports
if not cap.isOpened():
    raise SystemExit("no camera available")

ok, frame = cap.read()               # ok is False at end of stream or on failure
if ok:
    print("frame shape:", frame.shape)
    cv2.imwrite("frame.png", frame)

cap.release()                        # always release, or the camera stays locked

No output block for this one: what it prints depends on whether your machine has a camera and what resolution it reports, so any figure printed here would be wrong for most readers.

The two lines that matter are isOpened and the ok flag. Both fail by returning False rather than raising, exactly like imread. Every OpenCV input path works this way, so check every one.

Common mistakes

Not checking imread. A wrong path, a permission problem or an unreadable file all return None. Your next line then says something like 'NoneType' object has no attribute 'shape', which points at the wrong line entirely. Check for None immediately, every time.

Non-English characters in the file path. On Windows, cv2.imread fails on paths containing non-ASCII characters and returns None with no explanation. The workaround reads the bytes yourself:

python
import numpy as np, cv2

path = "फोटो.jpg"                      # any path your filesystem accepts
data = np.fromfile(path, dtype=np.uint8)   # Python opens it, not OpenCV
img = cv2.imdecode(data, cv2.IMREAD_COLOR)

No output block: this needs a real file at path on your own disk, so what it prints depends entirely on which image you point it at.

Mixing up (x, y) and [row, column]. Drawing functions and rectangles use (x, y). NumPy indexing and .shape use row first. There is no way around it; there is only checking.

Using cv2.imshow in a notebook or container. It needs a window server. Write the file out instead, or display with matplotlib after converting to RGB.

Enlarging with INTER_AREA or shrinking with the default. The right choice depends on direction. INTER_AREA to shrink, INTER_LINEAR or INTER_CUBIC to grow.

Installing more than one OpenCV package. opencv-python and opencv-python-headless overwrite each other's files. Uninstall both, then install exactly one.

Try it yourself

Change cv2.threshold(gray, 30, 255, ...) to a threshold of 200 and re-run analyse.py. Predict first: which of the two shapes survives?

The red rectangle is (0, 0, 255) in BGR. Grayscale conversion weights red at about 0.299, so it turns into a grey value near 76. The green circle weights at about 0.587, so it lands near 150. Neither clears 200, so you should find zero shapes. Run it and confirm.

Then drop the threshold to 100 and you get one shape, the circle alone. You have now used the grayscale weights from the previous lesson to predict a program's output before running it. That is the difference between using a library and understanding one.

What to learn next

Researcher — Mathematics and papers.

What OpenCV is, structurally

OpenCV is a C++ library with thin bindings for Python, Java and JavaScript. The Python binding is not a reimplementation. cv2 marshals arguments into cv::Mat and calls the same compiled code, so Python-level overhead is a function-call boundary rather than an algorithmic penalty.

The binding shares memory with NumPy rather than copying. A contiguous uint8 array passed into a cv2 function is wrapped as a Mat header over the same buffer. Returned Mat objects are wrapped as NumPy arrays over theirs. Two consequences follow. In-place operations with a dst= argument genuinely mutate the caller's array. And a non-contiguous array — the result of a slice with a step, or a transpose — forces a copy, silently.

History: released by Intel in 2000, C++ API from 2.0 in 2009, maintained since by OpenCV.org. The licence changed from 3-clause BSD to Apache 2.0 at version 4.5.0 in 2020, which matters for patent-grant reasons in commercial deployment.

Resampling, and a reproducibility hazard

cv2.resize selects a kernel by the interpolation flag:

FlagKernelCorrect use
INTER_NEARESTnearest samplelabel masks and index images only
INTER_LINEAR2 x 2 bilinearupscaling; the default
INTER_CUBIC4 x 4 bicubicupscaling, sharper, slower
INTER_AREAbox average over the source footprintdownscaling
INTER_LANCZOS48 x 8 windowed sinchigh-quality resampling

The hazard is specific and widely encountered. INTER_LINEAR does not adapt its kernel to the scale factor, so downscaling with it point-samples a bilinear interpolation and aliases badly. PIL's Image.resize with LANCZOS or BILINEAR does apply a scale-adaptive support and therefore antialiases. The two libraries produce measurably different arrays for the same nominal operation.

Parmar, Zhang and Zhu (2022), On Aliased Resizing and Surprising Subtleties in GAN Evaluation (arXiv:2104.11222), traced substantial FID discrepancies in the literature to exactly this. The practical rule for any vision pipeline: fix one resize implementation, use it at training and at inference, and record which one in the model card.

PyTorch's torch.nn.functional.interpolate gained an antialias=True flag to match PIL's behaviour, and torchvision.transforms.v2.Resize exposes it. It is off by default.

Separable filtering

A 2-D Gaussian is separable, so cv2.GaussianBlur applies two 1-D passes:

text
G(x, y) = G(x) * G(y)
cost per pixel:  2k  multiply-adds  instead of  k^2
  • k — kernel width in pixels.

At k = 31 that is 62 operations per pixel instead of 961. OpenCV also chooses a fixed-point path for small integer kernels and dispatches to SIMD, so measured throughput beats the operation count.

sigma and ksize interact: passing sigma = 0 makes OpenCV derive sigma from ksize as 0.3 * ((ksize - 1) * 0.5 - 1) + 0.8. Passing ksize = (0, 0) makes it derive the size from sigma. Specifying both inconsistently is a common source of results that do not match a reference implementation.

Canny, precisely

Canny (1986) defined edge detection as an optimisation over three criteria: good detection, good localisation and single response. From those he derived a filter well approximated by the first derivative of a Gaussian. The implementation has four stages:

  1. Gaussian smoothing at scale sigma.
  2. Gradient magnitude and orientation, via Sobel kernels.
  3. Non-maximum suppression along the gradient direction, thinning ridges to one pixel.
  4. Hysteresis thresholding with T_low and T_high: pixels above T_high are edges, pixels between the thresholds are edges only if connected to one.

cv::Canny implements stages 2 to 4. It does not perform stage 1 — it runs Sobel on the array you hand it. Supply the smoothing yourself. The L2gradient flag switches the magnitude from |Gx| + |Gy| to sqrt(Gx^2 + Gy^2); the default is the cheaper L1 form, which is anisotropic.

A serviceable heuristic for the thresholds is T_high at the 70th to 90th percentile of gradient magnitude, with T_low = 0.4 * T_high. Fixed absolute thresholds do not transfer across exposure changes.

Viola-Jones cascades

Viola and Jones (2001), Rapid Object Detection using a Boosted Cascade of Simple Features, made real-time face detection possible on 2001 hardware through three ideas:

Integral image. Precompute II(x, y) = sum of I(x', y') for x' <= x, y' <= y. Any axis-aligned rectangle sum is then four lookups, independent of rectangle size. This makes multi-scale Haar features constant-time.

AdaBoost feature selection. From roughly 180,000 candidate Haar features in a 24 x 24 window, boosting selects a few hundred that carry the discriminative signal.

Attentional cascade. Stages are ordered by cost. Early stages reject the overwhelming majority of windows with a handful of features. Expected cost per window is sum over i of (p_1 ... p_{i-1}) * n_i, where p_i is the pass rate of stage i and n_i its feature count. With pass rates near 0.5 and a high per-stage detection rate, average cost collapses to a few features per window.

The limitations are structural, not incidental. Haar features are axis-aligned and contrast-based, so the detector is trained per pose and does not generalise across viewpoint. And the training sets used for the shipped cascades were not demographically balanced, so per-subgroup error rates are both uneven and largely unmeasured. Buolamwini and Gebru (2018) established the disaggregated-evaluation methodology that any deployed face system should be held to; classical cascades predate it entirely.

For current work, OpenCV's dnn module runs a small ResNet-based SSD face detector, and cv2.FaceDetectorYN wraps YuNet — both are far more robust and still CPU-real-time.

Performance levers

LeverEffect
cv2.setUseOptimized(True)enables SIMD dispatch; on by default
cv2.setNumThreads(n)caps the internal thread pool; set to 1 inside a DataLoader worker
cv2.UMat(arr)routes through OpenCL to a GPU or integrated graphics if present
cv2.ocl.setUseOpenCL(False)disables the OpenCL path when it is slower than the CPU
Build with IPP / TBBIntel primitives and threading; enabled in the official wheels

The setNumThreads(1) case is the one that bites in practice. OpenCV's internal threading multiplied by a PyTorch DataLoader's worker count produces heavy oversubscription, and throughput drops while every core reads as busy.

Reading OpenCV's coordinate conventions

Three conventions coexist and none of them is wrong:

  • Mat and NumPy indexing: [row, column], origin at top-left.
  • Point, and drawing functions: (x, y), that is (column, row).
  • Size, and resize: (width, height), that is (columns, rows).

So img.shape[:2] is (height, width) while cv2.resize(img, (w, h)) takes them in the other order. This is a genuine API wart and the source of a persistent class of silent bugs. Write a wrapper in any project large enough to justify one.

References

  • Bradski, G. (2000). The OpenCV Library. Dr. Dobb's Journal of Software Tools.
  • Canny, J. (1986). A Computational Approach to Edge Detection. IEEE TPAMI 8(6).
  • Viola, P. & Jones, M. (2001). Rapid Object Detection using a Boosted Cascade of Simple Features. CVPR.
  • Bradski, G. & Kaehler, A. (2008). Learning OpenCV. O'Reilly. Dated on the API, still the clearest treatment of the algorithms.
  • Parmar, G., Zhang, R. & Zhu, J.-Y. (2022). On Aliased Resizing and Surprising Subtleties in GAN Evaluation. arXiv:2104.11222
  • Wu, W. et al. (2023). YuNet: A Tiny Millisecond-level Face Detector. Machine Intelligence Research 20.
  • Buolamwini, J. & Gebru, T. (2018). Gender Shades. PMLR 81.

Where OpenCV sits now

Its role has narrowed and deepened. Classical detection and segmentation inside OpenCV are largely of historical interest, displaced by learned models. What has not been displaced is everything around the model: decode, colour conversion, geometric transforms, camera and video I/O, calibration, and the classical geometry stack.

That geometry stack is the part worth knowing exists. Camera calibration, stereo rectification, solvePnP, homography estimation with RANSAC, and optical flow are mature, well-tested, and have no deep-learning replacement that wins on accuracy, speed and reliability together. Structure-from-motion and visual odometry pipelines still lean on them.

The dnn module reads ONNX and runs inference across CPU, OpenCL, CUDA and several vendor backends. It is a reasonable deployment target when you want one C++ dependency instead of a Python runtime. It also lags the training frameworks on operator coverage, so check your model converts before designing around it.

What to learn next