Detecting change from space
Change detection compares two satellite images of the same ground taken at different times and flags what is genuinely different, which is harder than it sounds because clouds, seasons and sun angle all look like change too.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Change detection compares two satellite pictures of the same place, taken months or years apart, and marks what genuinely changed.
Think of the "spot the difference" puzzle in a newspaper. Two pictures look almost identical, and you circle the five things that moved. A missing chimney, a repainted door, a tree that vanished. That is change detection, done by eye on two photographs.
Satellites do the same trick over entire countries, comparing a photo of a field from January against the same field in June.
Why it exists
Forests get cleared quietly, often in remote places nobody visits often. Floods appear and vanish within days. Illegal construction happens overnight, before an inspector can be sent out.
A person cannot watch every hectare of land, every week, for years. A satellite that revisits the same ground on a schedule can. Comparing two passes turns "is anything different here" into a question a computer can answer. No human visit is needed, only two saved images.
How it works
Image, January -> subtract, pixel -> Difference -> Threshold -> "Changed" map
by pixel map (how big a
Image, June -> change counts)Line the two images up first. Pixel (3, 7) in one photo must really be the same patch of ground as pixel (3, 7) in the other. This step is called co-registration — aligning two images so their pixel grids match exactly. Subtract one from the other. A patch of ground that has not changed produces a difference near zero. Bare soil that turned into a building produces a large difference.
Where you have already seen it
- Global Forest Watch, which flags deforestation almost as soon as satellite passes reveal it.
- News coverage of floods, where "before" and "after" satellite photos show submerged towns.
- Government tracking of illegal sand mining or unauthorised construction along riverbanks.
- Disaster relief agencies estimating how many buildings a cyclone destroyed, without sending anyone into the area first.
An honest warning
Two images of the same field can look different for reasons that have nothing to do with real change. A photo taken at 9 a.m. and one taken at 3 p.m. have different shadows. A photo from the dry season and one from right after monsoon have different soil brightness. A satellite reads slightly different sunlight and different sensor angle on every single pass.
Naive change detection confuses all of this with genuine change. That matters beyond accuracy scores. Flagging a farmer's field as "illegally built on" from a shadow artefact has real consequences for a real person. So does missing a real flood because the two photos were taken at different times of day. Serious systems normalise brightness across images and pick same-season pairs before comparing anything.
Remember this
- Change detection subtracts two aligned images of the same place and flags large differences.
- Misalignment, season and sun angle all fake the appearance of change — the hardest part is ruling those out first.
- Getting this wrong has consequences beyond a bad accuracy score, since these maps feed real decisions about land and disaster response.
What to learn next
- NDVI and vegetation indices — a more reliable signal than raw brightness for vegetation change.
- Working with satellite imagery — the pixel and band basics this lesson builds on.
- Anomaly detection — the general idea of flagging what looks unusual, applied over time instead of over a single pair of images.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
import numpy as np
# Two synthetic single-band images of the same 5x5 patch, six months apart
# Values stand in for reflectance brightness (higher = more reflective ground)
before = np.array([
[120, 118, 122, 60, 58],
[119, 121, 60, 59, 61],
[121, 60, 58, 60, 62],
[59, 61, 60, 58, 121],
[60, 59, 61, 120, 122],
])
after = np.array([
[121, 119, 121, 150, 148],
[120, 122, 149, 151, 152],
[122, 148, 150, 151, 60],
[149, 151, 150, 59, 119],
[148, 150, 151, 121, 120],
])
diff = after.astype(int) - before.astype(int)
print("difference grid:")
print(diff)
# a pixel counts as "changed" if brightness shifted by more than 30 units
changed = np.abs(diff) > 30
print("\nchanged pixels:")
print(changed)
print("\npercent of patch changed: {:.0f}%".format(100 * changed.mean()))difference grid: [[ 1 1 -1 90 90] [ 1 1 89 92 91] [ 1 88 92 91 -2] [90 90 90 1 -2] [88 91 90 1 -2]] changed pixels: [[False False False True True] [False False True True True] [False True True True False] [ True True True False False] [ True True True False False]] percent of patch changed: 56%
What actually happened
The "before" patch has two ground types mixed together, values near 120 and values near 60. In the "after" patch, most of the low-value ground rose to around 150 — a large real change, standing in for something like bare soil replaced by a bright roof or a dried-up field turning into infrastructure. A few pixels barely moved, standing in for ground that stayed the same.
.astype(int)matters. The arrays hold small integers by default, and subtracting one unsigned integer array from another can silently wrap around to a huge positive number instead of going negative. Casting first avoids this trap.- The threshold of 30 is a decision boundary, chosen here by looking at the data. Real systems tune this threshold against labelled examples of known change and no-change, the way any classifier threshold gets tuned.
changed.mean()works becauseTrue/Falsebehave as1/0in NumPy arithmetic, so the mean of a boolean array is exactly the fraction that isTrue.
Common mistakes
Comparing images that are not aligned. If the two images are shifted by even one pixel, every difference becomes noise from misalignment, not real change. Production pipelines co-register images before differencing, often to sub-pixel accuracy.
Ignoring the season. Comparing a monsoon-season image against a dry-season image will show a country's worth of "changed" vegetation that changed nothing but the calendar. Compare same-season pairs, or normalise for it explicitly.
Picking one fixed threshold for every scene. A threshold tuned on a desert scene will flag half a forest as "changed" when reused elsewhere. Brightness ranges differ by land type, sensor, and time of year.
Try it yourself
Change the threshold from 30 to 10 and rerun. Watch how many of the small, real-noise differences (the values around 1 or -2) get wrongly flagged as change. That trade-off between missing real change and flagging noise is the central tuning problem in every change-detection system.
What to learn next
- Anomaly detection — flagging unusual points more generally, with proper evaluation.
- NDVI and vegetation indices — a signal built to be more robust to lighting than raw brightness.
- Model evaluation — how to score a detector like this one honestly.
Researcher — Mathematics and papers.
Formal setup
Given two co-registered images I_1 and I_2 over the same spatial extent, define a change map C:
D(x, y) = f( I_2(x, y), I_1(x, y) )
C(x, y) = 1 if D(x, y) > T
0 otherwiseI_1,I_2— the earlier and later images, each a function from pixel location(x, y)to a value or vector of band values.f— a comparison function. Simple differencing usesf(a, b) = |a - b|. Band ratioing usesf(a, b) = a / b, which is less sensitive to multiplicative illumination change than a raw difference.T— a threshold, either fixed, or set statistically as some number of standard deviations above the mean ofDover a known no-change region.
Method families
Algebraic methods — image differencing, band ratioing, change vector analysis (CVA). CVA generalises differencing to multiple bands by treating each pixel's band values as a vector and measuring both the magnitude and direction of the vector's shift between dates; direction indicates the type of change (e.g. vegetation loss vs. new water), magnitude indicates how much.
Post-classification comparison — classify each date's image independently into land-cover classes (forest, water, built-up, bare soil), then compare the class labels. This avoids radiometric normalisation issues entirely, at the cost of compounding two classifiers' errors — an error rate of e per classification date yields a combined error rate approaching 2e - e^2 for the change map, ignoring correlation between the two classifiers' mistakes.
Deep learning approaches — Siamese convolutional or transformer encoders process I_1 and I_2 through shared weights, and a decoder head predicts the change map directly from the two feature representations. Trained end-to-end, these implicitly learn robustness to the confounds a hand-tuned threshold cannot handle, at the cost of needing labelled change/no-change training data, which is scarce and expensive relative to single-date land-cover labels.
Radiometric normalisation
Before differencing, dates are commonly normalised to reduce illumination and atmospheric differences unrelated to real change. Pseudo-invariant feature (PIF) normalisation regresses one date's reflectance against the other over pixels assumed stable (roads, exposed rock), then applies the fitted linear transform to the whole image before comparing.
Evaluation
Change detection is typically framed as pixel-wise binary classification, evaluated with precision, recall and the kappa coefficient against a human-labelled reference change map — see model evaluation for the general metrics. Kappa corrects for the class imbalance that dominates most real scenes, where the overwhelming majority of pixels do not change in any given interval and raw accuracy is misleadingly high by default.
Papers and current state
- Chen, J. et al. (2020), DASNet: Dual Attentive Fully Convolutional Siamese Networks for Change Detection, a widely cited deep architecture for the paired-image setting.
- Daudt, R. C. et al. (2018), Fully Convolutional Siamese Networks for Change Detection, ICIP — an early and still commonly used baseline architecture.
- The LEVIR-CD and OSCD datasets are the standard public benchmarks for building and general change detection respectively.
- Global Forest Watch's deforestation alerts (Hansen et al., University of Maryland, published continuously) are a production system built on this family of methods at country scale, run against Landsat and Sentinel data.
What to learn next
- Anomaly detection — thresholding and scoring unusual observations in general.
- Convolutional neural networks — the encoder architecture behind most learned change detectors.
- Model evaluation — precision, recall and the imbalance problem in depth.