NDVI and vegetation indices
NDVI turns the red and near-infrared bands into one number that tracks how alive and leafy a patch of ground is, because healthy plants reflect those two wavelengths in a very particular pattern.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
NDVI is a single number for how alive and leafy a patch of ground looks. It comes from comparing two colours of light.
Pick up a fresh green leaf snapped off a healthy plant, then a dry brown leaf that fell off weeks ago. The fresh leaf feels cool, full of water, and springs back when you bend it. The dry one is brittle and crumbles.
That difference is not only visible and touchable. It shows up strongly in near-infrared light too — a kind of light your eyes cannot see, covered in the previous lesson. A living leaf is full of water and the green pigment chlorophyll. It bounces back a lot of near-infrared light, while soaking up red light for photosynthesis. A dry or dead leaf loses that pattern almost completely.
Why it exists
An ordinary colour photo only tells you roughly how green something looks. Green is a weak, late signal. A crop shows stress in the infrared long before it turns visibly yellow to your eyes.
Scientists wanted one clean number for "how much like living plant tissue does this pixel behave". It had to be more reliable than colour alone. NDVI — the Normalized Difference Vegetation Index — was invented to be that number, from a comparison of red and near-infrared reflectance.
How it works
Near-infrared bounced back -> compare the two -> NDVI score
Red light bounced back roughly -1 to 1Every pixel gets compared this way, turning a two-band satellite image into a single grid of scores. High scores mean dense, healthy, leafy vegetation. Scores near zero mean bare soil or rock. Negative scores usually mean water, snow, or clouds.
Where you have already seen it
- Farm monitoring apps that show a field coloured green-to-brown, letting a farmer spot a stressed patch of a crop from a phone.
- Drought-monitoring dashboards run by agricultural departments, tracking how green a whole state is compared to a normal year.
- News reports on shrinking forest cover, which often use NDVI trends over years as evidence.
An honest warning
NDVI is a useful shortcut, not ground truth. It can be fooled. A patch of weeds looks every bit as "green" to NDVI as a healthy crop. The index cannot tell you what species is growing. It only tells you that something leafy and alive is there. Dense cloud shadow, wet soil, and some crops at some growth stages can also produce misleading scores.
Treat an NDVI map as a strong hint that tells you where to look closer, not as a verdict on its own.
Remember this
- NDVI compares near-infrared and red reflectance into one score, roughly from -1 to 1.
- High NDVI means healthy, leafy vegetation. Low or negative NDVI means soil, water, or non-living surfaces.
- It flags where something needs a closer look — it does not identify what the something is.
What to learn next
- Predicting crop yield — using NDVI trends as one input among several.
- Working with satellite imagery — the band and pixel concepts NDVI is built from.
- Feature engineering — the general skill of turning raw signals into a useful number.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
import numpy as np
# Synthetic red and near-infrared bands for a 4x4 patch
# Left half: healthy crop. Right half: bare soil / dry ground.
red = np.array([
[40, 42, 180, 185],
[38, 44, 178, 190],
[41, 39, 182, 188],
[43, 40, 179, 186],
], dtype=float)
nir = np.array([
[200, 210, 90, 85],
[205, 208, 88, 92],
[198, 202, 91, 87],
[207, 199, 89, 90],
], dtype=float)
ndvi = (nir - red) / (nir + red)
print("NDVI grid:")
print(np.round(ndvi, 2))
print("\nmean NDVI, left half (crop):", round(ndvi[:, :2].mean(), 2))
print("mean NDVI, right half (soil):", round(ndvi[:, 2:].mean(), 2))
# a common rough rule of thumb: NDVI above 0.3 suggests live green vegetation
healthy = ndvi > 0.3
print("\nlikely vegetation:")
print(healthy)NDVI grid: [[ 0.67 0.67 -0.33 -0.37] [ 0.69 0.65 -0.34 -0.35] [ 0.66 0.68 -0.33 -0.37] [ 0.66 0.67 -0.34 -0.35]] mean NDVI, left half (crop): 0.67 mean NDVI, right half (soil): -0.35 likely vegetation: [[ True True False False] [ True True False False] [ True True False False] [ True True False False]]
What actually happened
The formula is (nir - red) / (nir + red), applied to every pixel at once through NumPy's element-wise arithmetic — no loop needed.
- The healthy crop side has high NIR (around 200) and low red (around 40), so the difference is large and positive relative to the sum. The soil side has NIR and red close together, giving a small or negative score.
dtype=floaton both arrays matters. Left as integers,(nir - red) / (nir + red)performs integer division in older numeric code and silently loses the decimal part. Declaring float up front avoids that entire class of bug.- The 0.3 threshold used for
healthyis a common rough starting point in agricultural monitoring, not a universal law. Real thresholds are usually tuned per crop and per growth stage.
Common mistakes
Dividing by zero. A pixel where both NIR and red happen to be zero (common in "no data" regions) produces 0/0, which NumPy turns into nan with a runtime warning instead of crashing. Check for nan values before running statistics on a real NDVI map, with np.isnan(ndvi).sum().
Reusing the 0.3 threshold everywhere. A young crop early in the season legitimately scores lower than a mature one, without anything being wrong. Comparing a field's NDVI against its own history is usually more meaningful than comparing against a fixed number.
Assuming NDVI works the same on every satellite. Different sensors define their red and near-infrared bands with slightly different wavelength ranges, so raw NDVI values are not perfectly comparable across satellites without care.
Try it yourself
Add a nan value to one cell of the nir array (nir[0, 0] = np.nan) and rerun. Watch what happens to that cell's NDVI value and to ndvi.mean() — a single missing value quietly poisons an average unless you filter it out with np.nanmean(ndvi).
What to learn next
- Predicting crop yield — feeding an NDVI time series into a real model.
- Feature engineering — the general pattern of building a useful signal like this one.
- Time series features — turning an NDVI trend over a season into model inputs.
Researcher — Mathematics and papers.
Definition
NDVI = (NIR - RED) / (NIR + RED)NIR— surface reflectance in the near-infrared band, typically 760-900 nm.RED— surface reflectance in the red band, typically 620-690 nm.- The formula is bounded to
[-1, 1]wheneverNIR + RED > 0and both bands are non-negative reflectance values, since the numerator's magnitude cannot exceed the denominator.
Rouse et al. (1974) introduced NDVI for monitoring rangeland vegetation from early Landsat data. The physical basis: chlorophyll in living leaf mesophyll strongly absorbs red light for photosynthesis while leaf cell structure strongly scatters near-infrared light, producing the large NIR-RED contrast that healthy canopy exhibits and that senescent or absent vegetation does not.
Known failure modes and successor indices
NDVI saturates at high leaf area index — dense, closed canopies plateau near the top of the NDVI range and stop discriminating between "healthy" and "very healthy" vegetation. It is also sensitive to soil background reflectance in sparse canopies and to atmospheric aerosol contamination in the red band.
| Index | Formula | What it corrects for |
|---|---|---|
| SAVI (Soil-Adjusted Vegetation Index) | ((NIR - RED) / (NIR + RED + L)) * (1 + L) | Soil brightness in sparse canopy, via a soil-adjustment constant L, typically 0.5 |
| EVI (Enhanced Vegetation Index) | G * (NIR - RED) / (NIR + C1*RED - C2*BLUE + L) | Canopy background and atmospheric aerosol scattering, using a blue band term |
| NDRE (Normalized Difference Red Edge) | (NIR - RedEdge) / (NIR + RedEdge) | Saturation at high leaf area index, using the red-edge band available on Sentinel-2 |
| GNDVI | (NIR - GREEN) / (NIR + GREEN) | Chlorophyll concentration specifically, more sensitive than NDVI in later growth stages |
L, C1, C2 and G in SAVI and EVI are empirically fitted correction constants from the original papers (Huete, 1988; Huete et al., 1997), not universal physical values.
Cost
Computing NDVI over an image is O(n) in the number of pixels n, a single element-wise pass. The dominant cost in practice is I/O and atmospheric correction upstream of the index itself, not the arithmetic.
Papers
- Rouse, J. W. et al. (1974). Monitoring Vegetation Systems in the Great Plains with ERTS. NASA SP-351. Original NDVI definition.
- Huete, A. R. (1988). A Soil-Adjusted Vegetation Index (SAVI). Remote Sensing of Environment 25(3). Introduces SAVI.
- Huete, A. et al. (1997, 2002). Development of the Enhanced Vegetation Index (EVI) for MODIS.
- Gitelson, A. A. et al. (1996). Use of a Green Channel in Remote Sensing of Global Vegetation from EOS-MODIS. Introduces GNDVI.
Current state
NDVI remains the default first pass in operational agricultural monitoring because it needs only two bands, available on almost every optical satellite since Landsat 1. Research and commercial systems increasingly move to red-edge and multi-index feature stacks feeding a learned model rather than a single hand-derived index and threshold, especially for crop-specific applications — see predicting crop yield.
What to learn next
- Predicting crop yield — NDVI time series as model features.
- Feature engineering — the general theory of constructed features like this one.
- Time series features — extracting trend and seasonality from an NDVI series.