Predicting crop yield
Crop yield prediction turns a season's worth of satellite and weather signals into an estimate of how much a field will produce, and it is a regression problem wearing farming clothes.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Crop yield prediction estimates how much a field will produce before harvest, from satellite and weather signals gathered during the season.
Think about a grandmother judging a mango tree by how heavily it flowered this year. This happens weeks before a single fruit has formed. She has seen enough seasons to connect flowering with the eventual harvest, without measuring anything.
A yield model tries to formalise that same judgement, using satellite readings and weather records instead of a lifetime of memory.
Why it exists
Governments need to know roughly how much wheat or rice a state will produce. They need it months before harvest, to plan imports, storage and pricing. Insurance companies need an estimate to set fair premiums. Farmers benefit from an early warning that a field is trending toward a poor season, while there is still time to act.
Waiting until harvest to count sacks is too late for every one of those decisions. A model reads the signs mid-season instead — how green the crop is, how much rain has fallen. It gives an estimate while that estimate still matters.
How it works
Satellite readings Weather records Learned pattern from Estimated yield
through the season + through the season -> many past seasons -> for this field
(how green, how (a trained model) this season
much canopy)The model trains on many past field-seasons where the true final yield is already known, the way supervised learning always works. It learns the relationship between mid-season signals and the harvest that followed. Then it applies that relationship to a field whose harvest has not happened yet.
Where you have already seen it
- India's FASAL programme (Forecasting Agricultural output using Space, Agrometeorology and Land based observations), which produces national crop production estimates.
- Commodity price movements around harvest season, partly driven by early yield forecasts from agencies and private firms.
- Crop insurance schemes that use satellite-based yield estimates to help decide payouts across large areas, rather than inspecting every individual field.
An honest warning
A yield prediction is a statistical estimate built from limited history, not a promise. Weather can turn in the final weeks of a season in ways no mid-season signal captures. A late hailstorm, a sudden pest outbreak, an irrigation failure — none of these show up early. A model trained on one region's soil and crop varieties often performs badly elsewhere, without retraining.
These numbers can feed real payouts and real policy decisions. Treat a yield model's output as one input to a human decision, never the decision itself. A model that looks accurate on last year's data can still be badly wrong. A season with unfamiliar weather is exactly when that happens.
Remember this
- Yield prediction is regression: turning mid-season signals into a number, the way any other regression problem does.
- It trades early warning for uncertainty — it is never as accurate as counting the harvest after the fact.
- It feeds decisions with real financial weight, so it needs to stay one input among several, with a human in the loop.
What to learn next
- Linear regression — the general technique this lesson's model is built from.
- Why crop disease models fail in the field — a related model with sharper failure modes.
- Model evaluation — how to check whether a yield model is actually any good.
Developer — Code and libraries.
Setup
pip install scikit-learnMinimal runnable code
from sklearn.linear_model import LinearRegression
# Each row: [peak-season NDVI x100, total monsoon rainfall in mm] for one past field-season
# Synthetic numbers, illustrative only -- not real agronomic data
X = [
[72, 620], [68, 540], [75, 700], [55, 380], [61, 420],
[80, 760], [58, 400], [70, 610], [64, 480], [77, 690],
]
# yield in quintals per hectare for that same field-season
y = [42, 37, 46, 24, 29, 50, 26, 41, 32, 47]
model = LinearRegression()
model.fit(X, y)
# a new field: NDVI 0.66 at peak season, 560 mm of rain so far
new_field = [[66, 560]]
prediction = model.predict(new_field)
print("predicted yield (quintals/hectare):", round(prediction[0], 1))
print("NDVI coefficient:", round(model.coef_[0], 3))
print("rainfall coefficient:", round(model.coef_[1], 3))
print("R-squared on training data:", round(model.score(X, y), 3))predicted yield (quintals/hectare): 36.3 NDVI coefficient: 0.574 rainfall coefficient: 0.033 R-squared on training data: 0.997
What actually happened
LinearRegression finds the best straight-line relationship between the two features and the yield, the same technique covered in linear regression, applied to two inputs instead of one.
- The NDVI coefficient (0.574) is far larger than the rainfall coefficient (0.033), but this does not mean NDVI matters roughly 17 times more. The two features live on very different scales — NDVI as a small number near 70, rainfall in the hundreds — so raw coefficients cannot be compared directly without first scaling both features.
model.score(X, y)returns R-squared, the fraction of variation in yield the model explains on the data it was trained on. A number this close to 1.0, on only ten rows, is not a sign of a great model. It is a sign of too little data — see the mistake below.
Common mistakes
Trusting a near-perfect score from ten rows. Ten field-seasons is nowhere near enough to trust a model, and this example uses exactly that many on purpose. With few examples relative to the number of features, a linear model can fit the training data almost perfectly and still generalise badly to a field it has not seen. See overfitting and underfitting.
Never holding out a test set. This example trains and scores on the same ten rows, which always inflates the score. Any real yield model needs a proper split, covered in train-test split, ideally split by field or by year so the model is tested on seasons it genuinely never saw.
Applying a model trained on one region to a different one. Soil type, crop variety and typical rainfall all shift the relationship between NDVI and yield. A model trained on Punjab wheat fields is not safe to apply unmodified to Tamil Nadu paddy fields.
Try it yourself
Add three more rows with a poor-rainfall, high-NDVI field-season (something like [74, 300] paired with a lower yield than the NDVI alone would suggest). Watch how the rainfall coefficient changes once the model has an example where the two features disagree.
What to learn next
- Linear regression — the full mechanics of the model used here.
- Random forest — a model family that handles this kind of non-linear, multi-signal problem better in practice.
- Model evaluation — proper scoring instead of the training-set score shown above.
Researcher — Mathematics and papers.
Problem framing
Crop yield prediction is typically framed as regression over a feature vector aggregated per field-season:
y_hat = f(x_1, ..., x_p)y_hat— predicted yield, usually in mass per unit area (kg/ha, quintals/ha, bushels/acre).x_1, ..., x_p— features aggregated over the growing season: vegetation index statistics (mean, max, integral of NDVI or EVI over time), accumulated precipitation, growing degree days, soil moisture, and increasingly, raw multi-temporal satellite imagery consumed directly by a deep model instead of hand-engineered summary statistics.f— historically a linear or tree-based regressor; increasingly a recurrent or transformer-based sequence model consuming the full within-season time series rather than a handful of summary features.
Feature construction from time series
A common classical pipeline computes the integral of NDVI over the growing season (a proxy for total photosynthetic activity, related to accumulated biomass), plus growing degree days:
GDD = sum over days of max(0, (T_max + T_min)/2 - T_base)T_max,T_min— daily maximum and minimum temperature.T_base— a crop-specific base temperature below which growth is assumed negligible (commonly 10°C for maize).
Model families and current state
| Approach | Typical input | Note |
|---|---|---|
| Linear / ridge regression on aggregated features | Season-level summary statistics | Interpretable, strong baseline, common in agency reporting |
| Random forest / gradient boosting | Same, or per-timestep features | Handles non-linear interactions between rainfall and NDVI well |
| LSTM / temporal CNN over the raw NDVI series | Full within-season time series | Captures timing effects a summary statistic loses (e.g. a dry spell during flowering matters more than the same total rainfall deficit spread evenly) |
| CNN over raw multi-band, multi-date imagery | Stacked image tensor per field per date | Learns its own features rather than relying on NDVI; state of the art on large benchmark datasets, at the cost of far more data and compute |
You et al. (2017), Deep Gaussian Process for Crop Yield Prediction, combined a CNN feature extractor over MODIS imagery with a Gaussian process for uncertainty-aware yield prediction at county scale in the US Corn Belt, and remains a widely cited reference architecture.
Validation and the honest caveat
Held-out evaluation must respect the structure of the data or it silently overstates performance: a random row-wise train/test split lets highly correlated neighbouring-year or neighbouring-field observations leak between train and test, a case of group leakage. The defensible protocols are leave-one-year-out (testing generalisation to an unseen weather year) and leave-one-region-out (testing generalisation to unseen soil and management conditions), both of which typically show substantially worse error than a naive random split.
Reported errors in the literature (commonly 8-15% mean absolute percentage error at county or district scale, higher at individual-field scale) are specific to the crop, region and years studied, and do not transfer as a general accuracy guarantee to a new deployment. Any operational use feeding insurance payouts, subsidy decisions or food-security planning needs domain-expert agronomic validation against ground-truthed local harvest data before being trusted, not evaluation on held-out data from the same distribution alone.
Papers
- You, J. et al. (2017). Deep Gaussian Process for Crop Yield Prediction Based on Remote Sensing Data. AAAI.
- Jiang, Z. et al. (2020). Predicting County Level Corn Yields Using Deep Long Short Term Memory Models.
- Khaki, S. and Wang, L. (2019). Crop Yield Prediction Using Deep Neural Networks. Frontiers in Plant Science.
What to learn next
- Random forest — the tree-ensemble family widely used as a strong classical baseline here.
- LSTM for forecasting — the sequence-model approach to the within-season time series.
- Model evaluation — leave-one-year-out and leave-one-region-out validation in depth.