Climate and Energy AI

Modelling wildfire and flood risk

Wildfire and flood risk models combine weather and terrain data into a daily risk level, feeding early warning systems where a missed warning can cost lives.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. The honest note that matters most in this lesson
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A wildfire or flood risk model combines weather and land conditions into a single daily risk level for a place.

You already sense this without a model. Dry, crunchy grass on a hot, windy day feels like fire weather. Dark clouds stacking up over a river that is already full to the brim feels like flood weather. A risk model does the same kind of reasoning, combining many such signals precisely, instead of by a general feeling.

This is a genuinely high-stakes application. A risk model here can influence real decisions about evacuation and emergency response, where a missed warning risks lives, not only money. Read the honest note near the end of this lesson carefully — it matters more here than almost anywhere else on this platform.

Why it exists

Wildfire risk depends on several things at once. How hot and dry it is. How strong the wind is. How dry the vegetation itself has become, after a long dry spell. Flood risk depends on rainfall intensity, how saturated the ground already is, current river levels, and terrain. No single measurement tells the whole story on its own — risk comes from the combination.

Before models like this, risk assessment leaned heavily on fixed rule-of-thumb thresholds and expert judgement alone. A model that has learned from many past seasons can combine several weather and land signals into one risk estimate. Many of those signals are the same gridded weather data covered earlier in this section, updated daily, at the scale of an entire district or state.

How it works

WILDFIRE RISK                                FLOOD RISK
temperature, humidity, wind,          rainfall intensity, soil saturation,
days since last rain (dryness)        river level, terrain
        |                                     |
        v                                     v
   [ risk model ]                        [ risk model ]
        |                                     |
        v                                     v
    a risk LEVEL for a region and day  -->  feeds early warning systems,
                                             resource pre-positioning

Forest departments in fire-prone regions, and river authorities managing flood-prone basins, both use models built on this general pattern, alongside satellite monitoring and ground observation.

Where you have already seen it

India's Forest Survey of India runs a forest fire alert system that combines satellite fire detection with weather conditions. India's Central Water Commission issues flood forecasts for major river systems, combining upstream rainfall and river-level data. Both feed into public warnings and emergency planning, not only research reports.

The honest note that matters most in this lesson

A model like this makes two kinds of mistakes, and they are not equally costly. Missing a genuinely dangerous day — predicting "low risk" when conditions turn out to be severe — can mean people are not warned in time. A false alarm — predicting "high risk" on a day that turns out fine — wastes resources, and can erode public trust in future warnings. It does not put anyone in immediate danger. Because of that imbalance, real disaster risk systems are deliberately tuned to accept more false alarms, in exchange for catching more genuine danger. They are built, validated, and signed off by disaster management authorities and domain scientists. They are never deployed as a sole decision-maker, straight from a model like the one below.

Remember this

  • Risk comes from combining several signals at once — no single measurement tells the whole story for fire or flood risk.
  • A missed warning and a false alarm are not equally costly. Real systems are built with that imbalance in mind, not tuned for raw accuracy alone.
  • Public safety systems like this are validated by domain experts and disaster management authorities, not deployed from a model output alone.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Minimal runnable code

A simplified fire-danger classifier, combining temperature, humidity, wind, and days since last rainfall — the same basic ingredients real fire weather indices use. The key result here is not accuracy alone, but what happens to missed high-risk days at different decision thresholds.

fire_risk.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
N = 2000

temp = rng.normal(32, 6, N)
humidity = np.clip(rng.normal(45, 20, N), 5, 95)
wind = np.clip(rng.normal(15, 8, N), 0, None)
days_since_rain = np.clip(rng.exponential(6, N), 0, 40)

danger_score = 0.08*(temp-25) + 0.04*(60-humidity) + 0.05*wind + 0.10*days_since_rain
true_prob_high_risk = 1 / (1 + np.exp(-(danger_score - 3)))
is_high_risk = (rng.random(N) < true_prob_high_risk).astype(int)

X = np.stack([temp, humidity, wind, days_since_rain], axis=1)
X_train, X_test, y_train, y_test = train_test_split(X, is_high_risk, test_size=0.3, random_state=0)

model = LogisticRegression()
model.fit(X_train, y_train)
probs = model.predict_proba(X_test)[:, 1]

def evaluate(threshold):
    preds = (probs >= threshold).astype(int)
    missed = np.sum((y_test == 1) & (preds == 0))       # dangerous: unwarned high-risk days
    false_alarms = np.sum((y_test == 0) & (preds == 1))  # costly, but not dangerous
    return missed, false_alarms

total_high_risk = np.sum(y_test == 1)
for threshold in [0.5, 0.3, 0.15]:
    missed, false_alarms = evaluate(threshold)
    print(f"threshold={threshold:.2f}: missed {missed}/{total_high_risk} high-risk days, {false_alarms} false alarms")
Output
threshold=0.50: missed 116/242 high-risk days, 57 false alarms
threshold=0.30: missed 50/242 high-risk days, 159 false alarms
threshold=0.15: missed 6/242 high-risk days, 260 false alarms

What actually happened

At the default threshold of 0.5, the model misses nearly half of the truly dangerous days — 116 out of 242. In a warning system, that is a genuinely alarming failure rate, even though the model's overall accuracy across all days would look reasonable on paper.

Lowering the threshold to 0.15 catches all but 6 of those dangerous days — at the cost of nearly 5 times as many false alarms. That trade-off is not a flaw to be engineered away. It is the actual, unavoidable choice a real warning system has to make. That choice should be informed by how costly each kind of mistake really is — not by which threshold happens to look best on an accuracy chart.

Common mistakes

Optimising for overall accuracy alone. A model can be 90% accurate and still miss most of the rare, dangerous days, if those days are a small fraction of the data. That is exactly what happened at threshold=0.5 above. Always check recall on the dangerous class specifically, not overall accuracy.

Picking a threshold without involving domain experts. The "right" threshold depends on real costs — evacuation cost, warning fatigue, resource availability — that are not visible in a training dataset. This is a decision for disaster management professionals to make, informed by the model, not a number a data scientist should set alone.

Treating a synthetic, invented dataset like this one as representative of real fire or flood risk. The numbers above are illustrative only. Real fire-danger and flood models are built and validated on real historical data, reviewed by domain scientists, before they inform anything operational.

Try it yourself

Change rng.exponential(6, N) to rng.exponential(2, N) — a climate with shorter typical dry spells. Retrain, and recheck the same three thresholds. Watch how the trade-off between missed days and false alarms shifts as the underlying risk distribution changes.

What to learn next

Researcher — Mathematics and papers.

Established fire weather indices

Before machine learning, fire danger rating already had well-established, physically motivated formulations. The Canadian Forest Fire Weather Index (FWI) System (Van Wagner, 1987) computes a chain of intermediate moisture-content codes from temperature, relative humidity, wind, and precipitation — Fine Fuel Moisture Code, Duff Moisture Code, Drought Code. It combines them into indices of fire spread rate and fuel availability. This system, or regional adaptations of it, remains the operational baseline that most machine-learned fire risk models are compared against, and often ingest as input features rather than replace outright.

Flood forecasting structure

Operational flood forecasting typically chains two models. A rainfall-runoff hydrological model converts precipitation and soil state into streamflow. A hydraulic routing model then propagates that streamflow downstream through a river network, accounting for channel geometry and travel time. Machine learning increasingly targets specific stages of this chain — notably rainfall-runoff modelling. LSTM-based models (Kratzert et al., 2019) have shown strong benchmark performance here, including in ungauged basins with limited historical data, using transfer learning across many catchments simultaneously.

Cost-sensitive learning

The developer example's threshold trade-off is a specific case of the general cost-sensitive classification problem. Given a cost matrix C(y, y_hat) assigning an asymmetric cost to each combination of true and predicted class, the Bayes-optimal decision threshold on predicted probability p is:

threshold* = C(FP) / ( C(FP) + C(FN) )

where C(FP) and C(FN) are the costs of a false positive and false negative respectively. When C(FN) >> C(FP) — a missed dangerous day costing far more than a false alarm — the optimal threshold moves well below 0.5, exactly the pattern in the developer example's results. Eliciting realistic values for C(FP) and C(FN) in a life-safety context is a genuinely difficult applied problem. It generally requires structured input from domain experts and emergency management stakeholders — not a value a model developer sets unilaterally.

Evaluation for rare, high-stakes events

Standard accuracy is a poor metric when the dangerous class is rare, since a model predicting "low risk" always can still score well. Precision-recall curves, and metrics like the F-beta score (weighting recall more heavily than precision, for beta > 1) are more informative, alongside a fully specified, domain-reviewed cost analysis rather than a single summary number.

Current state and open problems

Machine-learned components are increasingly integrated into operational fire and flood forecasting, generally as one input alongside physically-based models and expert review, rather than as a standalone decision system. Genuine open challenges remain. Compound events — simultaneous or cascading extremes, like a heatwave followed immediately by extreme rainfall — are individually rare, and jointly rarer still, in historical training data. There is also the reliability of any statistically trained model under climate conditions genuinely unprecedented in the historical record — the same non-stationarity concern raised in downscaling climate projections.

Key references

  • Van Wagner, C. E. (1987). Development and Structure of the Canadian Forest Fire Weather Index System. Canadian Forestry Service.
  • Kratzert, F. et al. (2019). Toward Improved Predictions in Ungauged Basins: Exploiting the Power of Machine Learning. Water Resources Research.
  • Jain, P. et al. (2020). A Review of Machine Learning Applications in Wildfire Science and Management. Environmental Reviews.
  • Elkan, C. (2001). The Foundations of Cost-Sensitive Learning. IJCAI.

What to learn next

What to learn next

These follow on from what you just read.

  • Satellite and Agricultural AI

    Working with satellite imagery

    Satellite imagery is a grid of numbers per patch of ground, with more colour channels than your eyes have, so working with it means working with stacked arrays, not photographs.

  • Satellite and Agricultural AI

    Detecting change from space

    Change detection compares two satellite images of the same ground taken at different times and flags what is genuinely different, which is harder than it sounds because clouds, seasons and sun angle all look like change too.

  • Satellite and Agricultural AI

    NDVI and vegetation indices

    NDVI turns the red and near-infrared bands into one number that tracks how alive and leafy a patch of ground is, because healthy plants reflect those two wavelengths in a very particular pattern.