Modelling wildfire and flood risk
Wildfire and flood risk models combine weather and terrain data into a daily risk level, feeding early warning systems where a missed warning can cost lives.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A wildfire or flood risk model combines weather and land conditions into a single daily risk level for a place.
You already sense this without a model. Dry, crunchy grass on a hot, windy day feels like fire weather. Dark clouds stacking up over a river that is already full to the brim feels like flood weather. A risk model does the same kind of reasoning, combining many such signals precisely, instead of by a general feeling.
This is a genuinely high-stakes application. A risk model here can influence real decisions about evacuation and emergency response, where a missed warning risks lives, not only money. Read the honest note near the end of this lesson carefully — it matters more here than almost anywhere else on this platform.
Why it exists
Wildfire risk depends on several things at once. How hot and dry it is. How strong the wind is. How dry the vegetation itself has become, after a long dry spell. Flood risk depends on rainfall intensity, how saturated the ground already is, current river levels, and terrain. No single measurement tells the whole story on its own — risk comes from the combination.
Before models like this, risk assessment leaned heavily on fixed rule-of-thumb thresholds and expert judgement alone. A model that has learned from many past seasons can combine several weather and land signals into one risk estimate. Many of those signals are the same gridded weather data covered earlier in this section, updated daily, at the scale of an entire district or state.
How it works
WILDFIRE RISK FLOOD RISK
temperature, humidity, wind, rainfall intensity, soil saturation,
days since last rain (dryness) river level, terrain
| |
v v
[ risk model ] [ risk model ]
| |
v v
a risk LEVEL for a region and day --> feeds early warning systems,
resource pre-positioningForest departments in fire-prone regions, and river authorities managing flood-prone basins, both use models built on this general pattern, alongside satellite monitoring and ground observation.
Where you have already seen it
India's Forest Survey of India runs a forest fire alert system that combines satellite fire detection with weather conditions. India's Central Water Commission issues flood forecasts for major river systems, combining upstream rainfall and river-level data. Both feed into public warnings and emergency planning, not only research reports.
The honest note that matters most in this lesson
A model like this makes two kinds of mistakes, and they are not equally costly. Missing a genuinely dangerous day — predicting "low risk" when conditions turn out to be severe — can mean people are not warned in time. A false alarm — predicting "high risk" on a day that turns out fine — wastes resources, and can erode public trust in future warnings. It does not put anyone in immediate danger. Because of that imbalance, real disaster risk systems are deliberately tuned to accept more false alarms, in exchange for catching more genuine danger. They are built, validated, and signed off by disaster management authorities and domain scientists. They are never deployed as a sole decision-maker, straight from a model like the one below.
Remember this
- Risk comes from combining several signals at once — no single measurement tells the whole story for fire or flood risk.
- A missed warning and a false alarm are not equally costly. Real systems are built with that imbalance in mind, not tuned for raw accuracy alone.
- Public safety systems like this are validated by domain experts and disaster management authorities, not deployed from a model output alone.
What to learn next
- Classification — the general technique this lesson's risk model is built on.
- Choosing a threshold from costs — the formal version of the threshold trade-off demonstrated below.
- Weather and climate data formats — the underlying gridded data this kind of model is usually built from.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
A simplified fire-danger classifier, combining temperature, humidity, wind, and days since last rainfall — the same basic ingredients real fire weather indices use. The key result here is not accuracy alone, but what happens to missed high-risk days at different decision thresholds.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
N = 2000
temp = rng.normal(32, 6, N)
humidity = np.clip(rng.normal(45, 20, N), 5, 95)
wind = np.clip(rng.normal(15, 8, N), 0, None)
days_since_rain = np.clip(rng.exponential(6, N), 0, 40)
danger_score = 0.08*(temp-25) + 0.04*(60-humidity) + 0.05*wind + 0.10*days_since_rain
true_prob_high_risk = 1 / (1 + np.exp(-(danger_score - 3)))
is_high_risk = (rng.random(N) < true_prob_high_risk).astype(int)
X = np.stack([temp, humidity, wind, days_since_rain], axis=1)
X_train, X_test, y_train, y_test = train_test_split(X, is_high_risk, test_size=0.3, random_state=0)
model = LogisticRegression()
model.fit(X_train, y_train)
probs = model.predict_proba(X_test)[:, 1]
def evaluate(threshold):
preds = (probs >= threshold).astype(int)
missed = np.sum((y_test == 1) & (preds == 0)) # dangerous: unwarned high-risk days
false_alarms = np.sum((y_test == 0) & (preds == 1)) # costly, but not dangerous
return missed, false_alarms
total_high_risk = np.sum(y_test == 1)
for threshold in [0.5, 0.3, 0.15]:
missed, false_alarms = evaluate(threshold)
print(f"threshold={threshold:.2f}: missed {missed}/{total_high_risk} high-risk days, {false_alarms} false alarms")threshold=0.50: missed 116/242 high-risk days, 57 false alarms threshold=0.30: missed 50/242 high-risk days, 159 false alarms threshold=0.15: missed 6/242 high-risk days, 260 false alarms
What actually happened
At the default threshold of 0.5, the model misses nearly half of the truly dangerous days — 116 out of 242. In a warning system, that is a genuinely alarming failure rate, even though the model's overall accuracy across all days would look reasonable on paper.
Lowering the threshold to 0.15 catches all but 6 of those dangerous days — at the cost of nearly 5 times as many false alarms. That trade-off is not a flaw to be engineered away. It is the actual, unavoidable choice a real warning system has to make. That choice should be informed by how costly each kind of mistake really is — not by which threshold happens to look best on an accuracy chart.
Common mistakes
Optimising for overall accuracy alone. A model can be 90% accurate and still miss most of the rare, dangerous days, if those days are a small fraction of the data. That is exactly what happened at threshold=0.5 above. Always check recall on the dangerous class specifically, not overall accuracy.
Picking a threshold without involving domain experts. The "right" threshold depends on real costs — evacuation cost, warning fatigue, resource availability — that are not visible in a training dataset. This is a decision for disaster management professionals to make, informed by the model, not a number a data scientist should set alone.
Treating a synthetic, invented dataset like this one as representative of real fire or flood risk. The numbers above are illustrative only. Real fire-danger and flood models are built and validated on real historical data, reviewed by domain scientists, before they inform anything operational.
Try it yourself
Change rng.exponential(6, N) to rng.exponential(2, N) — a climate with shorter typical dry spells. Retrain, and recheck the same three thresholds. Watch how the trade-off between missed days and false alarms shifts as the underlying risk distribution changes.
What to learn next
- Choosing a threshold from costs — a rigorous treatment of exactly the trade-off shown above.
- Classification — the foundational technique this lesson builds on.
- Class weights — a common technique for training when the dangerous class is rare in the data, as it usually is here.
Researcher — Mathematics and papers.
Established fire weather indices
Before machine learning, fire danger rating already had well-established, physically motivated formulations. The Canadian Forest Fire Weather Index (FWI) System (Van Wagner, 1987) computes a chain of intermediate moisture-content codes from temperature, relative humidity, wind, and precipitation — Fine Fuel Moisture Code, Duff Moisture Code, Drought Code. It combines them into indices of fire spread rate and fuel availability. This system, or regional adaptations of it, remains the operational baseline that most machine-learned fire risk models are compared against, and often ingest as input features rather than replace outright.
Flood forecasting structure
Operational flood forecasting typically chains two models. A rainfall-runoff hydrological model converts precipitation and soil state into streamflow. A hydraulic routing model then propagates that streamflow downstream through a river network, accounting for channel geometry and travel time. Machine learning increasingly targets specific stages of this chain — notably rainfall-runoff modelling. LSTM-based models (Kratzert et al., 2019) have shown strong benchmark performance here, including in ungauged basins with limited historical data, using transfer learning across many catchments simultaneously.
Cost-sensitive learning
The developer example's threshold trade-off is a specific case of the general cost-sensitive classification problem. Given a cost matrix C(y, y_hat) assigning an asymmetric cost to each combination of true and predicted class, the Bayes-optimal decision threshold on predicted probability p is:
threshold* = C(FP) / ( C(FP) + C(FN) )where C(FP) and C(FN) are the costs of a false positive and false negative respectively. When C(FN) >> C(FP) — a missed dangerous day costing far more than a false alarm — the optimal threshold moves well below 0.5, exactly the pattern in the developer example's results. Eliciting realistic values for C(FP) and C(FN) in a life-safety context is a genuinely difficult applied problem. It generally requires structured input from domain experts and emergency management stakeholders — not a value a model developer sets unilaterally.
Evaluation for rare, high-stakes events
Standard accuracy is a poor metric when the dangerous class is rare, since a model predicting "low risk" always can still score well. Precision-recall curves, and metrics like the F-beta score (weighting recall more heavily than precision, for beta > 1) are more informative, alongside a fully specified, domain-reviewed cost analysis rather than a single summary number.
Current state and open problems
Machine-learned components are increasingly integrated into operational fire and flood forecasting, generally as one input alongside physically-based models and expert review, rather than as a standalone decision system. Genuine open challenges remain. Compound events — simultaneous or cascading extremes, like a heatwave followed immediately by extreme rainfall — are individually rare, and jointly rarer still, in historical training data. There is also the reliability of any statistically trained model under climate conditions genuinely unprecedented in the historical record — the same non-stationarity concern raised in downscaling climate projections.
Key references
- Van Wagner, C. E. (1987). Development and Structure of the Canadian Forest Fire Weather Index System. Canadian Forestry Service.
- Kratzert, F. et al. (2019). Toward Improved Predictions in Ungauged Basins: Exploiting the Power of Machine Learning. Water Resources Research.
- Jain, P. et al. (2020). A Review of Machine Learning Applications in Wildfire Science and Management. Environmental Reviews.
- Elkan, C. (2001). The Foundations of Cost-Sensitive Learning. IJCAI.
What to learn next
- Choosing a threshold from costs — the formal cost-sensitive framework introduced above.
- Class weights — practical techniques for the rare-positive-class setting common in disaster risk data.
- Weather and climate data formats — where the input data for models like this one originates.