Climate and Energy AI

Downscaling climate projections

Downscaling turns one coarse climate number covering a huge region into local detail, using known relationships like elevation, coastline, and terrain.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest note
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Downscaling turns one coarse climate number covering a huge region into detail for one specific place.

You have seen a weather map that shows one big blob of colour covering your entire state. But you know your cousin's house in the hills is far colder than your apartment in the city. The map draws no line between you at all. Downscaling fills in that difference. It turns "the whole region is about 27°C" into "the hill station is about 17°C, and the coastal city is about 29°C."

Why it exists

Global climate models simulate the entire planet's atmosphere and oceans, decades into the future. That is an enormous computation. So these models run at a coarse grid — commonly 50 to 250 kilometres per grid box. That is the kind of grid covered in weather and climate data formats. One box can easily contain a city, its suburbs, a river valley, and a range of hills, all reported as a single number.

Real decisions need detail at the scale of a few kilometres, not a hundred. Where to build a reservoir. How to plan a city's drainage. Which crops will still grow in a district by 2050. Downscaling bridges that gap.

How it works

Global climate model:  ~100km grid cells, whole planet, decades ahead
              |
              v
      DOWNSCALING
   (a learned or physics-based relationship between the coarse
    grid and local conditions: elevation, coastline, terrain)
              |
              v
Local projection:  ~1-10km grid cells, one region, same overall trend

There are two broad approaches. Dynamical downscaling runs a second, finer-resolution physics simulation over only the region of interest. It uses the coarse model's output as its boundary conditions — accurate, and expensive. Statistical downscaling learns a mapping from historical data, between what the coarse model showed and what actually happened locally. It applies that mapping to future coarse projections — much cheaper, and only as good as the assumption that the same relationship keeps holding.

Where you have already seen it

City-level climate adaptation plans are built from downscaled projections, not the raw global model output. How much monsoon rainfall intensity should a city plan drainage for? How many additional heatwave days should a district expect? These are the kinds of questions downscaling answers. India's state-level climate action plans use exactly this kind of localised data.

An honest note

A downscaled number carries two layers of uncertainty, stacked on top of each other. There is the original global model's uncertainty about the planet's future climate trajectory. Then the downscaling method adds its own approximation error on top. Statistical downscaling in particular assumes the historical relationship between the coarse and local scale keeps holding. That assumption gets shakier the further into an unprecedented climate future the projection reaches. Numbers used for real infrastructure or policy decisions should come from established climate science institutions. Their stated uncertainty ranges should stay attached, not get treated as a single precise prediction.

Remember this

  • Global climate models run coarse, because simulating the whole planet for decades is enormously expensive.
  • Downscaling adds local detail back in, using either a nested physics model or a learned statistical relationship.
  • A downscaled number inherits uncertainty from the global model, and adds its own on top.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Minimal runnable code

A simplified statistical downscaling model. Temperature genuinely drops with elevation, a real atmospheric effect called the lapse rate. A coarse regional model has no way to represent that within one grid box.

downscaling.py
import numpy as np
from sklearn.linear_model import LinearRegression

rng = np.random.default_rng(0)

LAPSE_RATE = 6.5  # degrees C cooler per km of elevation -- the real environmental lapse rate

def true_local_temp(coarse_regional_temp, elevation_km):
    return coarse_regional_temp - LAPSE_RATE * elevation_km

n_days = 300
coarse_temp = 25 + 4 * np.sin(np.linspace(0, 6 * np.pi, n_days)) + rng.normal(0, 1, n_days)
station_elevations = np.array([0.1, 0.5, 1.2, 2.0])  # km above sea level -- known stations

X_train, y_train = [], []
for day_temp in coarse_temp:
    for elev in station_elevations:
        observed = true_local_temp(day_temp, elev) + rng.normal(0, 0.5)
        X_train.append([day_temp, elev])
        y_train.append(observed)

downscaler = LinearRegression()
downscaler.fit(X_train, y_train)

# A brand-new hill station, at an elevation never seen in training.
NEW_ELEVATION = 1.6
test_coarse_temp = 27.0

downscaled = downscaler.predict([[test_coarse_temp, NEW_ELEVATION]])[0]
true_value = true_local_temp(test_coarse_temp, NEW_ELEVATION)

print(f"coarse regional forecast: {test_coarse_temp:.1f}C (same number for the whole region)")
print(f"true local temperature at {NEW_ELEVATION}km elevation: {true_value:.1f}C")
print(f"downscaled estimate at {NEW_ELEVATION}km elevation:    {downscaled:.1f}C")
print(f"error, using the coarse grid directly: {abs(test_coarse_temp - true_value):.1f}C")
print(f"error, after downscaling:              {abs(downscaled - true_value):.1f}C")
Output
coarse regional forecast: 27.0C (same number for the whole region)
true local temperature at 1.6km elevation: 16.6C
downscaled estimate at 1.6km elevation:    16.6C
error, using the coarse grid directly: 10.4C
error, after downscaling:              0.0C

What actually happened

Using the coarse grid's number directly is off by more than 10 degrees at this new hill station, because that number never accounted for elevation at all. The downscaling model was trained on other stations at different elevations. It learned the elevation relationship, and applied it correctly to a station it had never seen — a genuinely new elevation value, 1.6, that never once appeared during training.

  • The model never saw 1.6 as an elevation during training (station_elevations only contains 0.1, 0.5, 1.2, 2.0). It interpolated correctly because the relationship it learned — temperature drops steadily with elevation — is close to linear over this range. That lets it generalise smoothly between, and slightly beyond, training values.
  • This is a deliberately simple case. Real statistical downscaling also accounts for coastline distance, land cover, and urban heat effects, none of which follow as clean a relationship as elevation does.

Common mistakes

Applying a statistical downscaling model far outside the range it was trained on. The linear relationship used here works cleanly near training elevations. Extrapolating to a mountain three times higher than anything in training is a different, much less reliable claim.

Ignoring that downscaling does not fix a wrong coarse forecast. If the input global-model number itself is wrong, downscaling only adds local detail to that wrong number — it cannot correct an error that happened upstream.

Treating one downscaled run as the answer. Real climate projections use ensembles — many global models, many downscaling methods — and report a range, not one number. A single downscaled value, presented without its uncertainty range, misrepresents what the science actually supports.

Try it yourself

Add a fifth training elevation at 4.0 km, and generate its observations the same way. Retrain, and test the model at NEW_ELEVATION = 3.0. Compare how well it does now that a higher elevation was represented in training, versus how well it did purely extrapolating before.

What to learn next

Researcher — Mathematics and papers.

Dynamical downscaling

A Regional Climate Model (RCM) solves the same physical equations as a Global Climate Model (GCM), at higher spatial resolution, over a limited domain, using the GCM's output as time-varying lateral boundary conditions. This preserves physical consistency: energy and mass balance hold, and cloud and precipitation processes are simulated rather than statistically inferred. The cost is substantially higher than statistical methods, and it still inherits any systematic bias present in the driving GCM.

Statistical downscaling formulations

Given coarse-scale predictors X (large-scale temperature, pressure, circulation patterns) and local-scale observations Y, statistical downscaling fits:

Y = f(X) + epsilon

Common instantiations include:

  • Bias correction and spatial disaggregation (BCSD). Corrects the GCM's marginal distribution against observed local climatology, then spatially disaggregates using a fixed local pattern. It is cheap, and it assumes the local pattern of variability stays constant under climate change — which is not guaranteed.
  • Quantile mapping. Maps each quantile of the modelled distribution to the corresponding quantile of the observed distribution, correcting distributional bias more thoroughly than a simple mean shift.
  • Regression-based and machine-learned downscaling. The developer example's LinearRegression is the simplest member of this family; deep learning variants (often convolutional, treating this as an image super-resolution problem) are increasingly common in recent literature.

The stationarity assumption

Every statistical downscaling method implicitly assumes the statistical relationship between large-scale and local-scale climate — learned from a historical training period — remains valid under future forcing conditions outside that historical range. This is a non-stationarity risk: under a substantially warmer climate, local feedback processes (soil moisture depletion, altered atmospheric circulation patterns) can shift the very relationship the statistical model learned. Dynamical downscaling does not eliminate this concern either, since it inherits any structural bias in the driving GCM's representation of future forcing.

Uncertainty quantification

Downscaled projections are properly reported as ensembles, spanning multiple axes of uncertainty. Which GCM. Which emissions scenario — the Shared Socioeconomic Pathways, SSPs, in CMIP6. Which downscaling method. And internal climate variability, since the same forced trend produces a spread of possible individual weather realisations. The Coupled Model Intercomparison Project (CMIP) coordinates the multi-model ensembles that underlie most published downscaled products, specifically to make this uncertainty structure visible rather than hidden inside a single number.

Current practice and limits

Regional downscaling initiatives — CORDEX internationally, and national programmes building on it — provide standardised, peer-reviewed downscaled datasets for many regions, including South Asia, rather than leaving individual analysts to build ad hoc downscaling pipelines. Even so, downscaled precipitation extremes remain a documented weak point across both dynamical and statistical methods, because extreme events are, by definition, the sparsest part of any historical training record.

Key references

  • Wilby, R. L., Wigley, T. M. L. (1997). Downscaling General Circulation Model Output: A Review of Methods and Limitations. Progress in Physical Geography.
  • Maraun, D., Widmann, M. (2018). Statistical Downscaling and Bias Correction for Climate Research. Cambridge University Press.
  • Giorgi, F., Gutowski, W. J. (2015). Regional Dynamical Downscaling and the CORDEX Initiative. Annual Review of Environment and Resources.
  • Eyring, V. et al. (2016). Overview of the Coupled Model Intercomparison Project Phase 6 (CMIP6). Geoscientific Model Development.

What to learn next