Climate and Energy AI

Machine-learned weather forecasting

Instead of solving the physics of the atmosphere from scratch every time, a machine-learned weather model recognises patterns learned from decades of past weather.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest note
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A machine-learned weather model predicts tomorrow's weather by recognising patterns in decades of past weather, instead of solving the physics from scratch.

Your grandmother might tell you rain is coming because of how the crows are behaving, or a certain feel in the air. She is not solving equations. She is recognising a pattern she has seen play out many times before. Machine-learned weather forecasting works on a similar principle. It spots patterns in an enormous amount of past weather, rather than calculating the physics fresh each time.

Why it exists

Traditional weather forecasting, called numerical weather prediction (NWP), solves the physical equations of atmospheric motion directly — how air, heat, and moisture move and interact. This produces excellent forecasts, and it is genuinely expensive: a full global forecast can take hours on a supercomputer with thousands of processors.

A machine-learned model, once trained, skips that calculation entirely. It has learned, from years of gridded weather data, a direct mapping from "today's atmosphere" to "tomorrow's atmosphere." Producing a forecast becomes a single fast pass through a neural network. That takes seconds on one GPU, instead of hours on a supercomputer.

How it works

TRADITIONAL (physics-based)
atmosphere NOW  --[ solve fluid dynamics equations, hours, a supercomputer ]-->  forecast

MACHINE-LEARNED
atmosphere NOW  --[ a trained neural network, seconds, one GPU ]-->  forecast
                    (learned the pattern from decades of past weather)

Both approaches are still compared against the same yardstick: how close was the forecast to what actually happened.

Where you have already seen it

Research systems including Google DeepMind's GraphCast and Huawei's Pangu-Weather made international news in 2023. Published results showed them matching or beating traditional physics-based forecasts, on many standard accuracy measures, while running dramatically faster. Several national weather agencies now run machine-learned models alongside their traditional ones, comparing and sometimes blending the two.

An honest note

A model trained on historical weather has, by definition, seen mostly ordinary weather, because ordinary weather is what mostly happens. The rarest, most extreme events matter most for public safety warnings. They are also exactly the events a historical pattern-matcher has the least experience with. This is why forecasting agencies still combine machine-learned models with physics-based ones, and expert human judgement, rather than replacing either outright. That matters most for warnings that affect public safety.

Remember this

  • Traditional forecasting solves physics equations directly; machine-learned forecasting recognises patterns from historical data.
  • The main practical advantage is speed — seconds instead of hours, once the model is trained.
  • Rare, extreme weather is where a pattern-based model is on its least familiar ground.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy scikit-learn

Minimal runnable code

A simplified "weather-like" pattern: a warm front moving steadily across a small strip of grid points, with sensor noise. The model learns to predict tomorrow's grid directly from today's.

ml_weather.py
import numpy as np
from sklearn.linear_model import Ridge

rng = np.random.default_rng(0)

GRID = 20
STEPS = 400
SPEED = 0.6  # grid cells per timestep

def make_frame(t):
    center = (SPEED * t) % GRID
    x = np.arange(GRID)
    dist = np.minimum(np.abs(x - center), GRID - np.abs(x - center))  # wraps, like circling the globe
    return 20.0 * np.exp(-(dist ** 2) / 8.0) + rng.normal(0, 0.3, size=GRID)  # a "warm front" plus noise

frames = np.array([make_frame(t) for t in range(STEPS)])

X, y = frames[:-1], frames[1:]  # given today's grid, predict tomorrow's
split = 300
X_train, y_train = X[:split], y[:split]
X_test, y_test = X[split:], y[split:]

model = Ridge(alpha=1.0)
model.fit(X_train, y_train)
predicted = model.predict(X_test)

persistence = X_test  # the standard forecasting baseline: "tomorrow looks like today"

def rmse(a, b):
    return np.sqrt(np.mean((a - b) ** 2))

print(f"persistence baseline RMSE (predicting no change at all): {rmse(persistence, y_test):.3f}")
print(f"learned model RMSE:                                       {rmse(predicted, y_test):.3f}")
Output
persistence baseline RMSE (predicting no change at all): 1.824
learned model RMSE:                                       0.371

What actually happened

The persistence baseline — the simplest possible forecast, "tomorrow will look like today" — is a genuinely important comparison point in this field, not a strawman. Any real forecasting model has to beat it by a wide margin to be worth its cost. Here, the learned model's error is roughly a fifth of the persistence baseline's, because it picked up the moving-front pattern the persistence baseline has no way to see.

  • Ridge regression here plays the role of "the model," learning a fixed linear relationship between today's full grid and tomorrow's. Real systems like GraphCast use much larger, nonlinear architectures — but the training setup (predict the next state from the current one) is the same basic idea.
  • The % GRID wraparound in make_frame() is a small stand-in for the Earth being a sphere. Weather moving off one edge of a global grid re-enters from the other side.

Common mistakes

Comparing only against a strong baseline, and never a weak one, or vice versa. Reporting only "we beat persistence" is a low bar for a short forecast horizon, where persistence is often already good. Reporting only "we beat a state-of-the-art physics model" without ever mentioning persistence hides how much of the improvement is basic pattern-following.

Testing only on the kind of pattern seen in training. This example's test set is drawn from the exact same repeating pattern as training. A genuinely useful test set includes conditions meaningfully different from what training saw. That is the real-world equivalent of testing on a season, or a storm type, the model has not encountered.

Forgetting that weather is chaotic beyond a certain horizon. Even a perfect model faces a hard physical limit — small uncertainties in the current state grow over time, regardless of model quality. No forecasting approach, learned or physics-based, escapes this.

Try it yourself

Change SPEED to 1.5 so the pattern moves faster between snapshots. Retrain, and compare the new gap between persistence and the learned model. A faster-changing signal generally makes persistence worse and gives a real model more room to prove its worth.

What to learn next

Researcher — Mathematics and papers.

Problem formulation

Take the atmospheric state at time t on a discretised grid, X_t in R^{C x H x W}. C holds channels for variables like temperature, pressure, and wind components; H and W are spatial dimensions for latitude and longitude. A learned forecasting model approximates the flow map of the underlying PDE system directly:

X_hat_{t+dt} = f_theta(X_t)

where f_theta is typically a graph neural network, a vision transformer, or a Fourier neural operator, trained by minimising forecast error against reanalysis data (commonly ERA5) over a large historical period.

Notable architectures

  • GraphCast (Lam et al., 2023, published in Science) represents the global grid as a multi-mesh graph. It uses graph neural network message passing to model both local and long-range spatial interactions, trained autoregressively to roll forward multiple steps.
  • Pangu-Weather (Bi et al., 2023, Nature) uses a 3D Earth-specific transformer architecture operating directly over the pressure-level grid.
  • FourCastNet (Pathak et al., 2022) uses Adaptive Fourier Neural Operators, exploiting the efficiency of spectral convolution for global-scale spatial patterns.
  • Neural GCM (Kochkov et al., 2024, Nature) hybridises a traditional dynamical core with learned physics parameterisation, rather than replacing the physics-based approach outright.

All are trained and evaluated against ERA5 reanalysis, and benchmarked against ECMWF's operational HRES forecast — the reigning physics-based system — using standardised skill metrics.

Evaluation metrics

Standard metrics in this field include:

RMSE(t) = sqrt( mean( (X_hat_t - X_t)^2 ) )              -- root mean squared error at lead time t
ACC(t)  = corr( X_hat_t - climatology, X_t - climatology )  -- anomaly correlation coefficient

ACC is preferred at longer lead times. It measures skill relative to climatology — the historical average for that date and location — rather than raw error, which is dominated by seasonal variation that any reasonable model gets right without needing real skill.

Published results, and their limits

Lam et al. (2023) reported GraphCast outperforming ECMWF's HRES on the majority of 1,380 evaluated variable-and-lead-time combinations, at a small fraction of the computational cost. This is a genuine, published, peer-reviewed result — and it was measured on standard deterministic accuracy metrics, over the historical test period evaluated, predominantly reflecting typical rather than extreme conditions. Two things remain active, unsettled areas of ongoing research. One is performance on genuinely rare extreme events. The other is producing calibrated, physically consistent ensembles — multiple plausible future scenarios, not one number — rather than a single deterministic forecast.

Current state

Operational forecasting centres, as of current published practice, run machine-learned models as a complement to physics-based ensemble systems rather than a wholesale replacement, particularly for warnings tied to public safety. The chaotic nature of atmospheric dynamics imposes a hard predictability limit — roughly two weeks for any method, learned or physics-based — that no architecture has been shown to exceed.

Key references

  • Lam, R. et al. (2023). Learning Skillful Medium-Range Global Weather Forecasting. Science.
  • Bi, K. et al. (2023). Accurate Medium-Range Global Weather Forecasting with 3D Neural Networks. Nature.
  • Pathak, J. et al. (2022). FourCastNet: A Global Data-driven High-resolution Weather Model.
  • Kochkov, D. et al. (2024). Neural General Circulation Models for Weather and Climate. Nature.

What to learn next