Electricity load forecasting
Load forecasting predicts how much electricity a grid will need, hour by hour, because electricity has to be generated at nearly the same instant it is used.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Load forecasting predicts how much electricity a whole city or region will need, hour by hour.
You already know your own house uses more electricity on a hot afternoon, with every fan and air conditioner running. Far less flows at 3am, when everyone is asleep. You know a festival evening, lights strung up everywhere, pulls more power than an ordinary Tuesday. Grid operators need to know the same thing — not for one house, but for millions of them at once.
Why it exists
Electricity is unusual among things people buy: it has to be generated at almost the exact same instant it is used. Large-scale storage is limited, as the next lesson covers in detail. So a grid mostly cannot "make extra" electricity and save it for later.
That means supply has to be scheduled in advance to match expected demand, called the load. Power plants take real time to start up — some coal plants take hours. If a grid operator does not know demand is about to spike until it happens, there is no way to react fast enough.
How it works
past demand patterns + calendar (weekday/weekend, holiday, season) + weather forecast (temp -> AC/heating use)
|
v
[ model: predicted demand, in megawatts, for each coming hour ]
|
v
grid operators schedule power plants and reserves to matchTemperature is one of the strongest predictors. Hot weather drives air conditioning demand up sharply; cold weather does the same for heating, in regions that need it. A forecast that ignores weather entirely is ignoring one of the biggest sources of genuine, real variation in demand.
Where you have already seen it
The UK's National Grid famously tracks "TV pickup." Right after a major televised match ends, or at half-time, demand spikes sharply and suddenly, as millions of kettles switch on within minutes of each other. It is a well-documented, real example of how tightly linked electricity demand is to what people are actually doing, at that exact moment, together.
An honest note
Underestimating load is not only a cost problem. If actual demand exceeds available supply and reserves, the result can be a real blackout. That means consequences for hospitals, traffic systems, and water supply — not only inconvenience. This is why grid operators build reserve margins on top of any forecast. Those margins are sized to cover realistic forecast error, rather than trusting any single number completely.
Remember this
- Electricity generally has to be produced at nearly the same moment it is consumed, which makes accurate demand forecasting essential to keeping a grid stable.
- Weather, calendar, and shared human behaviour (a big match, a festival) all drive real, predictable swings in demand.
- Grids are built with reserve margin specifically because no forecast is perfect. The margin, not the forecast alone, is what prevents a shortfall from becoming a blackout.
What to learn next
- Battery dispatch and grid decisions — matching this demand forecast against available supply and storage.
- Forecasting solar and wind output — the matching supply-side forecast.
- Evaluating a forecast — the general standards this kind of forecast is judged against.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
A simplified city's electricity demand: a daily two-peak shape (morning and evening), lower on weekends, with extra demand from air conditioning once the temperature crosses 30°C.
import numpy as np
from sklearn.ensemble import RandomForestRegressor
rng = np.random.default_rng(0)
def daily_shape(hour, is_weekend):
morning = 30 * np.exp(-((hour - 9) ** 2) / 8)
evening = 40 * np.exp(-((hour - 19) ** 2) / 10)
if is_weekend:
return 60 + 0.6 * (morning + evening)
return 60 + morning + evening
def true_load(hour, is_weekend, temp_c):
ac_demand = np.clip(temp_c - 30, 0, None) * 3.0 # air conditioning above 30C
return daily_shape(hour, is_weekend) + ac_demand
n_days = 120
day_is_weekend = (np.arange(n_days) % 7) >= 5
day_temp = 28 + 4 * np.sin(np.linspace(0, 4 * np.pi, n_days)) + rng.normal(0, 2, n_days)
hours, weekends, temps, load = [], [], [], []
for d in range(n_days):
for h in range(24):
hours.append(h); weekends.append(day_is_weekend[d]); temps.append(day_temp[d])
load.append(true_load(h, day_is_weekend[d], day_temp[d]) + rng.normal(0, 1.5))
X = np.stack([hours, weekends, temps], axis=1).astype(float)
model = RandomForestRegressor(n_estimators=150, random_state=0)
model.fit(X, load)
# A heatwave weekday, hotter than anything typical in training.
test_hours = np.arange(24)
HEATWAVE_TEMP = 42.0
actual = true_load(test_hours, False, HEATWAVE_TEMP)
model_forecast = model.predict(np.stack([test_hours, np.zeros(24), np.full(24, HEATWAVE_TEMP)], axis=1))
# A common practical baseline: "load will look like the same hour, same day type, last week"
same_hour_last_week = true_load(test_hours, False, 28.0)
def rmse(a, b):
return np.sqrt(np.mean((a - b) ** 2))
print(f"'same hour last week' baseline RMSE: {rmse(same_hour_last_week, actual):.2f} MW")
print(f"weather-aware model RMSE: {rmse(model_forecast, actual):.2f} MW")
print(f"actual peak load during heatwave (7pm): {actual[19]:.1f} MW")
print(f"baseline forecast at 7pm: {same_hour_last_week[19]:.1f} MW")
print(f"model forecast at 7pm: {model_forecast[19]:.1f} MW")'same hour last week' baseline RMSE: 36.00 MW weather-aware model RMSE: 19.54 MW actual peak load during heatwave (7pm): 136.0 MW baseline forecast at 7pm: 100.0 MW model forecast at 7pm: 115.3 MW
What actually happened
The weather-aware model roughly halves the baseline's error. It can see the heatwave coming through the temperature input, while "same hour last week" has no way to know this week is unusually hot. Neither forecast is perfect — the model still underestimates the true evening peak by about 21 megawatts.
That gap matters. HEATWAVE_TEMP = 42.0 is hotter than almost anything in training (day_temp mostly stays in a much narrower range). Tree-based models like this one — the same limitation seen in imitation learning and behaviour cloning — do not extrapolate reliably past what they were trained on. A genuinely unprecedented heatwave can still catch a well-trained model off guard.
Common mistakes
Training only on "normal" weather. If historical training data rarely includes extreme heat, the model has little to learn the extreme-heat relationship from — exactly the gap shown above. Real load forecasting systems deliberately weight or oversample extreme conditions.
Ignoring holidays and special events explicitly. A public holiday can look like a weekday by the calendar, but behave like a weekend, or something else entirely, in demand. These need to be flagged as their own category, not inferred from day-of-week alone.
Reporting one national number instead of per-region forecasts. Weather, and the demand it drives, varies a great deal across even one country. A single nationwide forecast hides real regional swings that a grid operator managing a specific area needs to see.
Try it yourself
Add a is_holiday feature to the training data, and generate a few holiday days where demand follows a flatter, weekend-like pattern regardless of the actual day of week. Retrain, and check the forecast for a holiday that falls on what would otherwise be a weekday.
What to learn next
- Battery dispatch and grid decisions — how a forecast like this actually gets used to plan real-time operations.
- Forecasting solar and wind output — the supply-side forecast this demand number gets matched against.
- Multivariate forecasting — forecasting demand across many regions or feeders together.
Researcher — Mathematics and papers.
Problem structure
Short-term load forecasting (STLF, typically hours to a few days ahead) is a well-studied time series regression problem with strong, well-understood exogenous drivers. A standard formulation:
L_t = f( calendar_t, weather_t, L_{t-1}, L_{t-24}, L_{t-168}, ... ) + epsilon_tL_{t-24} (same hour, previous day) and L_{t-168} (same hour, previous week) are conventional lag features, capturing daily and weekly seasonality directly, alongside explicit calendar and weather covariates.
Temperature response modelling
The load-temperature relationship is characteristically nonlinear and asymmetric — a U-shaped or "hockey stick" response in most climates, flat over a comfortable range, rising steeply for both heating below and cooling above it. This is commonly modelled with piecewise-linear or spline terms in temperature. It can also be captured implicitly by tree-based ensembles, as in the developer example, and gradient boosting methods — both standard choices in the influential Global Energy Forecasting Competition series (Hong et al., 2016).
Hierarchical and probabilistic forecasting
Grid operators forecast at multiple levels simultaneously — individual substations, regional zones, and system-wide totals — which must be kept mutually consistent (regional forecasts summing to the system total). This is the same hierarchical forecast reconciliation problem covered generally elsewhere on this platform, applied to load. Probabilistic load forecasts predict a full distribution, or specific quantiles, rather than a single point estimate. These directly inform reserve margin sizing. A grid operator holds reserve calibrated to a chosen high quantile of the demand distribution, not only the expected value, because underestimating (a shortfall) costs far more than overestimating (idle reserve capacity).
Weather forecast uncertainty propagation
Weather is itself a forecast, not a known quantity, so load forecast uncertainty compounds two sources. One is the load model's own residual uncertainty, given perfect weather knowledge. The other is the weather forecast's own uncertainty, which grows with lead time. Rigorous operational systems propagate ensemble weather forecasts through the load model to produce a load distribution reflecting both sources, rather than treating the weather forecast as a fixed, certain input.
Evaluation
Alongside RMSE and MAPE (mean absolute percentage error, common in industry reporting since it is unit-free and interpretable to non-technical stakeholders), pinball loss at multiple quantiles is standard for probabilistic submissions, as used in GEFCom. Peak-hour accuracy is typically reported and weighted separately from off-peak accuracy, since forecast error at the daily peak carries disproportionate operational and financial consequence.
Current state
Gradient-boosted trees and, increasingly, deep learning sequence models (LSTM and transformer-based) are both in active operational and research use. Published comparisons show mixed results, depending on data volume, feature richness, and forecast horizon — no single architecture dominates across all settings. Extreme, rare weather events remain the hardest well-documented case across essentially every published method, consistent with the fundamental data-sparsity problem: extreme events are, by construction, rare in any historical training record.
Key references
- Hong, T., Fan, S. (2016). Probabilistic Electric Load Forecasting: A Tutorial Review. International Journal of Forecasting.
- Hong, T. et al. (2016). Probabilistic Energy Forecasting: Global Energy Forecasting Competition 2014 and Beyond. International Journal of Forecasting.
- Hyndman, R. J., Fan, S. (2010). Density Forecasting for Long-Term Peak Electricity Demand. IEEE Transactions on Power Systems.
What to learn next
- Battery dispatch and grid decisions — the operational use of a forecast distribution like this one.
- Hierarchical forecast reconciliation — keeping multi-level forecasts consistent, applied here to substations and regions.
- Prediction intervals for regression — the general framework behind reserve-margin-relevant quantile forecasts.