Retail, Demand and Supply Chain
Making store, region and national forecasts add up
Store forecasts, region forecasts and a national forecast rarely sum to the same number when built separately, and a business cannot plan around numbers that disagree with themselves.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
When you forecast a store, a region and a whole country separately, the numbers usually do not add up. A business cannot plan around numbers that contradict each other.
Think about planning a big family wedding budget. Your mother estimates the catering cost. Your uncle estimates the decoration cost. Separately, your father sits down and estimates the total cost for the whole event. When everyone's individual estimates are added up, the total does not match what your father predicted alone. Nobody made a mistake. Two honest ways of estimating the same thing landed on different numbers.
Retailers hit this constantly. A store forecasts its own sales. A regional office forecasts the region. Head office forecasts the whole country. Add up every store's forecast, and it rarely equals the regional forecast. Add up every region, and it rarely equals the national one.
Why it exists
Different levels of a business often use different forecasting methods, because different information is available at each level.
A single store's forecaster knows about a local road closing next week. Head office does not track road closures store by store — it looks at broad national trends instead. Both are reasonable. Both produce a number. The numbers were never built to agree with each other.
This becomes a real problem the moment the business tries to act on both numbers at once. Say the national plan assumes 10,000 units will move, but the sum of every store's order plan is only 9,200. Someone ordered too little somewhere — and nobody can tell where, from the mismatched totals alone.
How it works
NATIONAL forecast: 1,120 units
(made independently, by one team)
Store A: 490 Store B: 290 Store C: 200 Store D: 140
(made independently, store by store, by a different team)
Sum of stores = 490+290+200+140 = 1,120... except it usually is NOT
exactly equal, because the two forecasts were built differently.Reconciliation — making numbers at different levels agree — fixes this after the fact. Two common approaches:
- Bottom-up. Trust the store-level forecasts. Add them up to get the national number. The national figure is now defined as their sum, so it always matches.
- Top-down. Trust the national forecast. Split it across stores using each store's normal historical share of total sales.
Neither is universally "more correct" — the choice depends on which level of the business has the more reliable information.
Where you have already seen it
- A company's quarterly results call. Headline national revenue and the sum of regional numbers reported earlier sometimes need a footnote explaining a small gap.
- Election exit polls, where state-by-state predictions and the national seat-count prediction are reconciled by pollsters before being published together.
- A school's overall pass percentage. It should match the average of every section's pass percentage, weighted by section size — and often needs checking when it does not.
Remember this
- Forecasts made separately at different levels of a hierarchy rarely add up on their own.
- Bottom-up trusts the detailed, local forecasts. Top-down trusts the broad, aggregate one.
- A business cannot commit to inventory plans built on numbers that contradict each other.
What to learn next
- Turning a forecast into an order quantity — what a reconciled forecast is actually used for.
- Retail demand data — the table these forecasts are built from.
- Multivariate forecasting — forecasting series that relate to each other.
Developer — Code and libraries.
Setup
pip install numpy pandasMinimal runnable code
Four stores, each forecast independently. A separate national forecast, built with a different method. We check whether they agree, then reconcile both ways.
import numpy as np
import pandas as pd
rng = np.random.default_rng(11)
stores = ["store_A", "store_B", "store_C", "store_D"]
n_days = 30
base_rates = {"store_A": 50, "store_B": 30, "store_C": 20, "store_D": 15}
history = pd.DataFrame({s: rng.poisson(base_rates[s], n_days) for s in stores})
history["national"] = history[stores].sum(axis=1)
print(history.tail(5))
print()
# The store team forecasts with a plain 7-day average, per store.
store_forecast = history[stores].tail(7).mean()
# The head-office team forecasts the national number on its own, with a
# trend-weighted average -- a perfectly reasonable, different method.
national_forecast_independent = history["national"].ewm(span=7).mean().iloc[-1]
print("independent store forecasts:")
print(store_forecast.round(1))
print("sum of store forecasts: ", round(store_forecast.sum(), 1))
print("independent national forecast:", round(national_forecast_independent, 1))
print("mismatch:", round(national_forecast_independent - store_forecast.sum(), 1))
print()
# Bottom-up reconciliation: the national number is defined as the sum.
bottom_up_national = store_forecast.sum()
print("bottom-up reconciled national:", round(bottom_up_national, 1), "(now consistent by construction)")
print()
# Top-down reconciliation: split the national forecast using historical share.
historical_share = history[stores].sum() / history[stores].sum().sum()
top_down_stores = national_forecast_independent * historical_share
print("historical share of each store:")
print(historical_share.round(3))
print("top-down reconciled store forecasts:")
print(top_down_stores.round(1))
print("sum of top-down store forecasts:", round(top_down_stores.sum(), 1), "(now consistent by construction)")store_A store_B store_C store_D national 25 38 39 20 11 108 26 49 21 18 18 106 27 55 33 22 17 127 28 36 35 19 11 101 29 60 23 14 15 112 independent store forecasts: store_A 48.9 store_B 29.3 store_C 19.6 store_D 14.7 dtype: float64 sum of store forecasts: 112.4 independent national forecast: 111.1 mismatch: -1.4 bottom-up reconciled national: 112.4 (now consistent by construction) historical share of each store: store_A 0.437 store_B 0.256 store_C 0.179 store_D 0.129 dtype: float64 top-down reconciled store forecasts: store_A 48.5 store_B 28.4 store_C 19.8 store_D 14.3 dtype: float64 sum of top-down store forecasts: 111.1 (now consistent by construction)
What actually happened
The store team used a plain 7-day average. The national team used ewm — an exponentially weighted moving average, which leans more heavily on the most recent days than older ones. Both are defensible choices. They disagree by 1.4 units, small here, but the same gap at real retail scale can mean thousands of units of inventory planned against two different numbers.
Bottom-up reconciliation is the simpler line of code: it redefines the national number as store_forecast.sum(). Nothing about the store forecasts changes.
Top-down reconciliation keeps the national number and works backward, using historical_share — each store's typical fraction of total sales — to split it. Note the store forecasts changed slightly from the original independent ones, pulled toward consistency with the trusted national figure.
Common mistakes
Reconciling once and never again. Store shares drift as a business grows unevenly — a new store opening or an old one closing changes historical_share outright. Recompute it on a schedule, not once at setup.
Using bottom-up when the top level is more reliable. If national data is clean and store-level data is noisy or has many missing days, forcing everything bottom-up propagates the noisiest signal upward.
Ignoring the middle layer. Real retailers usually have three or more levels — store, region, national. Reconciling store-to-national directly and separately reconciling region-to-national can silently reintroduce the same disagreement one level down.
Treating the reconciled number as automatically more accurate. Reconciliation guarantees the numbers are consistent with each other. It does not guarantee either forecast was correct to begin with.
Try it yourself
Add a fifth store to base_rates with a very different scale, like 5. Re-run and watch how much historical_share — and therefore every top-down store forecast — shifts because of one new series.
What to learn next
Researcher — Mathematics and papers.
The hierarchy as a linear constraint
A forecasting hierarchy can be written as a summing matrix S mapping bottom-level series to every aggregate:
y_t = S * b_tb_t— the vector of bottom-level (most disaggregated) series at timetS— a 0/1 summing matrix encoding which bottom series roll up into which aggregatey_t— every series in the hierarchy, bottom and aggregate together
Independently produced base forecasts y_hat_t for every node in the hierarchy generally do not satisfy y_hat_t = S * b_hat_t — they are said to be incoherent. Bottom-up and top-down, described above, are the two simplest ways to force coherence, and both are special cases of a more general framework.
Optimal reconciliation (MinT)
Wickramasuriya, Athanasopoulos and Hyndman (2019), Optimal Forecast Reconciliation for Hierarchical and Grouped Time Series Through Trace Minimization, JASA, generalise bottom-up and top-down into a single linear projection:
y_tilde_t = S * (S' * W_h^{-1} * S)^{-1} * S' * W_h^{-1} * y_hat_ty_hat_t— the vector of independently produced base forecasts at every levelW_h— the covariance matrix of the base forecast errors at forecast horizonhy_tilde_t— the reconciled, fully coherent forecast vector
This is the MinT (trace minimisation) estimator: among all linear reconciliation methods, it is the one that minimises the trace of the reconciled forecast error covariance, given W_h. Bottom-up and top-down correspond to specific, structured choices of a projection matrix in this same family; MinT instead estimates the projection from the base forecasts' actual error covariance.
Why the error covariance matters
W_h is rarely known exactly and must be estimated. Common practical choices, from simplest to most data-hungry: an identity matrix (giving ordinary least squares reconciliation, ignoring correlation entirely), a diagonal matrix of each series' own error variance, and a full shrinkage estimator of the error covariance across all series. The shrinkage estimator is the one MinT is usually paired with in production, since a naively estimated full covariance matrix at retail scale is typically singular or near-singular.
Grouped, not strictly hierarchical, structures
Real retail hierarchies are usually grouped, not a single tree — a product rolls up by category and, separately, by store rolls up by region, and both aggregations must hold simultaneously. S extends naturally to this case (Hyndman, Ahmed, Athanasopoulos and Shang, 2011), and MinT applies unchanged once S correctly encodes every aggregation path.
Cost and practical scale
Computing (S' * W_h^{-1} * S)^{-1} exactly is O(n^3) in the number of nodes n in the hierarchy, which becomes prohibitive at retail scale (n in the hundreds of thousands including SKU x store combinations). Production implementations, such as the hierarchicalforecast package in the Nixtla ecosystem, use sparse matrix structure and structural shortcuts specific to strictly hierarchical (tree) topologies to avoid the dense inverse.
Key references
- Hyndman, R., Ahmed, R., Athanasopoulos, G. & Shang, H. (2011). Optimal Combination Forecasts for Hierarchical Time Series. Computational Statistics & Data Analysis 55(9).
- Wickramasuriya, S., Athanasopoulos, G. & Hyndman, R. (2019). Optimal Forecast Reconciliation for Hierarchical and Grouped Time Series Through Trace Minimization. Journal of the American Statistical Association 114(526).