Evaluating a forecast
A forecast error only means something next to a baseline and averaged over many origins, and the metric you pick quietly decides which model wins.
- 17 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A forecast error is meaningless on its own. It only means something compared to a baseline, and measured over many different starting points.
The analogy you have already lived
A bus arrives five minutes late.
On a ten-minute ride to the next stop, five minutes late is awful. You could have walked. On a six-hour intercity journey, five minutes is nothing at all.
The number five minutes did not change. What changed is what you compared it against.
Forecast errors work the same way. "Our model is off by forty units" tells nobody anything on its own. They need to know what forty units means for that series. They also need to know how badly a lazy guess would have done.
The first question, always
How would the dumbest possible method have scored?
Two dumb methods are worth keeping around forever.
The naive forecast says tomorrow will be the same as today. One line of code.
The seasonal naive forecast says next Saturday will be the same as last Saturday. Also one line.
If your careful model does not beat both, you do not have a model. You have a complicated way of being wrong, and you should ship the one-liner.
This is not a formality. Serious methods lose to these baselines all the time, on real data, in published work.
Three ways of measuring error
Average error in your own units. Add up how far off you were each day, ignoring direction, and take the average. If your unit is packets of milk, the answer is in packets of milk. Easy to explain to anybody.
Punishing big misses harder. Square the errors before averaging, then undo the squaring at the end. Being off by ten once hurts more than being off by two five times. Use it when a single big miss is genuinely worse than several small ones.
Percentage error. Express each miss as a share of the true value. This is the most requested and the most dangerous, and it gets its own section below.
Why percentage error will betray you
Percentage error looks like the friendly one. It has no units, so you can compare a bread series against an electricity series.
It has two flaws, and both are serious.
It explodes near zero. A shop that sells two packets on a quiet day, forecast at twelve, is off by ten packets. That is nothing in real terms. As a percentage it is five hundred percent, and it will drag your whole average up on its own. The developer section shows exactly that happening.
It is lopsided. Under-forecasting can never look worse than one hundred percent, because the true value is the ceiling. Over-forecasting has no ceiling at all. So percentage error quietly rewards a model that guesses low.
Some series sit near zero: daily sales of a slow item, hourly counts at night. For those, use a different measure.
The one that solves it
There is a measure built for this. Divide your model's average error by the average error of the seasonal naive baseline on the same data.
score below 1 -> better than the lazy baseline
score of 1 -> exactly as good as the lazy baseline
score above 1 -> worse than doing nothing cleverThe comparison is built into the number. You can put it in an email with no other context and it means something. It also handles zeros without trouble.
Never trust one split
Here is the mistake that survives even in careful teams.
You hold out the last month, score your model, and report the number. That number came from one particular month. A different month would have given a different answer, sometimes a very different one.
The fix is to move the cut-off and repeat.
train ................ | test |
train .................... | test |
train ........................ | test |
train ............................ | test |
each row: train on everything to the left, score on the block to the rightThen report the average across all of them, and also how much they varied. The developer section runs six of these. The winning model's error ranged from 29 to 58 across them. Reporting any single one would have been misleading.
What is honestly hard here
The metric you choose decides which model wins, and there is no neutral choice.
Average error favours a model that predicts the middle of the range. Squared error favours one that avoids big misses. Percentage error favours one that guesses low. These are not rounding differences; they change the ranking of your candidates.
So the honest approach is to choose the metric from the decision the forecast will feed. Suppose running out of stock costs you a customer, while holding stock costs you a little cash. Under-forecasting is then far worse than over-forecasting. No symmetric metric describes that. You need one that charges more for one direction than the other.
Pick the metric before you see the results. Choosing it afterwards, when you can see which one makes your model look best, is a way of lying to yourself.
Remember this
- Always report a naive baseline alongside your model.
- Percentage error breaks near zero and rewards guessing low.
- Score across many origins, and report the spread, not one number.
What to learn next
- Anomaly detection — using forecast errors as an alarm rather than a score.
- Monitoring and model drift — keeping these numbers honest after deployment.
- Model evaluation — the general case, so you can see exactly what time adds.
Developer — Code and libraries.
Setup
pip install numpy pandas statsmodelsWatch percentage error break
One quiet day is enough to destroy a MAPE.
import numpy as np
actual = np.array([100., 120., 90., 110., 2., 105.]) # one very quiet day
forecast = np.array([105., 110., 95., 100., 12., 100.])
err = forecast - actual
print("MAE :", round(np.abs(err).mean(), 2))
print("RMSE:", round(np.sqrt((err ** 2).mean()), 2))
print("MAPE:", round((np.abs(err) / actual).mean() * 100, 1), "%")
print("sMAPE:", round((200 * np.abs(err) / (np.abs(actual) + np.abs(forecast))).mean(), 1), "%")
print()
print("per-day percentage error:")
for a, f in zip(actual, forecast):
print(f" actual {a:6.1f} forecast {f:6.1f} -> {abs(f-a)/a*100:8.1f} %")MAE : 7.5 RMSE: 7.91 MAPE: 88.8 % sMAPE: 29.4 % per-day percentage error: actual 100.0 forecast 105.0 -> 5.0 % actual 120.0 forecast 110.0 -> 8.3 % actual 90.0 forecast 95.0 -> 5.6 % actual 110.0 forecast 100.0 -> 9.1 % actual 2.0 forecast 12.0 -> 500.0 % actual 105.0 forecast 100.0 -> 4.8 %
The MAE is 7.5 units. On a series running around 100, that is a good forecast.
The MAPE is 88.8 percent, which reads like catastrophe. Five of the six days are within 10 percent. The sixth day was off by ten units on a true value of two, giving 500 percent, and that single row carries the whole average.
Nothing is wrong with the forecast. The metric is wrong for this data.
MAPE is undefined when any actual is exactly zero. You get inf or a RuntimeWarning: divide by zero. Anyone who has forecast retail demand has met this.
The lopsidedness, in a table
actual = 100.0
print("forecast MAE MAPE sMAPE")
for forecast in (0.0, 50.0, 100.0, 150.0, 200.0, 400.0):
mae = abs(forecast - actual)
mape = mae / actual * 100
smape = 200 * mae / (abs(actual) + abs(forecast))
print(f"{forecast:8.0f} {mae:5.0f} {mape:6.1f}% {smape:6.1f}%")forecast MAE MAPE sMAPE
0 100 100.0% 200.0%
50 50 50.0% 66.7%
100 0 0.0% 0.0%
150 50 50.0% 40.0%
200 100 100.0% 66.7%
400 300 300.0% 120.0%MAPE's ceiling for under-forecasting is 100 percent. Predict zero when the truth is 100 and you are charged 100 percent — the worst possible. Predict 400 and you are charged 300 percent. A model tuned to minimise MAPE learns to aim low.
sMAPE has the opposite tilt. Forecasting 50 costs 66.7 percent while forecasting 150 costs 40.0 percent, for the same absolute miss. It was introduced to fix MAPE's asymmetry and introduced a different one. It is also bounded at 200 percent, which hides catastrophic errors.
This is why MASE exists.
A rolling-origin backtest, which is what you should actually run
import numpy as np
import pandas as pd
from statsmodels.tsa.holtwinters import ExponentialSmoothing
rng = np.random.default_rng(17)
n = 800
days = pd.date_range("2023-01-01", periods=n, freq="D")
y = pd.Series((900 + np.linspace(0, 200, n)
+ np.where(days.dayofweek >= 5, 200, 0)
+ rng.normal(0, 60, n)).round(), index=days).asfreq("D")
H, M = 14, 7 # forecast 14 days ahead; weekly season
def mase(actual, pred, train):
scale = np.abs(train.to_numpy()[M:] - train.to_numpy()[:-M]).mean()
return np.abs(actual - pred).mean() / scale
rows = []
for origin in range(600, 600 + 6 * H, H):
train, test = y[:origin], y[origin:origin + H]
snaive = np.resize(train.to_numpy()[-M:], H)
hw = ExponentialSmoothing(train, trend="add", seasonal="add", seasonal_periods=M,
initialization_method="estimated").fit().forecast(H).to_numpy()
rows.append({
"origin": str(y.index[origin].date()),
"snaive_MAE": round(float(np.abs(test - snaive).mean()), 1),
"hw_MAE": round(float(np.abs(test - hw).mean()), 1),
"snaive_MASE": round(float(mase(test, snaive, train)), 2),
"hw_MASE": round(float(mase(test, hw, train)), 2),
})
df = pd.DataFrame(rows)
print(df.to_string(index=False))
print("\nmean over folds: snaive MAE", round(df.snaive_MAE.mean(), 1),
" hw MAE", round(df.hw_MAE.mean(), 1))
print("spread over folds: snaive", round(df.snaive_MAE.std(), 1),
" hw", round(df.hw_MAE.std(), 1))
print("folds where Holt-Winters won:", int((df.hw_MAE < df.snaive_MAE).sum()), "of", len(df))origin snaive_MAE hw_MAE snaive_MASE hw_MASE 2024-08-23 61.9 37.9 0.92 0.56 2024-09-06 48.9 29.0 0.73 0.43 2024-09-20 47.2 36.8 0.71 0.55 2024-10-04 56.9 40.7 0.86 0.62 2024-10-18 59.8 41.5 0.91 0.63 2024-11-01 79.9 57.7 1.22 0.88 mean over folds: snaive MAE 59.1 hw MAE 40.6 spread over folds: snaive 11.8 hw 9.5 folds where Holt-Winters won: 6 of 6
Look at the hw_MAE column before anything else. It runs from 29.0 to 57.7. If you had held out only the fortnight starting 6 September, you would have reported 29.0. Held out only the fortnight starting 1 November, you would have reported 57.7. Same data, same model, double the number.
Six wins out of six is the result worth reporting, more than the mean. It says the improvement is consistent, not an average rescued by one lucky fold. With six paired folds, a sign test gives a two-sided p-value of 2/64 = 0.031, which is real evidence on a very small sample.
MASE reads without context. hw_MASE of 0.56 means the model made about 56 percent of the error the in-sample seasonal naive made. snaive_MASE of 1.22 in the last fold means the out-of-sample seasonal naive did worse there than its own in-sample scale, which is a signal that fold was unusually hard.
The scale in mase uses training data only. That is deliberate: the denominator must be knowable at forecast time, so the metric cannot be gamed by the test period.
The metrics, and when to reach for each
| Metric | Formula in words | Use it when | Avoid it when |
|---|---|---|---|
| MAE | mean absolute error | You want an answer in real units | You must compare across series of different sizes |
| RMSE | root of mean squared error | Big misses genuinely hurt more | The series has wild outliers you do not care about |
| MAPE | mean of absolute error over actual | Values are comfortably above zero and stakeholders demand percentages | Any actual is near or at zero |
| sMAPE | symmetric variant | You inherited a report that uses it | You have any other choice |
| MASE | MAE over the in-sample naive MAE | Comparing across series, or reporting one honest number | Nothing much — this is the sound default |
| Pinball | asymmetric, per quantile | The decision cares about direction, or you forecast a range | You need one number for a point forecast |
Pinball loss deserves a look if your forecast feeds a stock decision. It charges a different price for over-forecasting and under-forecasting, and minimising it produces a chosen quantile rather than the mean. That is usually what an inventory system actually needs.
Common mistakes
One holdout period. The single most common evaluation error in industry. Move the origin.
No baseline. A model with an MAE of 40 sounds fine until the one-line seasonal naive scores 38.
Averaging MAPE across series. A low-volume series produces enormous percentages and dominates the average. Use MASE.
Comparing MAE across differently scaled series. An MAE of 200 on electricity units and 3 on bread packets cannot be averaged.
Tuning on the backtest and reporting the backtest. If you chose a window size, a smoothing parameter or a feature set by looking at these folds, the number is optimistically biased. Hold out a final block you never look at until the end.
Comparing a one-step model against a fourteen-step model. Every method looks strong at horizon one. Fix the horizon that matches your decision and compare there.
Ignoring the interval. A point forecast with no range is an unfinished forecast. Check coverage: what fraction of actuals fell inside your 80 percent band? It should be near 80 percent.
Try it yourself
Change H from 14 to 1 in the backtest and rerun.
Predict what happens to both models before you run it. At a one-day horizon the seasonal naive has much less time to drift, so the gap should narrow. Then set H to 28 and watch it widen again. The lesson: a model's ranking is a function of the horizon, so quote the horizon whenever you quote a score.
What to learn next
- Anomaly detection — using forecast errors as an alarm rather than a score.
- Monitoring and model drift — keeping these numbers honest after deployment.
- Model evaluation — the general case, so you can see exactly what time adds.
Researcher — Mathematics and papers.
Notation
Let $y_t$ be the actual, $\hat{y}_t$ the forecast, $e_t = y_t - \hat{y}_t$ the error, $n$ the number of scored points, and $m$ the seasonal period.
$$ \mathrm{MAE} = \frac{1}{n}\sum |e_t|, \qquad \mathrm{RMSE} = \sqrt{\frac{1}{n}\sum e_t^2}, \qquad \mathrm{MAPE} = \frac{100}{n}\sum \frac{|e_t|}{|y_t|} $$
$$ \mathrm{MASE} = \frac{\frac{1}{n}\sum_{t}|e_t|}{\frac{1}{T-m}\sum_{t=m+1}^{T}|y_t - y_{t-m}|} $$
The denominator is the in-sample mean absolute error of the seasonal naive method over the training set of length $T$. Hyndman and Koehler (2006) proposed it precisely because it is scale-free, defined when $y_t = 0$, and has a finite mean under weaker conditions than percentage errors.
Which functional each loss elicits
This is the part that decides model rankings and is most often skipped.
A scoring function $S$ is consistent for a statistical functional $\mathcal{T}$ if $\mathbb{E}[S(\mathcal{T}(F), Y)] \le \mathbb{E}[S(x, Y)]$ for all $x$ and all $F$ in the class. Gneiting (2011), Making and evaluating point forecasts, JASA 106(494), sets this out:
| Loss | Optimal point forecast |
|---|---|
| Squared error | conditional mean |
| Absolute error | conditional median |
| Pinball at level $\tau$ | conditional $\tau$-quantile |
| MAPE-type | a quantity below the median, biased low |
| Relative error, $\mathbb{E}[ | e |
The consequence is concrete. On a right-skewed demand distribution the mean exceeds the median, so an MAE-optimal and an RMSE-optimal forecaster produce systematically different numbers from the same predictive distribution. Ranking models by MAE when the decision needs a mean is an error of specification, not of taste.
MAPE is not consistent for any standard functional. Its minimiser depends on the shape of the distribution in a way that has no clean interpretation. Kolassa (2020), Why the "best" point forecast depends on the error or accuracy measure, IJF 36(1), works through examples where MAPE-optimal forecasts for count data are absurd — for low-count Poisson demand the MAPE-optimal point forecast can be zero.
Proper scoring rules for distributions
Point metrics discard the uncertainty. For a predictive CDF $F$ and observation $y$, the continuous ranked probability score is
$$ \mathrm{CRPS}(F, y) = \int_{-\infty}^{\infty} \left(F(z) - \mathbf{1}{z \ge y}\right)^2 dz $$
CRPS is strictly proper: its expectation is uniquely minimised by the true predictive distribution. It reduces to MAE when $F$ is a point mass, so it is directly comparable with point metrics. It also has a closed form for Gaussian and for empirical-sample predictions, making it cheap on simulated forecast paths.
Gneiting and Raftery (2007), Strictly proper scoring rules, prediction, and estimation, JASA 102(477), is the reference. Gneiting, Balabdaoui and Raftery (2007) add the operational principle: maximise sharpness subject to calibration. Check calibration with a PIT histogram — under a correct predictive distribution the values $F_t(y_t)$ are uniform on $[0,1]$; a U-shape means intervals are too narrow, a hump means too wide.
Comparing two forecasters
Do not compare means of loss without a test. The Diebold-Mariano statistic (1995) tests $H_0: \mathbb{E}[d_t]=0$ for the loss differential $d_t = L(e_{1t}) - L(e_{2t})$:
$$ \mathrm{DM} = \frac{\bar{d}}{\sqrt{\widehat{\mathrm{LRV}}(d)/n}} $$
where $\widehat{\mathrm{LRV}}$ is a heteroskedasticity- and autocorrelation-consistent long-run variance estimate — necessary because $h$-step errors are MA($h-1$) correlated by construction. Harvey, Leybourne and Newbold (1997) give a small-sample correction and recommend a $t_{n-1}$ reference distribution.
Two constraints are routinely violated: DM is not valid for nested models (use Clark and West, 2007), and it is not valid across many series without a multiple-comparison adjustment. For ranking many methods over many series, the Friedman test followed by the Nemenyi post-hoc procedure (Demšar, 2006) is the standard, and it is what the M4 and M5 analyses used.
Backtest design
The estimator being computed is $$ \widehat{\mathrm{err}} = \frac{1}{|\mathcal{O}|}\sum_{t \in \mathcal{O}} L!\left(y_{t+h}, \hat{y}_{t+h|t}\right) $$ over a set of origins $\mathcal{O}$. Three design decisions determine its bias and variance:
- Expanding versus sliding window. Expanding uses all history and is preferred under stability; sliding adapts to regime change with higher variance.
- Origin spacing. Overlapping horizons make the fold errors strongly dependent, so the naive standard error across folds understates uncertainty. Spacing origins by at least $h$ (as the code above does) reduces but does not remove this.
- Purge and embargo when features have lookback windows or labels span time. See López de Prado (2018), chapter 7.
Bergmeir, Hyndman and Koo (2018), A note on the validity of cross-validation for evaluating autoregressive time series prediction, establishes a useful and counterintuitive result: for purely autoregressive models with uncorrelated residuals, standard $k$-fold cross-validation is asymptotically valid and more efficient than a single holdout. The validity fails as soon as residuals remain correlated, which is the common case, so rolling origin remains the safe default.
Reading
- Hyndman and Koehler (2006), Another look at measures of forecast accuracy, IJF 22(4) — the MASE paper.
- Gneiting (2011), Making and evaluating point forecasts, JASA 106(494).
- Gneiting and Raftery (2007), Strictly proper scoring rules, prediction, and estimation, JASA 102(477).
- Diebold and Mariano (1995), Comparing predictive accuracy, JBES 13(3).
- Kolassa (2020), Why the "best" point forecast depends on the error or accuracy measure, IJF 36(1).
- Bergmeir, Hyndman and Koo (2018), A note on the validity of cross-validation for evaluating autoregressive time series prediction, CSDA 120.
What to learn next
- Anomaly detection — using forecast errors as an alarm rather than a score.
- Monitoring and model drift — keeping these numbers honest after deployment.
- Model evaluation — the general case, so you can see exactly what time adds.