Exponential smoothing
Exponential smoothing keeps one running estimate and nudges it toward each new observation, giving recent data more weight without ever forgetting the past completely.
- 15 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Exponential smoothing keeps one running estimate and nudges it a little toward every new observation. Recent days matter most. Old days fade away instead of falling off a cliff.
The analogy you have already lived
Think about how you rate the restaurant down the road.
Last week's meal was disappointing, and that is most of your current opinion. The good meal a month ago still counts, a little. The excellent meal three years ago barely registers at all.
You never sat down and decided to remember exactly the last ten meals and forget the eleventh. Your memory fades smoothly. Recent experience carries the most weight, and everything older keeps shrinking without ever hitting zero.
That fading is exponential smoothing. You already run one in your head about almost everything.
Why it beats a plain moving average
A moving average has a hard edge. Take a window of seven days. The day that happened six days ago counts fully. The day before that counts nothing at all.
That edge causes an odd effect. A big spike sits inside your average for exactly seven days, then vanishes overnight. Your smooth line jumps for no reason connected to today.
Exponential smoothing has no edge. Everything counts, with weight fading as it gets older. Nothing ever drops out suddenly.
It is also cheap. A moving average needs the last seven values stored. Exponential smoothing needs one number — the current estimate. That mattered enormously when these methods were invented for warehouse inventory in the 1950s. It still matters when you track a million products.
How it works
There is one dial, called alpha, set between zero and one. It answers one question: how much do you trust today's new observation?
your new estimate -> take your old estimate
-> look at today's reading
-> move a step from the old one toward today's
how big is the step? alpha decides.
alpha near zero -> a tiny step. Barely listens to today.
Very smooth, very slow to react.
alpha in the middle -> half today, half everything before.
alpha near one -> a huge step. Mostly today.
Twitchy, follows every bump.Every day you take one step from your old estimate toward the new reading. A small alpha means a tiny step. A large alpha means a big one.
That single rule, applied over and over, produces the fading memory. Yesterday's estimate already contained the day before, which already contained the day before that.
Why the memory fades in a curve
Unwind the rule a few steps and a pattern appears. With alpha set to a third or so:
today gets about 33 out of 100 of the weight
yesterday gets about 22
two days ago gets about 15
three days ago gets about 10
... and so on, always shrinking, never reaching zeroAdd all of them up and you get one hundred. Nothing is lost, and nothing is thrown away abruptly.
The three levels of the method
Plain exponential smoothing tracks one thing: the current level, meaning roughly where the series sits now. That is enough for a flat, wobbling series.
It cannot handle a series that keeps climbing. So a second running estimate was added for the trend, meaning how much the level moves each step. That version is named after Charles Holt.
Then a third estimate was added for the repeating seasonal pattern. That version is called Holt-Winters, and it is still a genuinely strong forecaster today, decades later.
level -> where are we now?
trend -> which way are we heading, and how fast?
seasonal -> what does a typical Saturday look like?Each part gets its own dial, and each dial is learned from your data rather than guessed.
Where you have already seen it
- Warehouse stock systems decide reorder quantities this way, for millions of items.
- Your phone's battery-time-remaining estimate smooths recent drain rather than reacting to every second.
- Network speed meters show a smoothed rate, which is why the number glides instead of flickering.
- Weather "feels like" trends and many dashboard sparklines are exponentially smoothed.
What is honestly hard here
Choosing alpha by eye is a trap. A line that looks pleasing on a chart is not the same as a line that forecasts well. Let the software fit the dials by measuring forecast error, and resist the urge to overrule it.
The bigger honest point is this: exponential smoothing assumes tomorrow resembles a gently updated version of today. It cannot know that a festival is coming, that a competitor opened, or that a road closed. It reacts to change; it never anticipates change.
Remember this
- Exponential smoothing nudges one running estimate toward each new value.
- Older data fades smoothly instead of dropping off an edge.
- Holt-Winters adds a trend part and a seasonal part, and remains hard to beat.
What to learn next
- ARIMA — the other classical family, built on autocorrelation rather than smoothing.
- Evaluating a forecast — how to compare these fairly across many origins.
- Prophet — a modern tool that handles holidays and moving festivals directly.
Developer — Code and libraries.
Setup
pip install numpy pandas statsmodelsBuild it yourself first, in four lines
The whole method is one loop. Writing it once removes all the mystery from the library version.
import numpy as np
import pandas as pd
sales = pd.Series([120, 118, 125, 119, 121, 160, 122, 124],
index=pd.date_range("2025-05-01", periods=8, freq="D"), name="sales")
def ses(series, alpha):
level = series.iloc[0]
out = [level]
for y in series.iloc[1:]:
level = alpha * y + (1 - alpha) * level # nudge the level toward the newest reading
out.append(level)
return pd.Series(out, index=series.index)
print(pd.DataFrame({
"sales": sales,
"alpha=0.2": ses(sales, 0.2).round(2),
"alpha=0.8": ses(sales, 0.8).round(2),
"pandas ewm 0.2": sales.ewm(alpha=0.2, adjust=False).mean().round(2),
}))
print("\nhow much each past day counts, with alpha = 0.3:")
for lag in range(6):
print(f" {lag} days ago: {0.3 * 0.7 ** lag:.4f}")
print(" everything older:", round(0.7 ** 6, 4))sales alpha=0.2 alpha=0.8 pandas ewm 0.2 2025-05-01 120 120.00 120.00 120.00 2025-05-02 118 119.60 118.40 119.60 2025-05-03 125 120.68 123.68 120.68 2025-05-04 119 120.34 119.94 120.34 2025-05-05 121 120.48 120.79 120.48 2025-05-06 160 128.38 152.16 128.38 2025-05-07 122 127.10 128.03 127.10 2025-05-08 124 126.48 124.81 126.48 how much each past day counts, with alpha = 0.3: 0 days ago: 0.3000 1 days ago: 0.2100 2 days ago: 0.1470 3 days ago: 0.1029 4 days ago: 0.0720 5 days ago: 0.0504 everything older: 0.1176
Three things to take from that table.
The hand-written version matches pandas to the last decimal. adjust=False is the setting that makes ewm use the recursive form written above. The default adjust=True uses a different normalisation for the early rows, and the two disagree at the start of the series. Set it explicitly or you will chase a phantom bug.
Watch the spike on May 6. Sales jumped to 160 for one day. With alpha=0.2 the estimate moved to 128.38. With alpha=0.8 it lurched to 152.16 and then had to climb back down. Small alpha shrugs off a one-off; large alpha believes it.
The weights are a geometric sequence. Each lag gets 0.7 times the previous one. They sum to one, and everything older than five days still carries about 12 percent of the total. Nothing is discarded.
The three variants, scored honestly
Now the comparison that matters. Fit on the first 120 days, forecast the next 60 in one shot, and score against baselines computed from the same 120 days.
import numpy as np
import pandas as pd
from statsmodels.tsa.holtwinters import SimpleExpSmoothing, ExponentialSmoothing
rng = np.random.default_rng(21)
n, cut = 180, 120
days = pd.date_range("2025-01-01", periods=n, freq="D")
weekend = np.where(days.dayofweek >= 5, 150, 0)
visits = pd.Series((np.linspace(400, 560, n) + weekend + rng.normal(0, 30, n)).round(),
index=days, name="visits").asfreq("D")
train, test = visits[:cut], visits[cut:]
h = len(test)
def mae(pred):
return float(np.mean(np.abs(test.to_numpy() - np.asarray(pred))))
last_week = train.to_numpy()[-7:]
print("repeat the last value MAE:", round(mae(np.full(h, train.iloc[-1])), 1))
print("repeat the last week MAE:", round(mae(np.resize(last_week, h)), 1))
ses = SimpleExpSmoothing(train, initialization_method="estimated").fit()
print("simple exp smoothing MAE:", round(mae(ses.forecast(h)), 1))
hw = ExponentialSmoothing(train, trend="add", seasonal="add", seasonal_periods=7,
initialization_method="estimated").fit()
print("Holt-Winters (trend+week) MAE:", round(mae(hw.forecast(h)), 1))
print()
print("alpha (level) :", round(hw.params["smoothing_level"], 3))
print("beta (trend) :", round(hw.params["smoothing_trend"], 3))
print("gamma (season):", round(hw.params["smoothing_seasonal"], 3))repeat the last value MAE: 75.0 repeat the last week MAE: 39.9 simple exp smoothing MAE: 61.4 Holt-Winters (trend+week) MAE: 22.5 alpha (level) : 0.0 beta (trend) : 0.0 gamma (season): 0.0
Holt-Winters roughly halved the seasonal naive's error, from 39.9 to 22.5. The trend part keeps the forecast climbing over sixty days, and the seasonal part restores the weekend bump that a moving average would have flattened.
Simple exponential smoothing scored 61.4, worse than repeating the last week. That is not a bug. Simple exponential smoothing forecasts a flat line at the current level, which is hopeless against a series with both a trend and a weekly pattern. Match the model to the structure you actually have.
All three smoothing dials came out at zero, and that is meaningful. The optimiser searched and concluded that no updating at all beats any amount of updating. This series is a fixed straight trend plus a fixed weekly pattern plus pure noise, so the best strategy is to estimate that structure once and never chase a wobble. On real data you will see non-zero values; a fitted alpha near zero is a signal that your series is more stable than you assumed, and a fitted alpha near one is a signal that it is close to a random walk.
Choosing between the variants
| Your data | Model | Code |
|---|---|---|
| Wobbles around a level | Simple | SimpleExpSmoothing(y) |
| Has a trend, no repeating pattern | Holt | ExponentialSmoothing(y, trend="add") |
| Trend and a repeating pattern | Holt-Winters | ExponentialSmoothing(y, trend="add", seasonal="add", seasonal_periods=m) |
| Pattern grows with the level | Multiplicative | seasonal="mul" |
| Trend should flatten out far ahead | Damped | trend="add", damped_trend=True |
The damped option deserves a word. An undamped trend extrapolates in a straight line forever, which produces absurd numbers at long horizons. Damping pulls the slope toward zero as the horizon grows. Gardner and McKenzie (1985) introduced it, and empirical competitions have found damped trends beat undamped ones often enough that it is a sensible default for horizons beyond a few steps.
Common mistakes
Leaving adjust at its default in ewm. adjust=True and adjust=False produce different early values. If you compare a pandas result against a hand-written loop or another library, set adjust=False.
Setting seasonal_periods from the calendar instead of the data. For daily data with a weekly rhythm it is 7, not 365. Getting this wrong is the single most common Holt-Winters failure.
Fitting on too little data. Holt-Winters needs at least two full seasonal cycles to estimate the seasonal indices, and statsmodels will raise if you give it fewer. Two cycles is the bare minimum; three or four gives a usable estimate.
Using seasonal="mul" on data containing zeros. Multiplicative seasonality is undefined at zero and produces nan or wild values. Use additive, or add a small constant and remember to remove it.
Forecasting far past what the model can know. hw.forecast(365) from one year of daily data will happily return 365 numbers. The trend extrapolates linearly and the confidence interval widens without bound. Report the interval, or the point forecast will be believed.
Reusing a fitted model as new data arrives. Exponential smoothing is a recursive filter and can be updated cheaply, but statsmodels expects a refit. For a live system, refit on a schedule and monitor whether the fitted alpha drifts.
Try it yourself
Change seasonal_periods=7 to seasonal_periods=30 in the last script and rerun.
Predict what happens to the MAE first. Thirty days is not the true cycle of this data, so the model will try to estimate thirty seasonal indices from four noisy repetitions. Then set it back to 7 and change trend="add" to damped_trend=True alongside it, and see whether damping helps or hurts over a sixty-day horizon.
What to learn next
- ARIMA — the other classical family, built on autocorrelation rather than smoothing.
- Evaluating a forecast — how to compare these fairly across many origins.
- Prophet — a modern tool that handles holidays and moving festivals directly.
Researcher — Mathematics and papers.
The recursions
Simple exponential smoothing (SES). Component form:
$$ \hat{y}_{t+h|t} = \ell_t, \qquad \ell_t = \alpha y_t + (1-\alpha)\ell_{t-1} $$
- $\ell_t$ — the level at time $t$.
- $\alpha \in [0,1]$ — the smoothing parameter for the level.
- $h$ — the forecast horizon. The forecast function is flat in $h$.
Expanding gives the weighted-average form $\hat{y}{t+1|t} = \sum{j=0}^{t-1} \alpha(1-\alpha)^j y_{t-j} + (1-\alpha)^t \ell_0$, showing geometrically decaying weights and the persistent influence of the initial state $\ell_0$.
Holt's linear trend.
$$ \ell_t = \alpha y_t + (1-\alpha)(\ell_{t-1} + b_{t-1}), \qquad b_t = \beta^(\ell_t - \ell_{t-1}) + (1-\beta^)b_{t-1} $$ $$ \hat{y}_{t+h|t} = \ell_t + h\,b_t $$
- $b_t$ — the slope estimate at time $t$.
- $\beta^* \in [0,1]$ — the smoothing parameter for the slope.
Damped trend (Gardner and McKenzie, 1985) replaces $h$ by $\sum_{i=1}^{h}\phi^i$ and multiplies $b_{t-1}$ by $\phi$ in both recursions, with $\phi \in (0,1)$. As $h \to \infty$ the forecast converges to $\ell_t + \phi b_t/(1-\phi)$ rather than diverging.
Holt-Winters additive, with period $m$:
$$ \ell_t = \alpha(y_t - s_{t-m}) + (1-\alpha)(\ell_{t-1}+b_{t-1}) $$ $$ b_t = \beta^(\ell_t - \ell_{t-1}) + (1-\beta^)b_{t-1} $$ $$ s_t = \gamma(y_t - \ell_{t-1} - b_{t-1}) + (1-\gamma)s_{t-m} $$ $$ \hat{y}_{t+h|t} = \ell_t + h b_t + s_{t-m+h_m^+}, \qquad h_m^+ = \lfloor (h-1) \bmod m \rfloor + 1 $$
- $s_t$ — the seasonal index for the seasonal position of time $t$.
- $\gamma \in [0, 1-\alpha]$ — the seasonal smoothing parameter. The upper bound comes from the state-space stability region, not from the range of a probability.
The multiplicative form replaces the subtractions and additions involving $s$ with divisions and multiplications.
The ETS state-space framework
Hyndman, Koehler, Snyder and Grose (2002) put exponential smoothing on a statistical footing by writing each method as an innovations state-space model
$$ y_t = w(\mathbf{x}{t-1}) + r(\mathbf{x}{t-1})\,\varepsilon_t, \qquad \mathbf{x}t = f(\mathbf{x}{t-1}) + g(\mathbf{x}_{t-1})\,\varepsilon_t $$
with a single source of error $\varepsilon_t \sim \mathrm{N}(0,\sigma^2)$. The taxonomy is ETS(Error, Trend, Seasonal), each slot taking values from {A, M, N} plus damped variants — thirty combinations, of which fifteen are numerically stable in practice.
This buys three things the ad hoc recursions never had:
- A likelihood. Parameters and initial states are estimated jointly by maximum likelihood, and models are compared by AICc.
- Prediction intervals. Analytic variance expressions exist for most members; for the rest, simulation from the innovations model works.
- Automatic model selection.
AutoETSinstatsforecast, orets()in R's forecast package, searches the space by AICc.
Additive-error ETS(A,N,N) coincides with ARIMA(0,1,1); ETS(A,A,N) with ARIMA(0,2,2); ETS(A,A,A) with a restricted seasonal ARIMA. The multiplicative-error and multiplicative-seasonal members have no ARIMA equivalent, and non-linear members have no finite-order linear representation at all. The classes overlap; neither contains the other.
Parameter space and initialisation
The naive constraint $\alpha, \beta^, \gamma \in [0,1]$ is not the correct one. The admissible region is defined by requiring the roots of the model's characteristic equation to lie inside the unit circle, so that the weight on past observations decays. For ETS(A,A,N) the admissible region is $0 < \alpha < 2$, $0 < \beta^ < 4/\alpha - 2$, which extends well beyond the unit square. Hyndman et al. (2008), chapter 10, derives these.
Initialisation matters more than most practitioners expect on short series. initialization_method="estimated" treats $\ell_0, b_0, s_{-m+1:0}$ as free parameters in the likelihood; "heuristic" uses a decomposition-based starting guess. On series shorter than three or four seasonal cycles the two can differ materially.
Empirical standing
Exponential smoothing is not a historical curiosity. In M3 (Makridakis and Hibon, 2000) damped-trend exponential smoothing was among the top performers across 3003 series. In M4 (2018) the winning entry, Smyl's ES-RNN, was a hybrid that used Holt-Winters-style level and seasonality updates with a shared recurrent network learning the residual dynamics across series. In M5 (2020), on hierarchical retail data, simple exponentially-smoothed baselines were beaten mainly by gradient-boosted trees over engineered features — an approach covered in feature engineering for time series.
The recurring empirical finding across all of them: method complexity does not monotonically improve accuracy, and combinations of simple methods are unusually strong.
Reading
- Brown (1959), Statistical Forecasting for Inventory Control — the origin, motivated by naval inventory.
- Holt (1957/2004), Forecasting seasonals and trends by exponentially weighted moving averages, reprinted in IJF 20(1).
- Winters (1960), Forecasting sales by exponentially weighted moving averages, Management Science 6(3).
- Gardner and McKenzie (1985), Forecasting trends in time series, Management Science 31(10).
- Hyndman, Koehler, Ord and Snyder (2008), Forecasting with Exponential Smoothing: The State Space Approach, Springer — the definitive treatment.
- Hyndman and Athanasopoulos, Forecasting: Principles and Practice, 3rd ed., chapter 8 — otexts.com/fpp3/expsmooth.html.
What to learn next
- ARIMA — the other classical family, built on autocorrelation rather than smoothing.
- Evaluating a forecast — how to compare these fairly across many origins.
- Prophet — a modern tool that handles holidays and moving festivals directly.