Time Series and Forecasting

Moving averages

A moving average smooths a bumpy series by averaging a sliding window, and the window you choose decides how much noise you remove and how far behind you fall.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. The window size is the whole decision
  6. The trap that ruins forecasts
  7. Where you have already seen it
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A moving average replaces each value with the average of it and its neighbours. The bumps flatten out, and the underlying shape becomes visible.

The analogy you have already lived

Picture the shopkeeper at the corner deciding how many milk packets to order tomorrow.

He does not look only at today. Today a wedding party bought forty packets, and tomorrow nobody will. He does not look at last year either, because the colony has grown since then.

He thinks about roughly the last week and averages it in his head. Tomorrow he does the same thing, dropping the oldest day and adding the newest. The window slides forward with him.

That is a moving average, and he has been running one for twenty years without calling it that.

Why it exists

Raw daily data is loud. One rainy Tuesday, one festival, one server outage, and the chart looks like a heart monitor.

The question you actually want answered is quieter. Are we growing or shrinking? You cannot see that through the spikes.

A moving average turns down the volume. Individual odd days get averaged against their neighbours and shrink. Anything that lasts longer than the window survives.

How it works

   day     :   1     2     3     4     5     6
   visits  : 100   120    90   130   110   200

   window of 3, sliding along:

   [100 120  90]                  -> 103   (goes on day 3)
         [120  90 130]            -> 113   (goes on day 4)
               [ 90 130 110]      -> 110   (goes on day 5)
                     [130 110 200]-> 147   (goes on day 6)

   days 1 and 2 get nothing. There are not three values behind them yet.

Two facts fall out of that picture, and both matter.

You always lose the start. A window of three needs three values. A window of thirty needs thirty. On two years of monthly data, a twelve-month window costs you a whole year.

The average lags behind. Look at day 6. The real value jumped to 200, but the average only reached 147, because it is still carrying 130 and 110 from earlier. A moving average always arrives late to a change.

The window size is the whole decision

There is one dial, and it trades two things against each other.

A small window follows the data closely and reacts fast. It also keeps most of the noise, so you still see spikes.

A large window gives a beautifully smooth line. It also reacts slowly, so a genuine change takes a long time to show up.

There is no correct answer. There is only a choice about which mistake you would rather make. If you need to spot a problem within days, a thirty-day window will hide it from you for weeks.

The trap that ruins forecasts

Some tools offer a centred moving average. The window sits with equal amounts of past and future on either side.

A centred average is beautiful for looking at history. It has no lag at all, because the future is pulling it into place.

You can never use it to forecast. Computing today's centred average needs tomorrow's value, and tomorrow has not happened. Any model fed centred averages has been handed the answers, and its test score is worthless.

This is exactly the mistake described in what is time series data, wearing a different hat. Whenever a feature touches a future row, you have leaked.

Where you have already seen it

  • Cricket run rate over the last five overs is a moving average.
  • The daily case counts during the pandemic were reported as a seven-day average, because weekend reporting made raw counts jump around.
  • Your bike's fuel-efficiency readout averages the last stretch of riding.
  • Stock charts overlay a fifty-day and a two-hundred-day line on the price.
  • Fitness apps show a weekly average of steps rather than yesterday's number.

What is honestly hard here

A moving average is a description, not a prediction. It tells you where the series has been. Used as a forecast, it says "tomorrow will look like the recent past". That is a reasonable guess and rarely a good one.

It is also the wrong tool when your data has a repeating weekly or yearly pattern. Averaging a whole week together deletes the very Saturday spike you were trying to predict. The developer section shows this failing on purpose, with numbers.

Remember this

  • A moving average slides a window along and averages what is inside it.
  • Small window means fast and noisy. Large window means smooth and late.
  • Centred averages look at the future, so they are for studying history, never for forecasting.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas

Alignment is the whole game

Print a small table and stare at it. Ten rows is enough to check every number by hand, and hand-checking this once will save you a leaked feature later.

alignment.py
import pandas as pd

# Ten days of visits to a small blog, small enough to check the arithmetic by hand.
visits = pd.Series(
    [100, 120, 90, 130, 110, 200, 190, 105, 115, 95],
    index=pd.date_range("2025-03-01", periods=10, freq="D"), name="visits")

table = pd.DataFrame({
    "visits": visits,
    "trailing3": visits.rolling(3).mean(),               # today and the two days before
    "centred3": visits.rolling(3, center=True).mean(),   # yesterday, today, tomorrow
    "usable_forecast": visits.rolling(3).mean().shift(1),
})
print(table.round(1))
Output
            visits  trailing3  centred3  usable_forecast
2025-03-01     100        NaN       NaN              NaN
2025-03-02     120        NaN     103.3              NaN
2025-03-03      90      103.3     113.3              NaN
2025-03-04     130      113.3     110.0            103.3
2025-03-05     110      110.0     146.7            113.3
2025-03-06     200      146.7     166.7            110.0
2025-03-07     190      166.7     165.0            146.7
2025-03-08     105      165.0     136.7            166.7
2025-03-09     115      136.7     105.0            165.0
2025-03-10      95      105.0       NaN            136.7

Three columns, three different degrees of legality.

centred3 on 2025-03-02 is 103.3, which is the average of March 1, 2 and 3. It used March 3, a day that had not happened yet on March 2. Never feed this to a model.

trailing3 on 2025-03-03 is 103.3, the average of March 1, 2 and 3. No future is involved. But it contains today's own value, so it cannot be a forecast of today. It is a legitimate summary of today, useful as a feature for predicting tomorrow.

usable_forecast shifts everything down one row. On March 4 it reads 103.3, made only from March 1 to 3. That number was fully knowable on the evening of March 3, so predicting March 4 with it is honest.

The rule to memorise: .rolling(k).mean() then .shift(1). Windows first, then shift. This one line is the difference between a real result and a fantasy.

The lag, measured exactly

Take a series that climbs by exactly ten every day, with no noise. Now the lag is arithmetic, not opinion.

lag.py
import numpy as np
import pandas as pd

# A series that climbs by exactly 10 every day. No noise at all.
climb = pd.Series(np.arange(100, 200, 10), index=pd.date_range("2025-06-01", periods=10, freq="D"))

flat_weights = np.array([1, 1, 1, 1, 1]) / 5
recent_weights = np.array([1, 2, 3, 4, 5]) / 15      # yesterday counts most

def weighted(s, w):
    return s.rolling(len(w)).apply(lambda win: float(np.dot(win, w)), raw=True)

print(pd.DataFrame({
    "value": climb,
    "plain MA(5)": weighted(climb, flat_weights),
    "weighted MA(5)": weighted(climb, recent_weights),
}).round(1))
Output
            value  plain MA(5)  weighted MA(5)
2025-06-01    100          NaN             NaN
2025-06-02    110          NaN             NaN
2025-06-03    120          NaN             NaN
2025-06-04    130          NaN             NaN
2025-06-05    140        120.0           126.7
2025-06-06    150        130.0           136.7
2025-06-07    160        140.0           146.7
2025-06-08    170        150.0           156.7
2025-06-09    180        160.0           166.7
2025-06-10    190        170.0           176.7

The plain average sits exactly 20 below the truth, every single day. At ten units per day, that is a two-day lag — which is (5 - 1) / 2, the standard result for a window of five.

The weighted version sits 13.3 below, a lag of one and a third days. Leaning on recent values cuts the lag, at the price of a rougher line.

That trade-off has no free lunch in it, and the honest way to think about the next lesson is that exponential smoothing is what you get when you push the weighting idea to its logical end.

A moving average as a forecast, scored properly

Now the uncomfortable part. Here is a moving-average forecaster tested against two dumb baselines on data with a weekend pattern.

ma_forecast.py
import numpy as np
import pandas as pd

rng = np.random.default_rng(21)
n = 180
days = pd.date_range("2025-01-01", periods=n, freq="D")
weekend = np.where(days.dayofweek >= 5, 150, 0)
visits = pd.Series((np.linspace(400, 560, n) + weekend + rng.normal(0, 30, n)).round(),
                   index=days, name="visits")

test = visits.index >= days[120]          # score only the last 60 days

def mae(forecast):
    return (visits[test] - forecast[test]).abs().mean()

print("naive (yesterday)          MAE:", round(mae(visits.shift(1)), 1))
print("seasonal naive (last week) MAE:", round(mae(visits.shift(7)), 1))
for k in (3, 7, 14, 30):
    print(f"moving average of {k:2d} days   MAE:", round(mae(visits.rolling(k).mean().shift(1)), 1))
Output
naive (yesterday)          MAE: 59.4
seasonal naive (last week) MAE: 31.7
moving average of  3 days   MAE: 74.2
moving average of  7 days   MAE: 60.0
moving average of 14 days   MAE: 58.2
moving average of 30 days   MAE: 56.7

Every moving average lost to visits.shift(7), by a wide margin. The seasonal naive is a one-line forecast that copies the value from the same weekday last week. It scored 31.7 against the best moving average's 56.7.

The reason is structural. This series has a weekend that runs 150 visits high. Averaging seven days blends Saturdays into Tuesdays and produces one middling number for every day of the week. The pattern is deleted by the very thing meant to help.

The three-day average was the worst of all, at 74.2, worse than doing nothing. A three-day window straddles the Friday-to-Sunday boundary and mixes two different regimes.

The lesson generalises: when a strong seasonal pattern exists, smooth within the season, never across it. And always report a naive baseline, or your number means nothing. Evaluating a forecast makes that a rule.

Useful variants in pandas

python
s.rolling(7).mean()                      # trailing simple moving average
s.rolling(7, min_periods=3).mean()       # start producing values after 3 rows, not 7
s.rolling(7).median()                    # robust to a single wild spike
s.rolling("30D").mean()                  # 30 calendar days, correct with missing dates
s.expanding().mean()                     # average of everything so far
s.ewm(span=7).mean()                     # exponentially weighted, no hard window edge

s.rolling("30D") deserves attention. The integer form counts rows; the string form counts time. If your sensor dropped readings, thirty rows might span forty-five days. The string form is almost always what you meant.

Common mistakes

Forgetting .shift(1) on a rolling feature. The window includes the current row, so the feature knows today's target. Models trained this way score beautifully offline and collapse in production.

Using center=True in a feature pipeline. It is a direct look into the future. Keep it for charts and decomposition, never for features.

Using min_periods to make NaN go away. s.rolling(30, min_periods=1).mean() produces a value on row one computed from a single observation. Those early rows are wildly noisy and are not comparable with later rows. Drop them instead.

Counting rows when you meant days. Any gap in the index silently changes what a window covers. Check s.index.to_series().diff().value_counts() before trusting an integer window.

Choosing the window by trying values on the test set. That is tuning on your test data. Choose it on a validation split, then measure once.

Try it yourself

In the last script, replace visits.rolling(k).mean().shift(1) with a seasonal moving average: the average of the same weekday over the past three weeks. In pandas that is visits.shift(7).rolling(3).mean() combined with taking every seventh row.

Score it against the 31.7 from the plain seasonal naive. Averaging three same-weekdays should beat copying one, because it cancels noise without crossing the season boundary. Check whether it actually does.

What to learn next

Researcher — Mathematics and papers.

The simple moving average as a linear filter

The trailing SMA of order $k$ is

$$ \hat{T}t = \frac{1}{k} \sum{j=0}^{k-1} y_{t-j} $$

and the centred version for odd $k = 2q+1$ is

$$ \hat{T}t = \frac{1}{2q+1} \sum{j=-q}^{q} y_{t+j} $$

  • $k$ — the window length.
  • $q$ — the half-width, $(k-1)/2$.
  • $y_t$ — the observation at time $t$.

Both are finite impulse response filters with rectangular weights. Their behaviour is best understood in the frequency domain.

Frequency response, and why the rectangular window is poor

For a symmetric filter with weights ${a_j}_{j=-q}^{q}$, the transfer function is $A(\omega) = \sum_j a_j e^{-i\omega j}$. For the centred SMA this evaluates to the Dirichlet kernel

$$ A(\omega) = \frac{\sin(k\omega/2)}{k \sin(\omega/2)} $$

  • $\omega$ — angular frequency in radians per sample, $\omega \in [0, \pi]$.

Two consequences follow, and both are visible in practice:

  • Zeros at $\omega = 2\pi j / k$. A centred SMA of order $k$ annihilates any cycle whose period divides $k$ exactly. This is precisely why an SMA of order $m$ removes seasonality of period $m$, and it is the design principle behind classical decomposition.
  • Sidelobes that do not decay fast, and $A(\omega) < 0$ over some bands. A negative transfer means the filter inverts components at those frequencies. The smoothed series can show a peak where the data had a trough. This artefact is called the Slutsky-Yule effect: Slutsky (1937) and Yule (1926) showed that repeatedly averaging pure noise manufactures apparent cycles that were never in the data.

The trailing SMA has the same magnitude response but a linear phase delay of $(k-1)/2$ samples. That delay is exactly the lag measured empirically in the developer block, and it is unavoidable for any causal filter with non-negative weights.

Better kernels

Rectangular weights minimise variance for a locally constant signal and nothing else. Alternatives that trade a little variance for far better frequency behaviour:

  • Henderson filters (Henderson, 1916) minimise the sum of squared third differences of the smoothed output subject to reproducing cubic polynomials exactly. The 13-term Henderson MA is the trend filter inside X-11.
  • Loess fits a locally weighted low-degree polynomial with tricube weights, reproducing local linear or quadratic structure and behaving sensibly at the endpoints. This is the smoother inside STL, discussed in trend, seasonality and noise.
  • Savitzky-Golay (1964) fits a degree-$p$ polynomial by least squares in each window, which preserves peak height and width far better than an SMA — the standard choice in spectroscopy, available as scipy.signal.savgol_filter.
  • Hodrick-Prescott solves $\min_\tau \sum (y_t - \tau_t)^2 + \lambda \sum (\Delta^2 \tau_t)^2$. Widely used in macroeconomics and widely criticised: Hamilton (2018), Why you should never use the Hodrick-Prescott filter, shows it induces spurious dynamics and that the conventional $\lambda$ values are arbitrary.

Variance reduction and the bias-variance statement

For $y_t = \mu_t + \varepsilon_t$ with $\varepsilon_t$ white noise of variance $\sigma^2$,

$$ \mathrm{Var}(\hat{T}_t) = \frac{\sigma^2}{k}, \qquad \mathrm{Bias}(\hat{T}t) = \frac{1}{k}\sum{j=0}^{k-1} \mu_{t-j} - \mu_t $$

Variance falls as $1/k$. Bias grows with $k$ whenever $\mu_t$ is not locally constant; for a linear trend of slope $\beta$ the trailing bias is exactly $-\beta(k-1)/2$. Minimising mean squared error gives an optimal $k$ that scales as $\sigma^{2/3}$ against curvature — the same $O(n^{-1/5})$ bandwidth story as in kernel regression, and the reason no universal window exists.

The white-noise assumption also fails for autocorrelated errors. With AR(1) errors of parameter $\phi$, the variance of the mean of $k$ consecutive observations is inflated by roughly $(1+\phi)/(1-\phi)$, so a rolling mean over positively autocorrelated data is noisier than it looks.

The relationship to exponential smoothing

The SMA gives weight $1/k$ to the last $k$ observations and zero beyond. Exponential smoothing gives weight $\alpha(1-\alpha)^j$ to lag $j$, never reaching zero. Equating the average age of the weights,

$$ \frac{k-1}{2} = \frac{1-\alpha}{\alpha} \quad \Longleftrightarrow \quad \alpha = \frac{2}{k+1} $$

which is the standard span convention in pandas.Series.ewm(span=k). The two smoothers are comparable in responsiveness at that setting, but exponential smoothing requires storing one number instead of $k$, and produces no leading NaN block. That is developed in exponential smoothing.

Reading

  • Yule (1926), Why do we sometimes get nonsense-correlations between time-series?, JRSS 89(1).
  • Slutsky (1937), The summation of random causes as the source of cyclic processes, Econometrica 5(2).
  • Savitzky and Golay (1964), Smoothing and differentiation of data by simplified least squares procedures, Analytical Chemistry 36(8).
  • Hamilton (2018), Why you should never use the Hodrick-Prescott filter, Review of Economics and Statistics 100(5).
  • Hyndman and Athanasopoulos, Forecasting: Principles and Practice, 3rd ed., section 3.3 — otexts.com/fpp3/moving-averages.html.

What to learn next

What to learn next

These follow on from what you just read.

  • Time Series and Forecasting

    Exponential smoothing

    Exponential smoothing keeps one running estimate and nudges it toward each new observation, giving recent data more weight without ever forgetting the past completely.

  • Time Series and Forecasting

    ARIMA

    ARIMA forecasts a series from its own past values and its own past errors, after differencing away whatever kept moving.

  • Time Series and Forecasting

    Prophet

    Prophet forecasts by adding up a bendy trend, repeating seasonal shapes and named holiday effects, which makes moving festivals and missing days easy to handle.