Time Series and Forecasting

What is time series data?

Time series data is data where the order of the rows is part of the information, so you can never shuffle it into a random train and test split.

On this page 10
  1. The short answer
  2. The analogy you have already lived
  3. What makes it different from ordinary data
  4. Why this breaks the normal recipe
  5. How a forecast actually works
  6. The vocabulary, in plain words
  7. Where you have already seen it
  8. What is honestly hard here
  9. Remember this
  10. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Time series data is data where each row has a time stamp, and the order of the rows carries meaning.

The analogy you have already lived

Look at the back of your electricity bill. There is a little bar chart of the last twelve months of units used.

You already read that chart like a forecaster. The bars get taller in summer, because the fan and the cooler run all day. You can guess next month's bar within a rough range, without any maths.

You are doing three things at once. You see the long slow direction. You see the yearly summer bump. You ignore the one odd month when guests stayed over.

That is forecasting. The rest of this section gives names to what you already do.

What makes it different from ordinary data

Most machine learning data is a pile of independent rows. A thousand photos of cats. Ten thousand customer records. Shuffle them and nothing is lost.

Time series data is not a pile. It is a queue — a line where position matters.

Take the twelve bars on your bill and shuffle them into a random order. The total units stay the same. The average stays the same. The highest and lowest months stay the same.

But the shape is gone. You can no longer see that summer follows spring. The only thing you destroyed is the exact thing you wanted to predict.

This is the single most important sentence in this whole section. In time series work, the order is not metadata around the data. The order is the data.

Why this breaks the normal recipe

In ordinary machine learning you split your rows into a training pile and a testing pile at random. That is covered in train, test and validation splits, and for photos or customer records it is the right thing to do.

Do that with a time series and you have quietly cheated.

   RANDOM SPLIT  (wrong for time series)

   Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec
    T   E   T   T   E   T   E   T   T   T   E   T
                             T means train, E means test

   You trained on June and August.
   Then you asked the model to "predict" July.
   It was surrounded by answers on both sides.


   CHRONOLOGICAL SPLIT  (right for time series)

   Jan Feb Mar Apr May Jun Jul Aug Sep | Oct Nov Dec
    T   T   T   T   T   T   T   T   T  |  E   E   E

   You train only on the past.
   You test only on the future, which the model has never seen.

The random split gives you a wonderful score. That score is a lie, because in real life the future is not sitting there in your training data.

A model tested this way looks brilliant in your notebook and falls apart in production. This mistake is common, it is silent, and it has cost real companies real money.

How a forecast actually works

  past values you have          →   model   →   next values you do not have

  Jan Feb Mar ... Sep                          Oct  Nov  Dec
  410 395 480     620                          640  655  670
                                                (a guess, with a range)

Everything to the left of the arrow already happened. Everything to the right has not happened yet. The line between them is called the forecast origin, and it moves forward as new data arrives.

The vocabulary, in plain words

  • Frequency — how often a reading arrives. Hourly, daily, monthly.
  • Horizon — how far ahead you are predicting. Seven days ahead is a seven-step horizon.
  • Lag — a value from earlier. Yesterday's sales are the lag-one value for today.
  • Gap — a missing reading. The meter reader did not come that month.

Each of these gets used constantly in the lessons that follow.

Where you have already seen it

  • Your electricity bill — twelve monthly readings, with a summer bump.
  • A shop before Diwali — sales climb for two weeks, then fall off a cliff.
  • Weather apps — a temperature forecast for the next seven days, with a confidence band.
  • UPI and card statements — your spending, in order, month by month.
  • Website dashboards — visits per day, with a dip every weekend.

None of these make sense if you shuffle the rows.

What is honestly hard here

Forecasting has a ceiling that no model can break through.

Some series are close to unpredictable. Tomorrow's stock price is the famous example. A large part of the movement is genuinely random, and no amount of deep learning removes randomness.

Anyone promising accurate long-range forecasts of a noisy series is selling something. A good forecaster knows the limit and reports an honest range instead of a confident single number.

Remember this

  • Time series data is data in order, and the order is the information.
  • Never split it at random. Train on the past, test on the future.
  • Your goal is a sensible range, not a magic exact number.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy pandas scikit-learn

No GPU, no dataset download. Everything here is generated inline and runs in under two seconds.

A time series is a Series with a DatetimeIndex

In pandas, the thing that makes a table a time series is the index — the labels down the left side. When those labels are dates, pandas unlocks resampling, rolling windows and lag operations.

order_is_data.py
import numpy as np
import pandas as pd

rng = np.random.default_rng(7)

# 120 days of visits to a small recipe blog: a slow rise, a weekend bump, and noise.
days = pd.date_range("2025-01-01", periods=120, freq="D")
trend = np.linspace(400, 700, 120)
weekend_bump = np.where(days.dayofweek >= 5, 120, 0)
visits = pd.Series(
    (trend + weekend_bump + rng.normal(0, 25, 120)).round(),
    index=days, name="visits",
)

print(visits.head(8))
print()

# autocorr(1) asks: does today's value tell you anything about tomorrow's?
print("mean visits            :", round(visits.mean(), 1))
print("lag-1 correlation      :", round(visits.autocorr(1), 3))

shuffled = visits.sample(frac=1, random_state=0).reset_index(drop=True)
print()
print("after shuffling the rows")
print("mean visits            :", round(shuffled.mean(), 1))
print("lag-1 correlation      :", round(shuffled.autocorr(1), 3))
Output
2025-01-01    400.0
2025-01-02    410.0
2025-01-03    398.0
2025-01-04    505.0
2025-01-05    519.0
2025-01-06    388.0
2025-01-07    417.0
2025-01-08    451.0
Freq: D, Name: visits, dtype: float64

mean visits            : 580.3
lag-1 correlation      : 0.78

after shuffling the rows
mean visits            : 580.3
lag-1 correlation      : 0.013

Read those four numbers slowly, because they are the whole lesson.

The mean is identical before and after the shuffle. So is the standard deviation, the minimum and the maximum. Every summary statistic you learned in statistics survived untouched.

The lag-one correlation collapsed from 0.78 to 0.013. That number measures how much today's value tells you about tomorrow's. Shuffling destroyed it completely, and it is the only quantity a forecaster cares about.

freq: D in the output matters. pandas recorded that readings arrive daily with no gaps. Many time series functions rely on it. If your index has holes, freq becomes None and things quietly misbehave.

Now watch a random split lie to you

the_split_that_lies.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error

rng = np.random.default_rng(0)
n = 200
# A small shop's daily takings in rupees. It wanders up and down, like real takings do.
sales = 5000 + np.cumsum(rng.normal(0, 120, n))
X = np.arange(n).reshape(-1, 1)          # the only feature is the day number

def mae_of(X_train, y_train, X_test, y_test):
    model = RandomForestRegressor(n_estimators=50, random_state=0).fit(X_train, y_train)
    return mean_absolute_error(y_test, model.predict(X_test))

# WRONG: shuffle the days, the way you would shuffle photos
Xa, Xb, ya, yb = train_test_split(X, sales, test_size=0.25, random_state=0, shuffle=True)
print("random split   MAE:", round(mae_of(Xa, ya, Xb, yb), 1), "rupees")

# RIGHT: train on the first 150 days, test on the 50 days that came after
cut = 150
print("chrono split   MAE:", round(mae_of(X[:cut], sales[:cut], X[cut:], sales[cut:]), 1), "rupees")

print("day 150 takings:", round(sales[150]), "| day 199 takings:", round(sales[199]))
Output
random split   MAE: 75.4 rupees
chrono split   MAE: 478.9 rupees
day 150 takings: 6100 | day 199 takings: 5366

Same data. Same model. Same metric. The reported error is six times worse once the split respects time.

Here is the mechanism. Under the random split, day 91 lands in the test set while days 90 and 92 sit in the training set. A shop's takings barely move overnight, so the model looks up its neighbours and copies them. That is not forecasting. That is reading the answer sheet.

The chronological number, 478.9, is the honest one. It is what you would have experienced had you deployed on day 150.

Do not read the second number as a failure. These takings are a random walk, so a large part of the movement is genuinely unpredictable. The chronological error is close to the floor for this series. The random split's 75.4 was never achievable by anyone.

The correct splitting tool

scikit-learn ships a time-aware splitter. Use it instead of train_test_split for anything with a clock in it.

rolling_origin.py
import numpy as np
from sklearn.model_selection import TimeSeriesSplit

y = np.arange(12)                    # twelve months, in order
for i, (train_idx, test_idx) in enumerate(TimeSeriesSplit(n_splits=3).split(y), 1):
    print(f"fold {i}  train={list(train_idx)}  test={list(test_idx)}")
Output
fold 1  train=[0, 1, 2]  test=[3, 4, 5]
fold 2  train=[0, 1, 2, 3, 4, 5]  test=[6, 7, 8]
fold 3  train=[0, 1, 2, 3, 4, 5, 6, 7, 8]  test=[9, 10, 11]

Every test index is larger than every training index in the same fold. The training window grows as the origin moves forward. This is called rolling-origin evaluation, and it is the standard way to score a forecaster. Evaluating a forecast goes into it properly.

Common mistakes

Calling train_test_split without shuffle=False. The default is shuffle=True. That one default has broken more forecasting projects than any algorithm choice. If you use it at all, pass shuffle=False.

Scaling before splitting. Fitting a StandardScaler on the whole series computes a mean that includes future values, then bakes it into your training features. Fit the scaler on training data only, then apply it to the test data.

A string index instead of a datetime index. If dates load as text, "2025-1-9" sorts before "2025-10-01". Convert on read with pd.read_csv(path, parse_dates=["date"], index_col="date"), then check df.index.dtype says datetime64[ns].

Silent duplicate or missing timestamps. Run df.index.duplicated().sum() and df.index.to_series().diff().value_counts(). A single duplicated hour from a clock change will corrupt every rolling window downstream.

Shuffling inside cross-validation. cross_val_score uses KFold, which shuffles by default for regression in many setups. Pass cv=TimeSeriesSplit(5) explicitly.

Try it yourself

Take the visits series from the first script. Compute visits.autocorr(7) — the correlation with the value seven days earlier.

Predict whether it will be higher or lower than autocorr(1) before you run it. The weekend bump repeats every seven days, so think about what that does. Then explain the result to yourself in one sentence.

What to learn next

Researcher — Mathematics and papers.

Formal setup

A univariate time series is a realisation of a stochastic process ${Y_t}_{t \in \mathbb{Z}}$ observed at $t = 1, \dots, T$.

  • $Y_t$ — the random variable at time index $t$.
  • $y_t$ — the observed value, one realisation of $Y_t$.
  • $T$ — the number of observations available.
  • $h$ — the forecast horizon, the number of steps ahead being predicted.

The forecasting problem is to estimate the conditional distribution

$$ p!\left(Y_{T+h} \mid Y_{1:T} = y_{1:T}\right) $$

and the point forecast $\hat{y}{T+h|T}$ is a functional of it. Under squared-error loss the optimal point forecast is the conditional mean $\mathbb{E}[Y{T+h} \mid \mathcal{F}_T]$, where $\mathcal{F}_T$ is the sigma-algebra generated by observations up to $T$. Under absolute-error loss it is the conditional median. These differ for skewed series, which is why a metric choice is a modelling decision, not a reporting decision.

Why i.i.d. machinery does not transfer

Standard supervised learning theory assumes the training sample is independent and identically distributed. Time series violate both halves.

Dependence. $\mathrm{Cov}(Y_t, Y_{t-k}) \neq 0$ for small $k$ in almost every real series. The autocovariance function

$$ \gamma(t, t-k) = \mathbb{E}!\left[(Y_t - \mu_t)(Y_{t-k} - \mu_{t-k})\right] $$

is exactly the structure the model is meant to exploit, so it cannot be assumed away.

Non-identical distribution. $\mu_t$ and $\gamma$ typically vary with $t$. This is the subject of stationarity, and stationarity assumptions are what allow $\gamma(t, t-k)$ to collapse to a function $\gamma(k)$ of lag alone.

The consequence for validation is not a matter of taste. K-fold cross-validation estimates generalisation error under exchangeability. Time series are not exchangeable, so the estimator is biased, and the bias is optimistic. Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, quantifies this and gives the conditions under which a blocked variant recovers validity.

Rolling-origin evaluation

Tashman (2000), Out-of-sample tests of forecasting accuracy, formalises the standard protocol. For origins $t_0, t_0+1, \dots, T-h$ produce $\hat{y}{t+h|t}$ using only $y{1:t}$, then aggregate the errors $e_{t+h} = y_{t+h} - \hat{y}_{t+h|t}$.

Two design choices matter and are frequently conflated:

  • Expanding window — train on $y_{1:t}$, which grows. Preferred when the data-generating process is stable and data is scarce.
  • Sliding window — train on $y_{t-w+1:t}$ for fixed $w$. Preferred under regime change, at the cost of variance.

A purge gap between train and test is required whenever features have look-back windows that would otherwise straddle the boundary, and whenever the label itself spans time. López de Prado (2018), Advances in Financial Machine Learning, chapter 7, gives the purging and embargo construction for overlapping labels.

Decomposing the error floor

Write the $h$-step error as

$$ \mathbb{E}!\left[(Y_{T+h} - \hat{y}{T+h|T})^2\right] = \underbrace{\mathrm{Var}(Y{T+h} \mid \mathcal{F}T)}{\text{irreducible}} + \underbrace{\left(\mathbb{E}[Y_{T+h}\mid\mathcal{F}T] - \hat{y}{T+h|T}\right)^2}_{\text{reducible}} $$

For a random walk $Y_t = Y_{t-1} + \varepsilon_t$ with $\varepsilon_t \sim \mathrm{WN}(0, \sigma^2)$, the conditional mean is $y_T$ for every $h$, and the irreducible term is $h\sigma^2$. Forecast uncertainty grows linearly in the horizon and no model can shrink it. This is the formal version of the honesty warning in the beginner block, and it is why the naive forecast is a competitive benchmark on financial series rather than a straw man.

Benchmarks that must be beaten

Any proposed model should be reported against these, on the same rolling origins:

BenchmarkForecastWhen it is hard to beat
Naive$\hat{y}_{T+h\mid T} = y_T$Random-walk-like series
Seasonal naive$\hat{y}{T+h\mid T} = y{T+h-m}$Strong fixed seasonality, period $m$
Drift$y_T + h\frac{y_T - y_1}{T-1}$Steady trend, low noise
Historical mean$\bar{y}$Mean-reverting, weak dependence

The M-competitions (Makridakis et al., 1982 through M5 in 2020) repeatedly found that simple methods and combinations of them are strong. M4 (2018) was won by a hybrid of exponential smoothing and a recurrent network; the pure-machine-learning entries below it underperformed statistical baselines. Report the benchmark or the result is uninterpretable.

Reading

  • Hyndman and Athanasopoulos, Forecasting: Principles and Practice, 3rd edition — free at otexts.com/fpp3. The standard modern reference.
  • Box, Jenkins, Reinsel and Ljung, Time Series Analysis: Forecasting and Control, 5th edition, 2015.
  • Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, Information Sciences 191.
  • Makridakis, Spiliotis and Assimakopoulos (2018), The M4 Competition: Results, findings, conclusion and way forward.
  • Tashman (2000), Out-of-sample tests of forecasting accuracy: an analysis and review, International Journal of Forecasting 16.

What to learn next

What to learn next

These follow on from what you just read.

  • Time Series and Forecasting

    Trend, seasonality and noise

    Almost every time series is three things added together — a slow direction, a repeating pattern and leftover randomness — and separating them is the first thing you do.

  • Time Series and Forecasting

    Stationarity

    A stationary series keeps the same average and the same wobble forever, and most forecasting models quietly assume yours does too.

  • Time Series and Forecasting

    Moving averages

    A moving average smooths a bumpy series by averaging a sliding window, and the window you choose decides how much noise you remove and how far behind you fall.