What is time series data?
Time series data is data where the order of the rows is part of the information, so you can never shuffle it into a random train and test split.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 10
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Time series data is data where each row has a time stamp, and the order of the rows carries meaning.
The analogy you have already lived
Look at the back of your electricity bill. There is a little bar chart of the last twelve months of units used.
You already read that chart like a forecaster. The bars get taller in summer, because the fan and the cooler run all day. You can guess next month's bar within a rough range, without any maths.
You are doing three things at once. You see the long slow direction. You see the yearly summer bump. You ignore the one odd month when guests stayed over.
That is forecasting. The rest of this section gives names to what you already do.
What makes it different from ordinary data
Most machine learning data is a pile of independent rows. A thousand photos of cats. Ten thousand customer records. Shuffle them and nothing is lost.
Time series data is not a pile. It is a queue — a line where position matters.
Take the twelve bars on your bill and shuffle them into a random order. The total units stay the same. The average stays the same. The highest and lowest months stay the same.
But the shape is gone. You can no longer see that summer follows spring. The only thing you destroyed is the exact thing you wanted to predict.
This is the single most important sentence in this whole section. In time series work, the order is not metadata around the data. The order is the data.
Why this breaks the normal recipe
In ordinary machine learning you split your rows into a training pile and a testing pile at random. That is covered in train, test and validation splits, and for photos or customer records it is the right thing to do.
Do that with a time series and you have quietly cheated.
RANDOM SPLIT (wrong for time series)
Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec
T E T T E T E T T T E T
T means train, E means test
You trained on June and August.
Then you asked the model to "predict" July.
It was surrounded by answers on both sides.
CHRONOLOGICAL SPLIT (right for time series)
Jan Feb Mar Apr May Jun Jul Aug Sep | Oct Nov Dec
T T T T T T T T T | E E E
You train only on the past.
You test only on the future, which the model has never seen.The random split gives you a wonderful score. That score is a lie, because in real life the future is not sitting there in your training data.
A model tested this way looks brilliant in your notebook and falls apart in production. This mistake is common, it is silent, and it has cost real companies real money.
How a forecast actually works
past values you have → model → next values you do not have
Jan Feb Mar ... Sep Oct Nov Dec
410 395 480 620 640 655 670
(a guess, with a range)Everything to the left of the arrow already happened. Everything to the right has not happened yet. The line between them is called the forecast origin, and it moves forward as new data arrives.
The vocabulary, in plain words
- Frequency — how often a reading arrives. Hourly, daily, monthly.
- Horizon — how far ahead you are predicting. Seven days ahead is a seven-step horizon.
- Lag — a value from earlier. Yesterday's sales are the lag-one value for today.
- Gap — a missing reading. The meter reader did not come that month.
Each of these gets used constantly in the lessons that follow.
Where you have already seen it
- Your electricity bill — twelve monthly readings, with a summer bump.
- A shop before Diwali — sales climb for two weeks, then fall off a cliff.
- Weather apps — a temperature forecast for the next seven days, with a confidence band.
- UPI and card statements — your spending, in order, month by month.
- Website dashboards — visits per day, with a dip every weekend.
None of these make sense if you shuffle the rows.
What is honestly hard here
Forecasting has a ceiling that no model can break through.
Some series are close to unpredictable. Tomorrow's stock price is the famous example. A large part of the movement is genuinely random, and no amount of deep learning removes randomness.
Anyone promising accurate long-range forecasts of a noisy series is selling something. A good forecaster knows the limit and reports an honest range instead of a confident single number.
Remember this
- Time series data is data in order, and the order is the information.
- Never split it at random. Train on the past, test on the future.
- Your goal is a sensible range, not a magic exact number.
What to learn next
- Trend, seasonality and noise — the three parts every series is made of.
- Train, test and validation splits — the ordinary case, so you can see exactly what changes here.
- Evaluating a forecast — how to score one without fooling yourself.
Developer — Code and libraries.
Setup
pip install numpy pandas scikit-learnNo GPU, no dataset download. Everything here is generated inline and runs in under two seconds.
A time series is a Series with a DatetimeIndex
In pandas, the thing that makes a table a time series is the index — the labels down the left side. When those labels are dates, pandas unlocks resampling, rolling windows and lag operations.
import numpy as np
import pandas as pd
rng = np.random.default_rng(7)
# 120 days of visits to a small recipe blog: a slow rise, a weekend bump, and noise.
days = pd.date_range("2025-01-01", periods=120, freq="D")
trend = np.linspace(400, 700, 120)
weekend_bump = np.where(days.dayofweek >= 5, 120, 0)
visits = pd.Series(
(trend + weekend_bump + rng.normal(0, 25, 120)).round(),
index=days, name="visits",
)
print(visits.head(8))
print()
# autocorr(1) asks: does today's value tell you anything about tomorrow's?
print("mean visits :", round(visits.mean(), 1))
print("lag-1 correlation :", round(visits.autocorr(1), 3))
shuffled = visits.sample(frac=1, random_state=0).reset_index(drop=True)
print()
print("after shuffling the rows")
print("mean visits :", round(shuffled.mean(), 1))
print("lag-1 correlation :", round(shuffled.autocorr(1), 3))2025-01-01 400.0 2025-01-02 410.0 2025-01-03 398.0 2025-01-04 505.0 2025-01-05 519.0 2025-01-06 388.0 2025-01-07 417.0 2025-01-08 451.0 Freq: D, Name: visits, dtype: float64 mean visits : 580.3 lag-1 correlation : 0.78 after shuffling the rows mean visits : 580.3 lag-1 correlation : 0.013
Read those four numbers slowly, because they are the whole lesson.
The mean is identical before and after the shuffle. So is the standard deviation, the minimum and the maximum. Every summary statistic you learned in statistics survived untouched.
The lag-one correlation collapsed from 0.78 to 0.013. That number measures how much today's value tells you about tomorrow's. Shuffling destroyed it completely, and it is the only quantity a forecaster cares about.
freq: D in the output matters. pandas recorded that readings arrive daily with no gaps. Many time series functions rely on it. If your index has holes, freq becomes None and things quietly misbehave.
Now watch a random split lie to you
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
rng = np.random.default_rng(0)
n = 200
# A small shop's daily takings in rupees. It wanders up and down, like real takings do.
sales = 5000 + np.cumsum(rng.normal(0, 120, n))
X = np.arange(n).reshape(-1, 1) # the only feature is the day number
def mae_of(X_train, y_train, X_test, y_test):
model = RandomForestRegressor(n_estimators=50, random_state=0).fit(X_train, y_train)
return mean_absolute_error(y_test, model.predict(X_test))
# WRONG: shuffle the days, the way you would shuffle photos
Xa, Xb, ya, yb = train_test_split(X, sales, test_size=0.25, random_state=0, shuffle=True)
print("random split MAE:", round(mae_of(Xa, ya, Xb, yb), 1), "rupees")
# RIGHT: train on the first 150 days, test on the 50 days that came after
cut = 150
print("chrono split MAE:", round(mae_of(X[:cut], sales[:cut], X[cut:], sales[cut:]), 1), "rupees")
print("day 150 takings:", round(sales[150]), "| day 199 takings:", round(sales[199]))random split MAE: 75.4 rupees chrono split MAE: 478.9 rupees day 150 takings: 6100 | day 199 takings: 5366
Same data. Same model. Same metric. The reported error is six times worse once the split respects time.
Here is the mechanism. Under the random split, day 91 lands in the test set while days 90 and 92 sit in the training set. A shop's takings barely move overnight, so the model looks up its neighbours and copies them. That is not forecasting. That is reading the answer sheet.
The chronological number, 478.9, is the honest one. It is what you would have experienced had you deployed on day 150.
Do not read the second number as a failure. These takings are a random walk, so a large part of the movement is genuinely unpredictable. The chronological error is close to the floor for this series. The random split's 75.4 was never achievable by anyone.
The correct splitting tool
scikit-learn ships a time-aware splitter. Use it instead of train_test_split for anything with a clock in it.
import numpy as np
from sklearn.model_selection import TimeSeriesSplit
y = np.arange(12) # twelve months, in order
for i, (train_idx, test_idx) in enumerate(TimeSeriesSplit(n_splits=3).split(y), 1):
print(f"fold {i} train={list(train_idx)} test={list(test_idx)}")fold 1 train=[0, 1, 2] test=[3, 4, 5] fold 2 train=[0, 1, 2, 3, 4, 5] test=[6, 7, 8] fold 3 train=[0, 1, 2, 3, 4, 5, 6, 7, 8] test=[9, 10, 11]
Every test index is larger than every training index in the same fold. The training window grows as the origin moves forward. This is called rolling-origin evaluation, and it is the standard way to score a forecaster. Evaluating a forecast goes into it properly.
Common mistakes
Calling train_test_split without shuffle=False. The default is shuffle=True. That one default has broken more forecasting projects than any algorithm choice. If you use it at all, pass shuffle=False.
Scaling before splitting. Fitting a StandardScaler on the whole series computes a mean that includes future values, then bakes it into your training features. Fit the scaler on training data only, then apply it to the test data.
A string index instead of a datetime index. If dates load as text, "2025-1-9" sorts before "2025-10-01". Convert on read with pd.read_csv(path, parse_dates=["date"], index_col="date"), then check df.index.dtype says datetime64[ns].
Silent duplicate or missing timestamps. Run df.index.duplicated().sum() and df.index.to_series().diff().value_counts(). A single duplicated hour from a clock change will corrupt every rolling window downstream.
Shuffling inside cross-validation. cross_val_score uses KFold, which shuffles by default for regression in many setups. Pass cv=TimeSeriesSplit(5) explicitly.
Try it yourself
Take the visits series from the first script. Compute visits.autocorr(7) — the correlation with the value seven days earlier.
Predict whether it will be higher or lower than autocorr(1) before you run it. The weekend bump repeats every seven days, so think about what that does. Then explain the result to yourself in one sentence.
What to learn next
- Trend, seasonality and noise — the three parts every series is made of.
- Train, test and validation splits — the ordinary case, so you can see exactly what changes here.
- Evaluating a forecast — how to score one without fooling yourself.
Researcher — Mathematics and papers.
Formal setup
A univariate time series is a realisation of a stochastic process ${Y_t}_{t \in \mathbb{Z}}$ observed at $t = 1, \dots, T$.
- $Y_t$ — the random variable at time index $t$.
- $y_t$ — the observed value, one realisation of $Y_t$.
- $T$ — the number of observations available.
- $h$ — the forecast horizon, the number of steps ahead being predicted.
The forecasting problem is to estimate the conditional distribution
$$ p!\left(Y_{T+h} \mid Y_{1:T} = y_{1:T}\right) $$
and the point forecast $\hat{y}{T+h|T}$ is a functional of it. Under squared-error loss the optimal point forecast is the conditional mean $\mathbb{E}[Y{T+h} \mid \mathcal{F}_T]$, where $\mathcal{F}_T$ is the sigma-algebra generated by observations up to $T$. Under absolute-error loss it is the conditional median. These differ for skewed series, which is why a metric choice is a modelling decision, not a reporting decision.
Why i.i.d. machinery does not transfer
Standard supervised learning theory assumes the training sample is independent and identically distributed. Time series violate both halves.
Dependence. $\mathrm{Cov}(Y_t, Y_{t-k}) \neq 0$ for small $k$ in almost every real series. The autocovariance function
$$ \gamma(t, t-k) = \mathbb{E}!\left[(Y_t - \mu_t)(Y_{t-k} - \mu_{t-k})\right] $$
is exactly the structure the model is meant to exploit, so it cannot be assumed away.
Non-identical distribution. $\mu_t$ and $\gamma$ typically vary with $t$. This is the subject of stationarity, and stationarity assumptions are what allow $\gamma(t, t-k)$ to collapse to a function $\gamma(k)$ of lag alone.
The consequence for validation is not a matter of taste. K-fold cross-validation estimates generalisation error under exchangeability. Time series are not exchangeable, so the estimator is biased, and the bias is optimistic. Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, quantifies this and gives the conditions under which a blocked variant recovers validity.
Rolling-origin evaluation
Tashman (2000), Out-of-sample tests of forecasting accuracy, formalises the standard protocol. For origins $t_0, t_0+1, \dots, T-h$ produce $\hat{y}{t+h|t}$ using only $y{1:t}$, then aggregate the errors $e_{t+h} = y_{t+h} - \hat{y}_{t+h|t}$.
Two design choices matter and are frequently conflated:
- Expanding window — train on $y_{1:t}$, which grows. Preferred when the data-generating process is stable and data is scarce.
- Sliding window — train on $y_{t-w+1:t}$ for fixed $w$. Preferred under regime change, at the cost of variance.
A purge gap between train and test is required whenever features have look-back windows that would otherwise straddle the boundary, and whenever the label itself spans time. López de Prado (2018), Advances in Financial Machine Learning, chapter 7, gives the purging and embargo construction for overlapping labels.
Decomposing the error floor
Write the $h$-step error as
$$ \mathbb{E}!\left[(Y_{T+h} - \hat{y}{T+h|T})^2\right] = \underbrace{\mathrm{Var}(Y{T+h} \mid \mathcal{F}T)}{\text{irreducible}} + \underbrace{\left(\mathbb{E}[Y_{T+h}\mid\mathcal{F}T] - \hat{y}{T+h|T}\right)^2}_{\text{reducible}} $$
For a random walk $Y_t = Y_{t-1} + \varepsilon_t$ with $\varepsilon_t \sim \mathrm{WN}(0, \sigma^2)$, the conditional mean is $y_T$ for every $h$, and the irreducible term is $h\sigma^2$. Forecast uncertainty grows linearly in the horizon and no model can shrink it. This is the formal version of the honesty warning in the beginner block, and it is why the naive forecast is a competitive benchmark on financial series rather than a straw man.
Benchmarks that must be beaten
Any proposed model should be reported against these, on the same rolling origins:
| Benchmark | Forecast | When it is hard to beat |
|---|---|---|
| Naive | $\hat{y}_{T+h\mid T} = y_T$ | Random-walk-like series |
| Seasonal naive | $\hat{y}{T+h\mid T} = y{T+h-m}$ | Strong fixed seasonality, period $m$ |
| Drift | $y_T + h\frac{y_T - y_1}{T-1}$ | Steady trend, low noise |
| Historical mean | $\bar{y}$ | Mean-reverting, weak dependence |
The M-competitions (Makridakis et al., 1982 through M5 in 2020) repeatedly found that simple methods and combinations of them are strong. M4 (2018) was won by a hybrid of exponential smoothing and a recurrent network; the pure-machine-learning entries below it underperformed statistical baselines. Report the benchmark or the result is uninterpretable.
Reading
- Hyndman and Athanasopoulos, Forecasting: Principles and Practice, 3rd edition — free at otexts.com/fpp3. The standard modern reference.
- Box, Jenkins, Reinsel and Ljung, Time Series Analysis: Forecasting and Control, 5th edition, 2015.
- Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, Information Sciences 191.
- Makridakis, Spiliotis and Assimakopoulos (2018), The M4 Competition: Results, findings, conclusion and way forward.
- Tashman (2000), Out-of-sample tests of forecasting accuracy: an analysis and review, International Journal of Forecasting 16.
What to learn next
- Trend, seasonality and noise — the three parts every series is made of.
- Train, test and validation splits — the ordinary case, so you can see exactly what changes here.
- Evaluating a forecast — how to score one without fooling yourself.