Data Leakage and Results That Are Too Good
Using the future to predict the past
Shuffling time-stamped data lets your model peek at the future, so forecasting models must be evaluated strictly forward in time.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
If your data has timestamps, a random split lets the model study Tuesday while being examined on Monday.
Imagine predicting yesterday's cricket score while holding today's newspaper. The paper prints the result. You would be right every time, and it would prove nothing about your ability to predict tomorrow's match.
That is temporal leakage — the future sneaking into an exam about the past. In real life, prediction only ever runs one way: forward. An evaluation that mixes up time is testing a superpower your deployed model will never have.
Why it exists
The standard recipe — shuffle everything, split randomly — was built for data where rows are independent, like separate patients or separate photos. Time-stamped rows are not independent. Monday's sales resemble Tuesday's sales.
Shuffle a year of daily data, and your training set contains days from every month. When the model predicts a test day from 15 March, it has trained on 14 March and 16 March. It fills the gap between two known neighbours. Real forecasting never has the later neighbour.
How it works
random split (leaky):
train: Jan Feb [Mar] Apr [May] Jun ... ← knows both neighbours
test : [Mar] [May] ← gaps get filled, not forecast
forward split (honest):
train: Jan Feb Mar Apr ─────┐
test : └→ May Jun ← real forecastingThe honest split always draws a line in time. Training lives entirely before the line. Testing lives entirely after it.
A real example you have seen
Weather apps forecast tomorrow using data up to today — that is the honest task, and it is why they are sometimes wrong. Now think of stock-tip videos claiming a method that "would have earned lakhs last year". Many such methods secretly use each day's own future — buy signals computed from full-year statistics. Backwards-looking accuracy, forwards-facing uselessness.
Remember this
- Time-stamped data must be split by time, never shuffled.
- A shuffled evaluation tests gap-filling; deployment demands forecasting.
- Features count too: anything computed from later days poisons earlier rows.
What to learn next
- The same person in train and test — identity, the other axis your split must respect.
- Forecast evaluation — horizons, rolling origins and error metrics for honest forecasting.
- Stationarity — why the trend broke the forest, in proper vocabulary.
Developer — Code and libraries.
Setup
pip install scikit-learn pandasVerified with scikit-learn 1.7.2, pandas 2.2.3, numpy 1.26.4, CPU. Seeded, so your numbers should match.
The same model, two evaluations, opposite verdicts
Daily demand with a rising trend. Features are honest — yesterday's value and last week's value. Only the evaluation differs.
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import KFold, TimeSeriesSplit, cross_val_score
rng = np.random.default_rng(4)
n = 300
day = np.arange(n)
demand = 50 + 0.3 * day + 10 * np.sin(day / 7) + rng.normal(0, 3, n)
df = pd.DataFrame({"yesterday": np.roll(demand, 1),
"last_week": np.roll(demand, 7),
"day": day})
X, y = df.iloc[7:], demand[7:]
model = RandomForestRegressor(random_state=4)
shuffled = cross_val_score(model, X, y,
cv=KFold(5, shuffle=True, random_state=4), scoring="r2")
ordered = cross_val_score(model, X, y, cv=TimeSeriesSplit(5), scoring="r2")
print("shuffled folds R2:", shuffled.round(2), " mean", round(shuffled.mean(), 3))
print("time-ordered R2:", ordered.round(2), " mean", round(ordered.mean(), 3))shuffled folds R2: [0.97 0.98 0.97 0.96 0.97] mean 0.971 time-ordered R2: [ 0.17 0.09 -0.19 -0.66 -0.87] mean -0.291
Shuffled evaluation: a triumphant 0.97. Forward evaluation: negative — worse than predicting the average. Same model, same features, same data.
The walkthrough
R² of 0.97 versus −0.29 is not a small correction. R² is the share of variation explained; negative means the model's guesses are further from the truth than a flat average would be. The shuffled number said "ship it". The honest number said "this model cannot forecast at all".
Why so brutal? The series trends upward, and a random forest cannot predict values above anything it saw in training. In shuffled folds, training data surrounds every test day, so that weakness stays invisible. In forward folds, every test day sits above the training range, and the model's predictions saturate. The forward evaluation did not create this flaw — it revealed a flaw that shuffling concealed.
TimeSeriesSplit builds growing training windows with the test block always after them. Each fold is a rehearsal of deployment: train on the past, predict the next stretch. Details of horizon and gap handling live in forecast evaluation.
Features leak time too
Splitting forward is necessary and not sufficient. Every feature must also be computable from the past alone:
- A rolling average must end at yesterday, not be centred on today.
- "Customer's total purchases" must be totals as of that row's date.
- Normalising a series by its full-period average injects the future into every day.
The target leakage question — was this known at prediction time? — becomes: was this computable on that morning?
Common mistakes
Shuffling because "the model needs varied batches". Shuffle training batches if you like — never shuffle the train/test boundary across time.
Deduplicating, then still interpolating. Consecutive days are not duplicates, yet they carry nearly identical information. Add a gap between train and test (drop a few boundary days) when your features contain lags.
Grouping by month but ignoring order. Splitting "Jan–Oct train, Nov–Dec test" is right. Splitting "even months train, odd months test" reintroduces the sandwich.
Trusting one forward split. A single line in time is one rehearsal. TimeSeriesSplit gives several; report them all, as the honest output above does — including the ugly folds.
Try it yourself
Remove the trend: set demand = 50 + 10 * np.sin(day / 7) + rng.normal(0, 3, n). Rerun both evaluations and explain why the two verdicts move closer together. Then add the trend back and log-transform the target — how much does honest R² recover?
What to learn next
- The same person in train and test — identity, the other axis your split must respect.
- Forecast evaluation — horizons, rolling origins and error metrics for honest forecasting.
- Stationarity — why the trend broke the forest, in proper vocabulary.
Researcher — Mathematics and papers.
Filtrations and legitimate features
Index observations by time, and let $\mathcal{F}_t$ denote the information set (sigma-algebra) generated by everything observed up to time $t$. A forecasting evaluation is valid iff, for every test pair $(x_t, y_{t+h})$, the feature vector $x_t$ is $\mathcal{F}_t$-measurable and the model parameters are functions of ${(x_s, y_{s+h}) : s + h \leq t}$ only.
Where:
- $\mathcal{F}_t$ — the filtration: information available at time $t$.
- $h$ — the forecast horizon.
- $x_t, y_{t+h}$ — features at $t$, outcome $h$ steps ahead.
Random K-fold violates the parameter condition; future-computed features violate the measurability condition. Both are the same crime against the filtration.
When is K-fold on time series acceptable?
Bergmeir and Benítez (2012), On the use of cross-validation for time series predictor evaluation, Information Sciences, show K-fold can be approximately valid for purely autoregressive models with i.i.d. errors — the specific case where the model's inputs already summarise the past. Bergmeir, Hyndman and Koo (2018) extend this. The practical reading: the exception is narrow, is broken by trends, exogenous features, or misspecification, and the demo above is precisely such a breakage. Out-of-sample forward evaluation is the default that never betrays you; Tashman (2000), IJF, reviews the rolling-origin methodology.
Beyond the split: subtler temporal leaks
- Look-ahead in labels: "churned within 90 days" computed for rows near the dataset's end is right-censored; including them biases the label itself.
- Survivorship: building a stock dataset from today's index constituents removes companies that failed — the future filtered the past. Standard in quantitative finance folklore; the fix is point-in-time universes.
- Global statistics: detrending, standardising, or imputing with full-series statistics — the temporal case of preprocessing leakage.
- Embargoes: with overlapping labels (returns over $k$ days), adjacent train/test samples share outcome windows. Purged and embargoed CV — de Prado (2018), Advances in Financial Machine Learning — drops a buffer around each test block.
Backtest overfitting
Repeatedly tuning strategies against one historical period is test-set overfitting in temporal costume. Bailey et al. (2014), The Probability of Backtest Overfitting, and de Prado's deflated Sharpe ratio quantify how many strategy variants a backtest can absorb before its best result is expected noise. The general phenomenon and its bounds appear later in this section.
What to learn next
- The same person in train and test — identity, the other axis your split must respect.
- Forecast evaluation — horizons, rolling origins and error metrics for honest forecasting.
- Stationarity — why the trend broke the forest, in proper vocabulary.