Backtesting without fooling yourself
A backtest that tries enough parameter combinations will find a "winning" strategy in pure random noise — the fix is testing that winner on data it has never seen.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Backtesting means testing a trading idea on past data before risking real money. It is astonishingly easy to fool yourself with it.
Imagine a leaked exam paper from last year, with worked answers. Practice on it enough and you will score full marks. That score tells you nothing about how well you know the subject. You memorised the answers instead of learning the material. A real, unseen paper next week reveals the truth.
A backtest run carelessly is the same leaked paper. History is fixed and already known. Try enough strategies against it, and one will look brilliant by pure chance. Not because it found a real edge.
Why it exists
Before backtesting, traders tested ideas by trading real money and watching what happened. That was an expensive, slow way to learn. Backtesting lets you replay history instead, cheaply and quickly, to see how an idea would have performed.
That speed is also the trap. A backtest is a fixed, finite chunk of the past. Try 5 versions of a strategy against it, and one might do well by luck. Try 500 versions, and one doing brilliantly by luck alone is close to guaranteed. This is the same problem as flipping 500 coins. Someone is bound to be amazed that one flipped heads ten times in a row. With enough tries, someone always does.
Statisticians call this multiple comparisons. You test many things against the same data. You report only the best one, without accounting for how many you tried.
How it works
1,000 strategy variations
|
v
Test all 1,000 against the SAME 3 years of history
|
v
Pick the single best-looking one
|
v
"We found a strategy with a Sharpe ratio of 2.1!"
|
v
Run it on NEW, never-tested data -> performance collapsesA Sharpe ratio is a common score for a strategy. It is the average return divided by how bumpy the returns are. Higher generally means a smoother, more reliable profit — in theory.
The honest fix is simple to state and hard to practice. Split history into two parts. Search on one part. Never touch the other part until you have picked your one final strategy. That untouched part is your real answer key.
A real example you have seen
Financial media occasionally profiles a trader or a fund with an eye-popping backtested return. Some of those strategies go on to perform in the real world. Many quietly stop being mentioned a year later. The difference almost always comes down to one thing. Was the strategy validated on untouched data, with realistic costs, before anyone trusted it?
The honest part
A strategy can beat out-of-sample data honestly, and still fail once real money touches it. A backtest usually assumes trades happen instantly, at the exact price shown, with no fees. Real markets charge a spread. They take time to fill an order. They move the price when a large order lands. This gap between the backtest and the real world is called slippage. It eats returns that looked completely real on paper.
Remember this
- Trying many strategies against the same fixed history will eventually produce a "winner" by chance alone.
- Only trust a backtest number measured on data the search process never touched.
- Real trading costs — spread, fees, slippage — must be subtracted before a backtest number means anything.
- A promising backtest is a reason to paper-trade and get a risk team's sign-off next, never a reason to risk real capital straight away.
What to learn next
- Credit scoring — a different finance model, where the same discipline about honest evaluation applies.
- Overfitting and underfitting — the general version of "found a pattern in noise."
- Too good to be true — how to recognise this failure in any ML result, not only in finance.
Developer — Code and libraries.
Setup
pip install numpy pandasMinimal runnable code
We generate pure random price noise — by construction, there is no real trading edge in it. Then we search hundreds of parameter pairs for a moving-average crossover strategy and keep the best one. Watch what happens when that "winner" meets fresh data.
import numpy as np
import pandas as pd
rng = np.random.default_rng(14)
def make_prices(n, rng):
returns = rng.normal(0, 0.01, size=n) # pure noise, no real trend, on purpose
return pd.Series(100 * np.cumprod(1 + returns))
def crossover_sharpe(prices, fast, slow):
fast_ma = prices.rolling(fast).mean()
slow_ma = prices.rolling(slow).mean()
position = (fast_ma > slow_ma).astype(int) # 1 = holding, 0 = flat
daily_return = prices.pct_change()
strategy_return = position.shift(1) * daily_return # trade on the NEXT day's move
mean, std = strategy_return.mean(), strategy_return.std()
return 0.0 if std == 0 else (mean / std) * np.sqrt(252)
in_sample = make_prices(500, rng)
# Search 19 x 40 = 760 parameter pairs against the SAME 500 days
best = max(
((f, s, crossover_sharpe(in_sample, f, s)) for f in range(2, 20) for s in range(20, 60)),
key=lambda t: t[2],
)
print(f"tried {19*40} (fast, slow) pairs on the same 500 days")
print(f"best in-sample pair: fast={best[0]}, slow={best[1]}, sharpe={best[2]:.2f}")
# Test that exact winning pair on FRESH random data it never saw
out_of_sample = make_prices(500, rng)
oos_sharpe = crossover_sharpe(out_of_sample, best[0], best[1])
print(f"same exact pair, on fresh random data: sharpe={oos_sharpe:.2f}")tried 760 (fast, slow) pairs on the same 500 days best in-sample pair: fast=17, slow=23, sharpe=1.85 same exact pair, on fresh random data: sharpe=-0.33
What actually happened
The price data is generated by rng.normal(0, 0.01, ...) — plain random noise with zero built-in trend. There is no real pattern here, by construction.
And yet, searching 760 parameter pairs against that noise found one with an in-sample Sharpe ratio of 1.85 — a number that would excite most traders. That is multiple comparisons at work: with 760 attempts, some pair was always going to look good purely by chance.
The same exact pair, tested on a fresh batch of random data it had never seen, scored -0.33. The "edge" evaporated, because it was never real. This is what an overfit backtest looks like from the inside.
position.shift(1)matters. Without the shift, the strategy would use today's moving average to trade on today's return — which requires knowing today's close before the day has happened. Shifting by one day is what makes this a fair simulation.crossover_sharpereturns0.0whenstd == 0, guarding against a strategy that never traded (all zero returns, undefined Sharpe).- The random seed
14is fixed so this exact result reproduces. Try other seeds — the specific numbers change, but a large in-sample-to-out-of-sample gap keeps showing up, because the mechanism is structural, not a one-off unlucky draw.
Common mistakes
Reporting only the best backtest, not how many were tried. A Sharpe of 1.85 out of 1 attempt is very different information from a Sharpe of 1.85 out of 760 attempts.
Re-using the "out-of-sample" data after seeing it once. The moment you look at out-of-sample results and go back to tweak the strategy, that data has quietly become in-sample. You need a third, truly untouched slice for the final check.
Ignoring costs. Add even a small transaction cost per trade to the code above and watch how many "winning" strategies turn negative — moving-average crossovers trade often, and costs compound fast.
Testing on one historical period only. A strategy validated on 2015-2019 may quietly rely on conditions specific to that stretch. Walk-forward validation — repeatedly training on one window and testing on the next — is closer to how the strategy will actually be used.
Try it yourself
Add a fixed cost of 0.0005 (5 basis points) to strategy_return on every day the position changes from the day before. Re-run the search. Watch how many of the 760 pairs now show a negative in-sample Sharpe, once trading is not free.
What to learn next
- Credit scoring — a different model type, same discipline about what counts as a fair test.
- Too good to be true — general warning signs of an inflated result.
- K-fold cross-validation — the general-purpose tool this lesson's out-of-sample check is a cousin of.
Researcher — Mathematics and papers.
The multiple comparisons problem, formally
Suppose N independent strategies are tested against the same historical sample, and none has any true edge — each has a true Sharpe ratio of zero. Sample noise still produces a range of observed Sharpe ratios around zero. The maximum of N such draws grows with N, even though the true value never moves.
Bailey, Borwein, Lopez de Prado & Zhu (2014) give a closed-form expected value for the maximum Sharpe ratio observed under pure noise, as a function of the number of trials N and the variance of the Sharpe ratio estimator, and use it to define the Deflated Sharpe Ratio (DSR) — a corrected significance test that asks: given how many strategies were tried, is this Sharpe ratio still surprising?
DSR = Phi( ( SR_hat - SR_0 ) / sigma(SR_hat) )SR_hat— the observed Sharpe ratio of the selected strategySR_0— the expected maximum Sharpe ratio under the null hypothesis of no true skill, givenNtrialssigma(SR_hat)— the standard deviation of the Sharpe estimator, itself a function of sample skewness and kurtosisPhi— the standard normal CDF, turning the standardised gap into a probability
A DSR near 0.5 means the observed performance is exactly what pure luck, at that trial count, would produce.
Walk-forward validation
Rather than a single train/test split, walk-forward evaluation refits or reselects on an expanding or rolling window and tests only on the immediately following, never-seen period, then rolls forward:
Fit: [ 2015 ---- 2017 ] Test: [ 2018 ]
Fit: [ 2015 -------- 2018 ] Test: [ 2019 ]
Fit: [ 2015 ------------ 2019 ] Test: [ 2020 ]The final reported performance is the concatenation of the test-only segments, which is closer to a live production process — the model never trades on a period it selected its parameters using.
Combinatorial purged cross-validation
Standard k-fold cross-validation on financial time series leaks information across folds, because adjacent samples are temporally correlated (a return today is not independent of a return yesterday). Lopez de Prado (2018) proposes purging — removing training samples whose label windows overlap the test window — and embargoing — removing a further buffer immediately after the test window — before applying combinatorial cross-validation across many train/test splits.
Cost
Deflated Sharpe correction requires estimating the effective number of independent trials N, which is nontrivial when tested strategies are correlated with each other (most parameter sweeps are: fast=16 and fast=17 produce highly correlated strategies, so they are not 760 truly independent trials). Underestimating correlation between trials overstates N and makes the correction too lenient; overestimating it makes the correction too harsh.
Key references
- Bailey, D. H., Borwein, J., Lopez de Prado, M., & Zhu, Q. J. (2014). The Probability of Backtest Overfitting. Journal of Computational Finance.
- Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley — chapters 7 and 11 cover purged/embargoed cross-validation and the deflated Sharpe ratio in full detail.
- Harvey, C. R., Liu, Y., & Zhu, H. (2016). ...and the Cross-Section of Expected Returns. Review of Financial Studies — applies multiple-testing corrections to the published factor-investing literature itself.
Current state
Harvey, Liu & Zhu's survey of the published academic factor literature found that most claimed "anomalies" would not survive a proper multiple-testing correction — the same trap this lesson demonstrates in miniature applies to peer-reviewed finance research, not only to retail backtests. Any backtested finance-ML result presented without a trial count, a held-out period, and realistic costs should be treated as a starting hypothesis, not a validated conclusion.
What to learn next
- K-fold cross-validation — the general method purged cross-validation is adapted from.
- Too good to be true — recognising inflated results across any ML domain.
- Model risk management — the governance process that is supposed to catch exactly this before deployment.