# Overfitting

Also written over-fitting, backtest overfitting, curve-fitting, curve fitting.

*https://stockmarketstack.com/glossary/overfitting · next to Backtesting Frameworks & Algo Trading Libraries*

**Definition:** Choosing a trading rule on the same history it is then scored on, so that the backtest measures the search as much as the rule. Every configuration tried is another draw, and the best of many draws looks good by chance alone, so a high in-sample figure is evidence of an edge only once you know how many alternatives were tried before it. That count is the number a backtest almost never reports.

## How it works

In statistics the word means a model that has learned the noise in its sample along with the signal.
In this category it has a narrower and more useful meaning, which the literature calls backtest
overfitting: the rule looks good on a history because it was *selected* on that history, out of
several that were tried.

The mechanism is selection, not complexity. Take one price series and test a hundred variants of a
rule against it — different lookbacks, different thresholds, different exits. Even if none of them
has any real edge, their backtested results scatter, and the best of the hundred sits at the top of
that scatter. Report only that one and it looks like a discovery. Bailey, Borwein, López de Prado
and Zhu put it in one sentence: the higher the number of configurations tried, the greater the
probability that the backtest is overfit — and because analysts rarely report the number of
configurations tried, investors cannot evaluate the degree of overfitting in most claims.

The same paper makes the scale concrete. With only five years of data, no more than 45 independent
configurations should be tried, or the search is almost guaranteed to produce a strategy with an
annualised in-sample Sharpe ratio of 1 and an expected out-of-sample Sharpe ratio of zero. After
seven independent configurations on a two-year backtest, the same holds. That is not millions of
combinations from a generator; it is a few dozen genuinely different ideas tried against the same
five years. (Variants that differ only slightly are correlated and count as fewer independent
trials, which is why the paper calls its independence assumption conservative.)

Three consequences follow, and each one is a thing a reader gets wrong by reading the word loosely:

**It is a property of the search, not of the rule.** The same final rule is overfit or not
depending on how many alternatives were discarded on the way to it. Two people can publish an
identical strategy with identical backtests, and only the one who found it on the first try has
shown anything.

**A hold-out does not fix it by itself.** Setting data aside helps only while it has not been
looked at. The paper's point about the hold-out method is that it does not take into account the
number of trials attempted before selecting a model, so it cannot say whether the backtest is
representative. Check the out-of-sample result, adjust, check again, and the hold-out has become
part of the search.

**It is separate from bad data.** A search over a [survivorship-biased](https://stockmarketstack.com/glossary/survivorship-bias)
universe or restated fundamentals is wrong for a different reason, and fixing one does nothing for
the other.

## What the defences measure

Two published statistics turn "how many did you try" into a number, and both need the losers as
input — which is the practical point.

**The deflated Sharpe ratio.** Bailey and López de Prado's correction for selection bias under
multiple testing and for returns that are not normally distributed. As the number of independent
trials grows, so does the Sharpe ratio the best of them is expected to show by chance; the deflated
figure measures the winner against that raised hurdle rather than against zero. It needs the number
of trials and how widely their results varied, not only the winner's figure.

**The probability of backtest overfitting (PBO).** A cross-validation method from Bailey, Borwein,
López de Prado and Zhu that asks whether the selection process itself was conducive to overfitting,
in the sense that the strategies it selects tend to underperform the median of the trials out of
sample. It is non-parametric and works on any performance statistic, but by the authors' own
description it requires a large amount of information: the full record of every trial, not a
summary.

Walk-forward testing, Monte Carlo resampling and parameter-permutation tests are the other family:
they re-test the chosen rule under disturbance and reject the fragile ones. They are worth having and
they answer a different question — whether *this* rule is robust, not whether the search that found
it would have found something equally good in pure noise. See [walk-forward](https://stockmarketstack.com/glossary/walk-forward)
and [Monte Carlo simulation](https://stockmarketstack.com/glossary/monte-carlo-simulation) for what each actually resamples.

## Why it matters here

Every product in [backtesting frameworks](https://stockmarketstack.com/categories/backtesting-frameworks) produces a backtest,
and a backtest is exactly as informative as the count of attempts behind it. None of them prints
that count next to the equity curve by default.

It is sharpest for the strategy generators, where the search *is* the product.
[StrategyQuant X](https://stockmarketstack.com/tools/strategyquant-x) sells Monte Carlo, parameter permutation and walk-forward
tests above its Starter edition to shrink the problem, and its own documentation puts survival at
roughly one in a thousand profitable candidates. [Build Alpha](https://stockmarketstack.com/tools/build-alpha)'s own FAQ does
the arithmetic on an older signal library at around 2.6 × 10¹³ combinations, and its robustness
tests run on the same history the search already mined. [Adaptrade Builder](https://stockmarketstack.com/tools/adaptrade-builder)
holds back validation and test segments during the build and lets you penalise complexity. All
three cards say the same thing plainly: the tests reduce the risk and do not remove it.

The two statistics above are rare as shipped features. [ffn](https://stockmarketstack.com/tools/ffn), a free Python performance
library, includes both — `calc_deflated_sharpe_ratio` and `calc_prob_backtest_overfitting` — and its
documentation insists that the trial list cover all trials evaluated, not only the ones kept. That
requirement is the whole term in one line. [tradingview-mcp](https://stockmarketstack.com/tools/tradingview-mcp), an MCP server,
returns a walk-forward overfitting verdict on its fixed strategies; a verdict computed on nine
pre-chosen rules is a different claim from one computed on the search that chose them.

The question to take to any backtest, a vendor's or your own: how many configurations were tried on
this data before this one was shown, and is that number in the calculation?

## Where you will meet this

- [StrategyQuant X](https://stockmarketstack.com/tools/strategyquant-x.md)
- [Build Alpha](https://stockmarketstack.com/tools/build-alpha.md)
- [Adaptrade Builder](https://stockmarketstack.com/tools/adaptrade-builder.md)
- [ffn](https://stockmarketstack.com/tools/ffn.md)
- [tradingview-mcp](https://stockmarketstack.com/tools/tradingview-mcp.md)

## FAQ

### If a strategy passes an out-of-sample test, is it not overfit?

Only if that out-of-sample data was used once. Every time a result on the held-out period sends you back to change the rule, the held-out period has joined the search, and the test no longer tells you anything the in-sample result did not. What matters is how many configurations were tried on all of the data you have looked at, including the part you called out-of-sample.

## Sources

1. [Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance](https://www.davidhbailey.com/dhbpapers/backtest-pseudo.pdf) — Bailey, Borwein, López de Prado and Zhu (authors' pre-publication copy, peer-reviewed for the Notices of the AMS), 2014-04-01. A mathematical result about multiple testing, not a rule or a price; later work by the same authors builds on it rather than revising it.
2. [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://www.davidhbailey.com/dhbpapers/deflated-sharpe.pdf) — Bailey and López de Prado (authors' pre-publication copy, marked forthcoming in the Journal of Portfolio Management), 2014-07-31. A statistical method, not a rule or a price; the calc_deflated_sharpe_ratio function in ffn cites this article as the method it implements.
3. [ffn 1.2.2, ffn/core.py: docstrings of calc_deflated_sharpe_ratio and calc_prob_backtest_overfitting](https://github.com/pmorissette/ffn/blob/master/ffn/core.py) — ffn (open-source project), read 2026-09-27

*Last updated 2026-09-27. A reference page, corrected in place — not a dated post.*
