Back to the shelf
Cover of Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance

Read

Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance

Bailey et al.

Published in the Notices of the AMS in 2014, this paper by Bailey, Borwein, Lopez de Prado, and Zhu demonstrates mathematically that a relatively small number of trials is sufficient to identify an investment strategy with a spuriously high in-sample Sharpe ratio. The paper introduces the Minimum Backtest Length (MinBTL) concept, showing that the expected maximum in-sample Sharpe ratio grows with the number of trials attempted while out-of-sample performance remains near zero, and argues that failure to disclose the number of trials makes any backtest impossible to evaluate fairly.

Key takeaways

  • Backtest overfitting arises when parameters are selected to maximize in-sample performance across many trials, causing the chosen strategy to exploit noise rather than genuine signal, and leading to out-of-sample performance that regresses toward zero regardless of the in-sample Sharpe ratio reported.
  • The paper derives Minimum Backtest Length (MinBTL), the number of years of backtest data required to ensure that a selected strategy with a given in-sample Sharpe ratio is not merely a product of overfitting across N independent trials, with MinBTL growing approximately as 2 ln(N) divided by the square of the expected maximum Sharpe ratio.
  • Proposition 1 establishes that the expected maximum Sharpe ratio among N independent zero-true-SR strategies grows with N, approximated by a formula involving the Euler-Mascheroni constant and the inverse normal CDF, implying that more trials mechanically inflate the best observed in-sample result.
  • When there are no compensation effects (i.e., the financial process has no mean-reversion memory), overfitting does not induce negative out-of-sample performance on average, but the selected strategy still delivers near-zero OOS returns, making any high in-sample Sharpe ratio essentially uninformative about future returns.
  • Model complexity compounds the overfitting risk because each additional binary parameter doubles the number of possible configurations, making it straightforward to achieve high in-sample Sharpe ratios through brute-force search even without any genuine predictive insight.
  • A key practical implication is that any backtest submitted to investors or journals without disclosure of the number of trials N cannot be properly assessed for overfitting risk, and the authors argue that demanding this disclosure should be a standard requirement for evaluating investment research.

Reflections

No notes on this one yet. I add reflections as I finish or revisit a book.