Test 200 trading ideas that have no edge at all, and about ten of them will pass a standard significance test. Each one will come with a respectable backtest and a p-value under 5%. This is the single most common way quant research goes wrong, and it doesn't need anyone to make a mistake in their code. This lesson shows how big the effect is and what to do about it.
| Symbol | Meaning |
|---|---|
| SR | Annualised Sharpe ratio: mean return over volatility, per year |
| Length of the test, in years | |
| The t-statistic for "the true mean return is zero" | |
| The significance level, the false positive rate you accept (5%) | |
| The number of strategies (or variations) tested | |
| The inverse of the standard normal cumulative distribution function |
Testing one strategy
To ask whether a strategy's average daily return is really above zero, compare it with its standard error. For daily returns over years, the t-statistic works out to the Sharpe ratio times the square root of the number of years:
At the 5% level, one-sided, you call the result significant if . With one year of data that means a Sharpe ratio above 1.645. To show that a strategy with a true Sharpe of 1 is better than nothing, you need , about 2.7 years of data, and that's only to get the expected result over the line half the time.
A p-value below 5% means: if the strategy had no edge, results this good would happen less than 5% of the time. It does not mean there is a 95% chance the strategy works. That second number needs a prior, as the Bayesian updating lesson showed.
Testing many
Each test of a no-edge strategy has a 5% chance of a false positive. Run independent tests and the chance that at least one passes is
For that's 64%. For it's 99.4%. A researcher who tries a hundred ideas is almost certain to find something "significant", even in a world where nothing works.
Every strategy below is pure noise: daily returns with zero mean. Each is tested on one year of data (across) and then traded for a second year (up). Watch the white dots, the ones that passed:
- Strategies tried
- 0
- Passed the test
- 0
- Best Sharpe ratio
- –
- Next year, passers' average
- 0.00
They pass at about the expected 5% rate, and their next-year Sharpe ratios average about zero: in the second year, the winners are indistinguishable from the losers. Switch to the corrected threshold and the bar moves right as more strategies are tried. Almost nothing clears it, which is the right answer, because there is nothing to find.
Key idea. A significance test controls the false positive rate per test. Run enough tests and false positives are guaranteed. What matters is how many things were tried, not just how good the best one looks.
How good does luck look?
The best of no-edge strategies has an expected Sharpe ratio equal to the expected maximum of standard normals (for one year of data):
| Strategies tried | Expected best Sharpe ratio, from luck alone |
|---|---|
| 10 | 1.54 |
| 100 | 2.51 |
| 500 | 3.04 |
A Sharpe ratio of 2.5 would be an excellent real strategy. Here it is what you should expect from trying a hundred bad ones. Without knowing how many ideas were tried, a backtested Sharpe ratio can't be judged at all.
Corrections
The simplest fix is the Bonferroni correction: to keep the chance of any false positive at 5% across tests, test each one at . The threshold for the t-statistic becomes :
| Tests | Threshold |
|---|---|
| 1 | 1.64 |
| 20 | 2.81 |
| 100 | 3.29 |
| 500 | 3.72 |
Bonferroni is conservative: when the strategies are correlated, as variations of one idea usually are, it demands more than it needs to. Two refinements are worth knowing by name. The Holm procedure is uniformly better than Bonferroni and just as easy. The false discovery rate (Benjamini–Hochberg) controls the expected share of passing results that are false, which suits research where you will screen many candidates and accept a few mistakes. In strategy research, the same idea shows up as adjusting the Sharpe ratio itself for the number of trials, as in the "deflated Sharpe ratio".
Key idea. Raise the bar with the number of tries. With 100 independent tries, a t-statistic of about 3.3 is the new 1.64.
The tries you don't count
The hardest part is counting honestly. Every lookback window, entry threshold, universe of stocks, start date and "small fix" after seeing a backtest is another test. A researcher who tried one idea in fifty variations has run fifty tests, even if only the last one gets written up. This is often called the garden of forking paths.
Two habits help more than any formula. Keep a log of everything you tried, including the failures, so that is known. And keep some data you never look at until the very end: a genuine out-of-sample test, used once.
In code
from statistics import NormalDist
import numpy as np
rng = np.random.default_rng(7)
strategies, days = 200, 252
def sharpe(returns):
# Annualised Sharpe ratio of each row of daily returns.
return returns.mean(axis=1) / returns.std(axis=1, ddof=1) * np.sqrt(252)
# 200 strategies with no edge: daily returns are pure noise with 1% volatility.
year1 = rng.normal(0.0, 0.01, size=(strategies, days))
year2 = rng.normal(0.0, 0.01, size=(strategies, days))
tested, next_year = sharpe(year1), sharpe(year2)
passed = tested > 1.645 # one-sided 5% test, one year of data
print(f"passed: {passed.sum()} of {strategies}, best Sharpe {tested.max():.2f}")
print(f"passers' Sharpe: {tested[passed].mean():.2f} when tested, {next_year[passed].mean():.2f} the next year")
bonferroni = NormalDist().inv_cdf(1 - 0.05 / strategies) # 3.48 for 200 tries
print(f"passed after correction: {(tested > bonferroni).sum()}")
It prints passed: 13 of 200, best Sharpe 2.63. The 13 passers averaged a Sharpe ratio of 2.12
in the year they were tested and 0.01 the year after. After the Bonferroni correction, none
pass. Thirteen convincing backtests, and not one real strategy among them.
Where this shows up in quant work
- Strategy research. Firms ask how many variations were tried before a result was found, and discount accordingly. Being able to answer that question well in an interview stands out.
- Factor investing. Hundreds of published return "factors" have been found in the same historical data. Many fail to replicate, and the field now asks for t-statistics well above 2.
- Performance evaluation. Among many fund managers or traders, some will have great track records by luck alone. The best record in a large group is weak evidence of skill.
Exercises
- A colleague shows you a strategy with a backtested Sharpe ratio of 2 over one year, chosen from 30 variations. Using the tables above, how impressed should you be?
- How many years of data does a strategy with a true Sharpe ratio of 0.5 need before its expected t-statistic reaches 3?
- Modify the code to give 10 of the 200 strategies a real Sharpe ratio of 1. How many of them pass the uncorrected test? The corrected one? What does that say about the cost of correcting?
Key takeaways
- For a strategy, the t-statistic is about Sharpe × √years: one year of data needs a Sharpe ratio above 1.64 to be significant at 5%.
- Test enough no-edge strategies and some will pass; the expected best of 100 has a Sharpe of 2.5.
- Raise the threshold with the number of tests (Bonferroni, Holm, false discovery rate).
- Count every variation you tried, and keep untouched data for a final check.

