A new strategy has won 110 of its first 200 trades. Does it have an edge, or was it lucky? A confidence interval answers a slightly different question from the one you're asking. Bayes' rule answers yours directly: given these results, how likely is it that the hit rate is above 50%? This lesson shows how belief should change as evidence arrives, and why a sensible prior protects you from the backtests that look too good.
| Symbol | Meaning |
|---|---|
| The strategy's true hit rate (unknown) | |
| , | Wins, and total trades, observed so far |
| The probability of given | |
| A distribution on , used for beliefs about | |
| prior, posterior | Belief before, and after, seeing the data |
Bayes' rule
For two events, the definition of conditional probability rearranges into
Read as a hypothesis and as the data. The rule says how to turn the probability of the data under a hypothesis, which models give you, into the probability of the hypothesis given the data, which is what you want.
Here is why it matters before any formula about hit rates. Suppose that of all the strategy ideas a desk tries, 5% have a real edge. A backtest is a decent filter: it passes 80% of the real ones and only 10% of the fakes. A new idea passes. How likely is it to be real?
Seven passing strategies in ten are fakes, even though the test catches nine fakes in ten. The fakes win on numbers: there are nineteen of them for every real one.
Key idea. When real effects are rare, most positive results are false, however good the test. The base rate, the prior, is not optional.
Beliefs about a hit rate
For a hit rate we need a belief over a whole range of values, not one yes-or-no hypothesis. The Beta distribution is the standard choice. lives on , has mean , and gets narrower as grows. is flat: every hit rate equally plausible.
After wins in trades, the likelihood of the data is proportional to . Multiplying it by a prior, proportional to , gives
Updating is just counting: add wins to and losses to . You can read the prior's parameters as imaginary trades you've already seen.
Pick a true hit rate and a prior below and watch trades arrive. The posterior (green) starts as wide as the prior and closes in on the truth.
- Trades
- 0
- 0 wins
- Posterior mean
- 50.0%
- 95% credible interval
- 0.0% – 100.0%
- P(hit rate > 50%)
- 50.0%
Things to try:
- 53%, flat prior. For the first hundred trades the posterior is wide enough to include anything from a losing strategy to a great one. It takes around a thousand trades before the 95% interval clears 50%.
- 50%, flat prior. A strategy with no edge at all. Watch "P(hit rate > 50%)" swing around for a long time. Early runs of luck look exactly like skill.
- Sceptical prior. It acts like roughly 100 extra trades split evenly, so early lucky streaks barely move it. With enough data the two priors end up in the same place.
Shrinkage
The posterior mean is
It is a weighted average of the prior mean and the observed win rate, with the weight on the data growing as trades accumulate. This pull toward the prior is called shrinkage. With the sceptical prior and 110 wins in 200 trades, the estimate is , not the raw 55%.
Shrinkage is not timidity. Raw win rates from small samples are systematically too extreme: the strategies at the top of a ranking got there partly through luck, so their next results tend to be worse. Shrinking toward a sensible prior corrects that before you allocate money.
Key idea. A posterior is a compromise between what you believed and what you saw, weighted by how much of each you have. With little data, trust the prior; with a lot, trust the data.
Credible and confidence intervals
The shaded region is a 95% credible interval: given the prior and the data, the hit rate lies inside it with probability 95%. That is the statement people usually want, and wrongly read into a confidence interval (see How sure is an average?).
With a flat prior and plenty of data the two intervals are almost identical numerically. They differ when data are scarce, where the prior does real work, and in what they claim: the credible interval is a probability about ; the confidence interval is a property of the procedure.
In code
You don't need a statistics library. Put a grid over the possible hit rates, multiply prior by likelihood, and normalise:
import numpy as np
# A grid of possible hit rates, and a flat prior over them.
p = np.linspace(0.0005, 0.9995, 1000)
prior = np.ones_like(p)
def posterior(prior, wins, losses):
# Bayes' rule: prior times likelihood, then normalise to sum to 1.
unnormalised = prior * p**wins * (1 - p) ** losses
return unnormalised / unnormalised.sum()
post = posterior(prior, wins=110, losses=90)
cdf = np.cumsum(post)
mean = np.sum(p * post)
low, high = p[np.searchsorted(cdf, 0.025)], p[np.searchsorted(cdf, 0.975)]
print(f"after 110 wins in 200: mean {mean:.3f}, 95% interval [{low:.3f}, {high:.3f}]")
print(f"P(hit rate > 50%) = {post[p > 0.5].sum():.3f}")
# A sceptical prior: about as confident as having already seen 50 wins and 50 losses.
sceptical = p**49 * (1 - p) ** 49
post = posterior(sceptical, wins=110, losses=90)
print(f"sceptical prior: mean {np.sum(p * post):.3f}, P(> 50%) = {post[p > 0.5].sum():.3f}")
It prints a posterior mean of 0.550 with a 95% interval of and . So 110 wins in 200 is encouraging, but there is still about an 8% chance the strategy has no edge. The sceptical prior gives a mean of 0.533 and a probability of 0.876. The grid approach works for any prior and any likelihood, not only the Beta–binomial case.
Where this shows up in quant work
- Evaluating strategies and traders. A short track record is weak evidence. Bayesian estimates with a prior centred on "no edge" rank strategies more reliably than raw win rates.
- Research pipelines. The 30% calculation is why firms demand out-of-sample tests, economic reasoning and replication before trusting a backtest: each check raises the prior for the next.
- Parameter estimation. Shrinking noisy estimates (expected returns, betas, covariance matrices) toward a simple target is standard practice in portfolio construction.
Exercises
- In the screening example, what share of passing strategies is real if the base rate rises to 20%? What false positive rate would you need for a pass to mean "more likely real than not" at a 5% base rate?
- Starting from a flat prior, a strategy wins 6 of its first 8 trades. Write down the posterior and its mean. Why is the mean not 75%?
- Modify the code to find, for a strategy that wins exactly 53% of its trades, the smallest number of trades at which passes 95%.
Key takeaways
- Bayes' rule turns "how likely is this data if the hypothesis is true" into "how likely is the hypothesis given this data". When real effects are rare, most positive results are false.
- With a Beta prior, updating on wins and losses is counting: Beta(a, b) becomes Beta(a + wins, b + losses).
- The posterior mean shrinks the observed rate toward the prior, which corrects the luck in small samples.
- A small edge takes a lot of evidence to confirm: around a thousand trades for 53% against 50%.

