Skip to main content
QuantDXB

Probability and statistics · 13 min read

Bayesian updating

How belief should change as evidence arrives: Bayes' rule, why most good-looking backtests are false, and how many trades it takes to trust a small edge.

Before you start

  • Conditional expectation (this track)
  • Basic Python (NumPy arrays)

By the end you'll be able to

  • Apply Bayes' rule and explain why base rates dominate when effects are rare
  • Update a Beta prior on a hit rate with wins and losses
  • Explain shrinkage and why it improves estimates from small samples
  • Compute a posterior on a grid for any prior, with only NumPy

A new strategy has won 110 of its first 200 trades. Does it have an edge, or was it lucky? A confidence interval answers a slightly different question from the one you're asking. Bayes' rule answers yours directly: given these results, how likely is it that the hit rate is above 50%? This lesson shows how belief should change as evidence arrives, and why a sensible prior protects you from the backtests that look too good.

SymbolMeaning
ppThe strategy's true hit rate (unknown)
kk, nnWins, and total trades, observed so far
P(A∣B)\mathbb{P}(A \mid B)The probability of AA given BB
Beta⁡(a,b)\operatorname{Beta}(a, b)A distribution on [0,1][0, 1], used for beliefs about pp
prior, posteriorBelief before, and after, seeing the data

Bayes' rule

For two events, the definition of conditional probability rearranges into

P(H∣D)=P(D∣H) P(H)P(D).\mathbb{P}(H \mid D) = \frac{\mathbb{P}(D \mid H)\,\mathbb{P}(H)}{\mathbb{P}(D)}.

Read HH as a hypothesis and DD as the data. The rule says how to turn the probability of the data under a hypothesis, which models give you, into the probability of the hypothesis given the data, which is what you want.

Here is why it matters before any formula about hit rates. Suppose that of all the strategy ideas a desk tries, 5% have a real edge. A backtest is a decent filter: it passes 80% of the real ones and only 10% of the fakes. A new idea passes. How likely is it to be real?

P(real∣pass)=0.8×0.050.8×0.05+0.1×0.95=0.040.135≈30%.\mathbb{P}(\text{real} \mid \text{pass}) = \frac{0.8 \times 0.05}{0.8 \times 0.05 + 0.1 \times 0.95} = \frac{0.04}{0.135} \approx 30\%.

Seven passing strategies in ten are fakes, even though the test catches nine fakes in ten. The fakes win on numbers: there are nineteen of them for every real one.

Key idea. When real effects are rare, most positive results are false, however good the test. The base rate, the prior, is not optional.

Beliefs about a hit rate

For a hit rate we need a belief over a whole range of values, not one yes-or-no hypothesis. The Beta distribution is the standard choice. Beta⁡(a,b)\operatorname{Beta}(a, b) lives on [0,1][0, 1], has mean a/(a+b)a/(a + b), and gets narrower as a+ba + b grows. Beta⁡(1,1)\operatorname{Beta}(1, 1) is flat: every hit rate equally plausible.

After kk wins in nn trades, the likelihood of the data is proportional to pk(1−p)n−kp^k (1 - p)^{n - k}. Multiplying it by a Beta⁡(a,b)\operatorname{Beta}(a, b) prior, proportional to pa−1(1−p)b−1p^{a-1}(1-p)^{b-1}, gives

posterior  ∝  p a+k−1(1−p) b+n−k−1,sop∣data∼Beta⁡(a+k,  b+n−k).\text{posterior} \;\propto\; p^{\,a + k - 1}(1 - p)^{\,b + n - k - 1}, \qquad\text{so}\qquad p \mid \text{data} \sim \operatorname{Beta}(a + k,\; b + n - k).

Updating is just counting: add wins to aa and losses to bb. You can read the prior's parameters as imaginary trades you've already seen.

Pick a true hit rate and a prior below and watch trades arrive. The posterior (green) starts as wide as the prior and closes in on the truth.

True hit rate
Prior
Trades
0
0 wins
Posterior mean
50.0%
95% credible interval
0.0% – 100.0%
P(hit rate > 50%)
50.0%
Green: what you believe about the hit rate after the trades so far, with the shaded 95% credible interval. Dashed grey: the prior. A 53% strategy needs around a thousand trades before the interval clears 50%, and luck can make it many more.

Things to try:

  • 53%, flat prior. For the first hundred trades the posterior is wide enough to include anything from a losing strategy to a great one. It takes around a thousand trades before the 95% interval clears 50%.
  • 50%, flat prior. A strategy with no edge at all. Watch "P(hit rate > 50%)" swing around for a long time. Early runs of luck look exactly like skill.
  • Sceptical prior. It acts like roughly 100 extra trades split evenly, so early lucky streaks barely move it. With enough data the two priors end up in the same place.

Shrinkage

The posterior mean is

E[p∣data]=a+ka+b+n=a+ba+b+n⏟weight on prior⋅aa+b+na+b+n⏟weight on data⋅kn.\mathbb{E}[p \mid \text{data}] = \frac{a + k}{a + b + n} = \underbrace{\frac{a + b}{a + b + n}}_{\text{weight on prior}} \cdot \frac{a}{a + b} + \underbrace{\frac{n}{a + b + n}}_{\text{weight on data}} \cdot \frac{k}{n}.

It is a weighted average of the prior mean and the observed win rate, with the weight on the data growing as trades accumulate. This pull toward the prior is called shrinkage. With the sceptical Beta⁡(50,50)\operatorname{Beta}(50, 50) prior and 110 wins in 200 trades, the estimate is 160/300=53.3%160/300 = 53.3\%, not the raw 55%.

Shrinkage is not timidity. Raw win rates from small samples are systematically too extreme: the strategies at the top of a ranking got there partly through luck, so their next results tend to be worse. Shrinking toward a sensible prior corrects that before you allocate money.

Key idea. A posterior is a compromise between what you believed and what you saw, weighted by how much of each you have. With little data, trust the prior; with a lot, trust the data.

Credible and confidence intervals

The shaded region is a 95% credible interval: given the prior and the data, the hit rate lies inside it with probability 95%. That is the statement people usually want, and wrongly read into a confidence interval (see How sure is an average?).

With a flat prior and plenty of data the two intervals are almost identical numerically. They differ when data are scarce, where the prior does real work, and in what they claim: the credible interval is a probability about pp; the confidence interval is a property of the procedure.

In code

You don't need a statistics library. Put a grid over the possible hit rates, multiply prior by likelihood, and normalise:

python
import numpy as np

# A grid of possible hit rates, and a flat prior over them.
p = np.linspace(0.0005, 0.9995, 1000)
prior = np.ones_like(p)

def posterior(prior, wins, losses):
    # Bayes' rule: prior times likelihood, then normalise to sum to 1.
    unnormalised = prior * p**wins * (1 - p) ** losses
    return unnormalised / unnormalised.sum()

post = posterior(prior, wins=110, losses=90)
cdf = np.cumsum(post)
mean = np.sum(p * post)
low, high = p[np.searchsorted(cdf, 0.025)], p[np.searchsorted(cdf, 0.975)]
print(f"after 110 wins in 200: mean {mean:.3f}, 95% interval [{low:.3f}, {high:.3f}]")
print(f"P(hit rate > 50%) = {post[p > 0.5].sum():.3f}")

# A sceptical prior: about as confident as having already seen 50 wins and 50 losses.
sceptical = p**49 * (1 - p) ** 49
post = posterior(sceptical, wins=110, losses=90)
print(f"sceptical prior: mean {np.sum(p * post):.3f}, P(> 50%) = {post[p > 0.5].sum():.3f}")

It prints a posterior mean of 0.550 with a 95% interval of [0.480,0.617][0.480, 0.617] and P(p>0.5)=0.921\mathbb{P}(p > 0.5) = 0.921. So 110 wins in 200 is encouraging, but there is still about an 8% chance the strategy has no edge. The sceptical prior gives a mean of 0.533 and a probability of 0.876. The grid approach works for any prior and any likelihood, not only the Beta–binomial case.

Where this shows up in quant work

  • Evaluating strategies and traders. A short track record is weak evidence. Bayesian estimates with a prior centred on "no edge" rank strategies more reliably than raw win rates.
  • Research pipelines. The 30% calculation is why firms demand out-of-sample tests, economic reasoning and replication before trusting a backtest: each check raises the prior for the next.
  • Parameter estimation. Shrinking noisy estimates (expected returns, betas, covariance matrices) toward a simple target is standard practice in portfolio construction.

Exercises

  • In the screening example, what share of passing strategies is real if the base rate rises to 20%? What false positive rate would you need for a pass to mean "more likely real than not" at a 5% base rate?
  • Starting from a flat prior, a strategy wins 6 of its first 8 trades. Write down the posterior and its mean. Why is the mean not 75%?
  • Modify the code to find, for a strategy that wins exactly 53% of its trades, the smallest number of trades at which P(p>0.5)\mathbb{P}(p > 0.5) passes 95%.

Key takeaways

  • Bayes' rule turns "how likely is this data if the hypothesis is true" into "how likely is the hypothesis given this data". When real effects are rare, most positive results are false.
  • With a Beta prior, updating on wins and losses is counting: Beta(a, b) becomes Beta(a + wins, b + losses).
  • The posterior mean shrinks the observed rate toward the prior, which corrects the luck in small samples.
  • A small edge takes a lot of evidence to confirm: around a thousand trades for 53% against 50%.