Skip to main content
QuantDXB

Probability and statistics · 14 min read

How sure is an average?

Where error bars come from: the standard error, the central limit theorem, and what a 95% confidence interval really promises.

Before you start

  • Mean, variance and standard deviation
  • Basic Python (NumPy arrays)

By the end you'll be able to

  • Derive the standard error σ/√n of an average
  • Explain why averages become normally distributed (the central limit theorem)
  • Build a confidence interval and say correctly what it means
  • Check by simulation whether an error bar can be trusted

Almost every number a quant produces is an average: the mean return of a strategy in a backtest, the average fill price of an order, a Monte Carlo option price. An average on its own says very little. What matters is how far it could be from the truth. This lesson shows where that error bar comes from, why the bell curve shows up even when your data look nothing like one, and what a "95% confidence interval" does and doesn't promise.

SymbolMeaning
X1,…,XnX_1, \dots, X_nIndependent draws from the same distribution
μ\muThe true mean of that distribution (unknown in practice)
σ\sigmaIts true standard deviation
nnThe number of draws in one sample
Xˉn\bar{X}_nThe sample average 1n∑i=1nXi\tfrac{1}{n}\sum_{i=1}^{n} X_i
ssThe sample standard deviation, our estimate of σ\sigma
Φ\PhiThe standard normal cumulative distribution function

Averages settle down

The law of large numbers says that as you take more independent draws, the sample average gets closer and closer to the true mean:

Xˉn=1n∑i=1nXi  ⟶  μas n→∞.\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i \;\longrightarrow\; \mu \quad \text{as } n \to \infty.

That is reassuring, but it says nothing about how fast. With 30 observations, should you expect to be off by 1%, or by 50%? For that we need the spread of Xˉn\bar{X}_n itself.

How far off is an average?

Because the draws are independent, their variances add. The variance of a sum of nn draws is nσ2n\sigma^2, and dividing the sum by nn divides the variance by n2n^2:

Var⁡(Xˉn)=1n2⋅nσ2=σ2n,SE=Var⁡(Xˉn)=σn.\operatorname{Var}\big(\bar{X}_n\big) = \frac{1}{n^2}\cdot n\sigma^2 = \frac{\sigma^2}{n}, \qquad \mathrm{SE} = \sqrt{\operatorname{Var}\big(\bar{X}_n\big)} = \frac{\sigma}{\sqrt{n}}.

This standard deviation of the average has its own name, the standard error. It is the single most useful formula in this lesson.

Key idea. The error of an average shrinks like 1/n1/\sqrt{n}, not 1/n1/n. To make an estimate ten times more precise you need a hundred times more data.

The central limit theorem

The standard error tells us the size of the typical error. The central limit theorem tells us its shape. Whatever the distribution of the individual draws (as long as it has a finite variance), the standardised average is approximately standard normal when nn is large:

Xˉn−μσ/n  ≈  N(0,1).\frac{\bar{X}_n - \mu}{\sigma/\sqrt{n}} \;\approx\; \mathcal{N}(0, 1).

This is a strong statement. The draws can be flat, skewed or lumpy; their average still ends up bell-shaped. Try it below:

  • Start with the die roll at n=1n = 1. The bars are just the die: six equal bars, nothing like the curve.
  • Move to n=2n = 2. Averages of two dice already form a triangle, with 3.5 the most likely.
  • By n=10n = 10 the bars sit almost exactly under the curve.
  • Now try the rare event. It needs a much larger nn before the bell appears, because the source is so lopsided.

A fair six-sided die: every value from 1 to 6 equally likely. Flat, not bell-shaped.

1
Simulated versus theoretical mean and spread of the sample averages
MeasureSimulated (0)Theory
Mean of the averages–3.500
Spread (standard error)–1.708
Bars: averages of n draws, simulated. Curve: the normal distribution with mean μ and standard deviation σ/√n. Start at n = 1 to see the source itself, then drag n up and watch the bars take the curve's shape.

The table under the chart checks the other half of the story: the spread of the simulated averages matches σ/n\sigma/\sqrt{n} for every source and every nn.

Key idea. The normal distribution appears because averaging washes out the shape of the individual draws. How large nn must be depends on how skewed the draws are. "n≥30n \geq 30" is a rule of thumb, not a law.

Confidence intervals

Put the two results together. If the standardised average is approximately standard normal, then for z=1.96z = 1.96

P(−1.96≤Xˉn−μσ/n≤1.96)≈0.95,\mathbb{P}\left(-1.96 \leq \frac{\bar{X}_n - \mu}{\sigma/\sqrt{n}} \leq 1.96\right) \approx 0.95,

and rearranging the inequality for μ\mu gives the 95% confidence interval

Xˉn±1.96 σn.\bar{X}_n \pm 1.96\,\frac{\sigma}{\sqrt{n}}.

Other confidence levels use other values of zz, the point with Φ(z)=12+level2\Phi(z) = \tfrac{1}{2} + \tfrac{\text{level}}{2}:

Confidence levelzz
80%1.282
90%1.645
95%1.960
99%2.576

What does "95%" actually mean? The interval is random, because it is built from a random sample. The true mean is fixed. "95%" is a statement about the procedure: if you repeated the experiment many times, about 95% of the intervals you built would contain μ\mu. The chart draws one interval per fresh sample, so you can watch that happen.

Confidence level
Each line is one interval x̄ ± z·σ/√n from a fresh sample of 30 waiting times. Faint green lines contain the true mean of 1; bold white ones miss it. Over many samples the share that contain it settles at the confidence level.

Switch to 99% and the intervals get wider and miss less often; at 80% they are narrower and miss about one time in five. A confidence level is a trade between how wide your error bar is and how often it is wrong.

Key idea. Once you have computed one interval, it either contains μ\mu or it doesn't. The 95% is the long-run hit rate of the method, not a probability about this one interval.

When you don't know σ

In practice σ\sigma is unknown too, so you replace it with the sample standard deviation ss (dividing by n−1n - 1):

Xˉn±1.96 sn.\bar{X}_n \pm 1.96\,\frac{s}{\sqrt{n}}.

That adds a second source of error, and with small or skewed samples the interval becomes too narrow. Here is how often the 95% interval really contains the mean for exponential waiting times, from 20,000 simulated samples at each size:

Sample size nnUsing the true σ\sigmaUsing the estimate ss
1095.6%86.7%
3095.2%91.7%
10095.1%93.9%
1,00094.9%94.8%

With skewed data, a low sample average tends to come with a low ss, so the intervals that are furthest off are also the narrowest. For small samples from roughly normal data, the tt-distribution (a slightly larger multiplier than 1.96) fixes most of this. For skewed data, the honest answer is to collect more data, or check your intervals by simulation as below.

In code

Here is one confidence interval, then the same experiment repeated 10,000 times to check how often the interval contains the true mean of 1:

python
import numpy as np

rng = np.random.default_rng(0)

# One sample: 30 waiting times from an exponential distribution with mean 1.
sample = rng.exponential(scale=1.0, size=30)
mean = sample.mean()
se = sample.std(ddof=1) / np.sqrt(len(sample))
print(f"mean {mean:.3f}, 95% CI [{mean - 1.96 * se:.3f}, {mean + 1.96 * se:.3f}]")

# Repeat the experiment 10,000 times and count how often the interval contains the true mean 1.
samples = rng.exponential(scale=1.0, size=(10_000, 30))
means = samples.mean(axis=1)
ses = samples.std(axis=1, ddof=1) / np.sqrt(30)
covered = np.abs(means - 1.0) <= 1.96 * ses
print(f"coverage {covered.mean():.3f}")

Running it prints mean 1.184, 95% CI [0.716, 1.653] and then coverage 0.914, the shortfall from the table above. Checking coverage by simulation like this is a habit worth keeping: it tells you whether an error bar can be trusted before you rely on it.

Where this shows up in quant work

  • Monte Carlo pricing. A Monte Carlo price is an average of simulated payoffs, so its error bar is exactly 1.96 s/N1.96\,s/\sqrt{N}. That is the error bar in the option pricing lesson, and why it shrinks so slowly.
  • Backtests. A strategy's average daily return over a year is an average of about 250 numbers. Daily returns are noisy, so the standard error is often as large as the average itself: many "profitable" backtests are indistinguishable from zero.
  • Execution. Comparing two order-routing methods by their average slippage is a comparison of two averages, each with its own standard error.

Two warnings carry over to all of these. The formula σ/n\sigma/\sqrt{n} assumes the draws are independent: overlapping data or returns that trend together make the true error larger than it looks. And financial returns have heavy tails, so the central limit theorem needs more data to kick in than it does for a die.

Exercises

  • Using the chart, find roughly the smallest nn at which averages of the rare event look bell-shaped. Compare it with the die.
  • A strategy's daily returns have mean 0.05% and standard deviation 1%. How many trading days do you need before the 95% confidence interval for the mean excludes zero?
  • Modify the Python code to use the tt multiplier (scipy.stats.t.ppf(0.975, n - 1)) instead of 1.96 and measure the coverage again for n=10n = 10 and n=30n = 30.

Key takeaways

  • The standard error of an average is σ/n\sigma/\sqrt{n}: precision improves only with the square root of the data.
  • The central limit theorem makes averages approximately normal, whatever the source, once nn is large enough. Skewed sources need more.
  • A 95% confidence interval is a procedure that is right 95% of the time, not a 95% probability about one interval.
  • Estimating σ\sigma from small or skewed samples makes intervals too narrow. Check coverage by simulation.