Almost every number a quant produces is an average: the mean return of a strategy in a backtest, the average fill price of an order, a Monte Carlo option price. An average on its own says very little. What matters is how far it could be from the truth. This lesson shows where that error bar comes from, why the bell curve shows up even when your data look nothing like one, and what a "95% confidence interval" does and doesn't promise.
| Symbol | Meaning |
|---|---|
| Independent draws from the same distribution | |
| The true mean of that distribution (unknown in practice) | |
| Its true standard deviation | |
| The number of draws in one sample | |
| The sample average | |
| The sample standard deviation, our estimate of | |
| The standard normal cumulative distribution function |
Averages settle down
The law of large numbers says that as you take more independent draws, the sample average gets closer and closer to the true mean:
That is reassuring, but it says nothing about how fast. With 30 observations, should you expect to be off by 1%, or by 50%? For that we need the spread of itself.
How far off is an average?
Because the draws are independent, their variances add. The variance of a sum of draws is , and dividing the sum by divides the variance by :
This standard deviation of the average has its own name, the standard error. It is the single most useful formula in this lesson.
Key idea. The error of an average shrinks like , not . To make an estimate ten times more precise you need a hundred times more data.
The central limit theorem
The standard error tells us the size of the typical error. The central limit theorem tells us its shape. Whatever the distribution of the individual draws (as long as it has a finite variance), the standardised average is approximately standard normal when is large:
This is a strong statement. The draws can be flat, skewed or lumpy; their average still ends up bell-shaped. Try it below:
- Start with the die roll at . The bars are just the die: six equal bars, nothing like the curve.
- Move to . Averages of two dice already form a triangle, with 3.5 the most likely.
- By the bars sit almost exactly under the curve.
- Now try the rare event. It needs a much larger before the bell appears, because the source is so lopsided.
A fair six-sided die: every value from 1 to 6 equally likely. Flat, not bell-shaped.
| Measure | Simulated (0) | Theory |
|---|---|---|
| Mean of the averages | – | 3.500 |
| Spread (standard error) | – | 1.708 |
The table under the chart checks the other half of the story: the spread of the simulated averages matches for every source and every .
Key idea. The normal distribution appears because averaging washes out the shape of the individual draws. How large must be depends on how skewed the draws are. "" is a rule of thumb, not a law.
Confidence intervals
Put the two results together. If the standardised average is approximately standard normal, then for
and rearranging the inequality for gives the 95% confidence interval
Other confidence levels use other values of , the point with :
| Confidence level | |
|---|---|
| 80% | 1.282 |
| 90% | 1.645 |
| 95% | 1.960 |
| 99% | 2.576 |
What does "95%" actually mean? The interval is random, because it is built from a random sample. The true mean is fixed. "95%" is a statement about the procedure: if you repeated the experiment many times, about 95% of the intervals you built would contain . The chart draws one interval per fresh sample, so you can watch that happen.
Switch to 99% and the intervals get wider and miss less often; at 80% they are narrower and miss about one time in five. A confidence level is a trade between how wide your error bar is and how often it is wrong.
Key idea. Once you have computed one interval, it either contains or it doesn't. The 95% is the long-run hit rate of the method, not a probability about this one interval.
When you don't know σ
In practice is unknown too, so you replace it with the sample standard deviation (dividing by ):
That adds a second source of error, and with small or skewed samples the interval becomes too narrow. Here is how often the 95% interval really contains the mean for exponential waiting times, from 20,000 simulated samples at each size:
| Sample size | Using the true | Using the estimate |
|---|---|---|
| 10 | 95.6% | 86.7% |
| 30 | 95.2% | 91.7% |
| 100 | 95.1% | 93.9% |
| 1,000 | 94.9% | 94.8% |
With skewed data, a low sample average tends to come with a low , so the intervals that are furthest off are also the narrowest. For small samples from roughly normal data, the -distribution (a slightly larger multiplier than 1.96) fixes most of this. For skewed data, the honest answer is to collect more data, or check your intervals by simulation as below.
In code
Here is one confidence interval, then the same experiment repeated 10,000 times to check how often the interval contains the true mean of 1:
import numpy as np
rng = np.random.default_rng(0)
# One sample: 30 waiting times from an exponential distribution with mean 1.
sample = rng.exponential(scale=1.0, size=30)
mean = sample.mean()
se = sample.std(ddof=1) / np.sqrt(len(sample))
print(f"mean {mean:.3f}, 95% CI [{mean - 1.96 * se:.3f}, {mean + 1.96 * se:.3f}]")
# Repeat the experiment 10,000 times and count how often the interval contains the true mean 1.
samples = rng.exponential(scale=1.0, size=(10_000, 30))
means = samples.mean(axis=1)
ses = samples.std(axis=1, ddof=1) / np.sqrt(30)
covered = np.abs(means - 1.0) <= 1.96 * ses
print(f"coverage {covered.mean():.3f}")
Running it prints mean 1.184, 95% CI [0.716, 1.653] and then coverage 0.914, the shortfall
from the table above. Checking coverage by simulation like this is a habit worth keeping: it
tells you whether an error bar can be trusted before you rely on it.
Where this shows up in quant work
- Monte Carlo pricing. A Monte Carlo price is an average of simulated payoffs, so its error bar is exactly . That is the error bar in the option pricing lesson, and why it shrinks so slowly.
- Backtests. A strategy's average daily return over a year is an average of about 250 numbers. Daily returns are noisy, so the standard error is often as large as the average itself: many "profitable" backtests are indistinguishable from zero.
- Execution. Comparing two order-routing methods by their average slippage is a comparison of two averages, each with its own standard error.
Two warnings carry over to all of these. The formula assumes the draws are independent: overlapping data or returns that trend together make the true error larger than it looks. And financial returns have heavy tails, so the central limit theorem needs more data to kick in than it does for a die.
Exercises
- Using the chart, find roughly the smallest at which averages of the rare event look bell-shaped. Compare it with the die.
- A strategy's daily returns have mean 0.05% and standard deviation 1%. How many trading days do you need before the 95% confidence interval for the mean excludes zero?
- Modify the Python code to use the multiplier (
scipy.stats.t.ppf(0.975, n - 1)) instead of 1.96 and measure the coverage again for and .
Key takeaways
- The standard error of an average is : precision improves only with the square root of the data.
- The central limit theorem makes averages approximately normal, whatever the source, once is large enough. Skewed sources need more.
- A 95% confidence interval is a procedure that is right 95% of the time, not a 95% probability about one interval.
- Estimating from small or skewed samples makes intervals too narrow. Check coverage by simulation.

