Skip to main content
QuantDXB

Probability and statistics · 13 min read

Conditional expectation

A forecast is an average given what you know. The tower property, the best predictor, and how much of a return a signal can explain.

Before you start

  • Expected value and variance
  • How sure is an average? (this track)

By the end you'll be able to

  • Compute expectations conditional on events and on random variables
  • Estimate E[Y | X] from data by slicing and averaging, and judge the trade-off
  • Solve branching problems with the tower property
  • Split variance into the part a signal explains and the rest

A trading signal is a forecast, and a forecast is a conditional expectation: the average outcome given what you know. Given that buyers have outnumbered sellers for the last minute, what return should you expect over the next one? This lesson builds the idea from a die roll up to the two identities that every quant uses without thinking: the tower property and the law of total variance.

SymbolMeaning
E[Y]\mathbb{E}[Y]The expected value (long-run average) of YY
E[Y∣A]\mathbb{E}[Y \mid A]The expected value of YY given that event AA happened
E[Y∣X]\mathbb{E}[Y \mid X]The conditional expectation of YY given XX, a function of XX
m(x)m(x)That function: m(x)=E[Y∣X=x]m(x) = \mathbb{E}[Y \mid X = x]
Var⁡(Y)\operatorname{Var}(Y)The variance of YY

Conditioning on an event

Roll a fair die. Its expected value is 3.53.5. Now someone tells you the roll was even. Only 2, 4 and 6 are still possible, each with probability 13\tfrac{1}{3}, so

E[roll∣even]=13(2+4+6)=4.\mathbb{E}[\text{roll} \mid \text{even}] = \tfrac{1}{3}(2 + 4 + 6) = 4.

That is all conditioning is: throw away the outcomes the information rules out, rescale the probabilities of the rest so they add up to 1, and average again. The information changed the forecast from 3.5 to 4.

Conditioning on a random variable

Usually the information is not a yes-or-no event but a number you observe, like an order-flow imbalance XX (buy volume minus sell volume, scaled). For each value xx it could take there is an average outcome, and collecting those averages gives a function:

m(x)=E[Y∣X=x].m(x) = \mathbb{E}[Y \mid X = x].

Plugging the random XX into it gives E[Y∣X]=m(X)\mathbb{E}[Y \mid X] = m(X), which is itself a random variable: before you see the imbalance you don't know which forecast you'll make.

In the chart, each dot is one simulated minute: imbalance across, the next minute's return in basis points up. The true m(x)m(x) is the dashed curve: more buying pressure means a higher expected return, flattening out at the extremes. You never see that curve in real data. What you can do is slice the imbalance axis and average the returns inside each slice: the green bars.

6
Points
0
Variance explained
–
True share
19.4%
Each dot is one minute: imbalance across, next-minute return in basis points up. Green bars average the returns inside each slice. Too few slices blur the curve; too many leave each slice with too few points to average.

Try the slider. With one slice you get the overall average and learn nothing from XX. With 6 to 10 slices the bars trace the curve well. With 40, each slice holds only a handful of points, and the bars jump around with the noise. This is the same trade-off you will meet again in the overfitting lesson: a flexible estimate has less bias but more noise.

Key idea. E[Y∣X]\mathbb{E}[Y \mid X] is a function of what you know. Estimating it from data always means averaging over "similar" situations, and choosing how similar is a trade-off between bias and noise.

Why it is the best forecast

Suppose you forecast YY with some function g(X)g(X) and judge it by mean squared error. Write Y−g(X)=(Y−m(X))+(m(X)−g(X))Y - g(X) = \big(Y - m(X)\big) + \big(m(X) - g(X)\big) and expand the square. The cross term averages to zero, because Y−m(X)Y - m(X) has mean zero within every slice of XX. What is left is

E[(Y−g(X))2]=E[(Y−m(X))2]+E[(m(X)−g(X))2].\mathbb{E}\big[(Y - g(X))^2\big] = \mathbb{E}\big[(Y - m(X))^2\big] + \mathbb{E}\big[(m(X) - g(X))^2\big].

The first term doesn't depend on gg and the second is never negative, so the error is smallest when g=mg = m. No function of XX forecasts YY better, in the mean-squared sense, than the conditional expectation. Regression, gradient boosting and neural networks are all different ways of estimating it.

The tower property

If you average the conditional forecasts over all the situations you could be in, you get back the unconditional average:

E[E[Y∣X]]=E[Y].\mathbb{E}\big[\mathbb{E}[Y \mid X]\big] = \mathbb{E}[Y].

Check it on the die. Half the time the roll is even, with conditional mean 4; half the time it is odd (1, 3 or 5), with conditional mean 3. Then 12⋅4+12⋅3=3.5\tfrac{1}{2}\cdot 4 + \tfrac{1}{2}\cdot 3 = 3.5, the plain expected value.

The tower property is how you solve problems that branch. Roll a die; if it shows 6, roll again and add, and keep going as long as you roll 6s. Condition on the first roll: with probability 56\tfrac{5}{6} you stop, having rolled 1 to 5 (average 3), and with probability 16\tfrac{1}{6} you have 6 plus a fresh copy of the same game. So E=56⋅3+16(6+E)E = \tfrac{5}{6}\cdot 3 + \tfrac{1}{6}(6 + E), which gives E=4.2E = 4.2.

Key idea. To find an expectation that is hard to compute directly, condition on the first thing that happens, solve each branch, then average the branches.

How much does the signal explain?

The variance of YY splits into two parts, the law of total variance:

Var⁡(Y)=Var⁡(E[Y∣X])⏟explained by X+E[Var⁡(Y∣X)]⏟left over.\operatorname{Var}(Y) = \underbrace{\operatorname{Var}\big(\mathbb{E}[Y \mid X]\big)}_{\text{explained by } X} + \underbrace{\mathbb{E}\big[\operatorname{Var}(Y \mid X)\big]}_{\text{left over}}.

The first part measures how much the forecast moves around as XX changes; the second is the noise that remains even when you know XX. The ratio of the first to the total is the share of the variance XX explains, the R2R^2 of a perfect model.

In the simulated market, the true split is 11.16=2.16+9.0011.16 = 2.16 + 9.00: the imbalance explains about 19% of the variance of next-minute returns. The chart's readout estimates the same share from the dots. Notice it creeps above the true value when you use many slices: noise inside small slices gets counted as "explained". That is overfitting, measured.

Real markets are far less generous. For short-horizon returns the explained share of any public signal is usually tiny. A signal that explains even a small fraction of the variance can still be very profitable when it is traded thousands of times, because the forecast only needs to be right on average.

In code

Here is the slicing estimator with pandas, plus checks of both identities on 100,000 simulated minutes:

python
import numpy as np
import pandas as pd

rng = np.random.default_rng(1)
n = 100_000

# Simulated minutes: order-flow imbalance x, next-minute return y in basis points.
x = rng.standard_normal(n)
y = 2 * np.tanh(1.5 * x) + 3 * rng.standard_normal(n)
df = pd.DataFrame({"x": x, "y": y})

# Estimate E[Y | X] by averaging y within 12 equal-width slices of x.
df["slice"] = pd.cut(df["x"], bins=np.linspace(-3, 3, 13))
df["cond_mean"] = df.groupby("slice", observed=True)["y"].transform("mean")
print(df.groupby("slice", observed=True)["y"].mean().round(2).iloc[[0, 3, 6, 9, 11]])

# Tower property: the average of the conditional means is the overall mean.
inside = df.dropna(subset=["cond_mean"])
print(f"E[Y] {inside['y'].mean():.4f}   E[E[Y|X]] {inside['cond_mean'].mean():.4f}")

# Law of total variance: total = explained + unexplained.
total = inside["y"].var(ddof=0)
explained = inside["cond_mean"].var(ddof=0)
unexplained = ((inside["y"] - inside["cond_mean"]) ** 2).mean()
print(f"total {total:.3f} = explained {explained:.3f} + unexplained {unexplained:.3f}")
print(f"share explained {explained / total:.1%}")

The slice averages it prints run from about −1.8-1.8 to −2.0-2.0 on the selling side to about 2.02.0 on the buying side, tracking 2tanh⁡(1.5x)2\tanh(1.5x) (the outermost slices hold few points, so they are noisier). The two means agree exactly (−0.0096-0.0096), and the variance splits as total 11.194 = explained 2.117 + unexplained 9.078, an explained share of 18.9% against the true 19.4%.

Where this shows up in quant work

  • Signals. An alpha signal is an estimate of E[future return∣data]\mathbb{E}[\text{future return} \mid \text{data}]. Everything in the data and machine learning track is about estimating that function without fooling yourself.
  • Pricing. An option price is a conditional expectation of the discounted payoff given today's information, which is exactly what the Monte Carlo lesson estimates.
  • Fair prices are martingales. If a price already reflects all available information, its expected change given that information is zero: E[St+1∣information at t]=St\mathbb{E}[S_{t+1} \mid \text{information at } t] = S_t (after allowing for interest and risk). A predictable price would be an arbitrage. The random walks lesson builds on this.

Exercises

  • A card is drawn from a standard deck. What is the expected value of its rank (ace = 1, jack = 11, queen = 12, king = 13) given that it is a face card? Given that it is red?
  • You flip a fair coin until you see heads. Use the tower property (condition on the first flip) to show that the expected number of flips is 2.
  • In the Python code, replace the 12 equal-width slices with 12 slices containing equal numbers of points (pd.qcut). Which estimate is closer to the true curve at the extremes, and why?

Key takeaways

  • E[Y∣X]\mathbb{E}[Y \mid X] is the average of YY among situations with the same XX: a function of what you know, and the best mean-squared forecast of YY.
  • Estimating it means averaging over similar situations; how wide "similar" is trades bias against noise.
  • The tower property, E[E[Y∣X]]=E[Y]\mathbb{E}[\mathbb{E}[Y \mid X]] = \mathbb{E}[Y], solves branching problems by conditioning on the first step.
  • The law of total variance splits variance into what a signal explains and what it can't. Expect the explained part to be small in real markets.