Skip to main content
QuantDXB

Data and machine learning · 11 min read

Backtesting without look-ahead bias

How a backtest manufactures profits from noise: look-ahead, survivorship, and a checklist for the rest.

Before you start

  • Time series with pandas
  • Hypothesis tests and false discoveries

By the end you'll be able to

  • Explain why a backtest may only use information available at each decision
  • Find and fix look-ahead bias in pandas code
  • Recognise survivorship bias and the data needed to avoid it
  • Review a backtest against a checklist of common biases

A backtest answers one question: if I had followed these rules in the past, what would have happened? It can only answer honestly if, at every step, it uses nothing that wasn't known at that moment. Breaking that rule doesn't make a backtest a little optimistic; it can manufacture a spectacular strategy out of pure noise. This lesson shows the two most common ways it happens and a checklist for catching the rest.

TermMeaning
look-ahead biasUsing information in a decision before it was available
survivorship biasBuilding the universe from what exists today, dropping what disappeared
point-in-time dataData as it was known on each date, before later corrections or restatements
equity curveThe growth of money invested in the strategy over time
Sharpe ratioAnnualised mean return over volatility

Look-ahead: one missing day

Take a simple momentum rule: hold the stock if its average return over the last five days is positive, short it if negative. The signal is known at the close of day tt, so the earliest it can earn is day t+1t + 1's return. In pandas that is signal.shift(1) * returns, as the time series lesson showed.

Forget the shift and day tt's position earns day tt's return, a return that was part of the signal. A strategy that "knows" today's return before trading on it can't lose.

The chart below runs both versions on a random walk, a market where momentum has no edge at all:

Bias

A 5-day momentum signal on a random walk. Honest: trade it the day after it is known. Biased: trade it the same day, so the return it earns was part of the signal.

Annual return and Sharpe ratio of the honest and the biased backtest
BacktestAnnual returnSharpe ratio
Traded next day (honest)––
Traded same day (biased)––
Grey: the honest backtest. Green: the biased one. There is no edge to find in this market, so every bit of the green curve's profit comes from the bias.

The honest curve wanders around 1, as it should. The biased one climbs steadily for ten years. Nothing in its rules looks wrong; the bug is one line of alignment.

Key idea. If a backtest looks too smooth, assume look-ahead until you have proved otherwise. Real edges are noisy; an equity curve that rises almost every day usually means the strategy is seeing the answer.

Survivorship: the stocks that disappeared

Now switch the chart to survivorship. Two hundred stocks with no drift at all: on average, each neither rises nor falls. Some collapse and are delisted. An index that holds every stock while it trades earns about nothing, which is the truth.

Build the same index from the list of stocks that exist today, the obvious thing to download, and every collapse is silently removed. The survivors are, by construction, the stocks that did well enough not to collapse, and the index of survivors rises.

This bias is everywhere:

  • Index constituents today are winners; historical constituents included the companies that failed or were dropped.
  • Fund databases lose funds that closed, mostly the ones that did badly.
  • Price databases sometimes don't keep delisted symbols at all.

The fix is data that includes the dead: historical constituent lists, delisted securities with their final returns, and funds that closed.

Key idea. Ask of every dataset: what is missing, and why? Data that disappeared for a reason related to returns biases every result built on what is left.

The rest of the checklist

Look-ahead and survivorship are the famous two. A thorough review also asks:

  • Is every data point stamped with when it was known? Company earnings for a quarter are published weeks after it ends, and are sometimes restated later. Using the restated figure, dated to the quarter's end, is look-ahead. Point-in-time databases keep each value as first published.
  • Do the trade times line up? A signal computed from the close can't trade at that same close. Using the next open, or the next close, changes results more than people expect.
  • Are the statistics fitted on the future? Normalising a signal with the full-sample mean and standard deviation, or choosing parameters on the full history, leaks the future into every past decision. Use only data up to each point, as in the walk-forward lesson.
  • How many versions were tried? Every variation tested is a chance to fit noise, as the false discoveries lesson showed.
  • Are costs included? Spread, impact and fees, from the execution costs lesson. Fast strategies are the most affected.

In code

python
import numpy as np
import pandas as pd

rng = np.random.default_rng(3)

def sharpe(r):
    return r.mean() / r.std() * np.sqrt(252)

# 1. Look-ahead: 5-day momentum on a random walk (no edge exists).
returns = pd.Series(rng.normal(0, 0.01, 252 * 10))
signal = np.sign(returns.rolling(5).mean())  # known at the close of each day
honest = (signal.shift(1) * returns).dropna()  # trade it the next day
biased = (signal * returns).dropna()  # uses the return it is predicting
print(f"look-ahead: honest Sharpe {sharpe(honest):.2f}, biased Sharpe {sharpe(biased):.2f}")

# 2. Survivorship: 400 driftless stocks over 5 years; delisted below 30% of their start price.
days, stocks, vol = 252 * 5, 400, 0.03
log_paths = np.cumsum(rng.normal(-vol**2 / 2, vol, size=(days, stocks)), axis=0)
prices = pd.DataFrame(np.exp(log_paths))
delisted = prices.lt(0.3).cummax()  # once below the threshold, gone for good
daily = prices.pct_change().fillna(prices.iloc[0] - 1).mask(delisted.shift(1, fill_value=False))
survivors = ~delisted.iloc[-1]
all_stocks = daily.mean(axis=1)  # every stock that was trading that day
survivors_only = daily.loc[:, survivors].mean(axis=1)  # chosen knowing who survived
annual = lambda r: (1 + r).prod() ** (252 / len(r)) - 1
print(f"survivorship: {(~survivors).sum()} of {stocks} delisted; "
      f"all stocks {annual(all_stocks):+.1%} a year, survivors only {annual(survivors_only):+.1%} a year")

It prints an honest Sharpe ratio of 0.46, which for ten years of data is within the noise (a Sharpe estimate from ten years has a standard error of about 0.32), and a biased one of 5.98, a number no real strategy sustains. In the survivorship test 168 of the 400 stocks are delisted; the honest index returns +0.2% a year, the survivors-only index +13.9% a year. Neither bias needed a single wrong formula.

Where this shows up in quant work

  • Research reviews. Firms review backtests line by line for exactly these problems before any money is allocated. Spotting look-ahead in someone else's code is a common interview exercise.
  • Data engineering. A large part of quant data work is building point-in-time, survivorship-free datasets, because nothing built on top of biased data can be trusted.
  • Live versus backtest. When a strategy that looked excellent disappoints in live trading, bias in the backtest is the first suspect, ahead of "the market changed".

Exercises

  • In the code, change the momentum lookback from 5 days to 20. How does the biased Sharpe ratio change, and why does a longer lookback leak less?
  • Normalise the signal with (x - x.mean()) / x.std() over the full sample, then with an expanding window (x.expanding().mean()). Which version is a fair backtest?
  • Modify the survivorship example so delisted stocks lose a further 50% on their last day (as many real delistings do). How does the honest index change?

Key takeaways

  • A backtest may use only information available at each decision; shift(1) is the most important line in most backtests.
  • Look-ahead can turn a strategy with no edge into a spectacular one; too-smooth equity curves are a warning sign.
  • Survivorship bias removes the failures from history; use data that includes the dead.
  • Check point-in-time data, trade timing, full-sample fitting, the number of tries and costs.