Skip to main content
QuantDXB

Data and machine learning · 12 min read

Validation for financial machine learning

Why standard cross-validation finds signals in pure noise on time series, and how walk-forward validation with a gap fixes it.

Before you start

  • Overfitting (this track)
  • Backtesting without look-ahead bias (this track)

By the end you'll be able to

  • Explain how overlapping labels and slow features leak through shuffled folds
  • Set up walk-forward validation with an embargo
  • Interpret noisy validation scores on financial data
  • Keep preprocessing and tuning inside each training window

Cross-validation is the standard way to measure how well a model predicts data it hasn't seen: split the data into folds, train on some, test on the rest, repeat. Applied naively to financial time series, it reports skill that doesn't exist, and machine learning tools apply it naively by default. This lesson builds a dataset with nothing to predict, shows standard cross-validation finding a strong "signal" in it, and fixes it with walk-forward validation.

TermMeaning
foldOne of the parts the data is split into for cross-validation
k-fold (shuffled)Random folds; train on k − 1, test on the remaining one, in turn
leakageInformation about the test data reaching the model through the training data
overlapping labelsTargets that share data, like 20-day returns measured every day
walk-forwardTrain only on the past, test on the period after it, then roll forward
purging / embargoRemoving training days whose targets overlap or sit next to the test period

A dataset with nothing in it

Take a random walk: daily returns with no predictability at all. The target, as in many real models, is the return over the next 20 days, measured every day. The features are three slow-moving series of pure noise, unrelated to returns. Any model that finds a relationship here has found something that isn't there.

The model is kk-nearest neighbours: to predict a day, find the 5 training days whose features are closest and average their targets. It is simple, flexible, and a fair stand-in for the tree and neighbour-based models widely used on financial data.

Why shuffled folds leak

Shuffled kk-fold scatters the test days at random through the history. Every test day then has its own neighbours in time, the day before and the day after, in the training set. Two things make those neighbours a cheat sheet:

  • Overlapping labels. Day tt's target is the return over days t+1t + 1 to t+20t + 20; day t+1t + 1's target covers days t+2t + 2 to t+21t + 21. They share 19 of 20 days, so they are almost identical.
  • Slow features. Features that change slowly make adjacent days look alike, so the nearest neighbours in feature space are mostly the nearest neighbours in time.

The model "predicts" day tt by looking up day t+1t + 1, whose target it was trained on. Watch it happen:

Validation
Folds evaluated
0 / 5
Score (correlation)
–
There is no signal here: the target is a random walk's next 20 days, the features are unrelated noise. Shuffled folds still score well, because each test day's neighbours in time, with nearly the same target, are in the training set. Walk-forward folds train only on the past, with a 20-day gap, and score around zero.

With shuffled folds, the score comes out around 0.6, which would be an extraordinary signal for 20-day returns. Switch to walk-forward: each fold trains only on days before the test block, with a 20-day gap so no training target overlaps a test target, and the score falls to around zero, with some noise either way. That is the truth, because there is nothing to predict.

Key idea. When neighbouring observations share information, random train/test splits leak it. For time series, the test data must come after the training data, with a gap at least as long as the target's horizon.

Walk-forward validation

The procedure:

  1. Split time into consecutive blocks.
  2. For each block after the first, train on everything before it, minus a gap (embargo) at least as long as the label horizon, and test on the block.
  3. Collect predictions from all test blocks and score them together.

This mimics how a model would actually be used: fitted on the past and run on the future. It uses data less efficiently than shuffled folds (early blocks have little training data), and its scores are noisier, because overlapping targets mean far fewer independent observations than there are days. 1,500 daily observations of a 20-day target carry roughly the information of 75 independent ones. Both costs are real; they are the price of an honest estimate.

When even the gap isn't enough, because features use long lookback windows, for example, purging removes any training day whose information window overlaps the test period. The ideas are developed fully in the literature on combinatorial purged cross-validation.

The same leak elsewhere

The leak isn't specific to kk-nearest neighbours. It hits any model flexible enough to recognise "the same moment in time": tree ensembles, boosting, neural networks. It also appears whenever preprocessing uses the whole dataset: scaling features with the full-sample mean and standard deviation, selecting features by their correlation with the target over the full history, or tuning hyperparameters on the same folds used to report the score. All of these are versions of the look-ahead bias that backtests suffer from.

In code

python
import numpy as np

rng = np.random.default_rng(2)
days, horizon, window, k, folds, features = 1500, 20, 60, 5, 5, 3

returns = rng.normal(0, 0.01, days + horizon)
# Target: the next 20 days' return, which overlaps heavily from one day to the next.
target = np.array([returns[t + 1 : t + 1 + horizon].sum() for t in range(days)])
# Features: three slow-moving series of pure noise, unrelated to returns.
feature = np.column_stack([
    np.convolve(rng.standard_normal(days + window), np.ones(window) / window, mode="valid")[:days]
    for _ in range(features)
])

def knn_predict(train_idx, test_idx):
    # Average target of the k training days whose feature value is closest.
    distance = np.linalg.norm(feature[test_idx][:, None, :] - feature[train_idx][None, :, :], axis=2)
    nearest = np.argsort(distance, axis=1)[:, :k]
    return target[train_idx][nearest].mean(axis=1)

def score(splits):
    predicted, actual = [], []
    for train_idx, test_idx in splits:
        predicted.append(knn_predict(train_idx, test_idx))
        actual.append(target[test_idx])
    return np.corrcoef(np.concatenate(predicted), np.concatenate(actual))[0, 1]

# Shuffled k-fold: test days are scattered, so each has its neighbours in time in the training set.
order = rng.permutation(days)
shuffled = [(np.setdiff1d(order, part), part) for part in np.array_split(order, folds)]

# Walk-forward: train only on the past, with a gap of `horizon` days so no target overlaps.
blocks = np.array_split(np.arange(days), folds + 1)
walk_forward = [(np.arange(0, block[0] - horizon), block) for block in blocks[1:]]

print(f"shuffled k-fold correlation:  {score(shuffled):+.3f}")
print(f"walk-forward correlation:     {score(walk_forward):+.3f}")

It prints a shuffled correlation of +0.580+0.580 and a walk-forward correlation of −0.177-0.177. Running other seeds gives shuffled scores between about 0.55 and 0.65 every time, and walk-forward scores scattered around zero, between about −0.2-0.2 and +0.1+0.1. That scatter is itself a lesson: with overlapping targets, even an honest score of ±0.15\pm 0.15 can be noise.

Where this shows up in quant work

  • Model evaluation. Every machine learning result on financial data should come with how it was validated. "5-fold cross-validation" on daily data with multi-day targets is a red flag.
  • Feature research. Feature selection and hyperparameter tuning must happen inside each training window, never on the full dataset, or the leak comes back through the side door.
  • Interviews. Explaining why random cross-validation fails on time series, and what to use instead, is one of the most common machine learning questions for quant research roles.

Exercises

  • Change the target horizon from 20 days to 1 day (no overlap). How do the two scores change, and why does shuffled validation stop leaking so badly?
  • Remove the 20-day gap from the walk-forward splits. Does the score move? Try a 60-day feature window with a 0-day gap and a 60-day gap.
  • Replace kk-nearest neighbours with ordinary linear regression on the three features. Does shuffled validation still leak? What does that say about model flexibility and leakage?

Key takeaways

  • Shuffled cross-validation on time series leaks through overlapping labels and slow-moving features, and can report a strong signal in pure noise.
  • Walk-forward validation trains on the past and tests on the future, with a gap at least as long as the label horizon.
  • Honest scores on financial data are noisy, because overlapping targets leave few independent observations.
  • Do all preprocessing, feature selection and tuning inside each training window.