Skip to main content
QuantDXB

Data and machine learning · 12 min read

Overfitting

Why training error always falls and test error doesn't: bias, variance, the noise floor, and how to keep financial models honest.

Before you start

  • Conditional expectation
  • Thinking in arrays with NumPy

By the end you'll be able to

  • Distinguish training error from test error
  • Decompose test error into bias, variance and noise
  • Explain why financial data is especially prone to overfitting
  • Use holdout data, simplicity and regularisation to control it

Give a model enough freedom and it will fit any history perfectly, including all the noise in it. The fit then tells you nothing about the future. In finance this is the default outcome rather than an occasional risk, because the signal is weak and the noise is large. This lesson makes overfitting visible, measures it, and sets out the habits that keep it under control.

TermMeaning
training errorHow badly the model fits the data it was fitted on
test errorHow badly it predicts new data from the same source
biasError from a model too simple to capture the real pattern
varianceError from a model so flexible its fit changes with every sample of noise
noise floorThe error no model can beat: the variance of the unpredictable part
regularisationPenalising complexity so the fit prefers simple explanations

Fitting a curve

The chart below has 30 points from a smooth hidden signal plus noise. Fit a polynomial of degree dd by least squares and compare two numbers: the error on those 30 points (training error), and the error on fresh points from the same source (test error). A new sample of 30 points is drawn every few seconds, and the faint lines are the fits to earlier samples.

3
Training error
–
Test error
–
Noise floor
0.090
A new training sample is drawn every few seconds; faint lines are fits to earlier samples. Low degrees miss the shape (bias). High degrees chase the noise: training error keeps falling, test error explodes, and the fit changes completely from one sample to the next (variance).
  • Degree 1 is a straight line. It can't follow the curve, so both errors are high, and the fit barely changes between samples. That is bias.
  • Degree 3 follows the signal and ignores most of the noise. Test error is close to the noise floor, the 0.09 that no model can beat.
  • Degree 15 passes close to every training point. Training error keeps falling, but between the points, and especially near the edges, the curve swings wildly, and it is completely different for each new sample. That is variance, and test error explodes.

Key idea. Training error always falls as a model gets more flexible. Only error on data the model hasn't seen tells you whether the extra flexibility found signal or memorised noise.

Bias and variance

For squared error, the expected test error of a fitting method splits into three parts:

E[(y−f^(x))2]=(f(x)−E[f^(x)])2⏟bias2+Var⁡(f^(x))⏟variance+σ2⏟noise.\mathbb{E}\big[(y - \hat{f}(x))^2\big] = \underbrace{\big(f(x) - \mathbb{E}[\hat{f}(x)]\big)^2}_{\text{bias}^2} + \underbrace{\operatorname{Var}\big(\hat{f}(x)\big)}_{\text{variance}} + \underbrace{\sigma^2}_{\text{noise}}.

Simple models have high bias and low variance; flexible ones the reverse. Test error is lowest where the two balance, which is the U-shaped green curve in the chart's right panel. The noise term is a hard floor: in markets it is most of the variance of returns (the conditional expectation lesson measured it), which is why financial models overfit so easily.

Why finance is especially exposed

  • Weak signals. If a feature explains 1% of the variance of returns, a model flexible enough to explain 10% in a backtest is explaining at least 9% noise.
  • Few independent observations. Ten years of daily data is 2,500 points, but returns are noisy, regimes change, and overlapping targets reduce the effective sample further.
  • Many tries. Every parameter, feature and model choice tested against the same history adds flexibility, even if each model looks simple. The false discoveries lesson showed how fast this adds up.

A strategy with 12 tunable parameters fitted on the same history it is evaluated on is a degree-15 polynomial in disguise.

Keeping it under control

  • Hold out data. Fit on one period and test on a later one you haven't looked at. For time series, the test period must come after the training period; the walk-forward validation lesson covers how.
  • Prefer simple models. Fewer parameters, each with a reason to exist, and features with an economic story, not ones found by search.
  • Regularise. Ridge and lasso regression add a penalty on the size of coefficients, so the fit only uses flexibility the data strongly supports. The penalty's strength is itself chosen by validation.
  • Check stability. If the best parameters change a lot between subsamples, or performance collapses when they move slightly, the result is fitted noise.

In code

python
import numpy as np
from numpy.polynomial import chebyshev as C

rng = np.random.default_rng(0)

def truth(x):
    return np.sin(2.5 * x)

x_test = np.linspace(-1, 1, 400)
train_err, test_err = {}, {}
for trial in range(500):  # 500 independent training samples of 30 points
    x_tr = rng.uniform(-1, 1, 30)
    y_tr = truth(x_tr) + 0.3 * rng.standard_normal(30)
    y_te = truth(x_test) + 0.3 * rng.standard_normal(400)
    for d in (1, 3, 5, 9, 15):
        coef = C.chebfit(x_tr, y_tr, d)  # least squares in a numerically stable basis
        train_err.setdefault(d, []).append(np.mean((C.chebval(x_tr, coef) - y_tr) ** 2))
        test_err.setdefault(d, []).append(np.mean((C.chebval(x_test, coef) - y_te) ** 2))

for d in (1, 3, 5, 9, 15):
    print(f"degree {d:>2}: training error {np.mean(train_err[d]):.3f}, "
          f"test error median {np.median(test_err[d]):.3f}")
DegreeTraining errorTest error (median of 500)
10.1520.180
30.0780.103
50.0710.114
90.0600.182
150.04270.310

Training error falls at every step. Test error is lowest at degree 3, close to the noise floor of 0.09, rises again by degree 9, and at degree 15 is several hundred times worse than the simple model: a fit to 30 points with 16 coefficients is mostly noise. The test error uses the median because a few samples produce astronomically bad fits; the mean is dominated by them.

The fit uses the Chebyshev polynomial basis rather than plain powers 1,x,x2,…1, x, x^2, \dots, because high powers of xx are nearly identical on [−1,1][-1, 1] and make least squares numerically unstable, a floating point problem in its own right.

Where this shows up in quant work

  • Strategy research. Optimising parameters on a backtest is curve fitting. The more you optimise, the more of the backtest's return is fitted noise.
  • Machine learning. Gradient boosting and neural networks can fit almost anything; with financial data they need strong regularisation and careful validation to find anything real.
  • Interviews. "Your model has a great backtest. How do you know it isn't overfitted?" is a standard question. Holdout data, the number of tries, stability and economic sense are the answer.

Exercises

  • Using the chart, find the lowest degree at which training error is below the noise floor of 0.09. Is that degree a good model?
  • Add ridge regularisation to the Python code: minimise squared error plus λ∑ck2\lambda \sum c_k^2. How does test error at degree 15 change as λ\lambda grows?
  • Repeat the experiment with 300 training points instead of 30. Which degree is best now, and why does more data allow a more flexible model?

Key takeaways

  • Training error always falls with flexibility; only unseen data shows whether a model learned signal or noise.
  • Test error = bias² + variance + noise. Simple models err through bias, flexible ones through variance.
  • Finance has weak signals, limited data and many tries, so overfitting is the default.
  • Hold out data, keep models simple, regularise, and check that results are stable.