Skip to main content

Time Series and Walk-Forward Evaluation

Forecasting next quarter's sales requires reconstructing the data available at each prediction time. Splitting a historical table downloaded today into earlier and later rows can still put subsequent revisions into the past. An evaluation must establish both what was knowable and how far ahead the model must predict.

Data Splits and Leakage defines the general information boundary; pandas Dates and Time Series covers parsing, resampling, and windows. Here those boundaries become a sequence of forecast origins, baseline models, and error calculations.

Separate observation time, release time, and vintage​

ALFRED's guide distinguishes observation periods from historical data vintages: economic data are revised, and ALFRED preserves versions available on past dates. New values and revisions are typically added within one business day of release, and missing release dates may be replaced by other dates. Intraday forecasting therefore needs the actual release timestamp and the system's receipt time as well.

Time or versionQuestion it answersHypothetical monthly sales example
Observation periodWhen did the activity occur?January 2026
Publication timeWhen could someone see the value?First release of 100 on February 10
Data vintageWhich estimate was visible on a historical date?January revised to 104 on March 10
System receipt timeWhen did the forecasting system actually obtain it?May follow public release

At a February 28 origin, January's usable value is 100. Suppose February's value of 90 is not published until March 12; it is unavailable too. This example uses Python's standard-library date.fromisoformat to parse release dates, filters by the cutoff, then takes the latest eligible value for each observation period. The cutoff means the end of the specified day. It assumes immediate receipt and at most one release per observation period per day; multiple same-day versions require timestamps or an explicit release order.

from datetime import date

rows = [
("2026-01", "2026-02-10", 100),
("2026-01", "2026-03-10", 104),
("2026-02", "2026-03-12", 90),
]
cutoff = date(2026, 2, 28)
latest = {}
dated_rows = [(period, date.fromisoformat(released), value)
for period, released, value in rows]
for period, released, value in sorted(dated_rows, key=lambda row: row[1]):
if released <= cutoff:
latest[period] = value
print("as_of", cutoff.isoformat(), latest)
assert latest == {"2026-01": 100}

Output:

as_of 2026-02-28 {'2026-01': 100}

Retain release histories. Filter by availability before selecting a vintage; selecting today's latest values and then filtering observation periods reverses the required order. Training labels must also be available by the training cutoff. Decide in advance whether evaluation targets the first release or a revision at a specified date. The Macro Scenario Lab already preserves such constraints through information cutoffs.

Inspect the structure before differencing​

StructureMeaningModeling question
TrendA long-term rise or fall in levelExtrapolate levels or forecast changes?
SeasonalityRepetition at a fixed calendar periodShould quarterly data use m=4m=4?
Temporal dependenceValues at different times are relatedCan past values help, and are errors related too?
StationarityStatistical properties remain stable over timeCan historical relationships describe the current training interval?

FPP's chapter on stationarity and differencing explains how differencing can reduce trend or seasonality. With strong seasonality, try seasonal differencing first and avoid unnecessary additional differences. Berkeley's time-series lecture notes define weak stationarity: a constant mean, a finite constant variance, and covariance depending only on the lag. It allows temporal dependence.

First differences are Δyt=yt−yt−1\Delta y_t=y_t-y_{t-1}; seasonal differences are Δmyt=yt−yt−m\Delta_m y_t=y_t-y_{t-m}. In the hypothetical sales below, y1=100y_1=100, y2=120y_2=120, and y5=110y_5=110, giving Δy2=20\Delta y_2=20 and Δ4y5=10\Delta_4 y_5=10. These answer different questions: change since last quarter and change since the same quarter last year. Models such as Holt–Winters handle trend and seasonality directly without prior differencing. Apparent stationarity within training does not establish that the next period will retain the same relationships.

Build baselines from available lags​

FPP's simple forecasting methods provide two useful baselines: a naive forecast repeats the last observation, and a seasonal-naive forecast repeats the latest observation from the matching season. Let tt be the forecast origin and hh the horizon, assuming observations through tt are available. The seasonal baseline also requires a complete cycle, so t≥mt\ge m:

y^t+h∣tnaive=yt,y^t+h∣tseasonal=yt+h−m,1≤h≤m.\hat y_{t+h\mid t}^{\mathrm{naive}}=y_t,\qquad \hat y_{t+h\mid t}^{\mathrm{seasonal}}=y_{t+h-m},\quad 1\le h\le m.

Beyond one seasonal cycle, seasonal-naive forecasts keep repeating the last complete cycle; they must not read future actuals. Here the data are quarterly, m=4m=4, and h=1,2h=1,2.

Features might include past sales, a trailing four-quarter mean, and the target quarter's calendar information. If a training row is indexed by its target time t+ht+h, its lag-kk sales feature is yt+h−ky_{t+h-k}. Under immediate publication, it stays within the origin only when k≥hk\ge h. With release delays, availability must also be checked; shifting a column by one row is insufficient.

Construct historical training features using what was available at each row's original prediction time, and fit at the current origin only with labels already released. Fit mean imputation, scaling, feature selection, and differencing choices within that training interval. FPP's regression forecasting chapter distinguishes inputs known in advance from those observed later: future calendar information or schedules already known can be used, while realized future exchange rates, weather observations, and revisions released after the cutoff cannot be treated as known inputs. Revisions already released by the cutoff are usable. To reconstruct levels, add first-difference forecasts to the previous level and seasonal-difference forecasts to the level mm periods earlier. Start with levels known at or before the origin; use earlier forecasts for any post-origin levels needed in the reconstruction. Compare the restored levels with level baselines.

Move the forecast origin forward​

FPP's time-series cross-validation uses observations preceding each test point to construct forecasts and supports multi-step evaluation. Apply the release constraints above as well. Reserve the final period first, then use earlier origins for model comparison. Earlier validation observations may enter subsequent training after they become available.

Consider these synthetic quarterly sales, all in the same units, released immediately at quarter-end and never revised:

QuarterQ1Q2Q3Q4Q5Q6Q7Q8Q9Q10Q11Q12Q13Q14
Sales1001209013011013299143121145109157133160
OriginAvailable training intervalOne-step targetTwo-step targetRole
End of Q8Q1–Q8Q9Q10Validation
End of Q9Q1–Q9Q10Q11Validation
End of Q10Q1–Q10Q11Q12Validation
End of Q12Q1–Q12Q13Q14Final test after model selection is locked

This is an expanding window, retaining all available history. A fixed rolling window keeps only the most recent WW observations. Choose WW, minimum training length, and refitting frequency during validation. This example starts with two complete seasonal cycles and uses validation origins whose two-step targets both fall before Q13.

Alongside the baselines, fit a candidate called seasonal_drift: add the training mean of same-quarter annual changes, dtd_t, to the seasonal-naive forecast. It assumes a common absolute annual increment across quarters and is used here only at horizons one and two. The fit helper below prepares the state for all three candidates. Its input must contain complete, finite, regularly spaced quarterly values in time order, with t>mt>m so at least one annual change can be calculated. forecast accepts only the listed model names and integer horizons 1≤h≤m1\le h\le m.

dt=1t−m∑i=m+1t(yi−yi−m),t>m,y^t+h∣tseasonal_drift=yt+h−m+dt,1≤h≤m.d_t=\frac{1}{t-m}\sum_{i=m+1}^{t}(y_i-y_{i-m}),\quad t>m,\qquad \hat y_{t+h\mid t}^{\mathrm{seasonal\_drift}}=y_{t+h-m}+d_t,\quad 1\le h\le m.

At Q8, the annual changes are 10, 12, 9, and 13, averaging 11. The forecasts are therefore 110+11=121110+11=121 for Q9 and 132+11=143132+11=143 for Q10. Refitting at Q10 gives d10=68/6≈11.333d_{10}=68/6\approx11.333. The code refits at every origin and reads targets only after forecasting to calculate errors. Selection uses MAE with equal weights for the two horizons.

from math import sqrt
from statistics import mean

# Synthetic quarterly values; Q13-Q14 are reserved for the final test.
y = [100, 120, 90, 130, 110, 132, 99, 143,
121, 145, 109, 157, 133, 160]
m, H, dev_end = 4, 2, 12
names = ("naive", "seasonal", "seasonal_drift")

def fit(train):
if len(train) <= m:
raise ValueError("fit requires more than m observations")
change = mean(train[i] - train[i-m] for i in range(m, len(train)))
return train[-1], train[-m:], change

def forecast(state, name, h):
last, season, change = state
if name not in names:
raise ValueError("unknown model name")
if not isinstance(h, int) or not 1 <= h <= m:
raise ValueError("h must be an integer between 1 and m")
if name == "naive":
return last
return season[h-1] + (change if name == "seasonal_drift" else 0)

errors = {(name, h): [] for name in names for h in range(1, H+1)}
print("origin h actual naive seasonal seasonal_drift")
for t in (8, 9, 10):
assert t + H <= dev_end
state = fit(y[:t])
for h in range(1, H+1):
predictions = [forecast(state, name, h) for name in names]
actual = y[t+h-1] # Read the target only to score the forecasts.
print(t, h, actual, *(f"{p:.3f}" for p in predictions))
for name, p in zip(names, predictions):
errors[name, h].append(actual - p)

print("model h MAE RMSE")
for name in names:
for h in range(1, H+1):
e = errors[name, h]
print(name, h, f"{mean(abs(v) for v in e):.3f}",
f"{sqrt(mean(v*v for v in e)):.3f}")

# Selection rule: MAE with equal weights for the two horizons.
chosen = min(names, key=lambda name: mean(
abs(v) for h in range(1, H+1) for v in errors[name, h]))
print("selected", chosen)
state = fit(y[:dev_end])
print("final_origin h actual forecast error")
for h in range(1, H+1):
p = forecast(state, chosen, h)
actual = y[dev_end+h-1]
print(dev_end, h, actual, f"{p:.3f}", f"{actual-p:.3f}")

Output:

origin h actual naive seasonal seasonal_drift
8 1 121 143.000 110.000 121.000
8 2 145 143.000 132.000 143.000
9 1 145 121.000 132.000 143.000
9 2 109 121.000 99.000 110.000
10 1 109 145.000 99.000 110.333
10 2 157 145.000 143.000 154.333
model h MAE RMSE
naive 1 27.333 28.024
naive 2 8.667 9.866
seasonal 1 11.333 11.402
seasonal 2 12.333 12.450
seasonal_drift 1 1.111 1.388
seasonal_drift 2 1.889 2.009
selected seasonal_drift
final_origin h actual forecast error
12 1 133 132.500 0.500
12 2 160 156.500 3.500

Interpret errors by horizon​

FPP's point forecast accuracy chapter distinguishes training residuals from genuine out-of-sample forecast errors and defines MAE, RMSE, and percentage errors. For a fixed hh, let nhn_h be the number of forecasts at that horizon:

et,h=yt+h−y^t+h∣t,MAEh=1nh∑t∣et,h∣,RMSEh=1nh∑tet,h2.e_{t,h}=y_{t+h}-\hat y_{t+h\mid t},\qquad \mathrm{MAE}_h=\frac{1}{n_h}\sum_t|e_{t,h}|,\qquad \mathrm{RMSE}_h=\sqrt{\frac{1}{n_h}\sum_t e_{t,h}^2}.

MAE has the same units as sales; RMSE gives large errors more weight. MAPE is undefined for zero actuals and unstable near zero. MASE can compare series with different units, but its scaling denominator must come from training; it is undefined when all seasonal differences in that denominator are zero.

The one-step seasonal_drift errors are 0,2,−4/30,2,-4/3, giving MAE 10/9≈1.11110/9\approx1.111. The two-step errors are 2,−1,8/32,-1,8/3, giving MAE 17/9≈1.88917/9\approx1.889. Their equally weighted mean is 1.500, below 18.000 for naive and 11.833 for seasonal naive. Inspect each horizon before aggregating with business-relevant weights so a single average cannot conceal poor longer-range forecasts.

These errors cannot be assumed to be six independent trials. Q10 and Q11 each appear as targets from different origins. A direct derivation illustrates the issue: suppose yt=yt−1+εty_t=y_{t-1}+\varepsilon_t, with independent increments of mean zero and variance 0<σ2<∞0<\sigma^2<\infty. The two-step naive error is et,2=εt+1+εt+2e_{t,2}=\varepsilon_{t+1}+\varepsilon_{t+2}; at the next origin it is et+1,2=εt+2+εt+3e_{t+1,2}=\varepsilon_{t+2}+\varepsilon_{t+3}. The shared increment gives covariance σ2\sigma^2, each error has variance 2σ22\sigma^2, and their correlation is 1/21/2.

Compare model losses at matching origins and horizons, and account for temporal dependence when estimating uncertainty. Reducing overlap alone does not guarantee independent errors in real data. Three validation origins suffice to demonstrate the calculation but cannot support a reliable significance conclusion.

Connect forecast evaluation to a trading backtest​

A trading rule can lose money despite small forecast errors. Define when signals become available, the first executable trade, positions, turnover, and costs. A signal formed after the close cannot be assumed to have traded at that same closing price. The historical investment universe must also include eligible securities that later delisted; using only today's survivors creates survivorship bias.

Suppose a cash-funded trade has buy notional of 100,000 currency units and sell notional of 100,120 at pre-cost reference prices, giving a gross profit of 120. The cost model charges commission, spread, and slippage together at 8 basis points of the corresponding notional per leg. One basis point is 0.01%, so buying costs 100000×0.0008=80100000\times0.0008=80 and selling costs 100120×0.0008=80.096100120\times0.0008=80.096, totaling 160.096. Net profit is 120−160.096=−40.096120-160.096=-40.096. These fees are hypothetical inputs; add borrowing, financing, or market impact when they apply to the strategy.

Repeatedly choosing models, windows, or trading thresholds makes validation data part of strategy design. Bailey and colleagues' The Probability of Backtest Overfitting analyzes this selection effect and explains why a single holdout cannot assess bias from many preceding trials. Retain all candidates, including failed trials, compare results after costs, and finish selection before the final test.

Preserve a final period unused for selection​

Before the final test, lock features, windows, horizons, loss weights, refitting policy, and trading and cost rules. The program selects seasonal_drift using validation, then fits Q1–Q12, obtaining a mean annual change of 11.5. Its Q13 and Q14 forecasts are 132.5 and 156.5, with errors of 0.5 and 3.5. Final MAE is 2.000 and RMSE is 2.500. Only the selected model is scored here.

A real final test can continue rolling refits under a previously locked policy and use earlier test-period observations once available. Updating training data does not reopen model selection. Report errors at fixed horizons, along with date ranges, available vintages, and net trading results. If final results cause a model change, that period has become validation data; a new period unused for selection is needed. For familiar market history, an untouched period also requires that strategy design did not draw on remembered market outcomes. A new period recorded prospectively makes that condition clearer.

For the modeling questions in Quantitative Finance, first establish the data available at the time, then ask whether the model consistently beats baselines with the same information and horizons, and finally whether the improvement survives executable prices and full costs.

Explore connectionsOpen network