Time Series and Walk-Forward Evaluation
Forecasting next quarter's sales requires reconstructing the data available at each prediction time. Splitting a historical table downloaded today into earlier and later rows can still put subsequent revisions into the past. An evaluation must establish both what was knowable and how far ahead the model must predict.
Data Splits and Leakage defines the general information boundary; pandas Dates and Time Series covers parsing, resampling, and windows. Here those boundaries become a sequence of forecast origins, baseline models, and error calculations.
Separate observation time, release time, and vintage
ALFRED's guide distinguishes observation periods from historical data vintages: economic data are revised, and ALFRED preserves versions available on past dates. New values and revisions are typically added within one business day of release, and missing release dates may be replaced by other dates. Intraday forecasting therefore needs the actual release timestamp and the system's receipt time as well.
At a February 28 origin, January's usable value is 100. Suppose February's value of 90 is not published until March 12; it is unavailable too. This example uses Python's standard-library date.fromisoformat to parse release dates, filters by the cutoff, then takes the latest eligible value for each observation period. The cutoff means the end of the specified day. It assumes immediate receipt and at most one release per observation period per day; multiple same-day versions require timestamps or an explicit release order.
from datetime import date
rows = [
("2026-01", "2026-02-10", 100),
("2026-01", "2026-03-10", 104),
("2026-02", "2026-03-12", 90),
]
cutoff = date(2026, 2, 28)
latest = {}
dated_rows = [(period, date.fromisoformat(released), value)
for period, released, value in rows]
for period, released, value in sorted(dated_rows, key=lambda row: row[1]):
if released <= cutoff:
latest[period] = value
print("as_of", cutoff.isoformat(), latest)
assert latest == {"2026-01": 100}
Output:
as_of 2026-02-28 {'2026-01': 100}
Retain release histories. Filter by availability before selecting a vintage; selecting today's latest values and then filtering observation periods reverses the required order. Training labels must also be available by the training cutoff. Decide in advance whether evaluation targets the first release or a revision at a specified date. The Macro Scenario Lab already preserves such constraints through information cutoffs.
Inspect the structure before differencing
FPP's chapter on stationarity and differencing explains how differencing can reduce trend or seasonality. With strong seasonality, try seasonal differencing first and avoid unnecessary additional differences. Berkeley's time-series lecture notes define weak stationarity: a constant mean, a finite constant variance, and covariance depending only on the lag. It allows temporal dependence.
First differences are ; seasonal differences are . In the hypothetical sales below, , , and , giving and . These answer different questions: change since last quarter and change since the same quarter last year. Models such as Holt–Winters handle trend and seasonality directly without prior differencing. Apparent stationarity within training does not establish that the next period will retain the same relationships.
Build baselines from available lags
FPP's simple forecasting methods provide two useful baselines: a naive forecast repeats the last observation, and a seasonal-naive forecast repeats the latest observation from the matching season. Let be the forecast origin and the horizon, assuming observations through are available. The seasonal baseline also requires a complete cycle, so :
Beyond one seasonal cycle, seasonal-naive forecasts keep repeating the last complete cycle; they must not read future actuals. Here the data are quarterly, , and .
Features might include past sales, a trailing four-quarter mean, and the target quarter's calendar information. If a training row is indexed by its target time , its lag- sales feature is . Under immediate publication, it stays within the origin only when . With release delays, availability must also be checked; shifting a column by one row is insufficient.
Construct historical training features using what was available at each row's original prediction time, and fit at the current origin only with labels already released. Fit mean imputation, scaling, feature selection, and differencing choices within that training interval. FPP's regression forecasting chapter distinguishes inputs known in advance from those observed later: future calendar information or schedules already known can be used, while realized future exchange rates, weather observations, and revisions released after the cutoff cannot be treated as known inputs. Revisions already released by the cutoff are usable. To reconstruct levels, add first-difference forecasts to the previous level and seasonal-difference forecasts to the level periods earlier. Start with levels known at or before the origin; use earlier forecasts for any post-origin levels needed in the reconstruction. Compare the restored levels with level baselines.
Move the forecast origin forward
FPP's time-series cross-validation uses observations preceding each test point to construct forecasts and supports multi-step evaluation. Apply the release constraints above as well. Reserve the final period first, then use earlier origins for model comparison. Earlier validation observations may enter subsequent training after they become available.
Consider these synthetic quarterly sales, all in the same units, released immediately at quarter-end and never revised:
This is an expanding window, retaining all available history. A fixed rolling window keeps only the most recent observations. Choose , minimum training length, and refitting frequency during validation. This example starts with two complete seasonal cycles and uses validation origins whose two-step targets both fall before Q13.
Alongside the baselines, fit a candidate called seasonal_drift: add the training mean of same-quarter annual changes, , to the seasonal-naive forecast. It assumes a common absolute annual increment across quarters and is used here only at horizons one and two. The fit helper below prepares the state for all three candidates. Its input must contain complete, finite, regularly spaced quarterly values in time order, with so at least one annual change can be calculated. forecast accepts only the listed model names and integer horizons .
At Q8, the annual changes are 10, 12, 9, and 13, averaging 11. The forecasts are therefore for Q9 and for Q10. Refitting at Q10 gives . The code refits at every origin and reads targets only after forecasting to calculate errors. Selection uses MAE with equal weights for the two horizons.
from math import sqrt
from statistics import mean
# Synthetic quarterly values; Q13-Q14 are reserved for the final test.
y = [100, 120, 90, 130, 110, 132, 99, 143,
121, 145, 109, 157, 133, 160]
m, H, dev_end = 4, 2, 12
names = ("naive", "seasonal", "seasonal_drift")
def fit(train):
if len(train) <= m:
raise ValueError("fit requires more than m observations")
change = mean(train[i] - train[i-m] for i in range(m, len(train)))
return train[-1], train[-m:], change
def forecast(state, name, h):
last, season, change = state
if name not in names:
raise ValueError("unknown model name")
if not isinstance(h, int) or not 1 <= h <= m:
raise ValueError("h must be an integer between 1 and m")
if name == "naive":
return last
return season[h-1] + (change if name == "seasonal_drift" else 0)
errors = {(name, h): [] for name in names for h in range(1, H+1)}
print("origin h actual naive seasonal seasonal_drift")
for t in (8, 9, 10):
assert t + H <= dev_end
state = fit(y[:t])
for h in range(1, H+1):
predictions = [forecast(state, name, h) for name in names]
actual = y[t+h-1] # Read the target only to score the forecasts.
print(t, h, actual, *(f"{p:.3f}" for p in predictions))
for name, p in zip(names, predictions):
errors[name, h].append(actual - p)
print("model h MAE RMSE")
for name in names:
for h in range(1, H+1):
e = errors[name, h]
print(name, h, f"{mean(abs(v) for v in e):.3f}",
f"{sqrt(mean(v*v for v in e)):.3f}")
# Selection rule: MAE with equal weights for the two horizons.
chosen = min(names, key=lambda name: mean(
abs(v) for h in range(1, H+1) for v in errors[name, h]))
print("selected", chosen)
state = fit(y[:dev_end])
print("final_origin h actual forecast error")
for h in range(1, H+1):
p = forecast(state, chosen, h)
actual = y[dev_end+h-1]
print(dev_end, h, actual, f"{p:.3f}", f"{actual-p:.3f}")
Output:
origin h actual naive seasonal seasonal_drift
8 1 121 143.000 110.000 121.000
8 2 145 143.000 132.000 143.000
9 1 145 121.000 132.000 143.000
9 2 109 121.000 99.000 110.000
10 1 109 145.000 99.000 110.333
10 2 157 145.000 143.000 154.333
model h MAE RMSE
naive 1 27.333 28.024
naive 2 8.667 9.866
seasonal 1 11.333 11.402
seasonal 2 12.333 12.450
seasonal_drift 1 1.111 1.388
seasonal_drift 2 1.889 2.009
selected seasonal_drift
final_origin h actual forecast error
12 1 133 132.500 0.500
12 2 160 156.500 3.500
Interpret errors by horizon
FPP's point forecast accuracy chapter distinguishes training residuals from genuine out-of-sample forecast errors and defines MAE, RMSE, and percentage errors. For a fixed , let be the number of forecasts at that horizon:
MAE has the same units as sales; RMSE gives large errors more weight. MAPE is undefined for zero actuals and unstable near zero. MASE can compare series with different units, but its scaling denominator must come from training; it is undefined when all seasonal differences in that denominator are zero.
The one-step seasonal_drift errors are , giving MAE . The two-step errors are , giving MAE . Their equally weighted mean is 1.500, below 18.000 for naive and 11.833 for seasonal naive. Inspect each horizon before aggregating with business-relevant weights so a single average cannot conceal poor longer-range forecasts.
These errors cannot be assumed to be six independent trials. Q10 and Q11 each appear as targets from different origins. A direct derivation illustrates the issue: suppose , with independent increments of mean zero and variance . The two-step naive error is ; at the next origin it is . The shared increment gives covariance , each error has variance , and their correlation is .
Compare model losses at matching origins and horizons, and account for temporal dependence when estimating uncertainty. Reducing overlap alone does not guarantee independent errors in real data. Three validation origins suffice to demonstrate the calculation but cannot support a reliable significance conclusion.
Connect forecast evaluation to a trading backtest
A trading rule can lose money despite small forecast errors. Define when signals become available, the first executable trade, positions, turnover, and costs. A signal formed after the close cannot be assumed to have traded at that same closing price. The historical investment universe must also include eligible securities that later delisted; using only today's survivors creates survivorship bias.
Suppose a cash-funded trade has buy notional of 100,000 currency units and sell notional of 100,120 at pre-cost reference prices, giving a gross profit of 120. The cost model charges commission, spread, and slippage together at 8 basis points of the corresponding notional per leg. One basis point is 0.01%, so buying costs and selling costs , totaling 160.096. Net profit is . These fees are hypothetical inputs; add borrowing, financing, or market impact when they apply to the strategy.
Repeatedly choosing models, windows, or trading thresholds makes validation data part of strategy design. Bailey and colleagues' The Probability of Backtest Overfitting analyzes this selection effect and explains why a single holdout cannot assess bias from many preceding trials. Retain all candidates, including failed trials, compare results after costs, and finish selection before the final test.
Preserve a final period unused for selection
Before the final test, lock features, windows, horizons, loss weights, refitting policy, and trading and cost rules. The program selects seasonal_drift using validation, then fits Q1–Q12, obtaining a mean annual change of 11.5. Its Q13 and Q14 forecasts are 132.5 and 156.5, with errors of 0.5 and 3.5. Final MAE is 2.000 and RMSE is 2.500. Only the selected model is scored here.
A real final test can continue rolling refits under a previously locked policy and use earlier test-period observations once available. Updating training data does not reopen model selection. Report errors at fixed horizons, along with date ranges, available vintages, and net trading results. If final results cause a model change, that period has become validation data; a new period unused for selection is needed. For familiar market history, an untouched period also requires that strategy design did not draw on remembered market outcomes. A new period recorded prospectively makes that condition clearer.
For the modeling questions in Quantitative Finance, first establish the data available at the time, then ask whether the model consistently beats baselines with the same information and horizons, and finally whether the improvement survives executable prices and full costs.