Skip to main content

Data Splits and Leakage

Data leakage occurs when training, feature construction, selection, or evaluation uses information that would not legitimately be available for the target use, making results optimistic. High training accuracy alone is not leakage, and not every strong proxy is illegal information.

Define the Information Boundary​

unit prediction object
prediction_time when the output must be made
horizon how far ahead the target lies
available_info information genuinely available and allowed then
label_time when the label occurs, matures, and may be revised
split_key entities, groups, and times that must not cross sets

Legitimacy is relative to the deployment information set I(t)\mathcal{I}(t). A field existing in a database does not mean it existed at prediction time or is permitted for use.

Leakage Types​

TypeExampleDistortionControl
target/post-outcomedischarge billing predicts admission-time riskfeature follows outcomeaudit event times and lineage
temporalrandom time-series split or future revisionstraining sees the future regimeforward split and data vintages
entity/groupone patient/device appears on both sidesidentity or history is memorizedgroup, sometimes group+time split
duplicatecrops or template copies cross setstest is not a new casegroup near-duplicates before splitting
preprocessingscaler, imputer, PCA, selection fit globallytest distribution changes transformsfit inside every training fold
model selectiontest repeatedly changes model/thresholdtest becomes validationlock test or obtain a new holdout
benchmark contaminationtraining contains questions and answersmeasures reproductionprovenance, deduplication, source/time isolation

A deployment-available proxy may still be a shortcut, bias, or brittle correlate without being leakage. Use subgroup, intervention, and shift tests rather than relabeling every problem.

Split to Simulate Use​

Deployment questionPreferred splitDoes not answer
new independent same-regime casesstratified randomtemporal change or identity memory
new patients/users/devicesgroupfuture policy change
future periodtemporal forwardnew institution
new institution/regionexternal/group holdoutevery external environment
spatially dependent casesblocked spatialarbitrary unseen-region extrapolation

Stratification preserves label proportions; it does not prevent group, time, or duplicate overlap.

A held-out test set remains separate while five-fold cross-validation rotates the validation fold within the training data.Open full-size image

First separate the test set on the right. On the left, each row uses one blue fold for validation and the green folds for fitting. Fit preprocessing again inside each training fold. The test set stays outside model selection and is used only for final evaluation; five folds here illustrate the procedure rather than prescribe a universal split.

Pipeline Invariant​

For training indices ItrI_{tr} and validation indices IvaI_{va}, every learned transform TηT_\eta should satisfy

η^=fit⁡(T,DItr),Xva′=Tη^(XIva).\hat\eta=\operatorname{fit}(T,D_{I_{tr}}), \qquad X'_{va}=T_{\hat\eta}(X_{I_{va}}).

This includes scaling, imputation, vocabulary, PCA, feature selection, target encoding, oversampling, and learned augmentation. Cross-validation refits each transform per fold. Nested CV uses an outer loop for evaluation and an inner loop for selection; selecting features globally before CV still leaks.

A Tiny Train-Only Transformation​

The scikit-learn leakage guide uses the same rule: learn preprocessing on training data, then reuse it unchanged. Suppose training feature values are 2 and 4, while validation contains 100. Training-only centering learns 3, so validation becomes 97. Global centering would learn 106/3106/3 and let validation change the representation of the training rows.

train = [("A", 2.0), ("B", 4.0)]
valid = [("C", 100.0)]
assert {g for g, _ in train}.isdisjoint(g for g, _ in valid)
center = sum(x for _, x in train) / len(train)
assert [x - center for _, x in train] == [-1.0, 1.0]
assert [x - center for _, x in valid] == [97.0]

This toy split represents new entities. Predicting future visits of already known patients can legitimately use their available past histories; a group-disjoint test instead asks about unseen patients. For a future-new-patient test, choose a time cutoff, retain only earlier training rows whose labels have matured by that cutoff, and evaluate later rows from patients absent from training. Report the excluded returning patients; this restriction changes the population.

A pipeline prevents cross-fold fitting mistakes, not every within-training target leak. Target encoding can let a row's own label enter its feature; use out-of-fold or appropriately ordered training encodings. Resampling changes training data only, never validation/test prevalence.

Worked Example: 30-Day Readmission​

The system must predict at discharge whether a patient returns within 30 days. A flawed design randomly splits visits, uses billing states completed after the 30-day window, selects codes on all data, selects models by test ROC-AUC, and tunes thresholds by test recall.

A stronger design freezes the prediction-time field list; groups by patient and holds out a later cohort; fits imputation, vocabulary, and selection on training only; chooses model and threshold on validation; runs test once with hospital/time/subgroup slices; and excludes examples whose 30-day labels have not matured.

A lower score may reveal that the old system relied on identity or future information rather than that the corrected system became worse.

Leakage Probes​

  • record event, ingestion, and revision time for every feature;
  • train suspicious features alone and trace anomalous performance;
  • remove identifiers and high-cardinality proxies;
  • audit exact and near duplicates;
  • run label-permutation controls;
  • assert groups do not cross sets and transforms never fit on test;
  • acquire a fresh holdout after repeated benchmark inspection.

Permutation tests do not find every leak: if a direct label-coded feature is permuted together with the label, their relation may remain. Combine controls with lineage.

Distinctions​

Overfitting learns sample noise through a valid pipeline. Leakage crosses an information boundary. Distribution shift changes the deployment distribution. Confounding distorts a causal interpretation. All four can coexist.

prediction_time: explicit
label_horizon_and_maturity: explicit
unit_and_group_key: explicit
split_algorithm_and_seed: versioned
all_learned_transforms_fit_on_train_only: true
test_used_for_selection: false
duplicate_and_overlap_audit: complete
feature_availability_lineage: reviewed
known_residual_risks: listed

This note focuses on offline supervised prediction. Online experiments, RL, federated learning, and continual learning add intervention, feedback, and policy leakage.

Explore connectionsOpen network