Data Splits and Leakage
Data leakage occurs when training, feature construction, selection, or evaluation uses information that would not legitimately be available for the target use, making results optimistic. High training accuracy alone is not leakage, and not every strong proxy is illegal information.
Define the Information Boundary
unit prediction object
prediction_time when the output must be made
horizon how far ahead the target lies
available_info information genuinely available and allowed then
label_time when the label occurs, matures, and may be revised
split_key entities, groups, and times that must not cross sets
Legitimacy is relative to the deployment information set . A field existing in a database does not mean it existed at prediction time or is permitted for use.
Leakage Types
A deployment-available proxy may still be a shortcut, bias, or brittle correlate without being leakage. Use subgroup, intervention, and shift tests rather than relabeling every problem.
Split to Simulate Use
Stratification preserves label proportions; it does not prevent group, time, or duplicate overlap.
Open full-size imageFirst separate the test set on the right. On the left, each row uses one blue fold for validation and the green folds for fitting. Fit preprocessing again inside each training fold. The test set stays outside model selection and is used only for final evaluation; five folds here illustrate the procedure rather than prescribe a universal split.
Pipeline Invariant
For training indices and validation indices , every learned transform should satisfy
This includes scaling, imputation, vocabulary, PCA, feature selection, target encoding, oversampling, and learned augmentation. Cross-validation refits each transform per fold. Nested CV uses an outer loop for evaluation and an inner loop for selection; selecting features globally before CV still leaks.
A Tiny Train-Only Transformation
The scikit-learn leakage guide uses the same rule: learn preprocessing on training data, then reuse it unchanged. Suppose training feature values are 2 and 4, while validation contains 100. Training-only centering learns 3, so validation becomes 97. Global centering would learn and let validation change the representation of the training rows.
train = [("A", 2.0), ("B", 4.0)]
valid = [("C", 100.0)]
assert {g for g, _ in train}.isdisjoint(g for g, _ in valid)
center = sum(x for _, x in train) / len(train)
assert [x - center for _, x in train] == [-1.0, 1.0]
assert [x - center for _, x in valid] == [97.0]
This toy split represents new entities. Predicting future visits of already known patients can legitimately use their available past histories; a group-disjoint test instead asks about unseen patients. For a future-new-patient test, choose a time cutoff, retain only earlier training rows whose labels have matured by that cutoff, and evaluate later rows from patients absent from training. Report the excluded returning patients; this restriction changes the population.
A pipeline prevents cross-fold fitting mistakes, not every within-training target leak. Target encoding can let a row's own label enter its feature; use out-of-fold or appropriately ordered training encodings. Resampling changes training data only, never validation/test prevalence.
Worked Example: 30-Day Readmission
The system must predict at discharge whether a patient returns within 30 days. A flawed design randomly splits visits, uses billing states completed after the 30-day window, selects codes on all data, selects models by test ROC-AUC, and tunes thresholds by test recall.
A stronger design freezes the prediction-time field list; groups by patient and holds out a later cohort; fits imputation, vocabulary, and selection on training only; chooses model and threshold on validation; runs test once with hospital/time/subgroup slices; and excludes examples whose 30-day labels have not matured.
A lower score may reveal that the old system relied on identity or future information rather than that the corrected system became worse.
Leakage Probes
- record event, ingestion, and revision time for every feature;
- train suspicious features alone and trace anomalous performance;
- remove identifiers and high-cardinality proxies;
- audit exact and near duplicates;
- run label-permutation controls;
- assert groups do not cross sets and transforms never fit on test;
- acquire a fresh holdout after repeated benchmark inspection.
Permutation tests do not find every leak: if a direct label-coded feature is permuted together with the label, their relation may remain. Combine controls with lineage.
Distinctions
Overfitting learns sample noise through a valid pipeline. Leakage crosses an information boundary. Distribution shift changes the deployment distribution. Confounding distorts a causal interpretation. All four can coexist.
prediction_time: explicit
label_horizon_and_maturity: explicit
unit_and_group_key: explicit
split_algorithm_and_seed: versioned
all_learned_transforms_fit_on_train_only: true
test_used_for_selection: false
duplicate_and_overlap_audit: complete
feature_availability_lineage: reviewed
known_residual_risks: listed
This note focuses on offline supervised prediction. Online experiments, RL, federated learning, and continual learning add intervention, feedback, and policy leakage.