Skip to main content

Data and evaluation

Evaluation starts with the decision a result will support. Data provenance, prediction time and the cost of mistakes determine the split and metrics before a model is selected.

Start with dataset suitability and leakage-resistant splitting, then distribution shift. Calibration connects probabilities to thresholds and abstention. Local-model task replay and retrieval/generation evaluation bring these ideas into full systems; the evidence note explains how far a comparison supports a claim.

Reading Order​

StepArticleWhat it explains
1Datasets: Provenance, Leakage, and EvaluationA map from data-generating processes and provenance to leakage-resistant splits, distribution-shift evaluation, and authoritative public sources.
2Data Splits and LeakageBlock target, group, temporal, and preprocessing leakage by defining prediction time, entities, and train-only pipelines.
3Model Evaluation Under Distribution ShiftStart from deployment distribution, splits, metrics, thresholds, and uncertainty instead of treating one test score as universal ability.
4Probability Calibration, Decision Thresholds, and AbstentionSeparate prediction accuracy, reliable probabilities, and worthwhile actions using independent calibration and evaluation data.
5Evaluating Local Models and Agent TasksA reproducible protocol for measuring whether one exact local model configuration clears a real workflow threshold.
6Evaluating Retrieval and Generation: From Evidence to Correct AnswersMeasure evidence retrieval, ranking, faithfulness, and answer quality separately to locate a RAG system’s bottleneck.
7Judging Evidence in AI ExperimentsA practical way to label specifications, vendor claims, benchmark results, observations, and recommendations without pretending they prove the same thing.

Use What You Read​

Separate training, selection, calibration and final testing. Report the population and operating conditions along with a score, and preserve failures that would change the deployment decision.

Return to the AI reading paths to choose a neighboring topic.

Explore connectionsOpen network