Data and evaluation
Evaluation starts with the decision a result will support. Data provenance, prediction time and the cost of mistakes determine the split and metrics before a model is selected.
Start with dataset suitability and leakage-resistant splitting, then distribution shift. Calibration connects probabilities to thresholds and abstention. Local-model task replay and retrieval/generation evaluation bring these ideas into full systems; the evidence note explains how far a comparison supports a claim.
Reading Order
Use What You Read
Separate training, selection, calibration and final testing. Report the population and operating conditions along with a score, and preserve failures that would change the deployment decision.
Return to the AI reading paths to choose a neighboring topic.