Artificial Intelligence
Choose a reading path for using AI, understanding models, or building local systems across eight connected topics.
Choose a reading path for using AI, understanding models, or building local systems across eight connected topics.
Ordered categories, finite vocabularies, binning, and encoding decisions.
Safe joins, concatenation, cardinality checks, and unmatched-key diagnostics.
What makes a result trustworthy for its intended use?
A reproducible workflow for normalizing, parsing, validating, and reporting tabular data.
Practical patterns for importing delimited text, spreadsheets, JSON, and native R data while handling common parsing problems.
Block target, group, temporal, and preprocessing leakage by defining prediction time, entities, and train-only pipelines.
The labeled two-dimensional table and its core invariants.
Label, position, and MultiIndex selection without ambiguous semantics.
A map from data-generating processes and provenance to leakage-resistant splits, distribution-shift evaluation, and authoritative public sources.
Parsing, time zones, periods, offsets, resampling, and rolling windows in pandas.
Split-apply-combine with explicit output-shape and missing-key decisions.
Use the site’s Python examples in Jupyter to explore sample sizes and distributions, then rerun them without hidden state.
A defensive ingestion workflow for tabular files and external data.
A policy-driven approach to detecting, interpreting, and handling missing data.
Start from deployment distribution, splits, metrics, thresholds, and uncertainty instead of treating one test score as universal ability.
A practical NumPy reference covering array creation, indexing, shapes, vectorized operations, aggregation, and array comparison.
A task-oriented map for tabular data work with pandas.
Composable, vectorized patterns and a decision guide for map, apply, agg, transform, and pipe.
Boolean filtering and readable predicates for DataFrames.
A short map for learning R through data structures, import, transformation, visualization, and statistical work.
R expressions, vector indexing, missing values, data structures, control flow, functions, and packages.
A guide to pivot, pivot_table, melt, stack, unstack, and explode.
Explicit selection, vectorized operations, mapping, and alignment for Series.
The one-dimensional labeled array at the core of pandas.