Pandas
Pandas is a Python library for loading, selecting, cleaning, and summarizing tables that fit in memory. Rows and columns have labels, and different columns can store different data types, such as text, numbers, or dates. Use the table below to find a task or start with Series and DataFrame to understand the objects. These notes focus on common operations; the official documentation covers the full API and the limits of in-memory processing.
Working path
House rules
- Make labels, dtypes, units, time zones, and missing-value policy explicit.
- Use
.locfor labels and.ilocfor positions; do not rely on ambiguous[]behavior. - Prefer vectorized operations,
agg, andtransformbefore row-wiseapply. - Validate merge cardinality and inspect unmatched keys.
- Treat a DataFrame as an intermediate analytical object, not a database or a distributed compute engine.
Scope and versions
These notes cover small in-memory tables, from construction and CSV ingestion through selection, cleaning, joining, summarizing, reshaping, and time handling. They assume Python lists, dictionaries, slicing, and functions; they do not cover distributed execution or every storage backend. Examples use import pandas as pd, with local sample data on the operation pages. Check pd.__version__ when reproducing an older notebook.
The examples target pandas 2.2+ APIs and are also applicable to pandas 3.x with the stated dtype and parameter choices. Copy-on-Write is the only mode in pandas 3.0: writing to an independently selected object does not update its parent. In pandas 2.x this behavior is optional, so the notes use explicit assignment to the intended object rather than depending on view mutation. alias = df still binds the same object in either version; it is not a selection or copy.