Skip to main content

Pandas

Pandas is a Python library for loading, selecting, cleaning, and summarizing tables that fit in memory. Rows and columns have labels, and different columns can store different data types, such as text, numbers, or dates. Use the table below to find a task or start with Series and DataFrame to understand the objects. These notes focus on common operations; the official documentation covers the full API and the limits of in-memory processing.

Working path​

House rules​

  1. Make labels, dtypes, units, time zones, and missing-value policy explicit.
  2. Use .loc for labels and .iloc for positions; do not rely on ambiguous [] behavior.
  3. Prefer vectorized operations, agg, and transform before row-wise apply.
  4. Validate merge cardinality and inspect unmatched keys.
  5. Treat a DataFrame as an intermediate analytical object, not a database or a distributed compute engine.

Scope and versions​

These notes cover small in-memory tables, from construction and CSV ingestion through selection, cleaning, joining, summarizing, reshaping, and time handling. They assume Python lists, dictionaries, slicing, and functions; they do not cover distributed execution or every storage backend. Examples use import pandas as pd, with local sample data on the operation pages. Check pd.__version__ when reproducing an older notebook.

The examples target pandas 2.2+ APIs and are also applicable to pandas 3.x with the stated dtype and parameter choices. Copy-on-Write is the only mode in pandas 3.0: writing to an independently selected object does not update its parent. In pandas 2.x this behavior is optional, so the notes use explicit assignment to the intended object rather than depending on view mutation. alias = df still binds the same object in either version; it is not a selection or copy.

Source of truth​

Explore connectionsOpen network