Loading and Inspecting DataFrames
Before reading a table, decide which columns are required, what their values mean, and which entries count as missing. These expectations form the table’s schema. In the order example below, identifiers are read as strings so "001" keeps its leading zeros, while created_at is parsed as a date. After loading, check that the resulting DataFrame meets those expectations.
from io import StringIO
import pandas as pd
orders = pd.read_csv(
StringIO("order_id,customer_id,created_at,amount\n"
"001,C1,2026-01-02,120.50\n"
"002,C2,2026-01-03,0\n"),
usecols=["order_id", "customer_id", "created_at", "amount"],
dtype={"order_id": "string", "customer_id": "string"},
parse_dates=["created_at"],
na_values=["", "NA", "null"],
)
Ingestion checklist
- Identify delimiter, encoding, headers, decimal convention, and missing sentinels.
- Select required columns and specify semantic dtypes where inference is risky.
- Parse dates deliberately; localize or convert time zones explicitly afterward.
- Inspect
shape,head,info, duplicate keys, and missingness. - Assert the invariants the next stage relies on.
assert orders["order_id"].notna().all()
assert orders["order_id"].is_unique
assert orders["amount"].notna().all()
assert orders["amount"].ge(0).all()
Index choice
Keep the default range index unless a domain key genuinely benefits selection or alignment. A key can remain an ordinary column:
indexed = orders.set_index("order_id")
if not indexed.index.is_unique:
raise ValueError("order_id must be unique")
orders = indexed.reset_index()
Check indexed.index.is_unique when uniqueness is required; the verify_integrity argument to set_index is deprecated in pandas 3.0.
Do not use a non-unique business field as an index merely because it looks like an identifier.
Rename at the boundary
Normalize names once, close to ingestion:
orders = orders.rename(columns=str.strip)
orders.columns = orders.columns.str.lower().str.replace(" ", "_", regex=False)
Preserve a mapping when external names must remain traceable.
Scale boundary
Use chunksize when independent chunks can be aggregated incrementally. If the
workflow requires repeated full-table joins or shuffles beyond memory, changing
the execution engine is usually better than elaborate chunk bookkeeping.
Reading contracts
read_csv accepts a path or a file-like object such as the in-memory StringIO above. This example returns a (2, 4) DataFrame; orders["order_id"].tolist() is ["001", "002"], preserving leading zeros. A path is resolved from the process working directory. With chunksize, the return value is an iterator over DataFrames instead.
na_values adds to the default missing-token list. Use keep_default_na=False with an explicit list if a token such as "NA" is a valid identifier. parse_dates can leave an unparseable column as text; inspect its dtype or use pd.to_datetime(..., errors="raise") when conversion must succeed. Name normalization can create duplicate columns, so check orders.columns.is_unique afterward. Assertions suit examples, but Python can disable them with -O; use explicit exceptions at enforced boundaries.