Pandas Idioms
Readable pandas code exposes the sequence of table transformations and uses the most specific operation available.
Transformation style
import pandas as pd
orders = pd.DataFrame({
"customer_id": ["A", "A", "B"], "status": ["paid", "paid", "pending"],
"gross": [100, 50, 80], "fee": [10, 5, 8],
})
result = (
orders.loc[lambda x: x["status"].eq("paid")]
.assign(net=lambda x: x["gross"] - x["fee"])
.groupby("customer_id", as_index=False)
.agg(total_net=("net", "sum"), orders=("net", "size"))
.sort_values("total_net", ascending=False)
)
Break a chain into named stages when it becomes difficult to inspect, reuse, or debug. Method chaining is a readability tool, not a performance guarantee.
Choose the operation
def require_positive(frame, column):
if frame[column].isna().any() or not frame[column].ge(0).all():
raise ValueError(f"{column} must be present and non-negative")
return frame
clean = orders.pipe(require_positive, "gross")
Vectorize before apply
# Prefer
df = pd.DataFrame({"numerator": [10.0, 0.0], "denominator": [2.0, 0.0]})
df["ratio"] = df["numerator"].div(df["denominator"].replace(0, float("nan")))
# Reserve for genuinely row-dependent Python logic
def classify_row(row):
if pd.isna(row["ratio"]):
return "undefined"
return "high" if row["ratio"] >= 4 else "low"
df["label"] = df.apply(classify_row, axis="columns")
Row-wise apply constructs a Series per row and invokes Python repeatedly.
Benchmark only after choosing a correct, clear expression and testing on
representative data.
Mutation rule
Prefer expressions returning new objects and one-step .loc assignments. Avoid
chained assignment and avoid relying on inplace=True as an optimization.
Follow the returned object
The chain returns one row: customer A, total_net=135, orders=2; the original orders still has three rows and no net column. pipe returns whatever its function returns. Here require_positive only validates and returns its input frame, so clean.equals(orders) is true; no values are cleaned. Do not rely on object identity across pandas versions: pandas 3.0 passes a shallow copy into the function.
The ratio example explicitly maps a zero denominator to missing, giving [5.0, NaN] and labels high, undefined. Without a policy, nonzero divided by zero may produce infinity. This small classifier is only an illustration of the row contract; such simple logic can also be vectorized. DataFrame.map requires pandas 2.1 or newer. DataFrame.apply uses columns by default (axis=0) and rows with axis=1; result shape depends on what the callback returns.