Skip to main content

Categorical Data and Binning

A categorical dtype represents a finite vocabulary. It can reduce memory, make allowed values explicit, and preserve a meaningful order.

import pandas as pd

from pandas.api.types import CategoricalDtype

tasks = pd.DataFrame({"priority": ["high", "low", "urgent", None]})

priority_type = CategoricalDtype(
categories=["low", "medium", "high"],
ordered=True,
)
unknown = tasks["priority"].notna() & ~tasks["priority"].isin(priority_type.categories)
assert tasks.loc[unknown, "priority"].tolist() == ["urgent"]
tasks["priority"] = tasks["priority"].astype(priority_type)

Ordered categoricals support ordering and range comparisons according to the declared category sequence. Unordered categoricals represent membership without claiming that one category is greater than another.

Validate the vocabulary​

Values outside the declared categories become missing during conversion. Check for unknown values before or immediately after casting, and distinguish an unknown category from a genuinely missing observation.

Bin continuous values​

ages = pd.DataFrame({"age": [-1, 0, 17, 18, 35, 65]})
ages["band"] = pd.cut(
ages["age"],
bins=[0, 18, 35, 65, float("inf")],
labels=["child", "young_adult", "adult", "senior"],
right=False,
)

ages["quartile"] = pd.qcut(ages["age"], q=4, duplicates="drop")
  • cut uses value boundaries chosen from domain meaning.
  • qcut uses sample quantiles to target similarly populated bins.

Record interval closure and edge policy. Binning loses information and can make small input changes look discontinuous; keep the original numeric value.

Encode for models​

get_dummies creates indicator columns, but encoding belongs inside the model's training and inference pipeline when category vocabularies must remain identical. Do not infer production feature columns independently from each batch.

Edges and category codes​

The category example deliberately detects urgent before conversion; the converted Series then contains high, low, missing, missing. Casting declares the vocabulary but does not reject invalid rows by itself. Category codes are storage positions, with -1 for missing; they are not measurements and averaging them has no general meaning. Memory savings depend on repetition, not merely on choosing category.

With cut and right=False, intervals are left-closed and right-open: age 0 is child, 18 is young_adult, 35 is adult, and 65 is senior; -1 lies outside the bins and becomes missing. Passing a Series returns an index-aligned categorical Series; passing a list normally returns a Categorical. qcut can produce fewer than four bins when repeated quantile edges are dropped. Tied values cannot always be divided into equally sized bins. For train/test use, obtain training edges with retbins=True and reuse them with pd.cut(test_values, bins=training_edges, right=True, include_lowest=True) to preserve qcut’s boundary behavior; decide how out-of-range future values should be treated.

Source​

Explore connectionsOpen network