Skip to main content

Regular Expressions in Python

Regular expressions describe patterns in text. Use them for bounded lexical tasks such as extracting identifiers or validating a deliberately small format; use a parser for nested structure or a complete language grammar.

Write patterns as raw strings so Python string escaping does not obscure regex escaping:

import re

event = re.compile(
r"^(?P<level>INFO|WARN|ERROR)\s+user=(?P<user>[\w.-]+)$"
)

line = "INFO user=ada"
match = event.fullmatch(line)
if match:
level = match["level"]
user = match["user"]

Choose the matching contract​

  • search finds the first match anywhere.
  • match checks only at the beginning.
  • fullmatch requires the entire input to match and is usually clearest for validation.
  • finditer streams non-overlapping match objects; findall constructs a list whose shape changes with capturing groups.
  • sub replaces matches; a callable replacement can make context-sensitive changes without a second parsing pass.

Use named groups for fields with domain meaning and non-capturing groups (?:...) when grouping is only structural.

Meaning and escaping​

For str patterns, classes such as \w and \d use Unicode semantics by default. Add re.ASCII only when the format is explicitly ASCII. Case-insensitive matching is not locale-aware human-language comparison.

Escape untrusted text that must be treated literally:

user_supplied_prefix = "a.b"
literal_pattern = re.compile(re.escape(user_supplied_prefix))

Do not interpolate arbitrary text directly into a pattern or replacement template. Pattern escaping and replacement escaping have different rules.

Performance and validation boundaries​

Ambiguous nested repetitions and overlapping alternatives can cause extreme backtracking on adversarial input. Keep patterns simple, bound input size, test near misses, and isolate or replace regex work when a hard time limit matters. The standard re matching API does not provide a general per-match timeout.

A regex can check syntax but not ownership, deliverability, authorization, or business validity. For URLs, email addresses, dates, and structured formats, prefer a standard parser plus domain validation.

Use re.VERBOSE and comments when a pattern is complex enough to require maintenance. At some point a named parser is the more honest abstraction.

Read and test the pattern​

In the example, | selects a level, \s+ requires one or more whitespace characters, and [\w.-]+ accepts one or more word characters, dots, or hyphens. Named groups capture the selected fields. The anchors are redundant with fullmatch.

assert event.fullmatch("INFO user=ada").groupdict() == {"level": "INFO", "user": "ada"}
assert event.fullmatch("DEBUG user=ada") is None
assert event.fullmatch("INFO user=") is None
assert event.fullmatch("INFO\nuser=ada") is not None

That last match is intentional under \s: it includes newlines. For a single-line format that permits only spaces and tabs between fields, use [ \t]+ instead. Strip the record terminator deliberately before matching a line read from a file.

For literal replacement text, use a callable: re.sub(pattern, lambda match: replacement, text). Its returned string is inserted without interpreting backreferences; re.escape is for patterns, not replacement templates.

Open regex101 in Python mode and try the example pattern on a matching line and a near miss. Inspect the named captures and token explanations, then test a newline between fields to see the effect of \s+. Confirm the final fullmatch behavior with the Python example here.

Source​

Explore connectionsOpen network