Regular Expressions in Python
Regular expressions describe patterns in text. Use them for bounded lexical tasks such as extracting identifiers or validating a deliberately small format; use a parser for nested structure or a complete language grammar.
Write patterns as raw strings so Python string escaping does not obscure regex escaping:
import re
event = re.compile(
r"^(?P<level>INFO|WARN|ERROR)\s+user=(?P<user>[\w.-]+)$"
)
line = "INFO user=ada"
match = event.fullmatch(line)
if match:
level = match["level"]
user = match["user"]
Choose the matching contract
searchfinds the first match anywhere.matchchecks only at the beginning.fullmatchrequires the entire input to match and is usually clearest for validation.finditerstreams non-overlapping match objects;findallconstructs a list whose shape changes with capturing groups.subreplaces matches; a callable replacement can make context-sensitive changes without a second parsing pass.
Use named groups for fields with domain meaning and non-capturing groups
(?:...) when grouping is only structural.
Meaning and escaping
For str patterns, classes such as \w and \d use Unicode semantics by
default. Add re.ASCII only when the format is explicitly ASCII. Case-insensitive
matching is not locale-aware human-language comparison.
Escape untrusted text that must be treated literally:
user_supplied_prefix = "a.b"
literal_pattern = re.compile(re.escape(user_supplied_prefix))
Do not interpolate arbitrary text directly into a pattern or replacement template. Pattern escaping and replacement escaping have different rules.
Performance and validation boundaries
Ambiguous nested repetitions and overlapping alternatives can cause extreme
backtracking on adversarial input. Keep patterns simple, bound input size, test
near misses, and isolate or replace regex work when a hard time limit matters.
The standard re matching API does not provide a general per-match timeout.
A regex can check syntax but not ownership, deliverability, authorization, or business validity. For URLs, email addresses, dates, and structured formats, prefer a standard parser plus domain validation.
Use re.VERBOSE and comments when a pattern is complex enough to require
maintenance. At some point a named parser is the more honest abstraction.
Read and test the pattern
In the example, | selects a level, \s+ requires one or more whitespace characters, and [\w.-]+ accepts one or more word characters, dots, or hyphens. Named groups capture the selected fields. The anchors are redundant with fullmatch.
assert event.fullmatch("INFO user=ada").groupdict() == {"level": "INFO", "user": "ada"}
assert event.fullmatch("DEBUG user=ada") is None
assert event.fullmatch("INFO user=") is None
assert event.fullmatch("INFO\nuser=ada") is not None
That last match is intentional under \s: it includes newlines. For a single-line format that permits only spaces and tabs between fields, use [ \t]+ instead. Strip the record terminator deliberately before matching a line read from a file.
For literal replacement text, use a callable: re.sub(pattern, lambda match: replacement, text). Its returned string is inserted without interpreting backreferences; re.escape is for patterns, not replacement templates.
Open regex101 in Python mode and try the example pattern on a matching line and a near miss. Inspect the named captures and token explanations, then test a newline between fields to see the effect of \s+. Confirm the final fullmatch behavior with the Python example here.