Skip to main content

Regular Expressions

Build a pattern from a few readable pieces:

FormMeaning
.any character except a newline by default
^ / $start / end anchors
[abc] / [^abc]allowed / excluded character set
*, +, ?, {m,n}repetition
(…) / (?:…)capturing / non-capturing group
\bword boundary in Python's re syntax

In the example below, [A-Z] matches one uppercase ASCII letter, \d{3} matches three digits, and \b marks a word boundary, so codes becomes ["A-104", "B-208"]. The (?i) flag makes the replacement case-insensitive: updated becomes "item A-104, item B-208". is_code is True because the entire string "A-104" matches. The r prefix keeps Python’s string escaping from obscuring the regex backslashes:

import re

text = "Order A-104, order B-208"
codes = re.findall(r"\b[A-Z]-\d{3}\b", text)
updated = re.sub(r"(?i)\border\b", "item", text)
is_code = re.fullmatch(r"[A-Z]-\d{3}", "A-104") is not None

Choose the operation deliberately: search() finds the first match anywhere, fullmatch() requires the entire input, findall() collects matches, and sub() replaces them. Test representative matches, non-matches, empty input, Unicode, and unusually long input before using a pattern on untrusted data.

Regex syntax differs between Python, JavaScript, editors, and command-line tools. Treat the Python re documentation as authoritative only for Python; use RegExr as an interactive scratchpad, not as the specification for every engine.

In regex101, choose the Python flavor and try both matching and nonmatching strings from the example. Inspect the pattern explanation, add parentheses to capture the code’s letter and digits, then try a boundary case such as an extra digit. Selecting the engine matters: a successful match under another regex flavor may not reproduce Python’s behavior.

Explore connectionsOpen network