Strings and Unicode Text
Python str represents immutable Unicode text. bytes represents binary data,
which may be encoded text. Crossing between text and bytes requires an encoding.
text = "café"
payload = text.encode("utf-8")
restored = payload.decode("utf-8")
Keep text as str inside the program and encode/decode at file, network, or
subprocess boundaries.
Sequence operations
name = "Ada Lovelace"
first = name[:3]
last_character = name[-1]
contains_space = " " in name
Indexes and slices operate on Unicode code points, not user-perceived grapheme clusters. Some visible characters consist of multiple code points.
Because strings are immutable, transformations return string results without changing the original; a distinct object identity is not guaranteed:
normalized = raw.strip().casefold()
words = normalized.split()
line = ", ".join(words)
Repeated += in a large loop can create avoidable intermediate strings; collect
pieces and use join when construction cost matters.
Formatting
Use f-strings for local readable interpolation and format specifications:
message = f"{user}: {total:,.2f} CAD"
Formatting is not escaping. SQL, HTML, shell commands, regular expressions, and URLs each require context-specific APIs; never rely on an f-string for safety.
Unicode equality
Visually similar text can use different code-point sequences. Normalize when a
domain requires canonical comparison, and use casefold rather than lower for
aggressive caseless matching. Locale-aware collation and human name handling
need domain-specific libraries and rules.
Slices, cleanup, and canonical comparison
For the string methods,
indexes start at zero and negative indexes count from the end. A slice
text[start:stop:step] excludes stop; out-of-range slice bounds are clipped,
but an out-of-range single index raises IndexError. A zero slice step raises
ValueError.
assert "abc"[1:99] == "bc"
assert "abc"[::-1] == "cba"
assert "abc"[5:] == ""
assert " a b ".split() == ["a", "b"]
assert "a,,b".split(",") == ["a", "", "b"]
assert "abbaXba".strip("ab") == "X"
assert "report.txt".removesuffix(".txt") == "report"
strip(chars) removes any of those characters repeatedly from both ends, not
an exact prefix or suffix. find returns -1 when absent; index raises
ValueError. Use in when only presence matters. join requires strings;
it does not implicitly convert numbers.
The Unicode HOWTO distinguishes code-point equality from canonical equivalence:
import unicodedata
composed = "é"
decomposed = "e\u0301"
assert composed != decomposed
assert (len(composed), len(decomposed)) == (1, 2)
assert unicodedata.normalize("NFC", decomposed) == composed
assert "Straße".casefold() == "strasse"
assert len("café".encode("utf-8")) == 5
NFC does not merge every visually similar character or perform caseless matching.
Compatibility normalization such as NFKC can erase distinctions that a domain
needs. Decoding invalid UTF-8 raises UnicodeDecodeError by default; ignoring
errors silently loses data. Binary data need not represent text at all.