Skip to main content

Strings and Unicode Text

Python str represents immutable Unicode text. bytes represents binary data, which may be encoded text. Crossing between text and bytes requires an encoding.

text = "café"
payload = text.encode("utf-8")
restored = payload.decode("utf-8")

Keep text as str inside the program and encode/decode at file, network, or subprocess boundaries.

Sequence operations​

name = "Ada Lovelace"
first = name[:3]
last_character = name[-1]
contains_space = " " in name

Indexes and slices operate on Unicode code points, not user-perceived grapheme clusters. Some visible characters consist of multiple code points.

Because strings are immutable, transformations return string results without changing the original; a distinct object identity is not guaranteed:

normalized = raw.strip().casefold()
words = normalized.split()
line = ", ".join(words)

Repeated += in a large loop can create avoidable intermediate strings; collect pieces and use join when construction cost matters.

Formatting​

Use f-strings for local readable interpolation and format specifications:

message = f"{user}: {total:,.2f} CAD"

Formatting is not escaping. SQL, HTML, shell commands, regular expressions, and URLs each require context-specific APIs; never rely on an f-string for safety.

Unicode equality​

Visually similar text can use different code-point sequences. Normalize when a domain requires canonical comparison, and use casefold rather than lower for aggressive caseless matching. Locale-aware collation and human name handling need domain-specific libraries and rules.

Slices, cleanup, and canonical comparison​

For the string methods, indexes start at zero and negative indexes count from the end. A slice text[start:stop:step] excludes stop; out-of-range slice bounds are clipped, but an out-of-range single index raises IndexError. A zero slice step raises ValueError.

assert "abc"[1:99] == "bc"
assert "abc"[::-1] == "cba"
assert "abc"[5:] == ""
assert " a b ".split() == ["a", "b"]
assert "a,,b".split(",") == ["a", "", "b"]
assert "abbaXba".strip("ab") == "X"
assert "report.txt".removesuffix(".txt") == "report"

strip(chars) removes any of those characters repeatedly from both ends, not an exact prefix or suffix. find returns -1 when absent; index raises ValueError. Use in when only presence matters. join requires strings; it does not implicitly convert numbers.

The Unicode HOWTO distinguishes code-point equality from canonical equivalence:

import unicodedata

composed = "é"
decomposed = "e\u0301"
assert composed != decomposed
assert (len(composed), len(decomposed)) == (1, 2)
assert unicodedata.normalize("NFC", decomposed) == composed
assert "Straße".casefold() == "strasse"
assert len("café".encode("utf-8")) == 5

NFC does not merge every visually similar character or perform caseless matching. Compatibility normalization such as NFKC can erase distinctions that a domain needs. Decoding invalid UTF-8 raises UnicodeDecodeError by default; ignoring errors silently loses data. Binary data need not represent text at all.

Source​

Explore connectionsOpen network