Definition
Plain language
Rewriting text so that characters that look the same are stored the same way.
As stated in the literature
Mapping strings to a canonical form (NFC/NFD/NFKC/NFKD) so equivalent code point sequences compare equal; engines differ in whether and where they apply it, making it a usable behavioural probe.
Also called: Unicode-normalization, normalization
Why it matters: Without it, two strings that look identical on screen can fail to compare as equal — breaking searches, logins, and filters — and the fact that different systems normalize at different points makes it a handy way to tell them apart.
For example, the letter "é" can be stored either as one character or as an "e" followed by a separate accent mark, and normalization rewrites both into the same form so they match.
Heard on the show
“A tiling pattern that wins for matrix multiply often wins for normalization.”Episode 065 — One Loop to Optimize Them All: A Universal API for LLM-Driven Discovery