FM 1-04 Chapter 1 · Encoding
Unicode inspector
Show each code point with name, UTF-8 and UTF-16, find invisible, bidi and look-alike characters, normalise.
Local only Runs in your browser. Nothing you enter leaves this page.
Usable as a step in Chain- Code points
- 0
- Grapheme clusters
- 0
- UTF-8
- 0 bytes
- UTF-16
- 0 units
Cleaned
Normalisation
G is the grapheme cluster, what a reader sees as one character. Names cover ASCII, Latin-1, Latin Extended-A, basic Greek and Cyrillic, punctuation, spaces and invisible characters, and are computed for Hangul and CJK ideographs. Other names are left blank: the full name list is about 2 MB. The category comes from the browser's own Unicode data.
Look-alikes use a hand-picked subset of Unicode's confusables.txt (UTS #39): Cyrillic, Greek and Armenian letters that pass for Latin ones. They only matter in text meant to be Latin, such as domains and code. ZWJ and variation selectors inside emoji are kept by the cleaner, as they belong there. ZWNJ is removed, which is wrong for Persian and some Indic text. NFKC also folds fullwidth and mathematical letters to ASCII.