Home / Typography & Unicode Tools / Unicode Normalizer
Typography & Unicode Tools

Unicode Normalizer

Compare canonical and compatibility normalization forms, code-point counts, UTF-8 size, and optional combining-mark removal.

Unicode 17 referenceGrapheme ≠ code pointUTF-16 / UTF-8 offsetsLocal inspection
—Output equals input
—Input code points
—Output code points
—Input UTF-8 bytes
—Output UTF-8 bytes
ReadyNFC/NFD preserve canonical equivalence; NFKC/NFKD may replace compatibility characters such as ligatures or circled digits with simpler equivalents.

NFC / NFD / NFKC / NFKD comparison

Compare all four Unicode normalization forms side by side with code-point and UTF-8 byte counts so compatibility folding is visible before you replace source text.

FormChanged?Code pointsUTF-8 bytesOutput

Normalization and optional mark removal are separate operations

The selected UAX #15 normalization form remains distinct from optional combining-mark removal. Input/output metrics and exact code-point changes are reviewable before copy.

Unicode precision boundary

WebToolArc keeps grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes as separate measurements. Normalization follows the browser Unicode implementation; grapheme segmentation uses Intl.Segmenter when available. Property names shown by this tool are limited to the explicit local tables and JavaScript Unicode property escapes used by the page rather than claiming a complete Unicode Character Database name lookup.

Search by task, tool name, or category. Press Esc to close.
Start typing to find a tool.