Grapheme clusters are closer to what users see
Character counts can mean bytes, UTF-16 code units, code points or user-perceived graphemes. This counter focuses on extended grapheme clusters, which keeps many emoji and combining sequences together. Exact segmentation depends on the Unicode data implemented by the browser.
Four different ways to count text
Grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes are reported together. Cluster rows show their starting UTF-16 and UTF-8 offsets.
Unicode precision boundary
WebToolArc keeps grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes as separate measurements. Normalization follows the browser Unicode implementation; grapheme segmentation uses Intl.Segmenter when available. Property names shown by this tool are limited to the explicit local tables and JavaScript Unicode property escapes used by the page rather than claiming a complete Unicode Character Database name lookup.