Home / Typography & Unicode Tools / Grapheme Cluster Counter
Typography & Unicode Tools

Grapheme Cluster Counter

Segment emoji sequences, combining-mark sequences, and other extended grapheme clusters with the browser Internationalization API.

Unicode 17 referenceGrapheme ≠ code pointUTF-16 / UTF-8 offsetsLocal inspection
—grapheme clusters
—code points
—UTF-16 code units
—UTF-8 bytes
—Extended grapheme segmentation is intended to approximate user-perceived characters; language-specific tailoring can differ.

Grapheme, code-point, UTF-16 & byte evidence

Count user-perceived characters with Intl.Segmenter and compare them with code points, UTF-16 units and UTF-8 bytes. Export every grapheme cluster.

—Graphemes
—Code points
—UTF-16 units
—UTF-8 bytes
#GraphemeCode pointsBytes

Grapheme clusters are closer to what users see

Character counts can mean bytes, UTF-16 code units, code points or user-perceived graphemes. This counter focuses on extended grapheme clusters, which keeps many emoji and combining sequences together. Exact segmentation depends on the Unicode data implemented by the browser.

Four different ways to count text

Grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes are reported together. Cluster rows show their starting UTF-16 and UTF-8 offsets.

Unicode precision boundary

WebToolArc keeps grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes as separate measurements. Normalization follows the browser Unicode implementation; grapheme segmentation uses Intl.Segmenter when available. Property names shown by this tool are limited to the explicit local tables and JavaScript Unicode property escapes used by the page rather than claiming a complete Unicode Character Database name lookup.

Search by task, tool name, or category. Press Esc to close.
Start typing to find a tool.