Code units are storage, not characters
Surrogate pairs are grouped conceptually while every 16-bit unit remains inspectable. Unpaired surrogates are flagged instead of being called valid Unicode scalar values.
Unicode precision boundary
WebToolArc keeps grapheme clusters, Unicode code points, UTF-16 code units and UTF-8 bytes as separate measurements. Normalization follows the browser Unicode implementation; grapheme segmentation uses Intl.Segmenter when available. Property names shown by this tool are limited to the explicit local tables and JavaScript Unicode property escapes used by the page rather than claiming a complete Unicode Character Database name lookup.