Character diff is grapheme-aware
User-perceived grapheme clusters are compared where Intl.Segmenter is available, so emoji ZWJ sequences and combining text are not intentionally split into misleading per-code-point edits.
Inspect fine-grained text edits locally using browser grapheme segmentation when available.
Comparison runs locally in your browser. Interactive diff uses a bounded token count to avoid locking the page on extremely large inputs.
Locate the first mismatch, common prefix and suffix, grapheme lengths and edit concentration between two text strings.
User-perceived grapheme clusters are compared where Intl.Segmenter is available, so emoji ZWJ sequences and combining text are not intentionally split into misleading per-code-point edits.
Use the tool first, then apply these checks to verify inputs, interpret the result, and hand it off without displacing the primary workflow.
Unicode text can represent what looks like one character using multiple code points, and visually identical text can use different normalization forms. A grapheme-aware character comparison is often closer to what a reader sees, while source-code, protocol, or forensic work may require exact code-point or byte comparison. Choose the comparison level that matches the task and keep the untouched inputs when invisible differences could be significant.
Tabs, non-breaking spaces, zero-width characters, carriage returns, and trailing line breaks can cause mismatches that are hard to see in ordinary rendered text. Do not “clean” the inputs before investigating unless normalization is the intended operation. When a mismatch matters for code, CSV, identifiers, or digital signatures, inspect escaped or code-point representations so an invisible character cannot be mistaken for an application bug.
Unicode NFC or NFD normalization can make canonically equivalent strings compare the same, but it changes the underlying code-point sequence. Case folding, trimming, punctuation removal, and whitespace collapsing are even stronger transformations. Apply them only when the business rule says those distinctions are irrelevant. For passwords, cryptographic input, exact identifiers, and signed data, silent normalization can change meaning and should not be introduced casually.
A diff reports where two strings diverge, not why. Once the first mismatch or inserted/deleted region is identified, trace that position back to the file, database field, clipboard source, API response, or rendering pipeline that produced it. Recheck a small sample around the mismatch and confirm encoding. Fixing the upstream transformation is usually safer than repeatedly patching the displayed output after the fact.