Unicode escapes, code points, UTF-16 and UTF-8 are distinct representations
The review keeps scalar escapes separate from actual UTF-8 bytes and warns about unpaired UTF-16 surrogates in source strings.
Encode or decode Unicode scalar escapes locally, including supplementary-plane code points with \u{...} notation.
These tools operate on bytes and Unicode text locally. “Text to hex/binary/decimal” means the UTF-8 encoding of the text, not the numeric Unicode code point unless a tool explicitly says code point.
Inspect Unicode code points, UTF-16 units, UTF-8 bytes and JavaScript escape forms so supplementary characters and surrogate pairs are explicit.
The review keeps scalar escapes separate from actual UTF-8 bytes and warns about unpaired UTF-16 surrogates in source strings.
Use the tool first, then apply these checks to verify the inputs, interpret the result, and hand it off without displacing the primary workflow.
JavaScript, JSON, Python, Java, CSS and shell tools can represent Unicode differently. A sequence that looks familiar may have different escaping rules or may represent UTF-16 code units rather than a complete code point. Record the source language before converting.
Unicode code points, UTF-8 bytes and UTF-16 code units are different layers. A converter can show equivalent representations, but byte counts only make sense after an encoding is specified. Preserve both the text and encoding when comparing files or protocols.
Characters outside the Basic Multilingual Plane may require two UTF-16 code units, while visually accented characters can be composed from multiple code points. Compare code-point sequences rather than assuming one visible glyph always equals one scalar value.
For configuration, source code or data interchange, encode the original text, decode the result, and compare the recovered string. A successful round trip catches many copy/paste and escape-layer mistakes that a visually plausible output can hide.