Text Encoding & Byte Tools

UTF-8 Encoder

Encode Unicode text to exact UTF-8 bytes, inspect each code point and byte offset, compare transport representations, normalize deliberately, and strictly verify captured byte streams before you reuse them.

UTF-8 Byte & Round-Trip Studio

Turn Unicode text into exact UTF-8 bytes, inspect byte boundaries and transport representations, then strictly decode captured bytes to prove the round trip.

Browser-local · RFC 3629 · no upload

1. Enter the Unicode text you actually have

Pre-encoding policy

Local file import is strict UTF-8 and capped at 5 MB for browser responsiveness.
Ready.

Exact size metrics

—output bytes
—code points
—graphemes
—UTF-16 units
—ASCII-byte share
—bytes / code point

2. Compare the representations of the same bytes

Hex bytes
Decimal bytes
Binary bytes
\xHH byte escapes
%HH for every UTF-8 byte
Base64 of output bytes
URL component encoding
Unicode code-point escapes

3. Prove which character produced which bytes

Offsets are byte offsets in the exact output, including the optional three-byte UTF-8 BOM. The inspector follows Unicode code points; one visible grapheme can contain several rows.

#CharacterCode pointUTF-8 hexDecimalBinaryBytesByte offsetUTF-16 units

4. Strictly verify captured UTF-8 bytes

Fatal verification rejects malformed, truncated, overlong, surrogate and above-U+10FFFF sequences instead of silently inserting replacement characters.

5. Repeat-use and handoff

BOM is intentionally not added to each batch row.
LineSourceBytesHexBase64

Explicit local session

Nothing is auto-saved. Saving source requires this explicit button.

6. Export reproducible evidence

Truth boundary: UTF-8 is an encoding, not encryption. Unicode normalization can intentionally change the code-point sequence before encoding. A UTF-8 BOM is optional transport metadata, not part of the Unicode text. %HH byte notation is not the same thing as URL-component escaping because URL syntax leaves some ASCII characters unescaped. Base64 is another representation of the same bytes. This tool cannot infer the original character set of arbitrary unknown bytes.

How to verify UTF-8 instead of just generating it

Code point → UTF-8 bytes

UTF-8 encodes Unicode scalar values with one to four bytes. ASCII stays one byte; characters such as é, 한 and 😀 require two, three and four bytes respectively.

Byte count is not character count

Visible graphemes, Unicode code points, JavaScript UTF-16 units and UTF-8 bytes answer different questions. This Studio reports them separately so API or storage limits are not guessed from visible length.

Normalization changes bytes on purpose

NFC/NFD/NFKC/NFKD can change the code-point sequence before UTF-8 encoding. Two visually similar strings can therefore have different byte sequences. Preserve exact source unless normalization is a deliberate requirement.

BOM and newline policy are transport choices

UTF-8 does not require a BOM. CRLF versus LF also changes the byte stream. These controls are explicit because file, protocol and legacy-software requirements differ.

Strict decode catches real corruption

The reverse verifier rejects truncated sequences, isolated continuation bytes, overlong encodings, UTF-16 surrogate encodings and values above U+10FFFF instead of silently replacing them.

Percent bytes are not automatically a URL

%HH-for-every-byte is a diagnostic representation. URL component encoding follows URI syntax and leaves selected ASCII characters unescaped, so the two outputs are intentionally shown separately.

When this tool is useful

Use it for API payload debugging, byte limits, database/storage checks, source-code escapes, percent-encoded logs, Base64 handoff, hidden Unicode inspection and reproducing encoding bugs. The byte inspector makes the relationship between each code point and its UTF-8 sequence explicit.

Privacy and repeat use

The encoder, verifier, file check, batch conversion and exports run in the browser. Source text is not placed into share URLs and is only written to local storage when you explicitly press Save source locally.

Search by task, tool name, or category. Press Esc to close.
Start typing to find a tool.