What the identifier shows
For every character in your text (not just the first one — this is the biggest limitation of most competing tools), the identifier shows:
- Char — the visible glyph. Invisible characters (zero-width space, joiner, BOM, direction marks) are rendered as a striped label with the character's short name so you can see they exist.
- Code point — the standard U+XXXX Unicode notation.
- Name / category — the official Unicode name (LATIN SMALL LETTER A, GRINNING FACE, etc.) and general category (Lu, Ll, Nd, Mn, Cf, So, So-Emoji…).
- Dec / Hex — the code point in decimal and hexadecimal.
- UTF-8 — how many bytes UTF-8 needs for this character (1, 2, 3, or 4).
- HTML — the numeric HTML entity (
&#XXXX;) you can paste anywhere HTML is parsed.
Invisible & special characters
Copy-pasting from the web often brings hidden characters along for the ride — zero-width spaces (U+200B) inside email addresses, direction marks (U+200E, U+200F) inside filenames, byte-order marks (U+FEFF) at the start of a file. They look like nothing but they break URL matching, string comparisons, and search indices.
| Character | Code point | What it does |
|---|---|---|
| Zero-Width Space | U+200B | A break point for line wrapping without a visible space. |
| Zero-Width Non-Joiner | U+200C | Prevents ligature or contextual joining in Arabic / Indic scripts. |
| Zero-Width Joiner (ZWJ) | U+200D | Combines characters into a compound form — used heavily in emoji sequences. |
| LTR Mark / RTL Mark | U+200E / U+200F | Force text direction inside mixed left-to-right and right-to-left content. |
| Byte-Order Mark (BOM) | U+FEFF | Marks the start of a Unicode file as UTF-8/16/32. Invisible when in the middle of text. |
| Soft Hyphen | U+00AD | An invisible hyphenation hint; renders only when a line breaks there. |
| Non-Breaking Space | U+00A0 | Looks like a space but prevents line breaks and joining. |
Homoglyphs and phishing detection
The Unicode Confusables data set lists thousands of characters that look identical or nearly identical across scripts. Cyrillic а (U+0430) renders exactly like Latin a (U+0061). Attackers exploit this in phishing URLs (аpple.com with a Cyrillic 'а') and in scam emails. This identifier flags all common homoglyphs in red so you can spot suspicious strings at a glance.
Emoji sequences and ZWJ chains
A modern emoji like 👨👩👧 (Family: Man, Woman, Girl) is actually five separate Unicode characters joined by two Zero-Width Joiners. The identifier expands the sequence so you can inspect every code point involved — helpful when you're debugging text that renders differently across platforms.
When you need this
- Debugging string comparisons that silently fail (invisible characters in the input).
- Auditing phishing URLs (homoglyph detection).
- Understanding why an emoji renders differently on iOS vs Android (ZWJ sequences).
- Cleaning pasted text before saving to a database.
- Learning what a specific character is called (Unicode name lookup).
- Writing an HTML entity for a character you can't type.
Frequently asked questions
Does this identify multiple characters?
Yes — every character in your input gets its own row. Most competing tools limit you to one character at a time; this one has no cap.
Can it find invisible characters?
Yes. Zero-width spaces, joiners, BOMs, direction marks, and soft hyphens are shown as striped labels so you can see they're there. Use the Invisible / control filter to see only those rows.
Does it detect homoglyphs?
Yes. Cyrillic, Greek, and math-alphanumeric look-alikes are flagged in red.
What is a Unicode code point?
The numeric ID assigned to every character. Written as U+ followed by hex digits — U+0041 = A, U+2603 = SNOWMAN, U+1F600 = 😀.
How is UTF-8 byte size calculated?
1 byte for ASCII, 2 for characters up to U+07FF (Latin-1 supplement, Greek, Cyrillic), 3 for most of the Basic Multilingual Plane, 4 for characters above U+FFFF (most emoji). The tool uses the browser's native TextEncoder.
What is a ZWJ sequence?
A chain of emoji joined with Zero-Width Joiner (U+200D). It renders as a single compound emoji but is technically several code points — the identifier expands them all.
Is my text saved?
No. The identifier runs entirely in your browser. Nothing is uploaded.
Related tools
References
- Unicode Consortium — Full character database (name, category, code point).
- Unicode Technical Report #36 — Security considerations.
- Unicode Technical Standard #39 — Confusables data set.
- Unicode Emoji Charts — ZWJ sequences and emoji subcategories.