Free · Full-string · Invisible-char detection

Character Identifier — Unicode Inspector for Full Strings, Emoji & Invisible Characters

Paste any text and get a full character-by-character breakdown. For each character: code point (U+XXXX), Unicode name, general category, UTF-8 and UTF-16 byte size, decimal / hex / binary / HTML entity, homoglyph flag, and invisible-character warning. Search by name or code point. Handles emoji ZWJ sequences.
🔍 Full-string analysis👻 Invisible-char detection⚠ Homoglyph flags📥 CSV export

Character Identifier

Paste text, or search Unicode by name / code point.
0 characters · 0 code points
Insert test
Filter:
#CharCode pointName / categoryDecHexUTF-8HTML
Paste text above or insert a test character.

What the identifier shows

For every character in your text (not just the first one — this is the biggest limitation of most competing tools), the identifier shows:

  • Char — the visible glyph. Invisible characters (zero-width space, joiner, BOM, direction marks) are rendered as a striped label with the character's short name so you can see they exist.
  • Code point — the standard U+XXXX Unicode notation.
  • Name / category — the official Unicode name (LATIN SMALL LETTER A, GRINNING FACE, etc.) and general category (Lu, Ll, Nd, Mn, Cf, So, So-Emoji…).
  • Dec / Hex — the code point in decimal and hexadecimal.
  • UTF-8 — how many bytes UTF-8 needs for this character (1, 2, 3, or 4).
  • HTML — the numeric HTML entity (&#XXXX;) you can paste anywhere HTML is parsed.

Invisible & special characters

Copy-pasting from the web often brings hidden characters along for the ride — zero-width spaces (U+200B) inside email addresses, direction marks (U+200E, U+200F) inside filenames, byte-order marks (U+FEFF) at the start of a file. They look like nothing but they break URL matching, string comparisons, and search indices.

CharacterCode pointWhat it does
Zero-Width SpaceU+200BA break point for line wrapping without a visible space.
Zero-Width Non-JoinerU+200CPrevents ligature or contextual joining in Arabic / Indic scripts.
Zero-Width Joiner (ZWJ)U+200DCombines characters into a compound form — used heavily in emoji sequences.
LTR Mark / RTL MarkU+200E / U+200FForce text direction inside mixed left-to-right and right-to-left content.
Byte-Order Mark (BOM)U+FEFFMarks the start of a Unicode file as UTF-8/16/32. Invisible when in the middle of text.
Soft HyphenU+00ADAn invisible hyphenation hint; renders only when a line breaks there.
Non-Breaking SpaceU+00A0Looks like a space but prevents line breaks and joining.

Homoglyphs and phishing detection

The Unicode Confusables data set lists thousands of characters that look identical or nearly identical across scripts. Cyrillic а (U+0430) renders exactly like Latin a (U+0061). Attackers exploit this in phishing URLs (аpple.com with a Cyrillic 'а') and in scam emails. This identifier flags all common homoglyphs in red so you can spot suspicious strings at a glance.

Emoji sequences and ZWJ chains

A modern emoji like 👨‍👩‍👧 (Family: Man, Woman, Girl) is actually five separate Unicode characters joined by two Zero-Width Joiners. The identifier expands the sequence so you can inspect every code point involved — helpful when you're debugging text that renders differently across platforms.

When you need this

  • Debugging string comparisons that silently fail (invisible characters in the input).
  • Auditing phishing URLs (homoglyph detection).
  • Understanding why an emoji renders differently on iOS vs Android (ZWJ sequences).
  • Cleaning pasted text before saving to a database.
  • Learning what a specific character is called (Unicode name lookup).
  • Writing an HTML entity for a character you can't type.

Frequently asked questions

Does this identify multiple characters?

Yes — every character in your input gets its own row. Most competing tools limit you to one character at a time; this one has no cap.

Can it find invisible characters?

Yes. Zero-width spaces, joiners, BOMs, direction marks, and soft hyphens are shown as striped labels so you can see they're there. Use the Invisible / control filter to see only those rows.

Does it detect homoglyphs?

Yes. Cyrillic, Greek, and math-alphanumeric look-alikes are flagged in red.

What is a Unicode code point?

The numeric ID assigned to every character. Written as U+ followed by hex digits — U+0041 = A, U+2603 = SNOWMAN, U+1F600 = 😀.

How is UTF-8 byte size calculated?

1 byte for ASCII, 2 for characters up to U+07FF (Latin-1 supplement, Greek, Cyrillic), 3 for most of the Basic Multilingual Plane, 4 for characters above U+FFFF (most emoji). The tool uses the browser's native TextEncoder.

What is a ZWJ sequence?

A chain of emoji joined with Zero-Width Joiner (U+200D). It renders as a single compound emoji but is technically several code points — the identifier expands them all.

Is my text saved?

No. The identifier runs entirely in your browser. Nothing is uploaded.

References

  • Unicode Consortium — Full character database (name, category, code point).
  • Unicode Technical Report #36 — Security considerations.
  • Unicode Technical Standard #39 — Confusables data set.
  • Unicode Emoji Charts — ZWJ sequences and emoji subcategories.
Scroll to Top