Free · 6 encodings · DB fit check

Byte Counter — UTF-8, UTF-16, ASCII, Latin-1 with Database & SMS Fit Check

Counts bytes in six encodings — UTF-8, UTF-16 (BE / LE), UTF-32, ASCII, ISO-8859-1 (Latin-1), and Windows-1252 — and tells you whether the string fits inside common storage / messaging limits: MySQL VARCHAR, TINYTEXT, TEXT, MEDIUMTEXT, LONGTEXT, Redis key, DynamoDB item, SMS (GSM-7 / UCS-2), and a tweet. Byte distribution histogram and per-character breakdown included.
🔢 6 encodings🗄 Database limit check📊 1/2/3/4-byte histogram📥 CSV export

Byte Counter

Pick an encoding, paste text — bytes and fit-checks update instantly.
0 bytes · 0 characters

Fits in… (byte-based limits for common storage & messaging)

UTF-8 byte distribution

#CharCode pointUTF-8 bytesUTF-16 bytesHex (UTF-8)
Paste text above to see the per-character byte table.

What is a byte, exactly?

A byte is 8 bits — the smallest chunk of data most systems handle at once. When you type a character on a keyboard, the software encodes that character as one or more bytes according to a chosen encoding. Different encodings use different byte counts for the same character. The single character é is:

  • 2 bytes in UTF-8 (0xC3 0xA9)
  • 2 bytes in UTF-16 (0x00 0xE9 or 0xE9 0x00)
  • 4 bytes in UTF-32 (0x00 0x00 0x00 0xE9)
  • 1 byte in Latin-1 (0xE9)
  • Cannot be represented in 7-bit ASCII (falls back to ? or an escape)

This is why byte count and character count are different numbers — and why "how many bytes is my text?" depends entirely on the encoding you pick.

The 6 encodings this counter supports

EncodingBytes/charCharacter coverageNotes
UTF-81–4 (variable)All UnicodeDefault for the web. ASCII-compatible: 1 byte for A–Z 0–9 punctuation. Emoji: 4 bytes each.
UTF-16 BE/LE2 or 4All UnicodeDefault for JavaScript strings (in memory), Java char, Windows APIs. Above U+FFFF uses 4-byte surrogate pairs.
UTF-324 (fixed)All UnicodeSimple but wasteful. Rare outside of internal processing.
ASCII1 (7-bit)Only U+0000–U+007FAnything above 127 is lossy. Older protocols, cursed legacy systems.
Latin-1 (ISO-8859-1)1 (fixed)U+0000–U+00FFCovers Western European accents (é, ñ, ç). Anything above 0xFF is lossy.
Windows-12521 (fixed)~256 charactersSuperset of Latin-1 with curly quotes, em dash, euro sign, TM. Default for Windows Notepad.

⚠ When you switch to ASCII, Latin-1, or Windows-1252, the counter flags every character that cannot be represented. Those characters would be lost (or replaced with ?) if you actually wrote the text out in that encoding.

Database and messaging fit-checks

The Fits in panel checks your byte count against the byte-based limits of common storage systems. Character-based limits (like MySQL utf8mb4 VARCHAR(255) counting 255 characters) are honored on the character side; byte-based ones (like Redis, DynamoDB, and file uploads) are honored on the byte side.

TargetLimitApplies to
MySQL VARCHAR(255) — utf8mb4255 chars / max 1020 bytesCommon short-string column
MySQL TINYTEXT255 bytesSmall text blob
MySQL TEXT65 535 bytesArticle body, comments
MySQL MEDIUMTEXT16 777 215 bytesLarge documents
MySQL LONGTEXT4 294 967 295 bytesAnything huge
Redis key (recommended)< 100 bytesKeeps memory / lookup fast
DynamoDB item400 KBWhole record incl. attributes
SMS (GSM-7)160 chars / 140 bytesBasic Latin only
SMS (UCS-2)70 charsAny character; each is 2 bytes
Twitter / X post280 chars (weighted)Most emoji count as 2
URL (safe browser cap)~2 000 charsIE historical limit; modern browsers accept far more

How emoji byte-size works

Emoji surprise people because a single visible glyph often takes many bytes. 😀 (Grinning Face) is one Unicode code point (U+1F600) but takes 4 bytes in UTF-8. Modern compound emoji built with Zero-Width Joiners take more:

  • 😀 (Grinning Face) → 4 UTF-8 bytes.
  • 👨‍💻 (Man Technologist) → 11 UTF-8 bytes (Man + ZWJ + Laptop).
  • 👨‍👩‍👧 (Family: Man Woman Girl) → 18 UTF-8 bytes.
  • 🏳️‍🌈 (Rainbow Flag) → 14 UTF-8 bytes.

That means a 140-byte SMS can hold about 35 grinning faces… or one 👩🏽‍🚒 (Woman Firefighter: Medium Skin Tone) with a lot of room to spare — but not many family emoji.

When you need this

  • Checking whether a user-generated string fits inside a database column.
  • Estimating API request/response payload size.
  • Debugging why a "255-character" field is rejecting a 200-character emoji string (utf8mb4 uses up to 4 bytes/char).
  • SMS gateway pre-flight — knowing if you're going to be charged for 1 SMS or 3.
  • Anywhere character count and byte count disagree.

Frequently asked questions

What encodings does this counter support?

UTF-8, UTF-16 (BE / LE), UTF-32, ASCII, Latin-1 (ISO-8859-1), and Windows-1252. UTF-8/16/32 use the browser's native TextEncoder; the others use exact byte-mapping tables.

How many bytes is one character?

Depends on the character and the encoding. See the encodings table above — a single letter can be 1 byte (ASCII) or 4 bytes (UTF-32).

How many bytes is an emoji?

A basic emoji is 4 bytes in UTF-8. Compound emoji made with Zero-Width Joiners can be 11–18 bytes or more. See the emoji section.

Do spaces count as bytes?

Yes. Regular space (U+0020) is 1 byte in UTF-8. Non-breaking space (U+00A0) is 2 bytes.

Will my text fit in MySQL VARCHAR(255)?

The Fits in panel above answers exactly that for your text, whether you're on latin1, utf8, or utf8mb4.

Why does ASCII show fewer bytes than my character count?

It shouldn't. If it does, some characters in your text can't be encoded as 7-bit ASCII — the counter shows a warning banner and either drops or escapes them depending on your setting.

Is my text saved?

No. All processing happens in your browser. Nothing is uploaded.

References

  • Unicode Standard 15.1 — Encoding forms UTF-8, UTF-16, UTF-32.
  • RFC 3629 — UTF-8 encoding specification (byte-length ranges).
  • ISO/IEC 8859-1 (Latin-1) and Microsoft Windows-1252 codepage tables.
  • MySQL 8.0 Manual — VARCHAR, TINYTEXT, TEXT, MEDIUMTEXT, LONGTEXT byte limits.
  • GSM 03.38 — SMS GSM-7 / UCS-2 character set rules.
Scroll to Top