What is a byte, exactly?
A byte is 8 bits — the smallest chunk of data most systems handle at once. When you type a character on a keyboard, the software encodes that character as one or more bytes according to a chosen encoding. Different encodings use different byte counts for the same character. The single character é is:
- 2 bytes in UTF-8 (0xC3 0xA9)
- 2 bytes in UTF-16 (0x00 0xE9 or 0xE9 0x00)
- 4 bytes in UTF-32 (0x00 0x00 0x00 0xE9)
- 1 byte in Latin-1 (0xE9)
- Cannot be represented in 7-bit ASCII (falls back to ? or an escape)
This is why byte count and character count are different numbers — and why "how many bytes is my text?" depends entirely on the encoding you pick.
The 6 encodings this counter supports
| Encoding | Bytes/char | Character coverage | Notes |
|---|---|---|---|
| UTF-8 | 1–4 (variable) | All Unicode | Default for the web. ASCII-compatible: 1 byte for A–Z 0–9 punctuation. Emoji: 4 bytes each. |
| UTF-16 BE/LE | 2 or 4 | All Unicode | Default for JavaScript strings (in memory), Java char, Windows APIs. Above U+FFFF uses 4-byte surrogate pairs. |
| UTF-32 | 4 (fixed) | All Unicode | Simple but wasteful. Rare outside of internal processing. |
| ASCII | 1 (7-bit) | Only U+0000–U+007F | Anything above 127 is lossy. Older protocols, cursed legacy systems. |
| Latin-1 (ISO-8859-1) | 1 (fixed) | U+0000–U+00FF | Covers Western European accents (é, ñ, ç). Anything above 0xFF is lossy. |
| Windows-1252 | 1 (fixed) | ~256 characters | Superset of Latin-1 with curly quotes, em dash, euro sign, TM. Default for Windows Notepad. |
⚠ When you switch to ASCII, Latin-1, or Windows-1252, the counter flags every character that cannot be represented. Those characters would be lost (or replaced with ?) if you actually wrote the text out in that encoding.
Database and messaging fit-checks
The Fits in panel checks your byte count against the byte-based limits of common storage systems. Character-based limits (like MySQL utf8mb4 VARCHAR(255) counting 255 characters) are honored on the character side; byte-based ones (like Redis, DynamoDB, and file uploads) are honored on the byte side.
| Target | Limit | Applies to |
|---|---|---|
| MySQL VARCHAR(255) — utf8mb4 | 255 chars / max 1020 bytes | Common short-string column |
| MySQL TINYTEXT | 255 bytes | Small text blob |
| MySQL TEXT | 65 535 bytes | Article body, comments |
| MySQL MEDIUMTEXT | 16 777 215 bytes | Large documents |
| MySQL LONGTEXT | 4 294 967 295 bytes | Anything huge |
| Redis key (recommended) | < 100 bytes | Keeps memory / lookup fast |
| DynamoDB item | 400 KB | Whole record incl. attributes |
| SMS (GSM-7) | 160 chars / 140 bytes | Basic Latin only |
| SMS (UCS-2) | 70 chars | Any character; each is 2 bytes |
| Twitter / X post | 280 chars (weighted) | Most emoji count as 2 |
| URL (safe browser cap) | ~2 000 chars | IE historical limit; modern browsers accept far more |
How emoji byte-size works
Emoji surprise people because a single visible glyph often takes many bytes. 😀 (Grinning Face) is one Unicode code point (U+1F600) but takes 4 bytes in UTF-8. Modern compound emoji built with Zero-Width Joiners take more:
- 😀 (Grinning Face) → 4 UTF-8 bytes.
- 👨💻 (Man Technologist) → 11 UTF-8 bytes (Man + ZWJ + Laptop).
- 👨👩👧 (Family: Man Woman Girl) → 18 UTF-8 bytes.
- 🏳️🌈 (Rainbow Flag) → 14 UTF-8 bytes.
That means a 140-byte SMS can hold about 35 grinning faces… or one 👩🏽🚒 (Woman Firefighter: Medium Skin Tone) with a lot of room to spare — but not many family emoji.
When you need this
- Checking whether a user-generated string fits inside a database column.
- Estimating API request/response payload size.
- Debugging why a "255-character" field is rejecting a 200-character emoji string (utf8mb4 uses up to 4 bytes/char).
- SMS gateway pre-flight — knowing if you're going to be charged for 1 SMS or 3.
- Anywhere character count and byte count disagree.
Frequently asked questions
What encodings does this counter support?
UTF-8, UTF-16 (BE / LE), UTF-32, ASCII, Latin-1 (ISO-8859-1), and Windows-1252. UTF-8/16/32 use the browser's native TextEncoder; the others use exact byte-mapping tables.
How many bytes is one character?
Depends on the character and the encoding. See the encodings table above — a single letter can be 1 byte (ASCII) or 4 bytes (UTF-32).
How many bytes is an emoji?
A basic emoji is 4 bytes in UTF-8. Compound emoji made with Zero-Width Joiners can be 11–18 bytes or more. See the emoji section.
Do spaces count as bytes?
Yes. Regular space (U+0020) is 1 byte in UTF-8. Non-breaking space (U+00A0) is 2 bytes.
Will my text fit in MySQL VARCHAR(255)?
The Fits in panel above answers exactly that for your text, whether you're on latin1, utf8, or utf8mb4.
Why does ASCII show fewer bytes than my character count?
It shouldn't. If it does, some characters in your text can't be encoded as 7-bit ASCII — the counter shows a warning banner and either drops or escapes them depending on your setting.
Is my text saved?
No. All processing happens in your browser. Nothing is uploaded.
Related tools
References
- Unicode Standard 15.1 — Encoding forms UTF-8, UTF-16, UTF-32.
- RFC 3629 — UTF-8 encoding specification (byte-length ranges).
- ISO/IEC 8859-1 (Latin-1) and Microsoft Windows-1252 codepage tables.
- MySQL 8.0 Manual — VARCHAR, TINYTEXT, TEXT, MEDIUMTEXT, LONGTEXT byte limits.
- GSM 03.38 — SMS GSM-7 / UCS-2 character set rules.