Unicode converter
- What you would call characters
- 6
- Code points
- 6
- UTF-16 code units
- 7
- UTF-8 bytes
- 10
| Char | Code point | UTF-8 | UTF-16 | HTML | JS | Block |
|---|---|---|---|---|---|---|
| C | U+0043 | 43 | 0043 | C | \u0043 | Basic Latin (ASCII) |
| a | U+0061 | 61 | 0061 | a | \u0061 | Basic Latin (ASCII) |
| f | U+0066 | 66 | 0066 | f | \u0066 | Basic Latin (ASCII) |
| Γ© | U+00E9 | C3 A9 | 00E9 | é | \u00E9 | Latin-1 Supplement |
| β£ | U+0020 | 20 | 0020 |   | \u0020 | Basic Latin (ASCII) |
| π | U+1F600 | F0 9F 98 80 | D83D DE00 | 😀 | \u{1F600} | Emoji and pictographs |
Inspected in your browser Β· nothing is uploaded
Splits text into code points and gives each one its U+ notation, UTF-8 bytes, UTF-16 units, HTML entity, JavaScript escape and the part of Unicode it comes from. Above the table sit the four ways of counting the same string: graphemes, code points, UTF-16 code units and UTF-8 bytes, which agree only while the text stays plain English.
How to use the unicode converter
The four counts, and why they disagree
A grapheme is what a person means by a character β the thing one press of backspace removes. A code point is what the Unicode standard assigns a number to. A UTF-16 code unit is what JavaScript reports from .length, because JavaScript strings are stored as UTF-16. A UTF-8 byte is what a database column, a network payload and a file size actually measure. For unaccented English all four are the same number, which is exactly why nothing breaks until the first emoji or the first name with an Γ© in it.
π is one grapheme, one code point, two UTF-16 units and four UTF-8 bytes. So "π".length is 2, and slicing that string at position 1 leaves one half of a surrogate pair, which is not a character at all; that is where the replacement diamond in a truncated post comes from. π¨βπ©βπ§ goes further: one grapheme, five code points, eight UTF-16 units and eighteen bytes, because it is a man, a woman and a girl joined by two zero-width joiners whose only job is to say "draw these as one". Paste it in and the table shows all five rows, two of them with nothing visible in the character column. Those are the joiners, and they are also why deleting a family emoji sometimes takes several presses.
Which count your limit should be using
It depends what the limit is protecting. A database column is measured in bytes, so a VARCHAR(20) in a UTF-8 table holds twenty bytes and not twenty characters. A display limit, the kind that says "no more than 40 characters", wants graphemes, because that is what a reader perceives and what a text box appears to hold. An API limit is whatever the API documented, which is worth reading rather than assuming, since the three answers can differ by a factor of four on the same string.
The byte count is where the most common Unicode bug in production lives. MySQL has an encoding named utf8 that stores at most three bytes per character, and every emoji needs four, so the column has to be utf8mb4, where the mb4 is literally maximum bytes 4. An emoji that vanishes between a form submission and the page it renders on is nearly always this, and the character is genuinely gone rather than hidden.
What this page does not tell you
The Block column is a range lookup rather than a full character-database query, so it says "Latin-1 Supplement" and not the formal name LATIN SMALL LETTER E WITH ACUTE. It is enough to orient you, and the official code charts are linked below when the name is what you need. The table also stops at 300 characters and says so when it truncates, which is far more than anyone needs to diagnose one string and keeps a pasted document from freezing the page. The four counts above it always cover the whole input, however long it is.
The grapheme count comes from Intl.Segmenter, which every current browser has. Where it is missing the number falls back to the code-point count rather than guessing, so on a very old browser the first two rows read the same. One case is worth naming because these counts are what settle it: two strings that look identical on screen and do not compare equal are almost always a normalisation difference, where Γ© exists both as one code point and as a plain e followed by a combining acute. The table shows that immediately as one row against two, with the code-point count reading 1 against 2. The fix is a normalisation pass rather than anything on this page, and the text cleaner carries NFC and NFKC.
What people use it for
- Finding out why a string length disagrees with what is on the screen
- Tracking down an invisible character pasted in from a web page or a PDF
- Checking whether a name survived an import with its accents intact
- Getting the exact escape for one character to paste into source code
- Working out how many bytes a string will occupy in a database column
- Telling apart two identical-looking strings that refuse to match
Questions
The length property counts UTF-16 code units, and any code point above U+FFFF is stored as two of them. Slicing at position 1 splits the pair and leaves half a character. Spreading the string or iterating with for...of walks code points instead and fixes most of it.
What a person calls a character: the unit one press of backspace deletes. A family emoji is one grapheme and five code points, a flag is one grapheme and two, and a skin-tone thumbs-up is one grapheme and two.
A MySQL column declared utf8 stores three bytes per character and an emoji needs four. The column has to be utf8mb4. Newer MySQL treats utf8 as an alias for it, but plenty of schemas predate that.
U+ followed by the hexadecimal number Unicode assigned to that character. It names the character and says nothing about how it is stored; the UTF-8 and UTF-16 columns are the storage.
Because the character has no visible glyph: a zero-width joiner, a variation selector, a soft hyphen, a non-breaking space. The code point column still identifies it, which is usually the entire reason for looking.
The two 16-bit halves UTF-16 uses to store one code point above U+FFFF. Neither half is a character on its own, and one alone renders as a replacement glyph. The UTF-16 column shows both halves.
Because the four-digit \uXXXX cannot express a code point above U+FFFF, so anything that large needs \u{1F600} instead. The column picks the right form per character; writing the short one by hand is a silent bug rather than an error.
No. The Block column is an approximate range lookup, enough to tell you which part of Unicode a character comes from. For the formal name, the code charts are linked at the foot of this page.
As much as you like. The counts cover all of it and the table shows the first 300 characters, with a note when it has truncated.
The whole input as a space-separated list of U+ notations. That is the form to paste into an issue or a chat, because it survives every system between you and the person reading it.
Yes, and that is what the counts are for. Paste each in turn: the one whose code-point count is higher than its grapheme count is carrying a combining mark, a joiner or something with no glyph, and the table names it. Making them equal afterwards is a normalisation pass in the text cleaner.
Without the u flag a dot matches one code unit, so it takes half a surrogate pair. With /u it matches a code point. The UTF-16 column is what shows you the two halves it was choking on.
For plain text with no byte order mark, yes. Saving from an editor may add a BOM of three bytes and may rewrite line endings, and neither is counted here.
Not on this page, which reads in one direction. Paste the character itself and the table gives you every notation for it.
No. Everything happens in your browser; you can load the page, go offline, and it still works.