Developer Encoding

Unicode converter

What you would call characters
6
Code points
6
UTF-16 code units
7
UTF-8 bytes
10
CharCode pointUTF-8UTF-16HTMLJSBlock
CU+0043430043C\u0043Basic Latin (ASCII)
aU+0061610061a\u0061Basic Latin (ASCII)
fU+0066660066f\u0066Basic Latin (ASCII)
Γ©U+00E9C3 A900E9é\u00E9Latin-1 Supplement
␣U+0020200020 \u0020Basic Latin (ASCII)
πŸ˜€U+1F600F0 9F 98 80D83D DE00😀\u{1F600}Emoji and pictographs

Inspected in your browser Β· nothing is uploaded

Local Β· four counts, and a row for every character

Splits text into code points and gives each one its U+ notation, UTF-8 bytes, UTF-16 units, HTML entity, JavaScript escape and the part of Unicode it comes from. Above the table sit the four ways of counting the same string: graphemes, code points, UTF-16 code units and UTF-8 bytes, which agree only while the text stays plain English.

How to use the unicode converter

1 Paste the text that is misbehaving. A single character is enough.
2 Read the four counts first. Wherever they disagree is usually where the bug is.
3 Find the character in the table and take its code point, bytes or escape off the row.
4 Copy gives you the whole string as a list of U+ notations, which is the form worth pasting into a bug report.

The four counts, and why they disagree

A grapheme is what a person means by a character β€” the thing one press of backspace removes. A code point is what the Unicode standard assigns a number to. A UTF-16 code unit is what JavaScript reports from .length, because JavaScript strings are stored as UTF-16. A UTF-8 byte is what a database column, a network payload and a file size actually measure. For unaccented English all four are the same number, which is exactly why nothing breaks until the first emoji or the first name with an Γ© in it.

πŸ˜€ is one grapheme, one code point, two UTF-16 units and four UTF-8 bytes. So "πŸ˜€".length is 2, and slicing that string at position 1 leaves one half of a surrogate pair, which is not a character at all; that is where the replacement diamond in a truncated post comes from. πŸ‘¨β€πŸ‘©β€πŸ‘§ goes further: one grapheme, five code points, eight UTF-16 units and eighteen bytes, because it is a man, a woman and a girl joined by two zero-width joiners whose only job is to say "draw these as one". Paste it in and the table shows all five rows, two of them with nothing visible in the character column. Those are the joiners, and they are also why deleting a family emoji sometimes takes several presses.

Which count your limit should be using

It depends what the limit is protecting. A database column is measured in bytes, so a VARCHAR(20) in a UTF-8 table holds twenty bytes and not twenty characters. A display limit, the kind that says "no more than 40 characters", wants graphemes, because that is what a reader perceives and what a text box appears to hold. An API limit is whatever the API documented, which is worth reading rather than assuming, since the three answers can differ by a factor of four on the same string.

The byte count is where the most common Unicode bug in production lives. MySQL has an encoding named utf8 that stores at most three bytes per character, and every emoji needs four, so the column has to be utf8mb4, where the mb4 is literally maximum bytes 4. An emoji that vanishes between a form submission and the page it renders on is nearly always this, and the character is genuinely gone rather than hidden.

What this page does not tell you

The Block column is a range lookup rather than a full character-database query, so it says "Latin-1 Supplement" and not the formal name LATIN SMALL LETTER E WITH ACUTE. It is enough to orient you, and the official code charts are linked below when the name is what you need. The table also stops at 300 characters and says so when it truncates, which is far more than anyone needs to diagnose one string and keeps a pasted document from freezing the page. The four counts above it always cover the whole input, however long it is.

The grapheme count comes from Intl.Segmenter, which every current browser has. Where it is missing the number falls back to the code-point count rather than guessing, so on a very old browser the first two rows read the same. One case is worth naming because these counts are what settle it: two strings that look identical on screen and do not compare equal are almost always a normalisation difference, where Γ© exists both as one code point and as a plain e followed by a combining acute. The table shows that immediately as one row against two, with the code-point count reading 1 against 2. The fix is a normalisation pass rather than anything on this page, and the text cleaner carries NFC and NFKC.

What people use it for

  • Finding out why a string length disagrees with what is on the screen
  • Tracking down an invisible character pasted in from a web page or a PDF
  • Checking whether a name survived an import with its accents intact
  • Getting the exact escape for one character to paste into source code
  • Working out how many bytes a string will occupy in a database column
  • Telling apart two identical-looking strings that refuse to match

Questions

The length property counts UTF-16 code units, and any code point above U+FFFF is stored as two of them. Slicing at position 1 splits the pair and leaves half a character. Spreading the string or iterating with for...of walks code points instead and fixes most of it.

The Unicode Standard, latest versionUAX #29, Unicode text segmentation and grapheme clustersMySQL, the utf8mb4 character set
Was this tool any good?
Internal signal only Β· I use it to find the tools worth rebuilding