Unicode escape converter
- What you would call characters
- 11
- Code points
- 11
- UTF-16 code units
- 11
- UTF-8 bytes
- 13
Inspected in your browser · nothing is uploaded
Escapes every character above ASCII into JavaScript, HTML, CSS or URL notation, and decodes escapes back into characters at the same time. Both boxes update as you type, so a string that arrived already escaped shows you what it decodes to while you are still deciding how to re-escape it.
How to use the unicode escape converter
The four syntaxes are genuinely different
JavaScript gets \u00E9, and above U+FFFF the brace form \u{1F600}, because the four-digit form cannot express a code point that large. HTML gets a decimal numeric reference, é, rather than a named entity: there are only a few hundred named entities and they cover a fraction of Unicode, while the numeric form covers all of it and every parser has understood it since the beginning.
CSS gets \00E9 followed by a space, and that space is not decoration. A CSS escape is hexadecimal and variable length, up to six digits, so something has to mark where it ends; a single space does it and the parser then consumes it. Leave it out and \00E9a reads as one escape for U+0E9A, a Lao character, rather than an é followed by an a. The digits here are padded to four, which does not remove the need for the space, since four is still under the six a CSS escape may run to and the parser keeps reading.
URL is the odd one out, because it escapes bytes rather than characters. é becomes %C3%A9, two escapes for one character, because that is its UTF-8 representation. That is why a URL with accents in it looks so much longer than the text it came from, and why percent-decoding a string byte by byte gives nonsense unless you reassemble the UTF-8 afterwards.
The URL setting is not a URL encoder
It touches what is above ASCII and nothing else, so a space stays a space and an ampersand stays an ampersand. That is right for seeing what a piece of text costs in bytes and wrong for a whole query string, where the separators have to be encoded too or the parameters run into each other. Use the URL encoder for that; use this to check one value before it goes in.
Decoding, and what it will decode
The unescaped box runs continuously and takes JavaScript escapes in both forms along with HTML numeric references in decimal and hexadecimal. It does not take CSS escapes or percent-encoding back the other way. A CSS escape is ambiguous without the stylesheet around it, since the trailing space may belong to the escape or to the content, and percent-decoding belongs with the URL encoder, which handles the plus-for-space convention properly. So the round trip is exact for JavaScript and HTML and one-way for CSS and URL.
Escaping works by code point, not by storage unit, which matters for anything above U+FFFF: 😀 comes out as a single \u{1F600} or 😀 rather than a surrogate pair. Text below 128 passes through untouched in every mode, which is the point of the exercise. What comes out is safe to put in an ASCII-only file and still readable by a person.
Double escaping
This is the failure worth recognising on sight. Text that has been escaped twice shows \\u00E9 or é — decoding it once leaves an escape still on screen. Decode again to see the real text, then go and find the place in the pipeline that escaped an already-escaped value, because it will do the same thing tomorrow. Two escapes almost always means a template escaping something a framework had already escaped.
When not to escape at all
Most of the time you do not need any of this. With UTF-8 everywhere a JavaScript source file, an HTML page and a stylesheet can all hold é directly, and escaping only makes them longer and harder to read. Reach for it when something in the chain cannot carry the character itself: a Java properties file that has to be plain ASCII, a legacy system that mangles anything above 127, a CSS content value where the character would end the string, or a build step you do not control.
What people use it for
- Putting a non-ASCII string into a file that has to stay plain ASCII
- Writing a CSS content value for an arrow, a quote mark or an icon glyph
- Reading a JSON payload that arrived full of backslash-u escapes
- Checking what one value costs in bytes before it goes into a URL
- Untangling a string that has been escaped twice
Questions
CSS escapes are hexadecimal and variable length, so the space marks where the escape ends and is then consumed by the parser. Without it, \00E9a is read as a single escape for U+0E9A rather than as é followed by a.
For consistency with the other syntaxes and because it stays unambiguous. Four digits is still fewer than the six a CSS escape may run to, so the trailing space is required either way.
Percent-encoding escapes bytes rather than characters, and é is two bytes in UTF-8. A character outside the Basic Multilingual Plane becomes four.
No. It escapes only what is above ASCII, so a space stays a space and an ampersand stays an ampersand. A whole query string needs the URL encoder, which handles the separators as well.
Yes. The unescaped box decodes as you type, taking JavaScript escapes in both forms and HTML numeric references in decimal and hexadecimal.
A CSS escape is ambiguous on its own, because the trailing space may belong to the escape or to the text after it. Percent-decoding belongs with the URL encoder, which also handles plus as a space. Both directions are exact for JavaScript and HTML.
Only the brace form, \u{1F600}. The four-digit form cannot express it, and the escaped box picks the right form for each character automatically.
It escapes by code point, so 😀 comes out as one \u{1F600} or 😀 rather than as two surrogate halves.
No. Everything below 128 passes through unchanged in all four modes, which is what makes the output readable as well as safe.
It was escaped twice. Decode again, then find where in the pipeline the second escape happens; that is the actual bug.
Usually not. Numeric entities and backslash escapes for non-ASCII are a Latin-1 era workaround that now only adds bytes and hides what the string says.
Named entities cover a few hundred characters. The numeric form covers every code point and is understood everywhere, so a converter that has to handle arbitrary text should emit it.
For JavaScript and HTML, yes: escape then unescape returns the original string. CSS and URL are one-way here, by design.
The Unicode converter has a table with a row per character, giving the code point, both encodings and the escapes side by side.
No. Everything happens in your browser; you can load the page, go offline, and it still works.