Developer Encoding

Unicode escape converter

What you would call characters
11
Code points
11
UTF-16 code units
11
UTF-8 bytes
13

Inspected in your browser · nothing is uploaded

Local · escapes and decodes at the same time

Escapes every character above ASCII into JavaScript, HTML, CSS or URL notation, and decodes escapes back into characters at the same time. Both boxes update as you type, so a string that arrived already escaped shows you what it decodes to while you are still deciding how to re-escape it.

How to use the unicode escape converter

1 Paste text with accents, symbols or emoji in it.
2 Pick the syntax you are escaping for: JavaScript, HTML, CSS or URL.
3 Copy the escaped box, or paste escapes into the input and read the plain text out of the unescaped box.

The four syntaxes are genuinely different

JavaScript gets \u00E9, and above U+FFFF the brace form \u{1F600}, because the four-digit form cannot express a code point that large. HTML gets a decimal numeric reference, é, rather than a named entity: there are only a few hundred named entities and they cover a fraction of Unicode, while the numeric form covers all of it and every parser has understood it since the beginning.

CSS gets \00E9 followed by a space, and that space is not decoration. A CSS escape is hexadecimal and variable length, up to six digits, so something has to mark where it ends; a single space does it and the parser then consumes it. Leave it out and \00E9a reads as one escape for U+0E9A, a Lao character, rather than an é followed by an a. The digits here are padded to four, which does not remove the need for the space, since four is still under the six a CSS escape may run to and the parser keeps reading.

URL is the odd one out, because it escapes bytes rather than characters. é becomes %C3%A9, two escapes for one character, because that is its UTF-8 representation. That is why a URL with accents in it looks so much longer than the text it came from, and why percent-decoding a string byte by byte gives nonsense unless you reassemble the UTF-8 afterwards.

The URL setting is not a URL encoder

It touches what is above ASCII and nothing else, so a space stays a space and an ampersand stays an ampersand. That is right for seeing what a piece of text costs in bytes and wrong for a whole query string, where the separators have to be encoded too or the parameters run into each other. Use the URL encoder for that; use this to check one value before it goes in.

Decoding, and what it will decode

The unescaped box runs continuously and takes JavaScript escapes in both forms along with HTML numeric references in decimal and hexadecimal. It does not take CSS escapes or percent-encoding back the other way. A CSS escape is ambiguous without the stylesheet around it, since the trailing space may belong to the escape or to the content, and percent-decoding belongs with the URL encoder, which handles the plus-for-space convention properly. So the round trip is exact for JavaScript and HTML and one-way for CSS and URL.

Escaping works by code point, not by storage unit, which matters for anything above U+FFFF: 😀 comes out as a single \u{1F600} or 😀 rather than a surrogate pair. Text below 128 passes through untouched in every mode, which is the point of the exercise. What comes out is safe to put in an ASCII-only file and still readable by a person.

Double escaping

This is the failure worth recognising on sight. Text that has been escaped twice shows \\u00E9 or é — decoding it once leaves an escape still on screen. Decode again to see the real text, then go and find the place in the pipeline that escaped an already-escaped value, because it will do the same thing tomorrow. Two escapes almost always means a template escaping something a framework had already escaped.

When not to escape at all

Most of the time you do not need any of this. With UTF-8 everywhere a JavaScript source file, an HTML page and a stylesheet can all hold é directly, and escaping only makes them longer and harder to read. Reach for it when something in the chain cannot carry the character itself: a Java properties file that has to be plain ASCII, a legacy system that mangles anything above 127, a CSS content value where the character would end the string, or a build step you do not control.

What people use it for

  • Putting a non-ASCII string into a file that has to stay plain ASCII
  • Writing a CSS content value for an arrow, a quote mark or an icon glyph
  • Reading a JSON payload that arrived full of backslash-u escapes
  • Checking what one value costs in bytes before it goes into a URL
  • Untangling a string that has been escaped twice

Questions

CSS escapes are hexadecimal and variable length, so the space marks where the escape ends and is then consumed by the parser. Without it, \00E9a is read as a single escape for U+0E9A rather than as é followed by a.

CSS Syntax Module Level 3, escapingMDN, string literals and escape sequencesRFC 3986, percent-encoding
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding