Developer Formatting

HTML entity decoder and encoder

What to do

Processed in your browser · nothing is uploaded

Local · decoding is not sanitising

Turns HTML entities back into the characters they stand for, named ones like —, decimal ones like é and hexadecimal ones like é, and escapes in the other direction. Five characters need escaping in HTML: the ampersand, the two angle brackets, and both quote marks inside attribute values.

How to use the html entity decoder and encoder

1 Paste your input. The result appears immediately.
2 Pick the operation you want.
3 Copy the result, or save it as a file.

Double-encoding is the reason most people arrive here. A string that has been escaped twice shows & on the page, and decoding once gives &, which still looks wrong. The fix is decoding twice, and then finding where in the pipeline the value was escaped a second time, because the double encoding will come back tomorrow otherwise. Two escapes almost always means a template escaping a value that a framework had already escaped.

Going the other way, the encoder has two settings and the difference is worth understanding. The plain one escapes only what HTML actually requires: &, <, >, " and '. That is enough to make any string display as text rather than as markup. The non-ASCII option additionally turns every character above 127 into a numeric entity, which is a Latin-1 era workaround: with a UTF-8 page it only adds bytes, and it is worth reaching for only when something downstream mangles anything that is not plain ASCII.

Decoding is not sanitising, and the direction matters. Turning &lt;script&gt; back into <script> produces working markup, so decoded output must never be inserted into a page with innerHTML. Escaping is not sanitising either: it makes text display safely, and it does not clean up user-supplied markup you intend to render, which needs an allowlist sanitiser instead.

One detail: &nbsp; decodes to a non-breaking space (U+00A0), not an ordinary space. It looks identical and behaves differently; it will not wrap, and a string comparison against a normal space fails. That is the invisible cause of a great many "identical strings do not match" bugs.

What people use it for

  • Reading a feed or export whose markup arrived escaped
  • Escaping user text before it goes into a template
  • Untangling a value that was escaped twice
  • Putting a code sample on a page without it rendering

Questions

It was encoded twice. Decode again, then find where in the pipeline the second escape happens.

HTML Standard, named character references
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding