HTML entity decoder and encoder
Processed in your browser · nothing is uploaded
Turns HTML entities back into the characters they stand for, named ones like —, decimal ones like é and hexadecimal ones like é, and escapes in the other direction. Five characters need escaping in HTML: the ampersand, the two angle brackets, and both quote marks inside attribute values.
How to use the html entity decoder and encoder
Double-encoding is the reason most people arrive here. A string that has been escaped twice shows & on the page, and decoding once gives &, which still looks wrong. The fix is decoding twice, and then finding where in the pipeline the value was escaped a second time, because the double encoding will come back tomorrow otherwise. Two escapes almost always means a template escaping a value that a framework had already escaped.
Going the other way, the encoder has two settings and the difference is worth understanding. The plain one escapes only what HTML actually requires: &, <, >, " and '. That is enough to make any string display as text rather than as markup. The non-ASCII option additionally turns every character above 127 into a numeric entity, which is a Latin-1 era workaround: with a UTF-8 page it only adds bytes, and it is worth reaching for only when something downstream mangles anything that is not plain ASCII.
Decoding is not sanitising, and the direction matters. Turning <script> back into <script> produces working markup, so decoded output must never be inserted into a page with innerHTML. Escaping is not sanitising either: it makes text display safely, and it does not clean up user-supplied markup you intend to render, which needs an allowlist sanitiser instead.
One detail: decodes to a non-breaking space (U+00A0), not an ordinary space. It looks identical and behaves differently; it will not wrap, and a string comparison against a normal space fails. That is the invisible cause of a great many "identical strings do not match" bugs.
What people use it for
- Reading a feed or export whose markup arrived escaped
- Escaping user text before it goes into a template
- Untangling a value that was escaped twice
- Putting a code sample on a page without it rendering
Questions
It was encoded twice. Decode again, then find where in the pipeline the second escape happens.
Ampersand, less-than and greater-than always; double and single quotes inside attribute values.
Yes, both decimal (é) and hexadecimal (é).
Every numeric one, decimal and hexadecimal, plus the common named ones. Anything rarer is left untouched rather than guessed at.
The first escapes only what HTML requires. The second also turns every character above ASCII into a numeric entity, which a UTF-8 page does not need.
No. Decoding turns <script> back into working markup. Keep it encoded for display.
It makes text display safely. It does not sanitise user-supplied markup, which needs an allowlist sanitiser.
A non-breaking space, U+00A0, which looks like a space and is not one, and will not compare equal to it.
' is fine in XML and HTML5 but was not defined in HTML 4.01. The numeric form is understood everywhere.
Not with UTF-8. Numeric entities for non-ASCII are a Latin-1 era workaround that only adds bytes now.
No. Entities protect markup; percent-encoding protects a URL. A value going into an href needs both, percent-encoded first.
No. Everything runs in your browser, which matters when the file holds credentials or customer data.