Developer tools · HTML entity escaper
Decoding HTML entities without innerHTML: how a lookup-table decoder works
· How it works
html security encoding
The popular trick of decoding entities by assigning to innerHTML runs your input through the HTML parser, which is exactly what you do not want. This post explains the safer table-based approach and how it handles named, decimal and hex references.
The decoder that executed an <img onerror> — a concrete case where 'just decode it' became script execution
The common one-liner element.innerHTML = input does more than decode &. If the input also contains <img src=x onerror=...>, the browser creates an image element and an event-handler attribute. Depending on how that node is attached and loaded, this can turn a formatting shortcut into script execution. A pasted HTML-like string should remain data when the only job is to resolve character references, not be parsed into a DOM tree.
What innerHTML actually does with a string — parsing, element creation and event handler attributes, not just entity replacement
innerHTML invokes the HTML parser: tags become nodes, attributes acquire browser meaning, and a later textContent read strips markup out of the result. A <strong> tag you intended to preserve as literal input can vanish as text formatting. A detached element is not a general safety guarantee; code often re-inserts that subtree or uses the resulting HTML elsewhere. If you need to display untrusted input, assign textContent and sanitise only when you deliberately choose to render HTML.
The lookup-table approach — a regular expression for &name;, &#NNN; and &#xHHH;, and a map from names to characters
ToolAcre’s decoder uses a bounded regular expression to find a reference of the form &name;, { or {. Named references are looked up in a practical explicit table, including amp, lt, gt, quotes and common typography. An unknown name is left as written rather than guessed. This approach creates no elements and calls no HTML parser; it simply replaces recognized substrings in a string. The table is deliberately a subset, not all HTML named character references.
Handling numeric references — parsing decimal and hexadecimal code points and converting them to strings, including astral characters
For a numeric reference, parse decimal after &# or hexadecimal after &#x, then turn the numeric code point into a character with String.fromCodePoint. An astral value such as 0x1F600 yields an emoji, not two independent printable characters. The implementation also maps historic Windows-1252 control-range values as browsers do; zero, surrogate code points and values above U+10FFFF become a replacement character. That explicit error handling keeps an invalid number from crashing the decoder.
Worked example: decoding a string that mixes &, © and 😀 — each match resolved from the table or the number
Decode the literal input &, © and 😀: the first reference maps through the named table to &, decimal 169 becomes ©, and hex 1F600 becomes 😀. Include a raw <img onerror="alert(1)"> beside them. The decoder returns that tag-looking sequence as ordinary string characters; it does not create an image or execute an event. When putting the result into a real page later, use a safe text sink rather than taking the decoded string and assigning it back to innerHTML.
What the table approach will not do — legacy semicolon-less references and parser error-recovery quirks, unless deliberately implemented
The lookup approach intentionally does not reproduce the HTML parser’s legacy semicolon-less recovery rules. © without a semicolon may stay untouched. The fixed named table also omits many of the more than two thousand HTML5 named references. Those limitations are honest trade-offs for a small predictable decoder; accepting only explicit terminated references avoids treating arbitrary prose containing an ampersand as markup. Check the tool’s documented supported names if full browser compatibility is essential.
What this does not cover — sanitising HTML you intend to render, which is a different problem
Decoding references is not sanitising HTML for display. If the decoded text contains a <script> sequence, it remains dangerous if a different part of an application later inserts it as markup. Contexts such as an HTML attribute, a JavaScript string and a URL each need their own output encoding and policy. ToolAcre returns text; it cannot make a future unsafe sink safe.
Takeaway: treat input as data — how the HTML entity escaper decodes with a lookup table rather than the HTML parser, so your input stays text
Treat input as data. HTML entity escaper decodes with a table and code-point arithmetic, not an innerHTML trick, so markup-looking payloads remain inert characters within the tool. Try the three references, then inspect both the decoded text and how you plan to use it next: the safety boundary is lost if you reparse it as HTML.