Developer tools · HTML entity escaper
Numeric HTML entities: © vs © and why an emoji is one reference
· How it works
html unicode encoding
Numeric character references write a Unicode code point in decimal or hexadecimal. This post shows how to read them, why an emoji is one reference rather than two, and where the numbers come from.
The 😀 that appeared in a database export — reading a numeric entity without a lookup table
The 😀 that appeared in a database export — reading a numeric entity without a lookup table. A decimal reference exposes its value directly: 😀 asks for Unicode code point 128512. The decoder does not need a named-entry lookup when the number is present.
To verify html numeric character reference, construct the 128512 that appeared for a developer seeing 😀 in exported content. Preserve in a database export while numeric-reference arithmetic produces reading a numeric entity; identify where without a lookup table is consumed. The observation about numeric-reference arithmetic evidence belongs to HTML text only.
Decimal and hexadecimal forms — &#NNN; and &#xHHHH;, case rules for the x and the digits, and the semicolon
Decimal and hexadecimal forms — &#NNN; and &#xHHHH;, case rules for the x and the digits, and the semicolon. The accepted syntax is strict and terminated: one to seven decimal digits, or an x or X followed by one to six hexadecimal digits, then a semicolon. Missing terminators stay untouched.
A developer seeing 😀 in exported content can test decimal and hexadecimal forms by recording nnn and xhhhh case before the numeric-reference arithmetic pass. Compare rules for the x afterward and locate the parser responsible for and the digits and. This html numeric character reference result explains the semicolon, not executable contexts.
Code points, not bytes — why é is é regardless of the page's charset
Code points, not bytes — why é is é regardless of the page's charset. Numeric references identify Unicode code points rather than encoded bytes. Decimal 233 and hexadecimal E9 therefore resolve to é independent of how a surrounding file represents that character.
Isolate code points not bytes in a short numeric-reference arithmetic sample. Show why 233 is regardless as literal source, follow of the page s to its destination, and name the API reading charset. For html numeric character reference, numeric-reference arithmetic evidence remains parser-bound evidence.
Astral characters and UTF-16 — why 😀 is one reference in HTML but two code units in JavaScript
Astral characters and UTF-16 — why 😀 is one reference in HTML but two code units in JavaScript. The encoder iterates JavaScript strings with for-of, which yields a complete astral character. An emoji is encoded as 😀 or its decimal equivalent, not as two surrogate references.
Treat astral characters and utf as a boundary experiment. A developer seeing 😀 in exported content should retain 16 why x1f600 is, perform one numeric-reference arithmetic operation, and inspect one reference in html character by character before changing but two code units. The claim about in javascript stops at this HTML layer.
Worked example: converting between 😀, 😀, the character itself and the JavaScript string — all four forms
Worked example: converting between 😀, 😀, the character itself and the JavaScript string — all four forms. The computed chain is exact: 😀 is U+1F600, decimal 128512, hexadecimal 1F600, HTML 😀 or 😀, and a JavaScript string contains the same visible scalar value.
Reproduce worked example converting between with harmless input instead of customer material. Record x1f600 128512 the character, observe itself and the javascript, and count every intentional numeric-reference arithmetic pass. That html numeric character reference trail lets a developer seeing 😀 in exported content evaluate string all four forms and numeric-reference arithmetic evidence without guessing.
Invalid and disallowed values — surrogates, nulls and code points above 0x10FFFF, and how parsers replace them
Invalid and disallowed values — surrogates, nulls and code points above 0x10FFFF, and how parsers replace them. NUL, isolated surrogate values and values above U+10FFFF decode to U+FFFD. The tests cover �, � and �, so invalid references produce replacement characters instead of exceptions.
Place invalid and disallowed values, surrogates nulls and code, and points above 0x10ffff and side by side during the numeric-reference arithmetic review. A developer seeing 😀 in exported content can then decide whether how parsers replace them changed at conversion or downstream. Keep the html numeric character reference conclusion about numeric-reference arithmetic evidence out of generic security claims.
What this does not cover — the legacy Windows-1252 remapping of €–Ÿ, treated in a separate post
What this does not cover — the legacy Windows-1252 remapping of €–Ÿ, treated in a separate post. Values from decimal 128 through 159 are a deliberate compatibility exception handled by a Windows-1252 map. Their special mapping is separated from the ordinary code-point explanation here.
Define what this does not before running numeric-reference arithmetic. Save cover the legacy windows as a control, inspect the code points behind 1252 remapping of 128, and map 159 treated in a to the next interpreter. This makes separate post auditable for a developer seeing 😀 in exported content investigating html numeric character reference.
Takeaway: a numeric entity is a code point in disguise — how the HTML entity escaper's decoder turns references back into characters as data, never markup
Takeaway: a numeric entity is a code point in disguise — how the HTML entity escaper's decoder turns references back into characters as data, never markup. Numeric entities are code points written as text. The tool resolves them with arithmetic and String.fromCodePoint while never assigning the surrounding input to an HTML parser.
Connect takeaway a numeric entity to an observable numeric-reference arithmetic output. Keep is a code point beside the one-pass result, then verify where in disguise how the enters html entity escaper s. A developer seeing 😀 in exported content can now review decoder turns references back as a narrow html numeric character reference finding. The practical decision behind this article is specific: Numeric character references write a Unicode code point in decimal or hexadecimal. This post shows how to read them, why an emoji is one reference rather than two, and where the numbers come from. The reader action is equally concrete: Links to the HTML entity escaper and demonstrates decoding a numeric reference, pointing to the tool page's 'Supported input and output' section for the reference forms it accepts.