Developer tools · HTML entity escaper
Why © without a semicolon still decodes: HTML's legacy named references
· How it works
html browser-apis encoding
HTML parsers decode a small set of older named references even when the semicolon is missing, so a raw ¬ or © in text can turn into ¬ or ©. This post explains the legacy list, the longest-match rule, the attribute exception, and why escaping every ampersand avoids all of it.
The ¬ify that appeared in a page — a raw '¬ify' in text content, decoded to ¬ plus 'ify' by a legacy rule
The ¬ify that appeared in a page — a raw '¬ify' in text content, decoded to ¬ plus 'ify' by a legacy rule. A browser may interpret a semicolon-less legacy name inside prose, which explains surprising transformations such as a prefix of ¬ify. ToolAcre does not imitate that recovery behavior.
To verify html entity without semicolon, construct the ify that appeared for a developer whose page shows ¬ify where the text said ¬ify. Preserve in a page a while semicolon recovery boundary produces raw notify in text; identify where content decoded to plus is consumed. The observation about ify by a legacy belongs to HTML text only.
The legacy list — the older HTML 4 names that browsers must accept without a semicolon for compatibility
The legacy list — the older HTML 4 names that browsers must accept without a semicolon for compatibility. Legacy acceptance belongs to the browser tokenizer and varies by state. The repository decoder instead recognizes only explicit names in its 235-entry table followed by a semicolon.
A developer whose page shows ¬ify where the text said ¬ify can test the legacy list the by recording older html 4 names before the semicolon recovery boundary pass. Compare that browsers must accept afterward and locate the parser responsible for without a semicolon for. This html entity without semicolon result explains compatibility, not executable contexts.
How the tokenizer matches — consuming the longest name in the table, so ¬ wins inside ¬ify
How the tokenizer matches — consuming the longest name in the table, so ¬ wins inside ¬ify. Browser longest-match behavior can consume a known prefix before the reader expects it. The bounded regular expression in this tool avoids prefix guessing because its match must end at a semicolon.
Isolate how the tokenizer matches in a short semicolon recovery boundary sample. Show consuming the longest name as literal source, follow in the table so to its destination, and name the API reading not wins inside notify. For html entity without semicolon, semicolon recovery boundary evidence remains parser-bound evidence.
The attribute exception — why ©=2 inside an href survives, but © followed by & or the end of the value does not
The attribute exception — why ©=2 inside an href survives, but © followed by & or the end of the value does not. Attribute parsing adds exceptions involving equals signs and alphanumeric followers. Those browser rules are precisely why a compact utility should not claim parser-equivalent recovery.
Treat the attribute exception why as a boundary experiment. A developer whose page shows ¬ify where the text said ¬ify should retain copy 2 inside an, perform one semicolon recovery boundary operation, and inspect href survives but copy character by character before changing followed by or the. The claim about end of the value stops at this HTML layer.
Why new entities always require the semicolon — the compatibility line drawn when parsing was standardised in HTML5
Why new entities always require the semicolon — the compatibility line drawn when parsing was standardised in HTML5. In ToolAcre, © stays ©, © becomes ©, and ©x remains untouched because no matching table key exists. This differs intentionally from a forgiving browser parse.
Reproduce why new entities always with harmless input instead of customer material. Record require the semicolon the, observe compatibility line drawn when, and count every intentional semicolon recovery boundary pass. That html entity without semicolon trail lets a developer whose page shows ¬ify where the text said ¬ify evaluate parsing was standardised in and html5 without guessing.
Worked example: ©, ©, ©x and ©=2 in text and in an href — what a browser renders for each
Worked example: ©, ©, ©x and ©=2 in text and in an href — what a browser renders for each. New and obscure names should always carry semicolons. The utility makes that discipline observable: omission produces unchanged text rather than a guessed character or partial match.
Place worked example copy copy, copyx and copy 2, and in text and in side by side during the semicolon recovery boundary review. A developer whose page shows ¬ify where the text said ¬ify can then decide whether an href what a changed at conversion or downstream. Keep the html entity without semicolon conclusion about browser renders for each out of generic security claims.
What this does not cover — the full character reference state machine and error-recovery details
What this does not cover — the full character reference state machine and error-recovery details. This article does not reproduce the full character-reference state machine. It distinguishes browser legacy recovery from the tool’s strict contract so readers do not infer unsupported behavior.
Define what this does not before running semicolon recovery boundary. Save cover the full character as a control, inspect the code points behind reference state machine and, and map error recovery details to the next interpreter. This makes semicolon recovery boundary evidence auditable for a developer whose page shows ¬ify where the text said ¬ify investigating html entity without semicolon.
Takeaway: browsers accept some legacy omissions; ToolAcre deliberately requires semicolons
Takeaway: browsers accept some legacy omissions; ToolAcre deliberately requires semicolons. Escape every literal ampersand before inserting text into HTML. The encoder produces & deterministically, while its decoder requires terminated references and never promises legacy parsing.
Connect takeaway browsers accept some to an observable semicolon recovery boundary output. Keep legacy omissions toolacre deliberately beside the one-pass result, then verify where requires semicolons enters semicolon recovery boundary evidence. A developer whose page shows ¬ify where the text said ¬ify can now review semicolon recovery boundary evidence as a narrow html entity without semicolon finding. The practical decision behind this article is specific: HTML parsers decode a small set of older named references even when the semicolon is missing, so a raw ¬ or © in text can turn into ¬ or ©. This post explains the legacy list, the longest-match rule, the attribute exception, and why escaping every ampersand avoids all of it. The reader action is equally concrete: Links to the HTML entity escaper and demonstrates escaping a URL with query parameters so its ampersands become & before it goes into an href.