English

Developer tools · HTML entity escaper

Double-escaped HTML: how & and < happen and how to fix them

· How it works

html encoding debugging

Double-escaped HTML: how entity notationamp; and entity notationlt; happen and how to fix them shown as a browser-safe character-reference diagram
Original ToolAcre vector illustration

& on a page means text was escaped twice, usually once by a template and once by hand. This post explains the layers involved, how to spot the pattern in the source and how many decode passes are safe.

The product called 'Fish & Chips' on the live page — a concrete double-escape and how it slipped through review

The product called 'Fish & Chips' on the live page — a concrete double-escape and how it slipped through review. A page displaying Fish & Chips usually received text already escaped once and then escaped again. The second pass converted the first reference’s leading ampersand into &.

To verify html double escaping, construct the product called fish for a CMS editor seeing & in page titles. Preserve amp chips on the while double-escape diagnosis produces live page a concrete; identify where double escape and how is consumed. The observation about it slipped through review belongs to HTML text only.

Two layers that both escape — template auto-escaping plus manual escaping, or escaping on save and again on render

Two layers that both escape — template auto-escaping plus manual escaping, or escaping on save and again on render. Two boundaries often share responsibility: storage contains presentation-ready text, then a template auto-escapes it again. Another path manually escapes just before an already-safe rendering API.

A CMS editor seeing & in page titles can test two layers that both by recording escape template auto escaping before the double-escape diagnosis pass. Compare plus manual escaping or afterward and locate the parser responsible for escaping on save and. This html double escaping result explains again on render, not executable contexts.

Reading the signature — < ' and friends, and why the leading & is the tell

Reading the signature — < ' and friends, and why the leading & is the tell. Signatures such as <, " and ' reveal an encoded ampersand immediately before text that still resembles a reference. They indicate layers, not a new entity type.

Isolate reading the signature amp in a short double-escape diagnosis sample. Show lt amp 39 and as literal source, follow friends and why the to its destination, and name the API reading leading amp is the. For html double escaping, tell remains parser-bound evidence.

Worked example: decoding a doubly escaped string one pass at a time — what each pass reveals and when to stop

Worked example: decoding a doubly escaped string one pass at a time — what each pass reveals and when to stop. ToolAcre decodes one pass only. &amp;lt; becomes the literal text &lt; after the first click; a second explicit pass becomes <. Each intermediate value remains visible for diagnosis.

Treat worked example decoding a as a boundary experiment. A CMS editor seeing &amp; in page titles should retain doubly escaped string one, perform one double-escape diagnosis operation, and inspect pass at a time character by character before changing what each pass reveals. The claim about and when to stop stops at this HTML layer.

Fixing the pipeline — escape exactly once, at output time, and store the raw text

Fixing the pipeline — escape exactly once, at output time, and store the raw text. The durable repair is to store semantic text and encode at the final HTML output boundary. Marking pre-escaped values as trusted merely moves ambiguity and can later expose raw markup.

Reproduce fixing the pipeline escape with harmless input instead of customer material. Record exactly once at output, observe time and store the, and count every intentional double-escape diagnosis pass. That html double escaping trail lets a CMS editor seeing &amp; in page titles evaluate raw text and double-escape diagnosis evidence without guessing.

Common mistakes — decoding until nothing changes, which corrupts legitimate text about entities

Common mistakes — decoding until nothing changes, which corrupts legitimate text about entities. Repeatedly decoding until a value stops changing corrupts legitimate writing about entities and may reveal markup intended as text. Count architectural layers instead of using an open-ended loop.

Place common mistakes decoding until, nothing changes which corrupts, and legitimate text about entities side by side during the double-escape diagnosis review. A CMS editor seeing &amp; in page titles can then decide whether double-escape diagnosis evidence changed at conversion or downstream. Keep the html double escaping conclusion about double-escape diagnosis evidence out of generic security claims.

What this does not cover — sanitising rich HTML from users, which is a different discipline

What this does not cover — sanitising rich HTML from users, which is a different discipline. Removing duplicate escaping is not rich-HTML sanitization. If users may author markup, an allowlist sanitizer and a constrained rendering surface solve a different problem.

Define what this does not before running double-escape diagnosis. Save cover sanitising rich html as a control, inspect the code points behind from users which is, and map a different discipline to the next interpreter. This makes double-escape diagnosis evidence auditable for a CMS editor seeing &amp; in page titles investigating html double escaping.

Takeaway: escape once, at the edge — how the HTML entity escaper lets you decode one pass at a time and see each intermediate result

Takeaway: escape once, at the edge — how the HTML entity escaper lets you decode one pass at a time and see each intermediate result. The tool’s one-pass behavior is useful evidence: observe &amp;amp; becoming &amp;, then decide whether another decode belongs to a known layer. Do not automate “decode until clean.”

Connect takeaway escape once at to an observable double-escape diagnosis output. Keep the edge how the beside the one-pass result, then verify where html entity escaper lets enters you decode one pass. A CMS editor seeing &amp; in page titles can now review at a time and as a narrow html double escaping finding. The practical decision behind this article is specific: &amp;amp; on a page means text was escaped twice, usually once by a template and once by hand. This post explains the layers involved, how to spot the pattern in the source and how many decode passes are safe. The reader action is equally concrete: Links to the HTML entity escaper and demonstrates decoding an &amp;amp; sample twice, showing the intermediate and final text.