English

Developer tools · HTML entity escaper

Why HTML escaping is not enough inside script, onclick and href attributes

· Why it matters

html security url-encoding

Why HTML escaping is not enough inside script, onclick and href attributes shown as a browser-safe character-reference diagram
Original ToolAcre vector illustration

Entity escaping protects HTML text and attribute values, but a value inside a script block, an event handler or a URL is read by a different parser. This post explains each context and the encoding it needs.

The escaped value that still executed — a user string inside an onclick handler and why " did not save it

The escaped value that still executed — a user string inside an onclick handler and why " did not save it. An onclick attribute is parsed as HTML and then compiled as JavaScript. Encoding only the HTML delimiter can still leave a dangerous program after the browser decodes that attribute value.

To verify html escaping javascript context, construct the escaped value that for a developer that escaped everything and still got an XSS report. Preserve still executed a user while nested-parser context produces string inside an onclick; identify where handler and why quot is consumed. The observation about did not save it belongs to HTML text only.

One document, several languages — HTML, JavaScript, URLs and CSS each have their own parser and escape rules

One document, several languages — HTML, JavaScript, URLs and CSS each have their own parser and escape rules. A single document can invoke HTML, JavaScript, URL and CSS parsers. Each language assigns punctuation different meaning, so one generic “escaped” flag cannot describe every destination safely.

A developer that escaped everything and still got an XSS report can test one document several languages by recording html javascript urls and before the nested-parser context pass. Compare css each have their afterward and locate the parser responsible for own parser and escape. This html escaping javascript context result explains rules, not executable contexts.

Inside <script> — why entities are not decoded there and JSON or JavaScript string escaping is required instead

Inside <script> — why entities are not decoded there and JSON or JavaScript string escaping is required instead. Inside script raw text, HTML named references are not decoded like normal text. Serialize data as JSON with a proven serializer, avoid inline script interpolation, and account for closing-script sequences.

Isolate inside script why entities in a short nested-parser context sample. Show are not decoded there as literal source, follow and json or javascript to its destination, and name the API reading string escaping is required. For html escaping javascript context, instead remains parser-bound evidence.

Event handler attributes — decoded as HTML first and then run as JavaScript, so two layers of escaping apply

Event handler attributes — decoded as HTML first and then run as JavaScript, so two layers of escaping apply. Event-handler attributes combine two grammars and should be avoided. Binding behavior with addEventListener keeps data out of executable source and removes the need for nested HTML-plus-JavaScript escaping.

Treat event handler attributes decoded as a boundary experiment. A developer that escaped everything and still got an XSS report should retain as html first and, perform one nested-parser context operation, and inspect then run as javascript character by character before changing so two layers of. The claim about escaping apply stops at this HTML layer.

href and src — percent-encoding for the value plus entity escaping for the attribute, and the javascript: scheme problem

href and src — percent-encoding for the value plus entity escaping for the attribute, and the javascript: scheme problem. For href and src, first build and validate the URL, reject unwanted schemes, percent-encode components, then HTML-escape the final quoted attribute value. Entity conversion alone does not reject javascript: URLs.

Reproduce href and src percent with harmless input instead of customer material. Record encoding for the value, observe plus entity escaping for, and count every intentional nested-parser context pass. That html escaping javascript context trail lets a developer that escaped everything and still got an XSS report evaluate the attribute and the and javascript scheme problem without guessing.

Worked example: one value placed in text, in an attribute, in onclick and in href — the correct encoding for each

Worked example: one value placed in text, in an attribute, in onclick and in href — the correct encoding for each. The same value belongs in text through HTML escaping, in a data attribute through quoted-attribute handling, in code through structured serialization, and in a URL through URL construction plus scheme checks.

Place worked example one value, placed in text in, and an attribute in onclick side by side during the nested-parser context review. A developer that escaped everything and still got an XSS report can then decide whether and in href the changed at conversion or downstream. Keep the html escaping javascript context conclusion about correct encoding for each out of generic security claims.

What this does not cover — CSS injection and template-engine-specific helpers

What this does not cover — CSS injection and template-engine-specific helpers. CSS injection and template-engine-specific APIs are omitted because they introduce additional grammars and framework contracts. The safe choice is to use APIs that keep data separate from code.

Define what this does not before running nested-parser context. Save cover css injection and as a control, inspect the code points behind template engine specific helpers, and map nested-parser context evidence to the next interpreter. This makes nested-parser context evidence auditable for a developer that escaped everything and still got an XSS report investigating html escaping javascript context.

Takeaway: encode for the parser that will read it — how the HTML entity escaper handles the HTML layer, and the URL encoder & decoder in the same product handles the URL layer

Takeaway: encode for the parser that will read it — how the HTML entity escaper handles the HTML layer, and the URL encoder & decoder in the same product handles the URL layer. ToolAcre covers the HTML layer, and its URL panel covers representation of URL components. Neither panel grants authorization, sanitizes markup, validates schemes or protects SQL statements.

Connect takeaway encode for the to an observable nested-parser context output. Keep parser that will read beside the one-pass result, then verify where it how the html enters entity escaper handles the. A developer that escaped everything and still got an XSS report can now review html layer and the as a narrow html escaping javascript context finding. The practical decision behind this article is specific: Entity escaping protects HTML text and attribute values, but a value inside a script block, an event handler or a URL is read by a different parser. This post explains each context and the encoding it needs. The reader action is equally concrete: Links to the HTML entity escaper for the attribute-escaping step and demonstrates switching to the URL encoder & decoder panel for the href value without a page load.