Developer tools · HTML WYSIWYG editor
Sanitising Pasted HTML: What to Strip Before It Reaches Your CMS or Email
· How it works
html security text-cleanup
Explains what an HTML sanitiser does, from allowlisted tags and attributes to removed scripts and neutralised URLs, and how inspecting markup first makes sanitiser rules easier to write.
The pasted snippet that carried an onclick — opens with the hidden risk in rich-text input
A pasted anchor can hide an onclick beside an innocent destination. ToolAcre lowercases every attribute name and keeps only attributes explicitly listed for the accepted element, so onclick disappears even when its casing is altered. The removal report identifies that event handler instead of silently presenting unchanged source.
This boundary operates before rich paste is inserted, when source returns to visual mode, before copy, before plain-text extraction and again while building preview documents. Repetition reduces accidental bypass between actions, but the project still refuses to call the hand-written tokenizer a general-purpose XSS filter.
Why allowlists beat blocklists — explains that naming what is permitted is safer than trying to list every dangerous construct
An allowlist starts by naming permitted structure: paragraphs, headings, semantic inline elements, lists, description lists, quotations, code-like elements and anchors. An unknown ordinary wrapper loses its tag while retaining text. A dangerous container such as script, style, iframe, form, SVG or MathML loses its contents too.
A blocklist would need to anticipate every dangerous or unsupported construction. The allowlist rejects what it does not understand. That is a strong choice for one editor’s constrained output, yet it remains bounded by its tokenizer. Browser HTML5 error recovery can produce a different tree from a smaller parser when deliberately malformed input is involved.
Tags, attributes and URL schemes — covers the three layers a sanitiser filters, with typical decisions for each
Filtering occurs at three layers. Elements decide structural vocabulary. Per-element attributes permit only a few values such as link href and title, quotation cite, abbreviation title, and ordered-list start and type. URL inspection then decodes entities, strips controls and whitespace from a probe, and checks the resulting scheme.
Accepted schemes are http, https, mailto, tel and ftp, plus relative forms without an explicit scheme. JavaScript, data, file, blob, vbscript and about examples are refused in tests. Surviving links gain rel values, but that addition is not a substitute for destination policy or link review.
Styles: strip, allow or rewrite — discusses inline style handling and why many systems drop it entirely
Style attributes are removed wholesale. The module does not attempt to parse declarations, retain a safe subset or rewrite design tokens. That policy removes copied appearance and CSS-based request surfaces together. Class, id and data attributes also disappear, producing portable but deliberately less expressive markup.
Systems with a genuine style requirement need a different reviewed policy. Adding style to this allowlist without a CSS sanitizer would materially change its security surface. The current implementation avoids that problem rather than claiming to solve CSS safety for arbitrary hostile fragments.
Worked example: derive the rule from ToolAcre’s documented allowlist, not a Word-specific fixture
Start with `<div class="WordSection"><p style="color:red" onclick="x()">Notice <strong>today</strong></p></div>`. The div is unwrapped, class and style cannot survive, onclick is removed, and the paragraph plus strong element remain. The outcome follows generic policy without asserting which application produced the wrapper.
Add a javascript href and a script block. The anchor keeps its visible words but loses href; script and body disappear. Read the reported reasons. This exercise helps define a server policy, but copying ToolAcre’s exact subset blindly may omit elements your application requires or permit URLs your threat model forbids.
Server-side versus client-side sanitising — explains why the server must sanitise even if the browser already did
Client filtering improves local drafting but cannot be trusted by a server receiving user-controlled requests. Attackers can bypass the page, call an endpoint directly or exploit a parser difference. The server must parse and sanitize again with a maintained HTML5-aware implementation configured for its rendering context.
Output encoding remains separate too. HTML intended as text must be escaped by the template rather than inserted as markup. A fragment intentionally rendered as HTML needs sanitization before storage or output according to architecture. A sandboxed preview proves only that this preview grants no scripts, forms or same-origin access.
What this tool does cover — narrow editor output filtering, not general hostile-input sanitization
ToolAcre does filter its own output surface, contrary to the workbook’s claim that it is merely an inspection tool. The accurate correction is narrower: it is not a general-purpose XSS sanitizer for arbitrary hostile input. The source says this explicitly and documents a possible parser differential.
The iframe is defence in depth for rendering inside ToolAcre. Its empty sandbox attribute grants no script execution, form submission or same-origin access, and referrer policy is no-referrer. Once HTML is copied elsewhere, that frame no longer protects it. Publication safety belongs to the receiving system.
Takeaway: inspect locally, sanitise on the server — summarises the workflow and how the editor helps you see what a sanitiser will face
Inspect locally, sanitize on the server, and render according to context. Those are three distinct steps. ToolAcre helps reveal pasted baggage and offers a conservative draft subset, while removal notices make policy effects visible before a fragment reaches a CMS or email workflow.
Do not market a successful preview as proof against XSS. Use test payloads only in disposable content, preserve the raw source separately when investigation matters, and verify destination sanitization independently. Security claims should stop exactly where the code and rendering boundary stop.