Developer tools · HTML WYSIWYG editor
From DOM to Markup: How innerHTML Serialisation Shapes the HTML You Get
· How it works
html html-entities developer-workflow
Explains the rules browsers follow when turning a DOM tree back into an HTML string, from escaping and entities to void elements and attribute quoting, and why editor output looks the way it does.
Why <br> comes back as <br /> in this sanitizer
A source fragment containing `<br>` leaves this sanitizer as `<br />`. That result is not proof that every browser serializes HTML with a slash. It is a deliberate choice in ToolAcre’s string rewriter, whose void-element branch emits a space and slash for allowed br and hr elements.
The distinction prevents a common false explanation. The visual surface is read through innerHTML, but the copied source is then tokenized and rebuilt by application code. What users receive reflects both the browser-created DOM string and ToolAcre’s narrower output conventions, including aliases, filtered attributes and balanced tags.
Tokenising in, allowlist rewriting out — this is not browser innerHTML serialization
The tokenizer scans characters without constructing DOM nodes or calling DOMParser. It recognizes text, comments, declarations, start tags, end tags and raw-text element bodies. Sanitization iterates that flat sequence, keeping allowed structures, unwrapping ordinary unknown wrappers and suppressing dangerous containers with their contents.
Consequently, original spelling is not preserved. Mixed-case names become lowercase; b becomes strong, i becomes em, strike and del become s, and ins becomes u. The output is regenerated from accepted tokens rather than returned as the original substring, so source formatting and unsupported details disappear by design.
Escaping text and attributes while normalising known entities
Text is entity-decoded first and escaped again. Ampersands become `&`, less-than and greater-than signs become `<` and `>`, and prohibited control characters vanish. This order keeps an existing `&` stable across repeated passes instead of turning it into `&amp;`.
Attribute values follow a stricter escape path that also quotes double quotes, single quotes, angle brackets and ampersands. Known entities are decoded before accepted attributes are emitted. The tests pin idempotence for text and link queries, but the code does not promise preservation of each authorial entity spelling such as a named versus numeric reference.
Void elements and balanced closing tags in the custom rewriter
Only br and hr are both allowed and classified as void in the output set. They never receive closing tags. For non-void allowed elements, an open stack records nesting. A crossed close such as strong around em causes inner elements to close first, producing balanced markup rather than preserving malformed order.
Unclosed elements are closed at the end, while stray closing tags are ignored. This is the rewriter’s balancing policy, not a complete implementation of browser HTML5 error recovery. The source itself warns about parser differentials, which is why the output must not be marketed as a universal hostile-input sanitizer.
Attributes are allowlisted and quoted; source order alone is not the contract
No generic attribute bag survives. Paragraphs and headings permit none; links allow href and title; abbreviations allow title; quotations allow cite; ordered lists allow start and type. Every emitted value is double-quoted, and accepted links also gain `rel="noopener noreferrer nofollow"`.
Attribute order follows accepted tokens and the added rel field, but relying on order would be fragile. Meaning resides in the allowed names and values. A malformed integer start is dropped, unsafe href or cite schemes are refused, and unknown attributes generate removal reasons instead of silently becoming part of published markup.
Worked example: round-tripping a snippet through ToolAcre’s sanitizer
Try `<B class="loud">A & B<br></B><a href="javascript:alert(1)">open</a>`. The class is removed, B maps to strong, the existing ampersand remains one escaped ampersand, br becomes self-closed, and the unsafe link loses its href while retaining visible link text.
Pretty source mode then places supported block elements on lines and indents list children. It leaves inline content together and preserves text inside pre. Formatting is a readability pass over already rewritten output; it is not a second security parser and does not restore whitespace that earlier normalization removed.
What this does not cover — browser, XML or XHTML serialization
This article does not describe native browser serialization guarantees, XMLSerializer, XHTML syntax or optional HTML closing-tag rules in general. It documents the observable conventions in this repository. Claims about ` ` specifically are omitted because the local entity table, not a general browser serializer, determines recognized named references.
The preview builder sanitizes once more and wraps the fragment in a complete light-themed document. Copy HTML returns the fragment; Download HTML returns that document. Those are intentionally different artifacts. Read the action label before assuming a downloaded file is byte-identical to copied source.
Takeaway: the serialiser is consistent, your input was not — summarises how to predict editor output and how ToolAcre's HTML WYSIWYG editor shows what the browser actually produces
Predict ToolAcre by following its stages: tokenize the string, decode supported entities, apply element aliases and allowlists, inspect URLs, balance accepted tags, escape output and optionally pretty-print it. That sequence explains the actual result more accurately than saying innerHTML simply preserves or serializes the author’s source.
Use source mode to observe the rewrite with a small fixture before processing a long draft. The sanitizer refuses documents above 500,000 characters, so oversized input fails rather than freezing the tab. Keep the destination’s own HTML policy separate from this editor’s deliberately compact subset.