English

Developer tools · HTML WYSIWYG editor

Rich Text on the Clipboard: How Formatting Travels Between Applications

· Background

html clipboard contenteditable

HTML and plain-text clipboard representations taking separate paths into an editor
Original ToolAcre vector illustration

Explains the multiple formats a rich-text copy places on the clipboard, how each application chooses one, and why pasted markup is never quite what was copied.

Same copy, three different pastes — opens with content pasted into a browser, a terminal and a mail client

One copied paragraph can arrive as formatted content in a browser editor and unformatted characters in a terminal because destinations request different representations. ToolAcre’s paste handler asks the clipboard for HTML and plain text, then follows a fixed branch rather than trusting the browser’s default insertion.

If HTML is present, the original rich fragment never enters the surface directly. It passes through the sanitizer first. If HTML is absent, plain text is inserted. That local rule explains this editor’s behavior without claiming how every mail client, operating system or office suite prioritizes clipboard data.

This paste handler reads text/html and text/plain; broader clipboard flavours are outside the evidence

A clipboard can contain multiple representations of one selection. This implementation deals explicitly with `text/html` and `text/plain`. It does not enumerate RTF, images, file references or application-specific formats, so those should not be implied as supported inputs.

The two strings can encode different information. Plain text carries characters and line breaks, while HTML can carry paragraphs, lists, emphasis and links plus unwanted attributes. ToolAcre favors structure when available but narrows it through its element, attribute and URL policy before insertion.

Who writes the HTML flavour — explains that the source application generates it, which is why Word and Google Docs pastes look so different

The source application creates the HTML representation, which is why two visually similar copies can supply different wrappers and presentation. ToolAcre does not identify the producer. It applies one generic rule: unsupported wrappers and attributes are removed regardless of origin.

This prevents unsourced catalogs of Word or Google Docs internals from becoming product claims. The removal list tells the writer exactly what this parser rejected. If fidelity to a source application matters, capture a real fixture from the supported environment and test it rather than relying on a generic article example.

How the destination chooses — covers preference order and paste-as-plain-text shortcuts

The destination chooses by code. Here, non-empty HTML wins; otherwise plain text is used. There is no toolbar switch for “paste as plain text,” and repository sources do not define keyboard shortcut behavior. Browsers and operating systems may offer such gestures independently.

Drop handling mirrors the same idea with DataTransfer: prevent default, read HTML and plain values, sanitize HTML before `insertHTML`, or send plain characters through `insertText`. Unlike paste, the drop path does not display the same detailed removal count, though refresh still updates the visible document.

Platform fragment metadata is not parsed explicitly by this implementation

Some platforms wrap HTML clipboard fragments in metadata or markers, but ToolAcre contains no explicit parser for a named platform format. Its tokenizer sees the string returned by `getData("text/html")` and processes recognized markup. This article therefore omits header syntax, offsets and operating-system-specific guarantees.

If wrappers appear as comments, comments are removed. If metadata arrives outside recognized HTML, behavior should be tested. A qualitative possibility is not enough to claim support or rejection for a platform protocol the code never names.

Worked example: compare a rich clipboard payload with a plain-only payload

Copy a non-sensitive paragraph containing strong text, a short list and an https link. A rich paste can retain those allowed structures, remove source styling and add rel to the link. Inspect source and the removal warning. Then paste a plain-only representation through an environment that supplies no HTML.

The plain result contains characters inserted at the selection; any block structure is decided by the browser editing host rather than restored from the original HTML. Compare words and line boundaries separately from elements. The exercise demonstrates translation between representations, not byte-for-byte copying.

What this does not cover — clipboard APIs, images, platform shortcuts or permission matrices

This article does not cover navigator.clipboard programming, permission prompts, clipboard images, files, RTF, mobile shortcuts or a cross-platform preference matrix. It also does not claim that intercepted rich paste makes arbitrary hostile HTML safe to publish.

ToolAcre’s sanitizer is narrow and acknowledges parser differentials. Its sandbox blocks execution in the local preview, but copied HTML leaves that environment. A receiving server needs an appropriate parser-based sanitizer and output policy, while users should avoid pasting secrets into any unreviewed page.

Takeaway: the paste is a translation, not a copy — summarises the model and how ToolAcre's HTML WYSIWYG editor lets you inspect what the translation produced

A rich paste is a translation: the source writes representations, the destination chooses one, and ToolAcre filters the chosen HTML before insertion. Understanding that sequence explains why formatting can disappear without treating the clipboard as broken.

Use representative fixtures, inspect the resulting source and preserve original content when cleanup might remove meaning. Plain paste is useful for resetting presentation, while rich paste can retain structure. Neither path relieves the destination from sanitation, compatibility testing or editorial review.