English

Text & everyday tools · Text Toolkit

Invisible characters: non-breaking spaces, zero-width joiners and trim

· Background

text-cleanup unicode publishing

Visible text lines revealing hidden spaces and zero-width Unicode characters
Original ToolAcre vector illustration

Surveys the characters that look like nothing — non-breaking spaces, zero-width spaces and joiners, byte order marks — where they come from, which count as whitespace, and how to find and remove them.

Two identical lines that are not duplicates — how an invisible character defeats de-duplication and search

Two lines can look identical in an editor yet fail equality, search or de-duplication checks. One may contain an ordinary U+0020 space while the other contains U+00A0, a non-breaking space. A zero-width character can create the same problem without occupying any visible width, leaving a publisher with two apparently matching labels that remain distinct strings.

This is not merely cosmetic. Hidden code points can split analytics dimensions, defeat an exact find operation, preserve a duplicate line or make a copied identifier fail validation. Before rewriting content by eye, keep a copy of the source and compare the suspicious strings systematically. The visible rendering is evidence about appearance, not proof that their underlying Unicode sequences match.

Where they come from — web pages, word processors, emoji sequences and copy-paste from PDFs

Invisible characters commonly arrive through ordinary work. Web pages use non-breaking spaces to keep terms together, word processors preserve layout-oriented spacing, and PDF extraction reconstructs text from positioned glyphs. Copying between those systems can carry formatting characters into a CMS field even when the destination displays only familiar words and gaps.

Other invisible marks are intentional parts of writing systems or emoji. A zero-width joiner can combine emoji components into one displayed symbol, while joiners and non-joiners affect shaping in several scripts. That origin matters: “cannot see it” does not mean “safe to delete.” Cleanup should target a diagnosed artifact, not every character whose advance width happens to be zero.

The whitespace family — ordinary, non-breaking, narrow and ideographic spaces, and what JavaScript's trim considers whitespace

Unicode contains more than the everyday space. U+00A0 is the non-breaking space, U+202F is the narrow non-breaking space, and U+3000 is the ideographic space. JavaScript whitespace handling includes these characters, so `trim()`, `trimStart()` and `trimEnd()` can remove them when they occur at the relevant edge of a string.

The Text Toolkit’s line trimming calls those JavaScript methods on each line. It does not normalize interior spacing, so a non-breaking space between two words remains there. Its “characters without spaces” statistic removes characters matched by JavaScript `s`, which is broader than ASCII space. Use that number as a defined measurement, not as a promise that every invisible code point was excluded.

The zero-width family — zero-width space, joiner and non-joiner, word joiner and the byte order mark, which are not whitespace at all

The zero-width family follows different rules. U+200B zero-width space, U+200C zero-width non-joiner, U+200D zero-width joiner and U+2060 word joiner are format characters rather than JavaScript whitespace. Ordinary `trim()` therefore does not remove them. A line containing one of these marks is not blank merely because the screen shows no ink.

U+FEFF has two related histories: at the beginning of encoded data it can signal a byte order mark, while in text it has served as a zero-width no-break space. ECMAScript includes U+FEFF in its trim whitespace set, unlike U+200B through U+200D. Treating the whole zero-width family as one kind of whitespace would therefore contradict actual JavaScript behavior.

Detecting them — using the character counter to spot counts that do not match what you see, remembering that some joiners attach to a neighbour and stay hidden

A surprising count can alert you to hidden content, but it cannot identify every case. ToolAcre counts grapheme clusters with `Intl.Segmenter` when available. A joiner that participates in an emoji sequence may help several code points form one grapheme, so adding or removing it can change rendering without producing the simple one-character increase a code-point counter would show.

Use several clues together: failed exact matches, an unexpected “without spaces” result, copied text inspected in a Unicode-aware editor, or a regex search for specific code points. The fallback counter uses code points when Segmenter is unavailable, so results can also differ by browser capability. Counting narrows the investigation; it is not a hidden-character scanner or Unicode name viewer.

Detecting hidden characters: counts provide clues, not a complete inventory

For a confirmed copy-and-paste artifact, work on a duplicate and open Find and replace in regex mode. Searching for `[​]` and replacing with nothing removes zero-width spaces and embedded U+FEFF characters throughout the text. The tool compiles the pattern with JavaScript’s global flag, reports the replacement count, and leaves the text unchanged if the pattern is invalid.

Do not automatically broaden that pattern to `[​-‍]`. The range also removes U+200C and U+200D, which can change script shaping and split a joined emoji into separate symbols. Add those code points only when inspection proves they are unwanted. After replacement, run de-duplication and compare the output with the preserved original before publishing.

Worked example: remove only confirmed unwanted zero-width characters before de-duplicating

This cleanup does not address bidirectional control characters, homoglyphs or every security issue associated with confusing text. Directional marks can alter visual order, while homoglyph attacks use different visible characters that resemble one another. Neither problem is solved by deleting spaces, and an indiscriminate invisible-character regex can damage legitimate multilingual content while missing the actual risk.

The toolkit also does not expose regex flags such as Unicode or multiline mode, highlight each upcoming match, or step through replacements individually. Replace all operates across the full text, so the reported count and Undo action are important safeguards. For source code, legal copy or unfamiliar scripts, inspect each character with a dedicated Unicode tool before making bulk edits.

The takeaway — invisible does not mean harmless; the Text Toolkit's counter and regex find-and-replace expose and remove them without leaving your browser

Invisible does not mean empty, interchangeable or harmful. Non-breaking spaces are whitespace with layout semantics; zero-width joiners can carry shaping semantics; U+FEFF behaves differently again under JavaScript trimming. Accurate cleanup begins by naming the code point, understanding why it may be present and deciding whether the destination still needs its function.

Use the Text Toolkit as a controlled browser-side workbench: counts can expose a discrepancy, regex find-and-replace can remove a precisely defined set, and de-duplication can confirm that hidden differences are gone. Preserve an original, avoid deleting joiners by default, and verify the rendered result. A narrow, reviewable correction is safer than treating all unseen Unicode as debris.