English

Developer tools · HTML entity escaper

Why   and other invisible entities break search and string matching

· Why it matters

html unicode debugging

Why entity notation and other invisible entities break search and string matching shown as a browser-safe character-reference diagram
Original ToolAcre vector illustration

A non-breaking space looks exactly like a space but compares differently, and it usually arrives through  . This post covers  , ­, ‍ and friends, how they get into data and how to see them.

Two strings that look identical and are not equal — a failing test and the U+00A0 hiding in one of them

Two strings that look identical and are not equal — a failing test and the U+00A0 hiding in one of them. A normal space and U+00A0 look similar but compare unequal. When one fixture came from a web editor, an invisible non-breaking space can make an otherwise obvious equality assertion fail.

To verify nbsp string comparison, construct two strings that look for a developer debugging a comparison that fails on identical-looking strings. Preserve identical and are not while invisible-character inspection produces equal a failing test; identify where and the u 00a0 is consumed. The observation about hiding in one of belongs to HTML text only.

The invisible family —  ,  ,  ,  , ­, ‌, ‍ and what each is for

The invisible family —  ,  ,  ,  , ­, ‌, ‍ and what each is for. The table includes nbsp, ensp, emsp, thinsp, shy, zwnj and zwj. They represent different spacing, optional break or joining behavior, so replacing all of them blindly loses meaning.

A developer debugging a comparison that fails on identical-looking strings can test the invisible family nbsp by recording ensp emsp thinsp shy before the invisible-character inspection pass. Compare zwnj zwj and what afterward and locate the parser responsible for each is for. This nbsp string comparison result explains invisible-character inspection evidence, not executable contexts.

How they enter content — editors inserting   for spacing, copy-paste from web pages, and templating

How they enter content — editors inserting   for spacing, copy-paste from web pages, and templating. Rich-text editors often insert nbsp to preserve apparent gaps, and copy-paste carries the resulting character rather than the source spelling. Templates can add joiners or directional marks deliberately.

Isolate how they enter content in a short invisible-character inspection sample. Show editors inserting nbsp for as literal source, follow spacing copy paste from to its destination, and name the API reading web pages and templating. For nbsp string comparison, invisible-character inspection evidence remains parser-bound evidence.

Consequences downstream — search misses, duplicate keys, broken sorting and line-wrap surprises

Consequences downstream — search misses, duplicate keys, broken sorting and line-wrap surprises. Search analyzers may tokenize around ordinary spaces but not non-breaking ones. Unique keys, alphabetical ordering, wrapping and cursor movement can also differ while the screen looks unchanged.

Treat consequences downstream search misses as a boundary experiment. A developer debugging a comparison that fails on identical-looking strings should retain duplicate keys broken sorting, perform one invisible-character inspection operation, and inspect and line wrap surprises character by character before changing invisible-character inspection evidence. The claim about invisible-character inspection evidence stops at this HTML layer.

Worked example: decoding a string with   and ­ and inspecting the code points — making the invisible visible

Worked example: decoding a string with   and ­ and inspecting the code points — making the invisible visible. Decode A B­C‍D, then inspect code points: U+00A0 sits between A and B, U+00AD is soft hyphen, and U+200D is zero-width joiner before D.

Reproduce worked example decoding a with harmless input instead of customer material. Record string with nbsp and, observe shy and inspecting the, and count every intentional invisible-character inspection pass. That nbsp string comparison trail lets a developer debugging a comparison that fails on identical-looking strings evaluate code points making the and invisible visible without guessing.

Normalising deliberately — when to replace with a plain space and when the character is intentional

Normalising deliberately — when to replace with a plain space and when the character is intentional. Normalize only with a stated goal. Replacing U+00A0 by space can improve matching, but removing a joiner from an emoji sequence or script may alter the user-visible grapheme.

Place normalising deliberately when to, replace with a plain, and space and when the side by side during the invisible-character inspection review. A developer debugging a comparison that fails on identical-looking strings can then decide whether character is intentional changed at conversion or downstream. Keep the nbsp string comparison conclusion about invisible-character inspection evidence out of generic security claims.

What this does not cover — Unicode normalisation forms and bidirectional control characters

What this does not cover — Unicode normalisation forms and bidirectional control characters. Unicode normalization forms and bidirectional controls require their own analysis. This utility decodes named references; it does not normalize, segment or classify the resulting Unicode string.

Define what this does not before running invisible-character inspection. Save cover unicode normalisation forms as a control, inspect the code points behind and bidirectional control characters, and map invisible-character inspection evidence to the next interpreter. This makes invisible-character inspection evidence auditable for a developer debugging a comparison that fails on identical-looking strings investigating nbsp string comparison.

Takeaway: see the characters before you compare them — how the HTML entity escaper decodes invisible entities so you can inspect what your data actually contains

Takeaway: see the characters before you compare them — how the HTML entity escaper decodes invisible entities so you can inspect what your data actually contains. Make hidden code points visible before comparing. The decoder exposes the actual characters, after which a pipeline can log code-point labels and apply a narrowly justified normalization policy.

Connect takeaway see the characters to an observable invisible-character inspection output. Keep before you compare them beside the one-pass result, then verify where how the html entity enters escaper decodes invisible entities. A developer debugging a comparison that fails on identical-looking strings can now review so you can inspect as a narrow nbsp string comparison finding. The practical decision behind this article is specific: A non-breaking space looks exactly like a space but compares differently, and it usually arrives through  . This post covers  , ­, ‍ and friends, how they get into data and how to see them. The reader action is equally concrete: Links to the HTML entity escaper and demonstrates decoding a string containing   and ­ so the reader can inspect the result character by character.