English

Text & everyday tools · Text Toolkit

How whole-word search works: word boundaries and the cat|dog problem

· How it works

regex find-and-replace text-editing

The words cat and dog passing through boundaries while concatenate remains unchanged
Original ToolAcre vector illustration

Explains what a word boundary is in JavaScript regular expressions, why 'cat' must not match inside 'concatenate', and why an alternation like cat|dog has to be grouped before boundaries are added.

The rename that changed 'concatenate' — why a plain find-and-replace of a short word damages longer words

A product rename looks harmless until a short name appears inside unrelated words. Replacing every occurrence of `cat` with `lynx` changes `concatenate` into `conlynxenate`, even though no editor intended to rename part of that verb. A match count alone cannot protect the document when the search rule is too broad.

Whole-word matching narrows the rule before any replacement happens. In ToolAcre, enabling Whole word makes standalone `cat` eligible while leaving the same letters inside `concatenate` untouched. The distinction is structural rather than dictionary-based: the regular-expression engine checks the characters immediately beside each possible match.

What \b means — the zero-width position between a word character and a non-word character

In a JavaScript regular expression, `` marks a zero-width word boundary. It consumes no letter, space or punctuation. Instead, it succeeds at a position where one side is a word character and the other side is not, including the start or end of the text when the adjacent character qualifies as a word character.

JavaScript’s `w` category supplies the relevant word characters for this implementation: ASCII letters, digits and underscore. Therefore `cat` matches `cat` beside spaces, commas or line ends, but not the `cat` segment inside `concatenate`, because both sides of that segment continue through word characters and no boundary exists there.

Whole word in literal mode — the search text is escaped first, then wrapped in boundaries

With Regex switched off, ToolAcre first escapes every regular-expression metacharacter in the search text. A period remains a literal period, parentheses remain literal parentheses and a plus sign remains a plus sign. Only after that escaping step does Whole word wrap the resulting source, so user punctuation cannot silently become regex syntax.

The wrapper is `(?:… )` without the displayed space: the inner `(?:…)` is a non-capturing group. For a simple literal such as `cat`, the effect is equivalent to `cat`. The grouping still matters because one compilation path must safely handle simple terms and more complicated alternatives without changing capture-group numbering.

Whole word in regex mode — why cat|dog must be bounded as a group, or the boundary binds to one branch only

Regex mode leaves the entered pattern intact, but Whole word still groups it before adding boundaries. For `cat|dog`, ToolAcre compiles the conceptual form `(?:cat|dog)`. Both alternatives must therefore begin and end at word boundaries, so standalone `cat` and standalone `dog` match under the same rule.

Without that group, `cat|dog` would split into two branches: a left boundary would constrain only `cat`, while a right boundary would constrain only `dog`. Alternation has lower precedence than the surrounding sequence. Grouping is what applies both boundary assertions to the whole choice instead of distributing one assertion unevenly.

Worked example — renaming a product across a document with Whole word on, and inspecting the replacement count against expectations

Paste a representative passage containing `cat`, `dog`, `concatenate`, `dogged` and punctuation such as `cat, dog.` Enter `cat|dog`, turn on Regex and Whole word, and run Replace all with the new product name. The two standalone animal words should change; the longer words should remain exactly as they were.

Read the reported replacement count before accepting the edit. If the document has three known standalone references but the tool reports two or twelve, select Undo and inspect the sample contexts. ToolAcre counts matches before calling the standard string replacement, so the displayed total is a practical check on the search rule you actually compiled.

Where word boundaries surprise you — JavaScript's word characters are ASCII letters, digits and underscore, so accented and non-Latin words need testing

These boundaries do not implement a multilingual dictionary or Unicode text-segmentation algorithm. Because accented letters and non-Latin scripts fall outside the ASCII-style word category used here, Whole word can return zero or produce surprising edge positions for terms such as `café`. Punctuation inside a search term can cause the same mismatch.

Test the exact language and spelling before a bulk edit, especially for names containing accents, apostrophes or hyphens. A boundary assertion only compares adjacent character categories; it does not know whether typography forms one human word. If the sample fails, use a deliberately designed regex for those surroundings rather than trusting the checkbox.

Where JavaScript word boundaries surprise you: ASCII-style word characters are not linguistic words

The tool does provide a separate Case sensitive option, contrary to any implication that case handling is absent from the interface. Clearing it adds JavaScript’s case-insensitive flag. However, the implementation uses only global matching plus optional case insensitivity; it does not add Unicode, multiline, dot-all, sticky or other flags.

This article also does not turn `` into a Unicode-aware boundary or solve language-specific tokenization. JavaScript property escapes are not a substitute here because the tool does not compile patterns with the Unicode flag. For multilingual editorial work, establish explicit test cases and choose a pattern supported by the browser engine in use.

Case options exist in the tool; Unicode-aware word boundaries do not

Whole word is a composition rule, not a second replacement engine. Literal mode escapes the search; regex mode preserves it; then the same non-capturing group and two boundaries surround either result. That order protects literal punctuation and makes regex alternatives such as `cat|dog` behave as one bounded expression.

Use the option when standalone ASCII-style terms are the intended targets, then treat the replacement count and Undo control as safeguards rather than decoration. Open ToolAcre’s Find and replace, test the alternation against both standalone and embedded examples, and only apply it to the full document after those examples behave as expected.