Developer tools · Text comparison
Case Folding Is Harder Than It Looks: What 'Ignore Case' Means in Unicode
· Background
text-diff unicode case-sensitivity
Explains why ignoring case is trivial for ASCII and subtle for Unicode, with dotless i, sharp s and final sigma as examples, and what that means for any case-insensitive comparison.
Why two spellings of Istanbul may not match — opens with a real-world mismatch involving the Turkish dotted capital I that surprises English-speaking developers
Case-insensitive comparison sounds universal until identifiers, scripts and language rules disagree. ToolAcre’s actual rule is concrete: when Ignore case is selected, each line key receives JavaScript `toLowerCase()`, after any whitespace transformation. Strict equality then compares the resulting strings.
That mechanism is useful for headings or ordinary labels, but its name should not be expanded into locale-aware linguistic equality. The UI collects no locale and the function supplies none. Readers should inspect results whenever case carries identity or language meaning.
ASCII case: a single bit — explains the simple uppercase and lowercase relationship that made early tools' ignore-case flags easy
The workbook describes ASCII case as a single-bit relationship, but the source does not implement bit arithmetic or an ASCII-only branch. It invokes the platform string method on the complete line. Historical encoding explanations require sources beyond this repository.
For basic Latin text, `Hello` and `hello` become equal under the option, as the test suite proves. That test establishes one behavior, not every Unicode mapping. Avoid turning a passing English example into a promise about all scripts.
The implementation calls JavaScript toLowerCase; it does not expose an ASCII-specific bit operation
Unicode case folding is a comparison concept with its own data and status rules. ToolAcre does not call a named folding library, normalize strings or document a Unicode version. Calling this implementation “full case folding” would claim machinery that is absent.
The defensible description is lowercase-key comparison using the current JavaScript engine. Original rows remain unchanged for display. If the lowered strings differ in length or code points, the line remains a removal and addition rather than a forced match.
ToolAcre performs lowercase mapping, not a documented Unicode case-fold operation
The outline names Turkish I, German sharp s and Greek sigma. Those are legitimate topics for an externally sourced Unicode article, but this repository contains no table or tests guaranteeing outcomes for them. This article deliberately avoids declaring exact mappings.
Test the actual values in the target environment and consult authoritative Unicode and locale documentation before building identity rules. A convenience checkbox for visual review should not become a user-name comparator, security boundary or multilingual collation engine.
Named language exceptions require external Unicode sources and are not guaranteed here
Compare `Release Notes` with `release notes`. Strict mode reports a replacement; Ignore case reports the line identical and displays the original left text. Add a punctuation or accent change and observe that lowercasing alone does not erase every difference.
This controlled example demonstrates the option without pretending to settle script-specific behavior. When reviewing a glossary, run strict mode first and record case-only rows. Toggle the option afterward to separate capitalization noise from spelling or punctuation changes.
Worked example: test ordinary Latin variants without generalizing to every script
No locale selector appears in the UI, and `toLowerCase` receives no language parameter. Consequently the same transformation code runs regardless of the document’s intended language. The route does not know whether a line is Turkish, German, Greek or a programming identifier.
That absence is the practical warning. For locale-aware search, sorting or identity, use an API designed around those requirements and pin its behavior with appropriate data. Text diff is a reviewer’s lens over two strings, not a global naming policy.
No locale parameter is supplied by this comparison option
This article does not promise non-Latin equivalence, normalization, collation or security against confusable characters. Lowercasing cannot answer whether two names refer to the same person, product or filesystem object. Those systems may apply entirely different case rules.
Ignore whitespace can be combined with case, but doing so broadens what disappears. Toggle one option at a time to identify the cause. A result identical under both says only that transformed line keys match; it says nothing about byte equality or semantics.
Takeaway: ignore case for convenience, not for correctness — summarises when the option in ToolAcre's Text comparison is appropriate and when to inspect the results
Use Ignore case for convenience when capitalization is known to be presentation noise. Keep strict mode for code, paths, product names and any domain where case may distinguish values. Review both outputs instead of making the looser result the only record.
ToolAcre’s implementation is transparent enough to state precisely: optional whitespace handling, then JavaScript lowercase mapping, then strict line-key equality. That modest claim is more useful than an expansive Unicode promise the source cannot support.