Text & everyday tools · Text Toolkit
Why uppercase is not just A–Z: ß, the Turkish i and Unicode case mapping
· Background
unicode case-conversion localisation
Explains why upper- and lowercasing is defined by Unicode tables rather than arithmetic on letters, with the well-known exceptions — ß becoming SS, the Turkish dotless ı and Greek final sigma — and what that means for any browser-based converter.
STRASSE and the letter that grew — why uppercasing 'straße' changes the character count
Paste `straße` into a converter and uppercase it: JavaScript produces `STRASSE`. The result is visibly longer because the lower-case sharp s is mapped to two capital letters by the default operation used here. A validation rule that assumes case conversion preserves character count can therefore reject, truncate or misalign otherwise valid German text.
The reverse direction does not reconstruct the source. Lowercasing `STRASSE` yields `strasse`, not `straße`, because the capital sequence also represents an ordinary pair of s letters. Case conversion is consequently a text transformation, not a reversible encoding. Keep the original spelling whenever identity, display fidelity or exact comparison matters.
Case mapping as a table, not a formula — how Unicode defines simple and full case mappings for every cased letter
ASCII code often teaches a convenient pattern: upper- and lower-case Latin letters occupy predictable ranges, so a numeric offset appears to connect them. That model stops being useful once text contains letters outside those ranges or a mapping produces more than one output character. The converter therefore calls JavaScript string case methods instead of performing letter arithmetic.
The repository does not expose Unicode mapping tables or distinguish simple and full mappings, so those details should not be presented as features of this tool. Its observable contract is narrower: `toUpperCase()` and `toLowerCase()` delegate to the browser engine. Results should be tested in the environments where converted text will actually be consumed.
Case conversion uses engine-provided character mappings, not arithmetic or tables exposed by this tool
A length-changing result affects more than a counter beside an input. Stored offsets, selection ranges, highlights and cursor positions calculated before conversion may no longer point to the intended text afterward. The same risk applies to any workflow that allocates output space from input length or assumes one source character always corresponds to one result character.
Round-tripping can also lose distinctions even when a particular example keeps the same visible length. Several source spellings may converge on one upper-case form, leaving the lower-case operation without enough information to choose the original. Compare transformed values only when that loss is acceptable; otherwise retain originals and use a comparison strategy designed for the product requirement.
Length-changing mappings and why converting back cannot always restore the source
Greek sigma provides a different warning. Lowercase Greek uses sigma forms that can vary with position in a word, so changing case cannot always be decided from one isolated character. JavaScript receives the complete string, allowing its built-in lowercasing behavior to consider surrounding text rather than replacing every capital sigma with one fixed glyph.
Test sigma inside complete words, not only as a single-character sample. Punctuation, spaces and word boundaries are part of the input that the engine evaluates, and editing them can alter the lower-case result. For localisation review, preserve the whole phrase used by the interface and compare that phrase with an approved translation rather than inspecting code points alone.
Greek sigma shows that lowercasing can depend on surrounding text
Turkish distinguishes dotted and dotless i in ways that default, locale-independent case conversion does not model as Turkish-specific behavior. The ToolAcre config explicitly says its mapping is locale-independent and that Turkish dotted and dotless i are not special-cased. A plausible-looking result is therefore not evidence that names, addresses or interface strings were converted correctly for Turkish readers.
The practical response is not to add substitutions after a generic conversion. Such replacements can damage mixed-language text and still miss context. Use a locale-aware operation in software whose locale is known, and have Turkish or Azeri content reviewed in that setting. In this converter, treat the i examples as demonstrations of the documented limitation.
Turkish dotted and dotless i require locale-aware handling that this converter does not provide
ToolAcre implements UPPER and lower by calling the corresponding JavaScript string methods on the complete input. It does not pass a locale, load a language dictionary or request a server-side conversion. The output is therefore the browser engine’s default case result, which is useful for inspecting ordinary application behavior but not for certifying language-specific typography.
The distinction matters when reproducing a bug. Record the exact source string, the selected transform and the resulting string instead of reporting only that capitalization looked wrong. If production code uses the same default JavaScript methods, the converter offers a compact comparison. If production uses locale-aware APIs, matching this page is not the relevant acceptance test.
What this does not cover — title-casing rules for digraphs and the history of the capital ẞ
The Title Case button is another transformation with a deliberately limited contract. It finds runs of letters, combining marks or apostrophes, uppercases the first character and lowercases the rest. It has no language dictionary or style-guide exceptions, so acronyms and proper nouns can change, while small words such as `of` and `the` are capitalized like other words.
This article does not rely on the outline’s proposed history of a capital sharp s or on broader rules for title-casing digraphs, because neither is established by the inspected sources. Those subjects may be valuable with independent references, but they do not explain the shipped function. Here the reliable guidance is to review every converted heading and restore intentional spelling manually.
Title case behavior in this tool, without unsupported alphabet history
Try three focused checks in the case converter: uppercase `straße`, lowercase a Greek word containing capital sigma, and run dotted and dotless i samples through both directions. Observe the actual strings rather than predicting them from English. Then use Undo between experiments so each result begins from the original text instead of a previous lossy transformation.
The broader lesson is operational: case conversion can change length, collapse distinctions and depend on linguistic context that a default operation does not know. Use the tool to expose those behaviors, not to promise universal localisation correctness. Preserve source text, test complete real phrases, and apply locale-aware review wherever a language-specific result is required.