Text & everyday tools · Text Toolkit
Why sorted lists differ between computers: locale-aware sorting explained
· Why it matters
text-tools sorting locales
Explains the difference between code-point order and locale collation, why ä, é and uppercase letters land in different places on different machines, and how to share a sorted list without surprises.
The same list, two orders — a colleague's sort puts Ångström somewhere else, and neither of you is wrong
A team lead can send the same list of attendee names to colleagues in two countries and receive two plausible alphabetical orders. Names containing accented letters are the usual point of disagreement, because alphabetical placement is not determined by the visible base letter alone. The browser environment participates in the comparison.
Neither colleague needs to have made a manual mistake. The practical problem is that “sort alphabetically” does not identify one universal sequence for multilingual text. Before reviewing additions or removals, agree which completed order is authoritative; otherwise a line-by-line comparison will show movement caused only by separate sorting decisions.
Code-point order — why a naive sort puts every capital before every lowercase letter and pushes accented letters to the end
The outline describes a naive code-point sort that places capitals before lowercase letters and accented letters at the end, but that is not what the Text Toolkit implements. Its `sortLines` function copies the lines and compares each pair with `localeCompare`, so browser collation rules are consulted rather than raw character numbers.
Case handling is also explicit in the function options. Case-sensitive sorting requests `variant` sensitivity, while the default case-insensitive mode requests `base` sensitivity. That can make differently cased or accented forms compare as equivalent, leaving their relative placement dependent on the stable order supplied to the sorter rather than forcing capitals into a separate block.
ToolAcre does not use naive code-point order: it calls localeCompare for every line
The comparison passes `undefined` as the locale, which asks the runtime to use its default locale behavior. The source does not promise Swedish, German, English, or any other named collation, and the interface represented by this function has no locale argument. Publishing exact national orders here would therefore go beyond the implementation evidence.
This boundary explains why results can differ without proving exactly which setting caused a particular difference. Record the browser and environment when reproducibility matters, but do not infer a locale from one accented name’s position. The reliable fact is that the toolkit delegates ordering to `localeCompare` with the runtime default rather than pinning a shared language.
The browser supplies the default locale, but the toolkit offers no locale selector
Digit runs receive special treatment too. The function defaults `numeric` to true and forwards that option to `localeCompare`. Consequently, a list containing `item2` and `item10` is intended to place the smaller numeric run first in ascending order instead of comparing the character `1` in ten against the character `2` in two.
Zero-padding remains useful when data must pass through another sorter whose behavior is unknown, but it is not required to obtain numeric-aware ordering here. If an exact lexical order is desired instead, the underlying function supports turning numeric comparison off. The important lesson is to document the chosen option rather than assume every alphabetical command treats embedded digits alike.
Digit runs sort numerically by default, so item2 comes before item10
For a reviewable example, place `Ångström`, `Ana`, `Émile`, `item2`, and `item10` on separate lines, then sort once in the browser chosen by the team. Save that output as the reference. This article deliberately does not publish a second browser’s expected accented-name order because the repository contains no pinned locale fixture establishing it.
If a colleague has another order, compare the two completed texts with the line-based diff. The diff implementation aligns exact equal lines through a longest-common-subsequence table and labels unmatched rows as added or removed. It reports structural movement faithfully, but it does not explain which locale comparison placed a name at either position.
Worked example: compare one browser-produced order instead of inventing two locale results
The simplest collaboration rule is to sort once and share the resulting lines, not merely the unsorted input plus an instruction to press Sort. A plain-text result preserves the decision that was actually reviewed. Recipients can search, annotate, or diff that artifact without asking their own browsers to recreate an environment-dependent ordering.
Keep the original list as well when provenance matters. The sorted copy is a presentation for review, while the original may carry meaningful arrival or source order. Naming the reference file and recording whether sorting was ascending, case-sensitive, and numeric turns a vague alphabetical operation into a repeatable team handoff.
What this does not cover — numeric sorting options, reverse order and sorting by a specific column
The original outline says numeric sorting options and reverse order are not covered, but the source contradicts that limitation. `sortLines` accepts a numeric option and a direction of ascending or descending, while `reverseLines` can reverse the current line sequence. Those capabilities should not be described as absent from the toolkit’s text logic.
Sorting by a specific column is genuinely outside this function. Every complete line is the comparison value, so a row such as `Paris, Ana` is sorted from its beginning rather than by the name after the comma. Parse structured rows with a suitable data tool when fields matter; line sorting should not pretend to understand columns.
The toolkit does cover numeric sorting and reverse order, but not sorting by a column
Locale-aware sorting is useful because human alphabets cannot be reduced safely to one raw character-number rule. It also means that an unpinned default locale is not a cross-machine reproducibility contract. The Text Toolkit uses browser collation, defaults to case-insensitive and numeric-aware comparisons, and can reverse the requested direction.
For a team spanning environments, make the output the contract: choose one browser session, set the intended options, sort, and distribute that exact text. When another order appears, compare artifacts before debating which alphabet is correct. The source supports explaining the mechanism, while the saved result supplies the stable sequence the implementation itself does not guarantee globally.