English

Text & everyday tools · Text Toolkit

Trim, de-duplicate, sort: why the order of text cleanup steps matters

· Why it matters

text-cleanup duplicate-lines sorting

Three untidy membership lists passing through trim, empty-line removal, deduplication and sorting steps
Original ToolAcre vector illustration

Two lines that differ only by a trailing space are not duplicates until you trim them; this post explains why trim → remove empty → de-duplicate → sort is the right order and how to verify each step with counts.

The de-duplicated list that still had duplicates — how an invisible trailing space defeats a naive de-duplicate

A membership secretary merges three sign-up exports and still sees two entries for the same address after choosing Remove duplicates. One line ends immediately after the address, while the other carries a trailing space copied from a spreadsheet cell. They look identical on screen, but the comparison receives two different strings, so both survive.

Running the same cleanup buttons in another order can hide rather than solve that mismatch. Sorting first may place the near-duplicates together for visual review, yet it does not make them equal. The dependable sequence is to normalize line edges, discard lines with no content, remove exact repeats, and sort only after deciding that source order has no meaning.

Step one, trim lines — removing leading and trailing whitespace so identical entries become identical

Start with Trim lines. The implementation splits the text at CRLF, LF or lone CR line breaks, applies JavaScript trimming to each line, and joins the lines again with LF. Leading and trailing whitespace disappears from every row, so ` member@example.test ` and `member@example.test` become the same exact text before duplicate detection begins.

Trimming is deliberately narrower than general text repair. It does not change spacing inside a name, repair capitalization or decide that two differently written addresses belong to one person. That restraint makes this first pass reviewable: only line-edge whitespace changes, while every visible character in the entry remains available for the secretary to inspect.

Step two, remove empty lines — why blank lines should go before sorting, not after

Next choose Remove empty lines. A line is considered empty when trimming it would produce an empty string, so rows containing spaces or tabs disappear along with visibly blank rows. Doing this before sorting prevents blank entries from being moved to one end of the roster and mistaken for unexplained output from the sort operation.

This second pass also clarifies later counts. A blank separator between the three imported lists is useful while assembling the source, but it is not a member record. Once the lists are combined and their boundaries no longer matter, removing those separators makes each remaining line correspond to one candidate roster entry that can be compared.

Step three, remove duplicates — the first occurrence stays, later copies go, and the order of what remains is unchanged

Now run Remove duplicates. With the button defaults, comparison is case-sensitive and does not perform another trim. The function walks from top to bottom, stores each line it has seen, keeps the first occurrence and skips every later exact match. The relative order of all retained lines therefore remains the same after this operation.

Keeping the first copy matters when the source order carries a weak preference, such as placing the primary registration export before supplemental lists. It does not merge details from competing rows, and `Alex@example.test` remains distinct from `alex@example.test`. Review those differences manually instead of assuming that every case change is harmless identity normalization.

Step four, sort — only if order does not matter, to make the result easy to scan

Sort only after duplicates are settled, and only when alphabetical presentation is more valuable than import order. Sort A-Z copies the lines and compares them with the browser locale, case-insensitive sensitivity and numeric comparison enabled. Entries containing numbers can therefore follow natural numeric ordering rather than a simple character-by-character sequence.

The outline described sorting only as an easy scanning aid, but the implementation has specific comparison behavior worth recording. Results can depend on the browser locale, and equal-looking case variants may retain implementation-dependent relative placement. If list order records arrival time, priority or voting status, skip sorting and preserve the deduplicated sequence.

Step four, sort with locale-aware, case-insensitive and numeric comparison, but only when source order is disposable

Use the live Lines count as a running check, not as proof that every remaining person is unique. Record the count after the initial paste, after empty-line removal and after duplicate removal. Trimming normally leaves the count unchanged; blanks reduce it by the number of discarded rows; deduplication reduces it by the number of later exact copies.

A trailing line break creates a final empty line because line splitting preserves the text after that break. This can make the starting count look one higher than expected until Remove empty lines runs. Compare the actual roster rows as well as the number, since two different entries can still refer to one person and one identical shared address can represent two people.

Check line counts while remembering that a trailing line break creates an empty final line

Suppose source one contains `Mina@example.test`, source two contains the same address followed by spaces, and source three repeats the clean address below two whitespace-only rows. Paste all three blocks together and note the line count. Trim first so the spaced copy matches, then remove empty lines so separators no longer count as roster candidates.

Choose Remove duplicates next; the first clean Mina row stays and the later two copies disappear. Read the reduced count and scan nearby entries before choosing Sort A-Z. Every transformation pushes the previous editor text onto the history stack, so Undo can restore one step at a time when a count drops unexpectedly or source order proves important.

The takeaway — the Text Toolkit's buttons are single transforms applied in the order you choose, and the recommended order is documented on the page

The Text Toolkit exposes independent buttons rather than a bundled cleanup command. That design makes order the user's responsibility and also makes each change observable. Although the toolbar displays Remove duplicates before Remove empty lines and Trim lines, the safer workflow for mixed exports is Trim lines, Remove empty lines, Remove duplicates, then optional sorting.

Treat that order as a small data-cleaning pipeline: normalize what comparison sees, remove records with no content, collapse exact repeats, and rearrange only the final set. Keep counts and Undo available at every boundary. The result is not a verified membership database, but it is a cleaner, auditable roster ready for identity checks the text functions cannot perform.