Data & spreadsheets · CSV Cleaner
How Duplicate Row Removal Works: Exact Matches, Whitespace and Case
· How it works
csv data-cleaning duplicates
Two rows can look identical and still not match byte for byte. This post explains what 'duplicate' means to a cleaning tool, how whitespace, case and formatting create false non-matches, and how to prepare data so de-duplication catches what you intend.
'Remove duplicates' left obvious repeats behind — why a trailing space or a capital letter makes two rows different
A contact export can display two rows that appear identical while one email ends with a space or one surname uses a capital letter. The duplicate button does not apply human similarity. On the shipped page it compares the complete string value of every cell, with case and whitespace still significant at that moment.
That explains why an apparently obvious repeat may remain after the first click. The tool is avoiding a riskier decision: changing case or ignoring selected fields could merge records that only look related. Review the source, choose the normalization you actually want, and rerun exact removal after those edits.
What exact matching compares — every field and every character, including invisible ones such as non-breaking spaces and stray BOM bytes
Each row receives an identity built from all its cells. The implementation joins them with a null character that is not legal CSV text, avoiding the common mistake where one cell containing a comma collides with two separate cells. A second identical identity is skipped; the first row is retained unchanged.
The outline mentioned invisible non-breaking spaces and stray BOM bytes as though all receive special treatment. JavaScript trimming can remove surrounding Unicode whitespace when you choose Trim, but the parser’s explicit BOM rule only strips U+FEFF at the very start of the file. It does not scan every cell for arbitrary hidden bytes.
Exact comparison covers complete cell strings; only a leading file BOM is specially removed
Order matters. Trim whitespace removes leading and trailing space in every cell, while Collapse spaces also turns internal whitespace runs into one ordinary space. Applying either before duplicate removal can make previously distinct rows equal. Running removal first preserves both, because their original strings differ.
The panel does not expose case folding, Windows-1252 repair or Unicode normalization. Although the underlying library has an option for case-insensitive comparison, the route calls the default exact function. A claim that this page standardizes case or encoding would therefore exceed the controls a visitor can actually use.
Trim or collapse whitespace first; the panel does not normalize case or legacy encodings
Whole-row equality is not the same as identifying one customer. Two rows sharing an email but carrying different timestamps are both kept because at least one cell differs. The library can compare selected columns, yet the published duplicate page does not expose a key-column selector, so it cannot adjudicate that pair.
This is a valuable boundary rather than a missing magic trick. Choosing the newest record, merging fields or treating aliases as one person requires a business rule and often an audit trail. Export those conflicts for review or use the destination system’s documented merge workflow instead of disguising them as exact duplicates.
Order and survivorship — which copy is kept when duplicates are removed, and why that matters for 'last updated' data
When exact duplicates occur, the first appearance survives and later copies are removed. Surviving rows keep their original order. If a newer version appears later but differs in even one field, it is not a duplicate and remains; if it is byte-for-byte equal at the cell level, retaining either copy yields the same data.
The first-copy rule still matters operationally when row order records provenance outside the cells. Keep an untouched export before cleaning, and inspect counts after each action. ToolAcre’s undo stack retains twenty transformations in the current tab, but a downloaded original is the durable reference if the session closes.
Worked example — a contact export with near-duplicates, normalised so exact de-duplication catches them, with the rows that still need human review listed
Build a small test with Ada@example.com, Ada in one row, the exact same two cells in a second, and ada@example.com plus Ada followed by a space in a third. Exact removal drops only the second row. Trimming then removes the trailing space, but the lowercase email still keeps the third row distinct.
That outcome distinguishes mechanical certainty from human judgement. If the addresses are known to be case-insensitive in your system, apply that rule elsewhere and document it. The cleaner’s evidence is simpler: it can prove identical row arrays are reduced to their first occurrence without mutating the retained values.
What this does not cover — fuzzy matching, typo tolerance and merging conflicting records
Fuzzy names, typo tolerance, phonetic matching and conflict merging are outside this operation. It does not recognize that “Robert” and “Bob” might refer to one person, nor does it score two addresses for similarity. Those techniques can create false positives and require context unavailable in a flat export.
Key-column deduplication is also absent from this page despite support in the lower-level function. Articles should describe the route a reader can use, not latent parameters with no control. For identity resolution, select a purpose-built review process where match evidence and survivor rules are visible.
De-duplication is only as good as the normalisation before it — how the ToolAcre CSV Cleaner's duplicate removal fits after its encoding and header repairs
ToolAcre’s reliable sequence is inspect, normalize chosen whitespace, then remove exact repeats. Encoding repair and automatic header repair are not hidden preliminary stages. The parser removes a leading UTF-8 BOM, detects a delimiter and reports structural problems; it does not reinterpret damaged character encodings.
After removal, compare row counts and preview the survivors before download. What disappeared was provably equal across every cell under the strings then present. What remains may include legitimate near-duplicates, and keeping those uncertain records is safer than silently merging customer data on an assumption.