Data & spreadsheets · CSV Cleaner
How UTF-8 and Windows-1252 Get Confused: Repairing a Mojibake CSV Export
· How it works
csv encoding data-cleaning
When 'José' becomes 'José', the bytes are fine and the interpretation is wrong. This post explains how the two most common encodings collide, how to recognise the symptoms, and how re-decoding repairs them.
Accented names and curly quotes turned into symbol soup — the tell-tale patterns of a UTF-8 file read as Windows-1252, and the reverse
A customer name that becomes replacement diamonds after loading is not evidence that CSV Cleaner detected the wrong legacy code page. The configuration says the opposite: only UTF-8 is understood, and a Windows-1252 or Shift-JIS file is read as UTF-8. Invalid byte sequences can therefore already be replaced before the CSV parser sees characters.
The outline centered familiar mojibake such as José, but the browser read path uses `File.text()` and provides no encoding selector. This article corrects that promise. The actionable symptom inside this tool is replacement characters or otherwise damaged text, with the original file bytes preserved for recovery elsewhere.
A non-UTF-8 file reaches this tool as replacement characters, not a verified mojibake pattern
A delimited file is bytes on disk, while the parser operates on a JavaScript string. An encoding defines the mapping between those layers. CSV syntax names commas, quotes and record boundaries but carries no reliable on-disk declaration telling `File.text()` which legacy mapping created every non-ASCII byte.
Once decoding has produced U+FFFD replacement characters, later CSV operations receive those placeholders as ordinary text. Trimming or exporting cannot infer which original byte sequence or character belonged there. That is why the untouched source matters more than a find-and-replace list assembled from the damaged display.
The two usual suspects — UTF-8's multi-byte sequences and Windows-1252's single bytes, and why they produce predictable garbage when swapped
UTF-8 represents non-ASCII characters with multibyte sequences. Windows-1252 assigns many Western characters to individual byte values. Reading one convention under another can fail or create misleading text, but this route does not test alternative decoders, score plausible language or offer a Windows-1252 selection.
The only encoding-specific parser behavior is removal of a leading U+FEFF UTF-8 byte-order mark after text decoding. That prevents the marker from joining the first header. It is not general encoding detection and provides no support for Shift-JIS, UTF-16 or regional code pages mentioned nowhere in the implementation.
ToolAcre accepts UTF-8 text and does not compare Windows-1252 candidates
Replacement characters indicate that the text decoder could not map some input bytes under its chosen interpretation. Question marks may have been inserted by an earlier lossy export, in which case the original character could already be unavailable. A recognizable à sequence can arise in other workflows, but this page does not diagnose its history.
Do not decide the source encoding from one surname alone. Check the exporting application’s setting, file provenance and a byte-aware inspector that leaves the source untouched. The cleaner’s row warnings concern quote closure and column width; they are not evidence that character encoding is correct.
Re-decoding, not find-and-replace — why the fix is to read the bytes with the right encoding and write UTF-8, rather than patching characters one by one
The reliable repair is to return to original bytes and decode them once with the documented source encoding, then write UTF-8. That operation must occur before opening through a UTF-8-only text path. Replacing visible garbage fragments after decoding can corrupt legitimate occurrences and cannot distinguish several original characters that collapsed to one placeholder.
CSV Cleaner has no byte-level re-decoding control, so it cannot perform the outline’s promised conversion. Use a trusted source-aware conversion method, compare representative names with the source system, and then bring the UTF-8 result here for delimiter, quoting, whitespace and duplicate work.
Recover from the original bytes outside this tool; character replacement here cannot restore them
For a safe demonstration, create a tiny legacy-encoded file containing one accented name and retain a hexadecimal copy. Load it into the tool and observe whether replacement characters appear. That observation establishes the UTF-8 boundary; it does not establish the original code page merely because the expected name is known.
Next convert the untouched bytes with an explicitly selected decoder outside ToolAcre, save UTF-8, and load that result. The name should now arrive intact while the CSV parser handles separators normally. Comparing these two paths teaches the right lesson without claiming that the cleaner performed the recovery itself.
Worked example: demonstrate the UTF-8 boundary without claiming an unsupported repair
A double-encoded file may require reconstructing an earlier transformation, and data already saved with literal question marks may be irrecoverable without another source. This article does not prescribe a universal reversal because the implementation contains no encoding history or byte-preserving recovery function.
It also avoids claiming support for UTF-16, East Asian encodings or normalization forms. If those matter, choose a converter that names and tests them. A successful CSV parse only proves the delimiter state machine found rows; it says nothing about whether character decoding before that stage was faithful.
Fix the interpretation once — how the ToolAcre CSV Cleaner's encoding repair re-decodes and re-encodes the export on your device
ToolAcre can strip a leading UTF-8 BOM and serialize the resulting string as UTF-8 CSV through the browser’s download path. It cannot turn arbitrary legacy bytes into correct Unicode because those bytes have already crossed the browser’s fixed text-reading boundary without a user-selected decoder.
Treat replacement marks as a stop signal. Preserve the source, identify its encoding from the producer, convert once with an appropriate byte-aware tool, and verify important names. Only then use CSV Cleaner for the structural jobs its configuration actually promises.