English

Data & spreadsheets · CSV Cleaner

CSV Escaping Dialects: Doubled Quotes, Backslashes and Database Exports

· Background

csv parsing interoperability

A doubled-quote CSV path succeeding while a backslash escape remains literal text
Original ToolAcre vector illustration

RFC 4180 says to double a quote; several database tools export with backslashes instead. This post explains the escaping conventions in circulation, where each came from, and why a file that is valid in one dialect breaks in another.

A database export that a spreadsheet reads with stray backslashes — two escaping traditions colliding in one file

An export can contain a backslash followed by a quote because its producer follows a programming-language-style convention, while a spreadsheet-oriented parser expects doubled quotes. Feeding one dialect to the other does not create a neutral compromise. It changes where quoted state opens and closes or leaves escape characters in data.

ToolAcre supports one clear rule: fields can be wrapped in double quotes and an inner quote is represented by two quotes. Backslash has no escape role. Identifying that contract before loading is safer than asking the cleaner to guess which punctuation the producer intended.

The RFC convention — wrap in double quotes and double any inner quote, with no escape character at all

For a value such as `She said "hello", then left`, RFC-style CSV surrounds the field and writes the inner marks twice. The delimiter stays inside quoted state, and each doubled pair returns one literal quote. Serialization applies the same transformation in reverse.

Fields are quoted when they contain the current delimiter, quote, LF, CR or surrounding spaces, or whenever Quote every field is selected. There is no separate escape character. This small grammar is tested for delimiters, embedded record endings and quote pairs.

The C tradition — backslash-escaped quotes, newlines and delimiters inherited from programming languages and some database loaders

Some source systems can define backslash conventions, but ToolAcre has no switch for them. A backslash is appended to the current field like any ordinary character. A following quote is then interpreted according to whether the parser is in quoted state and whether it has a paired quote.

Consequently, a backslash export should be parsed by a tool configured for that dialect before conversion to doubled-quote CSV. Search-and-replace is risky because literal backslashes and escaped sequences can look alike. The source grammar must decide which ones are syntax.

Backslash escaping is an unsupported alternative dialect, not a mode ToolAcre can read

The workbook names database products and data-frame libraries, yet their defaults and configurable modes are not documented in this repository. This article does not attribute a dialect to any particular vendor. Consult the exact exporter or loader documentation used by the workflow.

Capture the command and options that produced the file. A product may support several modes, so its name alone is not a dialect. Reproducibility comes from configuration plus a sample containing delimiter, quote, newline and backslash edge cases.

Specific database and library defaults require their own documentation and are not claimed here

Under this parser, `"a \"quote\"",b` does not mean what a C-style reader might expect, because the slashes remain in the text and the quotes govern CSV state. Conversely, a doubled pair inside a quoted field becomes one quote rather than two literal characters.

A mismatch may trigger shifted columns or an unclosed-quote warning, or it may parse into wrong-looking strings without structural errors. Preview representative rows and compare them with the producer. Syntactic success alone cannot certify that escape semantics matched.

Under ToolAcre rules, backslashes remain literal and doubled quotes are the only quote escape

Write one cell containing both a comma and a quote as RFC CSV: `"She said ""hi"", then left"`. ToolAcre returns the exact value `She said "hi", then left`. Write the inner quote with backslashes instead and observe that the slash has no special protection.

The comparison demonstrates parser rules without depending on a database brand. After a correct parse, downloading produces doubled quotes consistently. After a mismatched parse, exporting merely standardizes the wrong cells, so resolve the source dialect before trusting the result.

What this does not cover — binary exports, fixed-width files and fully custom escape characters

Binary exports, fixed-width records and custom escape characters are outside the route. An unclosed quote is reported and consumes the remaining input as one field; the cleaner does not infer where a missing close belonged. That ambiguity cannot be repaired generically.

Nor does the interface convert a backslash dialect automatically. Use a parser that explicitly names the escape character, obtain a table of verified strings, and then serialize to the target convention. Each grammar should be applied once at a known boundary.

Know the dialect on both ends — how the ToolAcre CSV Cleaner's quoting repair normalises fields towards consistent quoting that standard parsers read

Know both endpoints. ToolAcre is appropriate when the source follows doubled-quote CSV or when another parser has already normalized a different dialect into verified cells. Its output offers consistent minimal or all-field quoting and doubles every embedded quote.

Do not mistake a cleaned download for dialect detection. The auto-detector chooses delimiters, not escape grammars. A reliable transfer records separator, quote rule, newline policy and encoding together, then tests an adversarial sample before production data moves.