English

Data & spreadsheets · CSV Cleaner

How a CSV Parser Handles Quotes, Embedded Commas and Newlines in Fields

· How it works

csv parsing data-formats

A parser state path keeping quoted commas, doubled quotes and embedded line breaks inside fields
Original ToolAcre vector illustration

Splitting on commas works until a field contains a comma, a quote or a line break. This post explains the small state machine a real CSV parser uses, why quoting rules exist, and what a repair tool does when quoting is broken.

An address field with a comma shifted every column to the right — why naive splitting is not parsing

Splitting a line at every comma fails as soon as a postal address, note or product description contains that character. The separator only ends a field while the parser is outside a quoted field. Inside quotes, the same byte belongs to the value and must reach the preview unchanged.

ToolAcre proves this with tests rather than assuming CSV is simple text. A semicolon file containing commas inside quoted notes is still detected as semicolon-delimited because each candidate is parsed into rows before its rectangular shape is scored. Raw punctuation counts never decide the dialect.

The quoting rules — fields containing the delimiter, quotes or newlines are wrapped in double quotes, and inner quotes are doubled

A field containing the chosen delimiter, a double quote, a line feed or a carriage return is quoted during serialization. A literal double quote inside a quoted field is represented by two adjacent quotes. On parse, that pair contributes one quote to the cell and does not close the field.

Quotes only open a quoted state when they appear at the start of an otherwise empty field. In an unquoted value such as 12" pipe, the quote remains literal data. That implementation choice is tested and prevents a measurement mark in the middle of text from swallowing subsequent columns.

A parser as a state machine — inside-quotes versus outside-quotes, and how one character at a time decides where a field ends

The core loop tracks a current field, current row and `inQuotes` state. Outside quotes, the delimiter pushes a field and CR, LF or CRLF pushes a row. Inside quotes, every character is appended except a quote, which either forms an escaped pair or exits quoted state.

This one-character progression explains why a simple regular expression or line split is fragile. Whether a newline ends a record depends on earlier input, and whether a quote closes the field depends on the next quote. The parser carries exactly that minimal state through the string.

Embedded newlines — why a row can span several lines and why line-based tools such as grep or wc miscount records

A quoted field may contain LF, CRLF or a lone CR, and those characters remain inside the value. Consequently one logical record can occupy several display lines in a text editor. Counting physical lines with a line-oriented command does not necessarily count data rows.

After the closing quote, the next delimiter or record ending resumes structural meaning. ToolAcre’s tests assert both embedded LF and embedded CRLF, then verify that only one row was produced. This behavior is not inferred from an RFC label; it is exercised directly by the shipped state machine.

Broken quoting — unbalanced quotes that swallow the rest of the file, and how a repair tool re-quotes fields consistently

The outline promised a quoting repair, but an unclosed quote is inherently ambiguous. ToolAcre reports `UNCLOSED_QUOTE` against the row where quoted state began and reads everything afterward into one field. It does not invent a closing position or re-quote the intended records.

That conservative failure preserves the consumed text for inspection and avoids pretending to know whether a later quote or newline was meant as data. Correct the source or regenerate the export, then reload. Serialization can normalize a successfully parsed table; it cannot recover row boundaries the malformed source made unknowable.

Broken quoting is reported but not repaired; an unclosed field consumes the remaining text

Use `name,note` as the header, then one row with Ada and a note containing a comma plus a line break, and another whose note is `She said "hi"`. In source CSV, quote the long note and write the inner quotation mark as `""hi""`. The parser returns two data rows and keeps both special characters.

When serialized with minimal quoting, plain names need no quotes, while the note fields are quoted because they contain structural characters. Parsing that output again yields the same header and cells even if redundant source quotes disappeared. Data equality, not byte-for-byte styling, defines this round trip.

What this does not cover — dialects that use backslash escapes, and heuristics for guessing intent when quoting is truly ambiguous

Backslash escape dialects are outside the parser. A backslash before a quote has no special state-machine meaning here; it remains text, while the quote is interpreted under the doubled-quote rules. The tool also declines to guess intent once an opening quote remains unmatched.

Those boundaries matter when importing a database dump configured for another escape convention. Identify the producer’s dialect before loading, or export with RFC-style quoting. A parser that silently recognizes several incompatible escape rules can turn literal backslashes into control syntax and conceal a mismatch.

Respect the quotes and rows stay intact — how the ToolAcre CSV Cleaner's quoting repair produces a file that standard parsers read correctly

Respecting quoted state keeps delimiters and record endings inside the cells where they belong. ToolAcre also handles three line-ending forms, removes a leading UTF-8 BOM and reports row-width mismatches. These behaviors are independent parts of one structural read rather than a collection of search-and-replace tricks.

After a valid parse, serialization doubles embedded quotes and quotes any field requiring protection, with CRLF output by default. After an invalid unclosed quote, stop and repair from source evidence. The tool can standardize known cells; it does not claim to reconstruct an unknowable table.