English

Data & spreadsheets · CSV Cleaner

RFC 4180 Explained: The Closest Thing CSV Has to an Official Standard

· Background

csv rfc-4180 data-formats

CSV records with comma separators, CRLF endings and doubled quotes aligned into a consistent output
Original ToolAcre vector illustration

CSV existed for decades before anyone wrote it down. This post walks through RFC 4180, what it actually says about delimiters, quoting, line endings and headers, and why real-world files still ignore it.

Two 'valid' CSV files that no single parser reads correctly — why the format grew for decades without a standard

Two files can both carry a .csv extension while one uses semicolons, another tabs and each applies quotes differently. A parser must know a dialect before punctuation becomes structure. ToolAcre begins with four candidates and scores parsed sample shapes rather than trusting the extension.

That flexibility explains why “valid CSV” is often an incomplete description. The cleaner can accept common deviations and then serialize a consistent result. Its behavior is defined by configuration and tests, giving a more useful contract than assuming every producer follows one memo perfectly.

What RFC 4180 is — an informational memo from 2005 documenting common practice, not a binding standard

The repository repeatedly names RFC 4180 as the output convention, but it is not a source for historical claims about the memo’s date or legal status. This article therefore focuses on the rules embodied by the implementation: record endings, quoting, header width and delimiter handling.

When external standards history matters, cite the primary publication separately. For using ToolAcre, the relevant evidence is local: code defaults to CRLF, quotes structural characters, doubles inner quotes and reports a row whose width differs from the header.

Repository evidence uses RFC 4180 as the target convention; publication history is outside the source set

The configured target describes comma-delimited output, CRLF line endings, an optional header concept and rectangular rows. On the shipped route the first row is not optional in practice: it is always consumed as the header. A headerless file loses its first record to that role.

Ragged input does not cause rejection. Short rows are padded with empty cells; long rows keep extras, and both receive row-numbered warnings. The serializer writes every retained row, so “clean output” does not mean the program guessed how to fix structural disagreement.

Quoting according to the memo — when fields must be quoted and how a literal double quote is doubled

A field must be quoted when it contains the chosen delimiter, a quote or a record-ending character. ToolAcre additionally quotes leading or trailing spaces so another parser does not silently trim them. Inside a quoted field, each literal double quote becomes two double quotes.

The Quote every field switch requests a uniform style; otherwise output is minimally quoted. Redundant source quotes may disappear while parsed values remain equal. This normalization is expected and is why tests compare the recovered table as well as exact output for selected examples.

The text/csv media type — the optional header and charset parameters, and why they rarely travel with a file on disk

A file on disk rarely arrives with media-type parameters attached. ToolAcre does not read a `header` or `charset` parameter from the selected file. It relies on its route contract: UTF-8 text, a first-row header and supported delimiter detection or manual selection.

The optional BOM export switch writes U+FEFF before the first header. That may help a spreadsheet recognize UTF-8, but the configuration warns it can break strict script or database import because the marker joins the first name when not stripped. Leave it off unless the consumer requires it.

Where the world diverges — semicolons, LF line endings, ragged rows and unlabelled encodings, and why the memo tolerates none of it

Real input may use semicolon, tab or pipe, LF or lone CR, uneven rows and a leading UTF-8 BOM. The parser intentionally accepts all those line endings and delimiter candidates. It does not “tolerate” ragged rows silently; it keeps data and attaches a warning.

An unclosed quote is another accepted-but-reported condition. Everything after the opening quote becomes one field, because no unambiguous closing boundary exists. Correct that source before trusting export. Permissive parsing is a preservation strategy, not certification that the input was conforming.

ToolAcre accepts several real-world deviations while reporting ragged rows instead of rejecting them

No rule here gives cells numeric, boolean or null types. The parser returns strings and the JSON converter preserves them as strings. Semantic conventions must be agreed by sender and receiver, including whether blank text is missing and whether a digit sequence is an identifier.

Formula-injection safety is an export policy layered above CSV syntax. When enabled, risky leading characters receive an apostrophe except ordinary numeric forms. That deliberately changes values for spreadsheet safety; it is not part of quote escaping or a general type system.

Aim for RFC-style output even when inputs ignore it — how the ToolAcre CSV Cleaner repairs files towards consistent delimiting and quoting

Aim for predictable output even when the source dialect varies. After parsing and resolving warnings, ToolAcre emits CRLF records and standard doubled-quote escaping. If the detected delimiter remains semicolon or tab, the download keeps that delimiter rather than always forcing comma.

Select comma manually when a receiving contract requires comma-separated output, and inspect the control because it displays the delimiter actually used. Consistency comes from an explicit selection plus reviewed row shape, not from the CSV extension alone.