Data & spreadsheets · CSV Cleaner
Tab-Separated Values vs CSV: Why TSV Exists and When to Use It
· Background
csv tsv data-formats
Tabs rarely appear inside data, which is exactly why TSV exists. This post explains the origins of tab-separated files, how they sidestep CSV's quoting problems, and where they still fall short.
Free text full of commas and quotes that keeps breaking CSV exports — the case that made tabs attractive
Free-text notes often contain commas, so comma-delimited output needs quotes around those cells. A tab may occur less often in that particular dataset, which can make a TSV file visually quieter. It does not eliminate the need for a parser or an agreement about embedded tabs and newlines.
ToolAcre treats tab as one of four separators. The same quoted-field state machine runs regardless of which separator is selected. This means the decision changes one structural character, not the safety model for every awkward value.
Where TSV comes from — the tab character's typewriter origins and its early use in database dumps and scientific data
The workbook requests a history from typewriters to database dumps, but those claims are not present in product configuration, implementation or tests. This module omits that chronology and focuses on behavior readers can reproduce with the shipped parser.
A tab in source text is simply ` ` to the state machine when selected as delimiter. Outside quotes it ends a field; inside quotes it remains part of the field. That rule is enough to compare formats without inventing an origin story.
Typewriter origins and early database history are outside repository evidence
TSV can reduce quoting when values contain commas but no tabs, quotes or record endings. The outline said most TSV dialects need no quoting at all; ToolAcre’s serializer is more precise. It quotes any field containing the selected delimiter, a quote, LF, CR or surrounding whitespace.
A note containing a real tab must therefore be quoted under tab output, and a literal quote is doubled inside that field. Free text also can contain line breaks, which remain quoted. Choosing an uncommon delimiter lowers collisions in a dataset; it never makes collision rules disappear.
ToolAcre still applies RFC-style quoting when tab-delimited data contains tabs, quotes or newlines
Tabs are visually ambiguous because editors can render them as variable-width space. Copying through software that converts tabs to spaces can destroy the separator while leaving the file readable to a person. Inspect invisible characters or rely on a parser preview before declaring a one-column result malformed.
ToolAcre writes the detected tab back into a labelled select and shows column counts, making the invisible choice visible. If the file is actually space-aligned text rather than TSV, tab detection will not create columns. Manual selection cannot manufacture separators that are absent.
Escaping conventions — backslash sequences in some TSV dialects versus RFC-style quoting in CSV
Some tabular dialects use backslash sequences for tabs or newlines. ToolAcre does not. Its parser recognizes RFC-style doubled double quotes and quoted fields under every selected delimiter. A backslash is ordinary cell text and does not escape the next character.
That mismatch can leave stray slashes or premature boundaries when loading an export configured for a different dialect. Identify the producer’s escape rules first. Converting safely requires parsing the original dialect, not replacing slash-letter pairs without knowing whether they were literal.
Backslash TSV dialects are unsupported; ToolAcre uses doubled quotes for every selected delimiter
Take headers id,note and two values: one note says `red, small`, while another contains a real tab. In comma output, the first note is quoted and the tab may remain unquoted because it is not the delimiter. In tab output, the comma is ordinary but the tab-bearing value is quoted.
Parse each output with its matching delimiter and both tables recover the same strings. This exercise shows that neither format universally avoids special handling. The chosen separator merely changes which character triggers quoting in addition to quotes and record endings.
What this does not cover — fixed-width formats and binary tabular formats
Fixed-width and binary tabular formats are outside the route. It does not infer columns from visual alignment or decode a columnar container. Inputs must be textual CSV, TSV or plain text using comma, semicolon, tab or pipe, or supported JSON in the reverse direction.
The route also offers no custom escape character. If a scientific or database workflow defines its own TSV grammar, use a parser configured for that grammar before bringing normalized text here. An extension is not enough evidence of dialect compatibility.
Choose the delimiter for the data and the consumer — how the ToolAcre CSV Cleaner's delimiter repair produces a consistently delimited file for the system you are feeding
Choose a delimiter according to both data and consumer, then state it explicitly. ToolAcre can detect tab and comma from a rectangular ten-row sample, or accept a manual override. It does not examine semantic content and decide that notes deserve tabs.
After parsing, resolve row warnings and serialize with the same or another supported separator. Standard doubled-quote handling protects collisions. The result is consistent because the delimiter is known, not because TSV or CSV is inherently immune to punctuation in data.