Text & everyday tools · Text Toolkit
CR, LF and CRLF: where line endings come from and why pasted text breaks
· Background
line-endings text-cleanup data-exports
Explains the teletype origins of carriage return and line feed, why operating systems chose different conventions, and how those choices show up as doubled blank lines and stray characters in pasted text.
The export with a blank line after every row — why a file from another system pastes with twice as many lines
You paste a three-row export and see a blank line after every record. It is tempting to blame Windows CRLF immediately, but a standards-conforming text area normally presents line breaks in a normalised form. Doubled rows more often mean that an earlier conversion treated carriage return and line feed as two independent separators, or inserted an extra newline between records.
Keep the original until you know which stage changed it. Compare the source in an editor that can reveal control characters, then compare the pasted line count. ToolAcre deliberately accepts CRLF, LF and lone CR as one boundary each, so an untouched three-record file should produce three lines rather than six merely because it came from Windows.
An extra blank row is a symptom, not proof that CRLF alone caused it
The names describe physical actions on printing terminals. A carriage return moved the carriage to the start of the current line, while a line feed advanced the paper to the next line. They were separate controls because either motion could be requested independently. ASCII retained them as control characters CR at decimal 13 and LF at decimal 10.
Modern screens no longer move paper, yet the byte values survived in files, protocols and programming interfaces. The history explains why CR and LF are not interchangeable punctuation marks and why CRLF is a two-character sequence. It does not mean every modern application processes them separately; parsers commonly recognise the pair as one logical line ending.
Three conventions — LF on Unix and modern macOS, CRLF on Windows, CR on classic Mac OS, and why each seemed reasonable
Unix and Unix-like systems conventionally use LF for a line ending, and modern macOS follows that convention. Windows text files conventionally use CRLF. Classic Mac OS used CR alone, but Mac OS X adopted the Unix foundation and LF. Those choices remain visible when tools exchange plain text without agreeing on normalisation.
No convention makes the words themselves different. Trouble appears at a boundary whose reader expects only one representation or naively splits on every control character. A robust line parser checks CRLF as a pair before checking lone CR or LF. ToolAcre does exactly that with the ordered pattern ` | | `, then joins transformed output with LF.
What happens in a browser text box — how line breaks are generally normalised on paste and where stray characters still appear
HTML defines special handling for line breaks in text controls. In a textarea value, browsers normalise CRLF and lone CR to LF in the exposed value, while form submission can apply the form-data rules for line breaks. Consequently, a paste that looks correct in the box can still be serialised differently by another layer or copied into software with another convention.
Stray CR can still surface when text bypasses a normal textarea, when escaped bytes such as the literal characters `\r` are displayed, or when a parser splits only on LF and leaves CR attached to each field. The browser is one stage in the path, not a universal repair service for files, clipboard producers, APIs and command-line consumers.
Browsers normalise textarea line breaks, but clipboard and downstream formats still differ
Blank rows and trailing carriage returns require different diagnoses. Remove Empty Lines deletes rows whose content is empty or whitespace after ToolAcre has recognised all three line-ending styles. Trim Lines removes leading and trailing whitespace from every row. Because ToolAcre’s splitter consumes a real CR separator, trimming is not normally needed merely to erase that separator.
Use trimming only when spaces, tabs or a literal unconsumed character remain, and inspect meaningful indentation before changing it. If the text contains visible `^M`, determine whether the viewer is rendering an actual CR or those two printable characters. A blanket replacement can damage intended content, whereas checking counts before and after gives a reviewable result.
Remove empty lines fixes empty rows; trimming is usually unnecessary after ToolAcre splits CR correctly
For a reproducible example, start with three names whose damaged intermediate representation contains an empty row between each name: `Ada`, blank, `Grace`, blank, `Linus`. The Word and Character Counter reports five lines. This is deliberately doubled input; a genuine CRLF sequence alone would be recognised as one boundary and would not create those empty rows in ToolAcre.
Choose Remove Empty Lines and the output becomes three lines joined with LF. The counter should now report three. If imported values also carry padding, run Trim Lines separately and review the result. Separating those operations proves which defect each action fixed instead of attributing every cleanup to a vague conversion from Windows text.
Worked example: clean a deliberately doubled export and verify the line count
Line-ending cleanup does not repair hard wrapping, where one logical sentence was deliberately broken at a column width. Removing every newline from that material would also join real paragraphs and list items. Decide whether the boundary represents a record, a paragraph or visual wrapping before applying a line operation across the whole block.
It also does not diagnose character encoding. A UTF-8 byte order mark, replacement diamonds, mojibake and decoding failures concern how bytes become characters, not whether CR or LF separates those characters into rows. Preserve the original file and identify its encoding with an appropriate file-aware tool before treating odd visible symbols as line endings.
The takeaway — line endings are history you can see; the Text Toolkit's line tools and counter let you fix the symptoms in seconds
CR, LF and CRLF are historical controls with present-day compatibility consequences. The safest mental model is one logical line boundary with several physical representations. Count records, inspect the source convention and identify the stage that introduced empty rows or retained control characters before deleting anything.
For pasted material, open the Text Toolkit at `/tools/text/`, note the initial line count, apply Remove Empty Lines only when empty rows are genuinely unwanted, and use Trim Lines only for surrounding whitespace. Recheck the count and sample records afterward. That short audit turns an invisible formatting problem into a controlled, reversible text transformation.