English

Developer tools · Text comparison

A Short History of Unix diff: Hunt, McIlroy and the Line-Based Model

· Background

text-diff algorithms software-history

A historical timeline stopping at a verified modern LCS comparison panel
Original ToolAcre vector illustration

Traces diff from Bell Labs in the 1970s to today's tools, explaining why comparing text by lines became the standard and how later algorithms refined it.

Why every picture of 'what changed' looks like Unix diff — opens with the observation that code hosting sites, review tools and browser tools all inherit the same model

Modern review screens often share a visual grammar of unchanged context, removals and additions. It is tempting to turn that resemblance into a complete lineage story, but this repository is evidence for ToolAcre’s current implementation rather than a primary archive for Unix history.

The safe article therefore distinguishes two tasks. It explains what the present source proves in detail and marks the workbook’s historical names, dates and motives as requiring external authoritative references. Familiarity is not a citation, especially for algorithm authorship.

Bell Labs and the problem of comparing files — describes the early-Unix need to compare source files and the work of Douglas McIlroy and James Hunt

The outline attributes early work to Bell Labs figures. None of the four Text diff source files documents that history, and the task forbids invented or uncited claims. This module consequently does not repeat biographical details, quotations or dates as settled facts.

A future historical article should cite papers, manuals or institutional archives directly. Until those sources are part of the evidence set, the correction is explicit: ToolAcre can demonstrate a line-diff model today, but it cannot authenticate every step in the model’s origin story.

Bell Labs authorship and early-Unix history need external primary sources not present in this repository

Current code splits on normalized newline conventions and compares arrays of lines. That fact explains today’s granularity. It does not prove teletypes or storage constraints caused a historical design choice, because no historical material is imported into the implementation.

The distinction matters beyond scholarship. “This tool works on lines” is reproducible by tests. “Lines were chosen for this historical reason” is a causal claim needing documents from the period. Good technical writing does not use present code as retroactive evidence of past intent.

The repository proves ToolAcre compares lines; it does not prove why historical systems chose them

`toUnifiedText` emits `---` and `+++` labels followed by every row prefixed with space, minus or plus. It is patch-shaped and useful for download, but it contains no `@@` hunk headers. Calling it a complete unified diff would overstate the format implemented here.

The route also does not generate normal or context formats. Historical evolution among those formats may be worth explaining with external manuals, yet the shipped output supports a narrower description: a readable full-row stream with familiar markers and caller-supplied labels.

ToolAcre exports a patch-shaped stream, not full historical diff formats

The outline says modern engines use Myers, but ToolAcre does not. Its file header names plain LCS dynamic programming and documents O(n·m) time and memory over the differing middle. The implementation fills a `Uint32Array` table and reconstructs one path through it.

This is not a cosmetic correction. Algorithm names carry specific complexity and tie behavior. Prefix and suffix trimming plus a 2,000-line cap make this design suitable for bounded pasted differences. Readers should not transfer a Myers complexity claim onto code with a quadratic matrix.

ToolAcre uses plain LCS dynamic programming, not Myers

Line-based comparison still cannot identify semantic equivalence, moved blocks or formatting intent. Those limitations follow directly from ordered string keys. A compiler may consider two source fragments equivalent while a line diff reports replacements, or the reverse under aggressive normalization.

The original model remains useful precisely because its result is inspectable. Each row has text, type and side-specific line numbers. The reviewer can disagree with an alignment while still account for every line, which is harder with an opaque “meaning changed” score.

What this does not cover — binary comparison, delta compression and version-control internals

Binary deltas, compression, repository object storage and merge internals are outside both this route and the evidence set. The article also omits the requested historical chronology rather than filling it from memory. Missing a claim is preferable to publishing an unverified one.

For those topics, collect primary specifications and source from the systems being described. Do not infer a file format from visual similarity or an algorithm from the word “minimal.” ToolAcre’s config uses that word, while the implementation supplies the exact mechanism that qualifies it.

Takeaway: fifty years of one good idea — summarises the lineage and notes that ToolAcre's Text comparison applies the same line-based model in a browser tab

The durable lesson is methodological: separate history from implementation. This browser tool proves normalized line splitting, LCS reconstruction, explicit options, a differing-line cap and local rendering. It does not prove who invented the surrounding conventions or why.

Use the route to inspect present text and use primary sources to tell the past. That boundary may make the historical article less sweeping, but it makes every included statement auditable. Precision beats a polished lineage assembled from uncited recollection.