English

Data & spreadsheets · CSV Cleaner

How a Browser Cleans a Large CSV in a Web Worker Without Uploading It

· How it works

csv browser-processing privacy

A local CSV file passing into a browser worker and returning as a download without a server
Original ToolAcre vector illustration

'No upload' sounds like marketing until you see how it works. This post explains how a browser reads a file locally, why a Web Worker keeps the page responsive, what limits file size, and how to verify nothing is sent.

A large export and a web tool — what usually happens to your file, and why an upload is not a technical necessity

Choosing a file on a web page can resemble an upload, but selection and transmission are separate browser operations. CSV Cleaner receives a File object after the visitor chooses it, validates the allowed type and configured size, then reads its text inside the page. No upload is necessary for delimiter detection or row cleanup.

The authoritative product record says rows are parsed on the device, while the tool record narrows the claim: its code makes no request carrying file data, pasted text or generated output. Disclosed analytics script requests are separate page traffic, so privacy language must distinguish them from the CSV processing path.

Reading a file without sending it — how the browser's File API gives a page access to bytes on your device

The implementation calls `file.text()` rather than the FileReader API named in the outline. That promise resolves with a JavaScript string available to the page. A JSON file takes a separate conversion branch; delimited text is sent as a message to the worker together with an optional delimiter choice.

Reading locally does not grant the site general access to the drive. The page receives only the file selected through the browser control. It stores a base name for the eventual download, keeps the current table in tab memory and clears that state when the visitor requests a reset.

The page uses File.text() to read selected bytes without an upload

Parsing runs in a Web Worker created lazily on the first call. The worker detects the delimiter, scans characters through the same state machine used by tests, and returns the completed header, rows, warnings and BOM flag. Progress labels are sent around those stages so the main interface can keep updating.

Offloading the scan prevents that synchronous character loop from monopolizing the main UI thread. It does not make parsing free or streamed; the source text and returned table still occupy memory. On pagehide, the client aborts an inflight call and terminates the worker so a backgrounded tab does not retain it indefinitely.

Memory, not an upload cap, is the limit — why the ceiling depends on your device and browser rather than a server quota

The outline said memory rather than an upload cap was the limit, but the shipped interface explicitly validates a maximum of 50 MB. That threshold is in both configuration prose and panel code. Device and browser resources can still fail below a theoretical maximum, yet the article must not erase a deliberate product guard.

Nor can repository evidence supply a universal row ceiling, memory multiplier or processing time. File content changes the number and length of strings created. The defensible expectation is bounded input plus worker parsing, not a performance figure that was never measured across the visitor’s hardware.

The interface enforces a 50 MB file limit in addition to device constraints

To verify the narrow no-upload claim, open developer tools after the page loads, clear the Network list, choose a harmless distinctive sample and run a cleanup. Inspect new request URLs, methods and bodies for that marker. A CSV download may appear as a local browser action rather than a request to a conversion endpoint.

Source inspection complements the observation: the load path uses file text, the worker imports local parser code and downloadText creates the output. Runtime inspection catches deployment additions that source alone may miss. Neither proves what browser extensions or the operating system do, so keep the conclusion scoped to ToolAcre’s processing path.

Worked example — cleaning a large export while watching the network panel and memory usage, with what to expect on a low-memory device

Use a generated sample comfortably below the documented limit, with enough rows to make progress visible but no customer data. Watch the worker activity and preview, apply Trim whitespace, then download. Record whether any request contains the unique marker; do not publish fabricated megabytes-per-second or memory charts.

On a constrained device, the only safe prediction is that available resources differ. The configured limit prevents larger files from entering this path, while ordinary browser memory pressure remains possible. If a test tab becomes unstable, close other work, reduce the sample and preserve the original rather than promising a universal workaround.

Worked example: verify a harmless sample without inventing memory or timing figures

Local cleaning cannot retract a file previously sent to another service, and it cannot process an input rejected by the size guard. It also does not isolate the page from extensions, clipboard history or malware on the machine. Those are distinct threat boundaries outside the parser and worker implementation.

The tool retains an undo history of up to twenty table states during the session, which also consumes local memory. Clearing resets the table and controls; leaving the page aborts and terminates the worker. This behavior is useful operational detail, but it is not a secure-erasure guarantee for browser or device memory.

On-device processing is checkable — how the ToolAcre CSV Cleaner runs in your browser and never uploads your file

The complete verified path is selection, `File.text()`, worker parsing, pure row transforms, serialization and browser download. Configuration identifies the real 50 MB admission limit and the privacy statement explicitly separates disclosed analytics requests from spreadsheet contents. That is stronger than a slogan because each boundary can be inspected.

Run the network check with non-sensitive data whenever deployment details matter to your organization. Evidence from the current source and observed session supports a precise claim: ToolAcre’s CSV operation has no upload endpoint and sends no table content. Precision keeps the statement useful without pretending the whole browser is offline.