English

Documents · PDF Toolkit

How a Browser Tab Reads and Rewrites a PDF Without Uploading It

· How it works

pdf privacy web-workers

A document moving through browser memory and returning as a local download
Original ToolAcre vector illustration

'Runs in your browser' is a claim you can understand and test. This post follows a PDF from the file picker into memory, through a Web Worker, and back out as a download, explaining what each step does and why no server is involved.

The file picker is not an upload — the moment of confusion when choosing a file looks like sending it

A file chooser resembles an upload control because many sites send the chosen file immediately. Selection alone does not require that transfer. In this toolkit, the browser grants the page access to the user-selected File object, and the controller validates its type and size before reading PDF bytes into tab memory.

The boundary is observable: choosing a document changes local interface state, displays its name, size, and page count, and enables the operation. It does not post the file to an application endpoint. The PDF cap is 50 MB and the image cap is 30 MB, enforced before expensive processing begins.

From disk to memory — how the File API hands the page a byte array the site's own JavaScript can read

For PDFs, `arrayBuffer()` supplies bytes and a `Uint8Array` holds them for worker calls. Image batches are handled differently to avoid an unnecessary second read: raw File objects are retained until browser image preparation. These are memory operations inside the current browser context, not remote storage or account history.

The first PDF is inspected through an `info` call in the worker so a large document does not block painting while page count is determined. The tab can temporarily hold source bytes, parser state, prepared assets, and eventual output together. Clearing or closing matters because local processing still consumes real device resources.

Most PDF transforms use the toolkit worker; PDF parsing and canvas encoding have a separate split

Merge, split, extract, delete, reorder, rotate, watermark, and final image-to-PDF assembly call pdf-lib through the dedicated module worker. Progress and cancellation messages cross that boundary while computational work stays away from the primary interface thread. Clearing files terminates the worker instead of waiting for garbage collection.

PDF-to-image is the important exception. pdf.js uses its own worker for parsing, but browser canvas encoding must remain on the main thread. The renderer yields between pages so controls and progress can repaint. Saying every transform runs wholly in one worker would contradict the shipped implementation.

Parsing and writing are local stages, but image preparation and canvas encoding can run on the main thread

The general pipeline is read, interpret, transform, and encode, but each operation chooses its own concrete machinery. Page-copy operations ask pdf-lib to build a fresh document. Rotation changes additive page metadata. Watermarks append drawings. PDF-to-image renders pages, while images-to-PDF prepares browser-decoded images before worker assembly.

Those distinctions affect fidelity. Copying pages retains selectable content; rendering them to PNG or JPEG turns the page into pixels and loses its text layer. Non-PNG/JPEG images may be decoded and re-encoded as PNG before PDF assembly. “Local” describes data movement, not one universal transformation algorithm.

The download is a local object — how a Blob and an object URL give you a file that never existed on any server

Each operation returns bytes or a ZIP that the interface wraps in a Blob with the appropriate media type. The result action calls the shared `downloadBlob` helper, which creates the browser download rather than navigating to a server file. The generated artifact existed in memory before the user saved it to normal device storage.

Organizer thumbnails also use Blob object URLs, but their lifecycle is explicit: old URLs are revoked before a new load and all are revoked on destroy or clear. That distinction prevents a local privacy claim from hiding a memory leak. Once downloaded, the file follows ordinary device backup and sharing rules.

Why the limit is your device's memory — the original, the parsed structure and the rewritten copy all live in RAM at once

Local work is bounded by RAM and browser policies as well as explicit input caps. A PDF can exist simultaneously as source bytes, a parsed object model, transferred worker buffers, previews, and serialized output. Raster pages add large canvases, so the renderer measures a pixel budget and may reduce scale before allocation.

A phone can struggle sooner than a desktop even below 50 MB because compressed file size says little about decoded page imagery. Close unrelated tabs, process fewer pages, or use a larger device for demanding jobs. The hard cap remains 50 MB per PDF and 30 MB per image; memory is not an excuse to claim no fixed limit.

Other ToolAcre products may use the network; consented production analytics is also separate

The repository states that two other ToolAcre products contact the network by design, so local behavior must be verified per product rather than generalized across the site. The PDF operation code has no endpoint carrying documents, and an automated isolation test checks that invariant during processing.

Site-wide analytics can run only on the canonical production host after consent. Its ToolAcre event allow-list excludes file names, contents, pasted text, URLs, and exact file sizes, while Google scripts remain third-party code described by the privacy policy. Static asset or analytics requests are distinct from a document upload.

Takeaway — every step happens on your device, and the PDF Toolkit's page and the Network panel let you confirm it

The complete lifecycle is visible: choose a File, validate and read it, transform locally with the appropriate worker or rendering path, wrap the output in a Blob, download it, then clear buffers and object URLs. No application server is required to perform those page operations.

Treat “in browser” as an architecture you can inspect rather than a slogan. Watch requests, read the operation-specific note, and use Clear files and free memory when finished. That control releases organizer images, drops references, and terminates the worker, reducing both memory pressure and the lifetime of sensitive document material in the tab. Download first, verify the saved file, and only then clear; local processing intentionally provides no remote fallback for a result discarded too early.