Documents · PDF Toolkit
What Converting a PDF Actually Means: Rasterising vs Re-encoding
· Background
pdf image-conversion rasterization
'Convert' covers very different operations: drawing a page to pixels, wrapping an image in a page, or reconstructing editable text. This post explains what each involves and why some conversions are exact while others are approximations.
The 'converted' file that lost its text — why a page turned into an image can no longer be searched or selected
A converted page can look identical and still lose its most useful property. PDF-to-image renders each page into pixels, so words can no longer be selected, searched, or read as a text layer. The ZIP contains pictures of pages, not smaller PDFs or editable document content.
That loss is appropriate when a slide, chat, catalog, or preview needs an image. It is harmful when accessibility, search, copy, or later editing matters. Choose conversion based on the required representation rather than the reassuring similarity of the first visual inspection.
Rasterising — rendering a page's vector instructions to a bitmap at a chosen resolution, and what is gained and lost
Rasterizing asks pdf.js to interpret vector drawing instructions, fonts, and images, then paint a canvas at a selected scale. ToolAcre offers PNG for crisp document edges and JPEG for smaller photographic output. Each page becomes a separately named image inside one ZIP.
The renderer checks a per-device pixel budget before allocation and can reduce oversized pages below the requested scale. PNG and JPEG preserve the rendered appearance at that pixel grid, not the source vectors or text semantics. Zooming beyond the chosen resolution eventually reveals the finite raster.
Wrapping — placing an image inside a new page so it becomes a PDF, and why this is lossless for the image itself
Images-to-PDF travels the other direction by placing one prepared image on each new PDF page. Match each image sizes the page to the picture plus margins; A4 and US Letter presets contain and center the whole image without cropping. The result is still image content, not reconstructed text.
JPEG and PNG bytes are embedded directly when possible. WebP, AVIF, GIF, BMP, HEIC, and HEIF depend on browser decoding and are re-encoded as PNG; animated GIF contributes only its first frame. Large pictures may be reduced to fit a device pixel budget, and every input image is capped at 30 MB.
Reconstructing — why turning a PDF into an editable document requires guessing at paragraphs, columns and tables
Reconstructing an editable word-processing document would require inferring reading order, paragraphs, columns, tables, and styles from page appearance. ToolAcre does not offer that operation or OCR. PDF-to-image deliberately discards semantic structure, while images-to-PDF deliberately wraps pictures in pages.
The absence of an editable conversion is important scope, not a missing toggle. A scan has no text unless another system recognizes it, and recognition still does not recover the original authoring model perfectly. Use a dedicated OCR or document-conversion workflow when editable structure is the actual requirement.
Resolution, colour and size — the settings that decide whether a rasterised page prints cleanly or looks soft
Resolution controls output pixel dimensions and memory use. Screen, Good, High, and Very high correspond to increasing scales in the interface, while the renderer may downscale a page that exceeds its budget. JPEG uses fixed quality near 92 percent; there is no quality slider.
PNG suits small text and line work because it avoids JPEG edge artifacts, although it can be larger. JPEG suits photographic pages when smaller delivery matters. Mixed-size source pages produce mixed-size images at one scale, and a ZIP uses stored entries rather than recompressing already compressed image data.
PDF parsing uses a worker, while canvas encoding and image preparation require main-thread browser APIs
Local conversion does not mean every step runs in one worker. pdf.js parses through its own bundled worker with document JavaScript evaluation disabled, while canvas encoding must stay on the main thread. The loop yields between pages so progress and controls continue to paint.
For images-to-PDF, browser image preparation can decode and draw unsupported formats on the main thread, then pdf-lib assembles prepared PNG or JPEG data in the toolkit worker. These boundaries explain the implementation accurately and replace the outline’s broader claim that rendering and re-encoding simply occur in a Web Worker.
The toolkit specifically offers PDF-to-PNG/JPEG and images-to-PDF
The shipped conversions are specific. PDF to Image accepts one unencrypted PDF up to 50 MB and returns PNG or JPEG pages in a ZIP. Images to PDF accepts supported browser image formats up to 30 MB each and returns one `images.pdf` with one image per page.
The toolkit does not convert PDF into editable office formats, run OCR, stitch all pages into one long image, or preserve text through a PDF-image-PDF round trip. The supported-input sections and operation controls state these exact paths, so there is no need to leave capability ambiguous.
Takeaway — know which kind of conversion you need before you start, then use the PDF Toolkit for the ones it supports
“Convert” can mean rendering structured pages into pixels or wrapping pixels inside newly structured pages. Those directions may look reversible, but they are not: once page text becomes a raster, putting that raster into another PDF does not recreate selectable characters or vector drawing instructions.
Use PDF-to-image when the required artifact is genuinely a picture, and images-to-PDF when the required container is genuinely a paged PDF. Both run locally with explicit worker and main-thread responsibilities, hard input caps, and memory guards. Choosing the correct representation before starting prevents a visually successful but functionally wrong result. Check the result note for automatic reduction or format conversion, because those reported changes may affect print suitability even when the preview appears acceptable. Preserve the structured PDF whenever future search, accessibility, or editing might matter, and create raster derivatives as disposable delivery assets rather than replacements for the source. Name those derivatives with their format and scale so nobody mistakes a screen-resolution image bundle for the document intended for print or archival use.