English

Text & everyday tools · QR & Barcode Toolkit

How much data can a QR code hold? Versions 1 to 40 and capacity explained

· Background

qr-code encoding usability

QR grids increasing from a sparse version to a dense maximum matrix
Original ToolAcre vector illustration

Explains the 40 QR versions, how grid size, encoding mode and error correction combine to set capacity, and why the theoretical maximum is rarely the practical one.

The code that became an unreadable grey square — what happens when you encode a paragraph instead of a link

A long note can turn a compact symbol into a dense field of tiny modules even before the encoder rejects it. That visual change matters because a fixed printed width leaves fewer pixels or printer dots for each added row and column.

Density is visible in both the module counter and physical proof. The generator may accept a long note, yet the fixed-size preview now contains many smaller cells and a later print may blur them. Capacity answers whether a matrix can be constructed, not whether a chosen printer and phone can resolve it. Reduce the payload before reducing margin or cell size, because those changes attack the scanner’s visual evidence.

Versions 1 to 40 — 21×21 modules growing by four per side to 177×177, and what each step adds

The implementation supports QR versions one through forty: 21 modules per side at the first version, then four more per side until 177. Automatic selection chooses the smallest version that fits the UTF-8 byte payload and correction level.

Version dimensions provide a simple observable ladder: 21 modules for version one, then four additional modules per side at each step until 177. The tests verify the `4n + 17` dimension rule across correction levels. ToolAcre asks the dependency for automatic selection rather than exposing a version field, so users should read the resulting module count instead of forcing a table lookup.

Capacity in this toolkit is UTF-8 byte-mode capacity; all-mode published maxima are outside the implementation

Generic capacity tables separate numeric, alphanumeric, byte and kanji modes, but ToolAcre deliberately uses byte mode for every payload. Reproducing other-mode maxima would misdescribe this tool, so capacity here is measured in encoded UTF-8 bytes.

UTF-8 byte length explains why equal character counts can behave differently. Forty ASCII `x` characters use fewer bytes than forty Japanese characters, and the test confirms that the multibyte payload needs a matrix at least as large. The panel displays bytes rather than characters for this reason. Generic numeric or alphanumeric maxima would not predict ToolAcre because it sends every payload through byte mode.

Capacity depends on the selected correction level; use the implementation’s exact byte limits

The verified ceilings are 2,953 bytes at L, 2,331 at M, 1,663 at Q and 1,273 at H. Those are implementation constants, not a promise about practical print quality, and stronger correction leaves less room for payload bytes.

The exact configured ceilings also define the error path: 2,953 bytes at L, 2,331 at M, 1,663 at Q and 1,273 at H. These are valid implementation figures to show in the interface. They should not be converted into character counts because accents and emoji vary in UTF-8 length. When a value exceeds capacity, shortening it is safer than silently dropping content.

Practical ceilings — camera resolution, print size and scan distance shrink the usable range well below version 40

A matrix that technically fits can still be a poor physical design when printed too small or viewed too far away. Camera resolution, quiet zone, contrast and substrate reduce the practical ceiling, which is why a short URL is often preferable to a whole record.

Practical range depends on the whole symbol. Stronger correction can make the matrix larger for identical text, as the 200-character test demonstrates between L and H. At a fixed print width, that shrinks every module. Raising correction is therefore not automatically safer: redundancy may tolerate some damage while increased density makes clean capture harder. Choose a level and then test the resulting physical code.

Worked example: compare generated module counts instead of predicting exact versions from memory

Generate a short URL, a supported vCard and a long plain-text note, then compare returned module counts and scan proofs. The source does not expose a guaranteed version estimator for arbitrary text, so observe the actual output instead of guessing.

For the example set, generate a short HTTPS URL, a supported vCard and a 500-character note at the same correction level. Record byte count and module dimensions instead of predicting exact versions. The vCard adds field labels and separators around the visible contact data, so its encoded length is not merely the sum of what appeared in the form.

What this does not cover — structured append across multiple codes and Micro QR

Structured append and Micro QR are not implemented. ToolAcre also does not split an oversized payload across symbols; it reports that the content is too long and suggests shortening it or choosing a lower correction level.

ToolAcre does not divide data across multiple codes and does not offer Micro QR. An oversized payload returns an error rather than a partial image. If a record is too large, host it behind a stable URL or select specialist tooling whose multi-symbol format and reader support meet the requirement. Manually cutting text into unrelated QR images creates an assembly problem for the scanner and user.

The takeaway — encode a pointer rather than the payload where you can, and let the QR & Barcode Toolkit pick the version for the content you enter

Encode a pointer when the destination can host the larger record, keep long-lived URLs under your control, and let the generator select the matrix. Capacity is a byte budget, while reliable use is a physical-system test.

The operational takeaway is a hierarchy: preserve correct content, remove unnecessary bytes, choose an appropriate correction level, observe the generated matrix, then size and test it. A theoretical maximum is the final boundary of the encoder, not a design target. Short pointers usually leave more room for robust modules and let the destination content change without replacing the print.