English

Text & everyday tools · QR & Barcode Toolkit

Why accented text sometimes scans wrong in QR codes: character sets and ECI

· Background

qr-code encoding browser-processing

UTF-8 bytes entering a QR grid and decoding into accented and Japanese text
Original ToolAcre vector illustration

Explains why the QR standard's default byte interpretation is not UTF-8, what the Extended Channel Interpretation mechanism does, and why some readers show mojibake for accented or non-Latin text.

The name that scanned as 'é' — what mojibake looks like in a decoded QR code and why it happens

Mojibake such as é appears when UTF-8 bytes are interpreted under another character mapping. ToolAcre addresses that failure before matrix generation by using TextEncoder, and its tests round-trip accented, Japanese and emoji examples.

The corruption is silent because the QR can remain structurally valid. A scanner decodes bytes, applies a different character interpretation and shows the wrong text, so finder patterns and error correction all appear to work. ToolAcre’s regression tests compare decoded output with original samples such as `café`, Japanese text and emoji. That catches semantic corruption a visual snapshot of black modules could never detect.

ToolAcre pre-encodes UTF-8 bytes to avoid the dependency’s Latin-1 default behaviour

The underlying QR dependency treats its byte-mode string as Latin-1 pass-through data. ToolAcre converts the intended text to UTF-8 bytes first and maps each byte to one code unit, so the library receives the correct octets instead of corrupting characters.

The conversion wrapper creates a `Uint8Array`, processes it in chunks and makes a binary string whose code units equal the UTF-8 byte values. The library’s Latin-1 pass-through then preserves those values instead of re-encoding the original JavaScript characters. Chunking avoids passing an excessive number of arguments to `String.fromCharCode`, while avoiding a global library mutation keeps other callers isolated.

ECI is background only; this implementation does not claim to emit an ECI header

Extended Channel Interpretation can label character encoding in QR systems, but no ECI emission appears in this implementation. This article therefore does not promise an ECI header or describe one as the mechanism behind ToolAcre’s UTF-8 support.

ECI would be a separate signal to a decoder, but ToolAcre does not request or expose one. Its compatibility strategy is correct UTF-8 bytes plus device testing, not an advertised encoding header. This distinction matters in support: a successful repository round trip proves byte preparation and matrix recovery; it cannot prove every external reader chooses the same character interpretation in every payload context.

Repository tests prove matrix round trips, not behaviour across named third-party camera applications

The repository decodes generated matrices in tests and proves its own byte round trip. It does not test every camera application, so claims about readers guessing UTF-8 or failing on specific platforms require separate device evidence.

The unit decoder used in tests is controlled and valuable for regression, but it is not a catalogue of camera applications. Record results from the devices the audience actually uses, including the decoded string rather than “scan succeeded.” Two apps can both recognise a code while one displays mojibake. Report such a difference as reader compatibility evidence instead of changing ToolAcre’s bytes without understanding the decoder.

Reducing the risk — keeping payloads to ASCII where possible, URL-encoding non-ASCII paths, and testing on more than one phone

Keep payloads concise, prefer ordinary HTTPS URLs when they can represent multilingual content on a web page, and test direct non-ASCII text on supported devices. URL encoding may change a URL’s bytes and must preserve the destination semantics.

A stable URL often reduces this risk because non-ASCII presentation can live on the destination page while the QR payload remains a concise ASCII address. If a URL path contains international characters, preserve its correct encoded destination and test it; blindly percent-encoding or transliterating can change routing. For direct contact or plain text, keep the test matrix small and scan with more than one supported reader.

Worked example: verify UTF-8 round trips in the implementation and test external readers separately

Encode café, 日本 and an emoji in separate test codes, confirm the repository’s decoder returns the original text, then scan exported images with the real applications your audience uses. Record differences instead of generalising from one phone.

Use three separate payloads—`café`, a short Japanese phrase and an emoji—then decode each with the repository test path and selected phone applications. Compare exact Unicode characters, not screenshots or visual similarity. If an app fails, retain the exported code and decoded bytes for diagnosis. Regenerating repeatedly from identical input should produce the same matrix and will not fix a reader interpretation difference.

What this does not cover — kanji mode's Shift JIS specifics and font rendering on the scanning device

Kanji-mode Shift JIS details and font rendering after decoding are outside the implementation. QR stores bytes; the scanner and destination interface decide how decoded characters are presented to the person holding the phone.

Kanji mode, Shift JIS and post-decode font selection are outside the implementation. Even correct Unicode may render with a missing glyph on a device that lacks a suitable font, which is different from receiving wrong characters. Separate byte corruption, decoder interpretation and font display when documenting a failure; they occur at different stages and require different remedies.

The takeaway — test any non-ASCII payload before printing; the QR & Barcode Toolkit generates locally so you can iterate quickly

ToolAcre’s verified claim is strong but bounded: it prepares UTF-8 bytes correctly and round-trips multilingual test strings. A print release with non-ASCII payloads still deserves representative reader testing.

The release criterion for multilingual print is exact recovery on representative readers. ToolAcre supplies verified UTF-8 preparation and local generation, and its byte counter reflects multibyte cost. The publisher must still preserve the tested artifact, avoid unreviewed payload edits and disclose reader requirements where compatibility is narrow. Correct encoding is necessary, but the user experiences the entire decoding, interpretation and display chain.