Developer tools · JSON formatter & validator
Invisible characters that break JSON: BOM, smart quotes and NBSP
· How it works
json developer-workflow validation
When the validator reports an error at the very first character and the file looks perfect, an invisible character is usually to blame. This post explains byte order marks, typographic quotes and non-breaking spaces, and how each one is reported.
Line 1, column 1, nothing wrong to see
Line 1, column 1, nothing wrong to see — a JSON document can begin with a character that occupies a real position but renders with no visible glyph. The opening brace then appears to be first even though a byte-order mark or zero-width character precedes it. A strict parser encounters that hidden code point before it reaches `{`, so reporting the first column is accurate rather than vague. The display and the character sequence simply tell different stories.
Do not delete a brace that looks correct merely because the caret appears beside it. Inspect the code point at the reported offset, enable visible whitespace or switch to a hexadecimal view. ToolAcre does not silently remove a leading BOM before parsing, and its scanner reports the first unexpected character.
The UTF-8 byte order mark
The UTF-8 byte order mark — the byte sequence EF BB BF decodes to U+FEFF at the beginning of a file. Byte order is not ambiguous in UTF-8, so the mark is unnecessary, but some editors and export tools still add it as an encoding signature. RFC 8259 says JSON generators must not add a BOM to networked JSON, although parsers may choose to ignore one for interoperability. That tolerance cannot be assumed across tools.
In a JavaScript string, the BOM is one character even though its UTF-8 representation uses three bytes. ToolAcre reports positions in string characters, so a leading mark appears at line 1, column 1. Configure the editor to save UTF-8 without BOM or remove U+FEFF before distributing the file.
Smart quotes from word processors
Smart quotes from word processors — typographic opening and closing marks look polished in prose, but JSON recognizes only the ASCII quotation mark U+0022 as a string delimiter. U+201C and U+201D are ordinary Unicode characters. Outside a string, they cannot begin a property name or value, so the validator reports the smart quote itself. Auto-correction in chat, email or a document editor often introduces the change after the JSON was originally valid.
Replace delimiters with straight double quotes, then examine apostrophes and quotation marks that belong to the value. Curly quotes are perfectly legal as content inside a correctly delimited JSON string, such as `"She said “go”"`; they fail only when asked to perform the delimiter’s grammatical job.
Non-breaking spaces and zero-width characters
Non-breaking spaces and zero-width characters — JSON whitespace is a deliberately short list: ordinary space U+0020, tab U+0009, line feed U+000A and carriage return U+000D. A non-breaking space U+00A0 may look identical to a normal space between a colon and value, but it is not on that list. A zero-width space U+200B shows nothing at all, yet it remains an unexpected character outside a quoted string.
Web pages use non-breaking spaces to keep words together, and messaging systems may insert zero-width characters for wrapping or script handling. Copying formatted snippets can carry them into configuration. Replace structural NBSP characters with ordinary spaces and remove unintended zero-width characters, guided by the reported offset.
Worked example: a config copied from a chat message
Worked example: a config copied from a chat message — suppose the visible text resembles `{"mode": "safe"}`, but validation fails at the start. A hex view reveals EF BB BF before the brace. Removing that BOM advances the next report to the quote before `mode`, which is actually U+201C. Replacing both smart delimiters with U+0022 then exposes a U+00A0 between the colon and the value.
Change that structural non-breaking space to U+0020 and validate once more. The accepted result can now be formatted normally. This sequence shows why repairing only what the screen appears to show is unreliable: several invisible or look-alike characters can occupy different grammatical positions. Follow each line and column, identify the actual code point, make one intentional replacement and rerun validation.
How to see the invisible
How to see the invisible — enable an editor's render-whitespace option to distinguish tabs from spaces and reveal unusual gaps, then use a Unicode inspector or hex view for characters that still look identical. A UTF-8 BOM appears as EF BB BF, a non-breaking space as C2 A0 and a zero-width space as E2 80 8B. Smart opening and closing quotes appear as E2 80 9C and E2 80 9D.
Match the diagnostic’s coordinate system before counting. ToolAcre scans a JavaScript string, so its columns count UTF-16 code units rather than UTF-8 bytes. A byte-oriented hex editor can therefore show a larger numeric offset after non-ASCII characters. Use the reported line to narrow the search, inspect the neighboring code points and translate only as needed.
What this does not cover
What this does not cover — mojibake such as `café` can be completely valid JSON. The parser sees an ordinary sequence of string characters and has no evidence that UTF-8 bytes were previously decoded as another encoding. Likewise, a non-breaking space or zero-width character inside a quoted value is syntactically valid. Validation catches characters that violate JSON grammar; it cannot decide whether valid Unicode content matches the author’s intention.
Repair encoding corruption at the boundary where bytes become text, using knowledge of the original and mistaken encodings. Do not repeatedly encode and decode a JSON string until it looks better, because that can damage already-correct characters. Application-level normalization is also a separate decision: visually identical Unicode sequences may compare differently while remaining valid.
Takeaway: trust the reported column even when the line looks clean
Takeaway: trust the reported column even when the line looks clean — invisible characters and look-alike punctuation still occupy precise positions in the source. A leading BOM, curly delimiter, non-breaking space or zero-width mark can prevent a parser from reaching the brace or quote that appears correct. Reveal whitespace, inspect code points or bytes and replace the character whose identity conflicts with its grammatical role rather than editing nearby visible JSON at random.
Remember that positions may count characters while a hex tool counts encoded bytes, so compare the surrounding text instead of expecting every offset number to match. Remove a BOM only at the document boundary, convert smart delimiters to U+0022 and replace invalid structural spacing without erasing legitimate Unicode inside strings.