English

Developer tools · JSON formatter & validator

How a JSON validator finds the exact line and column of an error

· How it works

json validation developer-workflow

A caret positioned under the unexpected key in a small JSON document
Original ToolAcre vector illustration

Browser engines report JSON.parse failures differently, and some give only a character offset. This post explains how a validator turns that into a line and column, and why the position marks where parsing stopped rather than where you made the mistake.

The error message that tells you nothing — why 'Unexpected token in JSON at position 1432' is useless in a 400-line file

An error such as “Unexpected token” is frustrating in a long configuration because it offers no location you can open in your editor. JSON.parse is the browser’s authoritative parser, but its diagnostic text differs among JavaScript engines and versions. ToolAcre does not guess the location by matching an unstable English error string. If JSON.parse fails, a separate strict scanner walks the original text to identify the first character the JSON grammar cannot accept.

What a JSON parser actually does as it reads — a walk through tokenising and the recursive-descent grammar that consumes one value at a time

JSON has six structural characters—braces, brackets, colon and comma—and values that can be strings, numbers, arrays, objects, true, false or null. A scanner needs to know whether it is inside a quoted string before calling a comma a separator: {"note":"A,B"} has one value, not two. It steps through a value or an object member and checks what can legally follow next. RFC 8259 defines this grammar, and unlike JavaScript object literals it does not permit comments or trailing commas.

From character offset to line and column — counting newlines up to the failure offset, and why CRLF endings and multi-byte characters complicate the count

A scanner usually starts with a zero-based offset in the original JavaScript string. To make that useful, count line breaks before the offset and find how far the failure lies from the last break. CRLF should be treated as one visual line ending, not two lines; positions in JavaScript strings count UTF-16 code units, not UTF-8 bytes on disk. A non-BMP emoji can occupy two code units in an editor that visually shows one glyph. The UI reports a line, column and excerpt so you can check the caret against the file you pasted.

Where parsing stops is not where the mistake is — a missing comma is reported at the next key, and a stray quote can push the error many lines down

The first impossible token is often after the original mistake. In an object, forgetting a comma after true makes the quote beginning the next property illegal: the parser was expecting a comma or closing brace. An unterminated string may cause the error to appear at a later line break or end of input. Read backward from the reported point to find the missing delimiter; do not assume the character under the caret must be deleted.

Worked example: a config with one missing comma — the reported position, the surrounding tokens, and how to walk backwards to the real cause

Try the literal three-line document {"name":"demo", followed by "enabled":true on line two and "port":8080} on line three, with no comma after true. ToolAcre reports line 3, column 1, offset 31: it expected a comma or } after the preceding property, and it shows a caret under the first quote of "port". Insert a comma at the end of line two, then validate again. This is a diagnostic of the first syntax obstacle, not a judgment that the word “port” is wrong.

How browser engines differ — V8, SpiderMonkey and JavaScriptCore phrase the same failure differently, which is why a consistent line-and-column report helps

V8, SpiderMonkey and JavaScriptCore have used different wording and sometimes different contextual snippets for the same JSON.parse failure. ToolAcre’s scanner supplies its own structural reason and location when the native parser rejects the value. If the scanner ever disagrees with JSON.parse, the tool returns the engine error rather than fabricating a position. That fallback is safer than confidently pointing at a guessed character.

What this does not cover — semantic problems such as wrong types, missing fields or schema violations, which a syntax validator will never flag

A syntactically valid object can still be wrong for your application: a missing required field, an age written as text, two duplicate keys or a reference to a nonexistent file are not automatically invalid JSON. RFC 8259 says member names should be unique for interoperability, but mere parsing does not enforce your API schema. Validate syntax here and validate semantic constraints in the program that consumes the document.

Takeaway: read the position as 'the first token the grammar could not accept' — and how the JSON formatter & validator reports that line and column without uploading the text

Treat the reported position as “the first token this grammar could not accept.” Work backward to the cause, fix one issue and rerun. JSON formatter & validator does this locally without uploading a pasted configuration. Do not paste real production credentials into any public website if an offline editor can diagnose the file instead.