Developer tools · JSON formatter & validator
JSON string escapes explained: \n, \uXXXX and control characters
· How it works
json developer-workflow validation
A raw newline inside a JSON string is invalid, and so is a tab. This post covers the eight escape sequences, how \u escapes and surrogate pairs work, and why a pasted paragraph can invalidate a whole file.
The paragraph that broke the payload
The paragraph that broke the payload — pasting a visible line break into a quoted description inserts a control character directly into the JSON string. The first line looks complete, but the opening quote still expects string content or a closing quote. When the parser reaches the raw line feed, it stops there because JSON strings cannot span physical lines that way. The text must contain the two-character escape `\n` wherever the decoded value needs a newline.
Tabs copied from a document or spreadsheet cause the same class of failure, even though an editor may render them as harmless spacing. Replace a literal tab with `\t`, a carriage return with `\r` and other forbidden controls with their named or Unicode escapes.
The eight escapes JSON allows
The eight escapes JSON allows — after a backslash, the short forms are `"`, `\\`, `\/`, `\b`, `\f`, `\n`, `\r` and `\t`. They represent a quotation mark, backslash, slash, backspace, form feed, line feed, carriage return and horizontal tab. A slash may also appear unescaped; `\/` exists mainly for compatibility with contexts that once treated a closing script sequence specially.
No other letter may follow a JSON backslash. Sequences familiar from programming languages, such as `\v`, `\0`, `\x41` or a backslash followed by a physical newline, are invalid here. Use `\u` followed by exactly four hexadecimal digits when no short escape exists. This small fixed vocabulary keeps JSON strings portable: a consumer does not need JavaScript, Python or shell-specific escape rules to determine the characters represented by valid text.
Why a literal tab is invalid but a literal é is fine
Why a literal tab is invalid but a literal é is fine — JSON forbids unescaped code points from U+0000 through U+001F inside strings. That range contains tabs, newlines and other controls whose invisible effects can disrupt framing or display. The letter `é` is U+00E9, well outside the control range, so UTF-8 JSON may include it directly between quotes. The same is true for most scripts, symbols and emoji.
Escaping ordinary Unicode is therefore optional, not a cleanliness requirement. `"café"` and `"caf\u00e9"` decode to the same sequence of characters. Direct text is usually easier for people to read, while escapes can help an ASCII-only transport or make a specific code unit visible. Control characters are different: their escape is mandatory.
How \uXXXX escapes work
How `\uXXXX` escapes work — the `u` must be followed by exactly four hexadecimal digits, using 0–9 or A–F in either case. `\u00E9` represents the UTF-16 code unit for `é`, and `\u000A` represents a line feed. Fewer digits, braces such as `\u{1F600}` or a non-hexadecimal letter make the JSON invalid, even if another programming language accepts that notation.
Characters above U+FFFF are represented in this escape form as a surrogate pair. The emoji 😀 can be written as `\uD83D\uDE00`: the high and low surrogate combine after parsing into one Unicode scalar value. JSON grammar can carry an unpaired surrogate escape, but downstream encoders and applications may reject or replace it because it does not identify a complete Unicode character.
Worked example: escaping a Windows path and a snippet of HTML
Worked example: escaping a Windows path and a snippet of HTML — the intended path `C:\Temp\report.txt` needs each backslash doubled in JSON source: `"C:\\Temp\\report.txt"`. Without doubling, `\T` is an invalid escape and sequences such as `\r` or `\t` can silently become control characters instead of path separators. Build the JSON from the intended value, not by guessing which displayed slashes already belong to an outer language.
An HTML fragment such as `<a title="Report">Open</a>` may keep its angle brackets and slash literally, but the attribute quotes must become `\"` inside the JSON string. If a newline separates two tags, encode it as `\n`. The resulting JSON member can be validated and parsed back to the original HTML text.
Where escapes get doubled
Where escapes get doubled — every enclosing text grammar gets its own chance to interpret backslashes. A JSON document containing the decoded string `line1\nline2` must escape that backslash, producing `"line1\\nline2"`. If that JSON text is itself stored as a JSON string, its quotes and both backslashes need another layer of escaping. The apparent clutter records multiple representations, not a special extended form of JSON.
Shells and programming-language literals add their own quoting rules before a JSON parser sees the argument. Diagnose from the inside out: write the exact decoded value first, encode it as JSON once, then encode that complete JSON text for the surrounding shell or source language. At each boundary, inspect what bytes or characters the next parser actually receives.
What this does not cover
What this does not cover — HTML entities such as `"` and URL percent-encoding such as `%20` are separate transformations for separate syntactic contexts. A JSON parser does not decode either form. The string `"""` contains six literal characters after parsing, not a quotation mark, and `"%20"` contains a percent sign followed by two digits, not a space. Apply those encodings only when data crosses into HTML or a URL component.
This discussion also does not replace output encoding. Valid JSON received from an untrusted source can still contain HTML, script-like text or terminal control sequences as ordinary string data. The application that later renders or executes a command must handle that destination safely. JSON escaping protects JSON structure; it is not universal sanitization.
Takeaway: escape what the grammar forbids, nothing more
Takeaway: escape what the grammar forbids, nothing more — double quotes, backslashes and code points below U+0020 need attention inside JSON strings. Ordinary Unicode can remain readable, while `\uXXXX` provides an exact four-digit alternative and surrogate pairs represent characters above U+FFFF. A diagnostic on an apparently blank position often identifies a literal newline, tab or other control character that must be replaced with its textual escape.
Count encoding layers instead of counting slashes by sight. Start from the value the application should receive, encode it once for JSON and only then quote the resulting document for any outer shell, source file or second JSON string. Validate the text presented to the JSON parser and, when correctness matters, inspect the decoded string afterward.