Developer tools · URL encoder & decoder
WHATWG URL Standard vs RFC 3986: why browsers and libraries disagree
· Background
url-encoding standards developer-tools
There are two living definitions of a URL, and they disagree on purpose. This post explains why the WHATWG wrote its own standard, where the two differ in encoding and parsing, and which one your code follows.
The URL that a strict library rejects and the browser happily loads — one string, two verdicts
A string with a backslash in JavaScript might be interpreted smoothly by your browser as part of the URL path. The same string reaches a Python library backend, and it refuses to parse it because backslashes are not allowed. One URL, two different outcomes. Neither is wrong—they follow different standards. The WHATWG standard describes what browsers actually do with real-world URLs, including how they handle malformed input. RFC 3986 defines a formal grammar that URLs should ideally conform to. Many backend libraries built on RFC 3986 enforce that grammar strictly and reject anything outside it.
This divergence matters when you move data between environments. A URL browsers accept might fail validation in a backend tool. Understanding which standard your code implements prevents debugging phantom problems—URLs working fine in one place but mysteriously failing somewhere else for no apparent reason.
Why the WHATWG started over — describing what browsers really do with malformed input rather than what is valid
The WHATWG Working Group formed in 2004 to standardize how browsers actually handle URLs in practice, rather than defining stricter formal rules that browsers would not follow. RFC 2396 described a formal grammar specification, but browsers in practice never followed it exactly correctly. Real-world browsers developed practical rules for tolerating spaces, handling escaped characters, and recovering from malformed input that the RFC did not anticipate or expect.
RFC 3986 arrived in 2005 with formal grammar for well-formed URLs and strict requirements. Browsers implement WHATWG; backend libraries often implement RFC 3986.
Encode sets versus reserved characters — how the URL Standard's per-component lists relate to RFC 3986's categories
RFC 3986 divides characters into three categories: reserved, unreserved, and everything else that must be encoded. Reserved characters like colon, slash, question mark and hash have structural meaning in URLs. Unreserved characters are letters, digits, hyphen, underscore, period and tilde; these are always safe. Everything else gets percent-encoded as bytes. The standard provides one clear rule: know which category your character belongs to.
The WHATWG URL Standard takes a component-based approach. It specifies different encoding rules for scheme, authority, path, query and fragment separately instead of using global categories. An ampersand might be encoded in a path but left alone in a query string. A space is always encoded, but the exact representation varies by context. This per-component design matches browser behavior much better but requires knowing which part of the URL you are encoding.
Error tolerance: spaces, backslashes and tabs — inputs one standard rejects and the other repairs
Spaces must become %20 under both standards, but browsers silently convert literal spaces. Backslashes are forbidden by both standards, yet some browsers treat them as path separators. Tabs, newlines and control characters are prohibited. WHATWG specifies lenient parser behavior: convert or ignore them.
Non-ASCII characters like é or 中 must be percent-encoded using UTF-8 encoding. RFC 3986 does not actually specify the character encoding step itself; it assumes bytes exist but does not say how to get them from text. The WHATWG standard explicitly requires UTF-8: turn the string into UTF-8 bytes first, then percent-encode them. Both standards reach the same encoding result, but they start from different underlying assumptions and are not explicit about the same things.
Worked example: parsing a URL with a backslash and a space in both models — the outputs compared
Take the example string "https://example.com/café\ search". A browser encounters the backslash and sees it as a path character; it sees the space and encodes it to %20, producing something like https://example.com/café%5C%20search. An RFC 3986 parser rejects the entire URL immediately because backslashes are forbidden and spaces are forbidden. The browser continues parsing; the strict parser stops completely. Try another example: "https://user@example.com:80/path?q=a&b=c". Both standards identify the userinfo, host, port, path and query clearly. They agree completely on this structured URL. The disagreement happens only on unusual or malformed inputs.
Open the URL encoder & decoder and compare RFC 3986 mode with browser behavior. Paste a string with spaces, backslashes or other edge cases. The tool shows you exactly how each standard transforms the same input differently. You see immediately which one is stricter and what each one does.
Which one your environment uses — browsers and Node follow the URL Standard; many server libraries follow the RFC, described generally
In browsers, JavaScript uses the WHATWG URL Standard by default. The URL API implements it exactly. Node.js also uses WHATWG. Python libraries tend to implement RFC 3986; urllib follows it closely. Java libraries vary; java.net.URL tends toward RFC 3986. Rust's url crate follows WHATWG. Go's net/url is influenced by WHATWG. This is a general pattern, not an absolute rule.
When you build URLs programmatically and they move between browser and backend, pick one standard and stick to it. Use the browser's URL API for WHATWG. If your backend library is stricter, it is not a contradiction but a design choice.
What this does not cover — hostname parsing, IPv6 literals and IDNA processing
Hostname parsing involves IDNA, punycode and registrar rules that go beyond URL parsing itself entirely. IPv6 addresses, special schemes like mailto: or data:, and empty components are separate topics distinct from percent-encoding completely. Domain length limits and hostname validity vary by registrar and are not relevant to this discussion. Also excluded: relative references and scheme-specific parsing rules. This post focuses only on encoding and parsing differences.
This discussion focuses on the encoding and parsing differences that distinguish these standards. Excluding hostname rules, DNS rules and scheme-specific behavior prevents confusion about percent-encoding rules.
Takeaway: the same URL is valid in one world and an error in another — how the URL encoder & decoder gives you the plain RFC 3986 encoding so you can see what the browser normalised away
The same URL string can be valid under one standard and invalid under the other. Both are correct within their own design goals. When encoding URL components programmatically, use the right tool for your environment. WHATWG describes what browsers actually do; RFC 3986 defines formal grammar. The URL encoder & decoder shows RFC 3986 rules alongside browser behavior so you can see exact differences and choose which fits your situation.
Troubles appear most often when URLs cross browser-to-backend boundaries. Understanding this difference means handling that crossing intentionally rather than accidentally or by mistake.