English

Developer tools · URL encoder & decoder

What the URL constructor encodes for you: the browser's percent-encode sets

· How it works

url-encoding javascript whatwg web-apis

The WHATWG URL Standard encoding sets applied differently across URL path, query, and fragment components
Original ToolAcre vector illustration

The URL API silently percent-encodes some characters and leaves others alone, depending on which part of the URL they land in. This post explains the WHATWG encode sets and how to predict the output.

The space that became %20 and the | that stayed — a concrete case where new URL() partially encoded a path

When new URL("https://example.com/hello world") runs, the space becomes %20 silently. But new URL("https://example.com/hello|world") leaves the pipe untouched. This difference is not random. The WHATWG URL Standard defines separate character sets to encode for each URL component: path, query, fragment, and userinfo each have their own rules. Understanding these sets means predicting what the constructor will do.

The space needs percent-encoding because it is unsafe over HTTP and breaks readability. The pipe is different: not a reserved character that splits structure, so the browser leaves it alone. The line between safety and readability is drawn by WHATWG, not guesswork. Testing "hello world" shows encoding; testing "hello|world" reveals the boundaries for each URL part.

One URL, several encode sets — path, query, fragment and userinfo each have their own list of characters to escape

One URL contains multiple regions, each with its own encoding rules. The path follows one set, query another, fragment a third, userinfo a fourth. A space becomes %20 in path and query. An equals sign stays in query, where it separates keys and values, but encodeURIComponent turns it to %3D. The URL constructor knows its context and applies the correct rules for each part.

The encode sets are precise and small. Path has its own character list; query has a similar but different list. This reflects which characters have structural meaning. A forward slash splits path segments, so encodeURIComponent encodes it as %2F. In the fragment, a forward slash can exist without breaking anything. Understanding WHATWG rules means predicting output without running code.

Why percent-encoding is one-way: what stays encoded stays that way

The URL constructor performs one-way normalisation. Pass "%20" to new URL and it produces %20 unchanged. The constructor recognizes it as already encoded and leaves it alone. This is why double-encoding matters: encode once, pass through the constructor, and the encoding sticks. The constructor does not decode, re-interpret, and re-encode; it reads forward.

This one-way property affects applications trusting URL.href as canonical. If you concatenate user input with your path, the input gets normalised but not decoded. A value like "my+file" stays as is, or becomes "my%2Bfile" in some contexts. Later code using decodeURIComponent might read plus as space if from form data. The constructor normalises once; after that, your value is fixed.

Worked example: passing the same messy string through new URL() and reading href, pathname and searchParams — three different views

Take "hello world&foo=bar|test#anchor" and put it through new URL with different components. The space becomes %20 everywhere. The ampersand in the path stays (no structural meaning there), but in query it stays too (separates parameters, so normalising would lose the boundary between "q=" and "foo=bar"). The pipe and hash behave differently by location.

Reading href, pathname, and searchParams shows three different views. pathname shows the encoded path without scheme, host, or query. searchParams gives decoded parameters, so "hello+world" from form data becomes a space. The search property preserves the literal string. href shows the complete normalised URL. These coexist on one object; which to use depends on your next step.

URLSearchParams and the form-encoding rule — why it produces + for spaces while pathname produces %20

URLSearchParams applies form-encoding: space becomes plus, not %20. new URLSearchParams({q: "hello world"}) produces "q=hello+world", not "q=hello%20world". This is the historical application/x-www-form-urlencoded rule. But if you pass this string as raw query to new URL, the plus stays plus; only URLSearchParams decodes it as space. The constructor is faithful to what it sees.

This plus-sign difference causes common bugs. A URL from the address bar uses %20 for spaces. Form data uses plus. If you decode with decodeURIComponent (which reads plus literally) instead of URLSearchParams.get, spaces become plus characters. The URL encoder & decoder shows both: paste "hello+world" and compare component and form modes to see where space appears.

Comparing with encodeURI — where the two agree and where they diverge

The URL constructor and encodeURIComponent are different tools. encodeURIComponent encodes almost everything except unreserved letters, digits, and - _ . ! ~ * ' ( ). It assumes no context. The URL constructor parses an actual URL and applies WHATWG rules per component. encodeURIComponent turns "hello/world" into "hello%2Fworld"; new URL sees slashes as path separators. Same input, different output.

Use encodeURIComponent when building a URL by concatenating pieces. Use URLSearchParams or the URL constructor for complete or partial URLs. Do not use encodeURIComponent on a whole URL; you will mangle the scheme. Compare the result with intent. The browser enforces URL structure opinions, and new URL implements them. The URL encoder & decoder shows both views side by side.

What this does not cover — host parsing, IDNA and special versus non-special schemes

The WHATWG URL Standard is the source of truth, though reading it requires patience. The encode sets are defined in algorithm fragments, not plain lists. In practice, understanding the principle matters more than memorising sets. Path allows more characters (slashes are structural); query has its own rules; fragment has fewest restrictions (handled client-side, never sent to servers). Each component has its own rules; knowing this tells you where to look.

Normalisation and validation are different boundaries. The constructor normalises: cleans percent-encoding, applies component rules, gives canonical form. It does not validate: invalid characters throw, but empty hosts are accepted. The constructor is strict about format but lenient about interpretation. For exact spec compliance, read WHATWG's percent-encoded bytes section. For everyday building, use URLSearchParams, the URL API, and real examples.

Takeaway: the parser has opinions — how the URL encoder & decoder shows you the plain percent-encoding of a value or an address so you can compare it with what the browser produced

Unsupported WHATWG features here include host parsing with IDNA conversion (international domain names to ASCII) and special versus non-special scheme handling. File: URLs use double-slash authority; data: URLs do not. The constructor enforces these rules. Converting hostnames and determining special status belongs to spec-reading, not percent-encoding. This matters when building URLs across different schemes.

Test your URL construction by comparing browser interpretation with expectations. Build with new URL, read the properties that matter: href for complete form, pathname for path, search for raw query, searchParams for decoded. If output surprises you, paste into URL encoder & decoder and follow transformation step by step. The tool shows normalised output alongside raw encoding, revealing the difference. Understanding WHATWG sets means understanding browser choices and how to work with them.