English

Developer tools · URL encoder & decoder

RFC 3986 reserved and unreserved characters: what the URI standard says

· Background

url-encoding rfc3986 percent-encoding

URI characters categorized into reserved gen-delims, reserved sub-delims and unreserved sets
Original ToolAcre vector illustration

RFC 3986 divides characters into reserved, unreserved and everything else, and that split explains every percent-encoding rule you have met. This post reads the relevant sections plainly.

RFC 3986 reserved and unreserved characters—what matters when you build URLs

RFC 3986 divides characters into three categories: unreserved, reserved, and everything else that must be encoded. Unreserved characters never need encoding—these are letters, digits, hyphen, period, underscore, and tilde. The RFC lists these explicitly in section 2.3, stating they are safe to leave unencoded in any URI context. Testing in URL encoder & decoder with these characters shows they pass through unchanged. Reserved characters subdivide into gen-delims (: / ? # [ ] @) and sub-delims (! $ & ' ( ) * + , ; =), each with structural meaning in different URL components.

When does a character need encoding? Reserved characters must be percent-encoded only where they create ambiguity. A slash marks path segments; in a query value it must be %2F. An ampersand separates parameters; & in a value requires %26. Unreserved characters never need encoding—a hyphen stays hyphen. The URL standard ensures correct parsing. Testing with URL encoder & decoder: entering "hello/world" with encodeURIComponent produces "hello%2Fworld"; with encodeURI it preserves the slash.

Unreserved: letters, digits, hyphen, period, underscore and tilde — the characters that never need encoding and should never be encoded

Percent-encoding uses %HH where HH is hexadecimal notation. ASCII letter A (code 65) becomes %41. Non-ASCII é requires UTF-8 encoding: é (U+00E9) becomes %C3%A9. Modern standards specify UTF-8 uniformly across browsers.

Complete URLs need structural syntax intact; query values need internal reserved characters harmless. Query parameter ?q=R&D should encode the & as %26 if manual, otherwise ampersand becomes separator. Values with forward slashes become %2F in component mode. Component encoding (encodeURIComponent) handles this by encoding everything except unreserved letters, digits and - _ . ! ~ * ' ( ). Testing shows the difference between methods clearly.

Reserved: gen-delims and sub-delims — the two groups, their members and their structural roles

Query strings demonstrate why reserved characters matter. Ampersand separates key=value pairs: ?utm_source=email&utm_campaign=sale means two parameters. Inside a value, unescaped ampersand ends the pair. Equals separates keys from values. Parsing happens at multiple layers; each applies same rules.

Characters requiring encoding in query values include ampersand, equals, hash, question mark, spaces and non-ASCII letters. Hash is sneakiest: #anything becomes fragment identifier, never sent to server. Campaign names ending with hash lose everything after it before request leaves browser. Spaces must become %20. Testing with URL encoder & decoder shows component and form modes. Understanding position determines encoding need.

When reserved characters must be encoded — only where they would be mistaken for a delimiter, component by component

Percent-encoding persists in RFC 3986. Unreserved set stays small ensuring portability. Percent-encoded unreserved characters can decode without meaning change. Decoding %41 to A is correct because A is unreserved. Decoding %2F to / changes meaning when slash is data, not separator. RFC 3986 normalization section 6 covers syntactic approaches.

Reserved characters in different positions have different roles. Colon in scheme marks scheme:authority boundary; colon in userinfo is data. Question mark opens query section; slash in query value is literal. Hash marks fragment start. Position determines encoding need. Query strings carry values that are themselves URIs. Encoding a redirect URL like https://example.com/page?param=value as a parameter requires encoding slashes and colons to %2F and %3A. Context defines safe characters always.

Worked example: classifying every character of a real URL — unreserved, reserved-as-delimiter, reserved-as-data

RFC 1738 (1994) considered many characters unsafe. As deployments standardized on UTF-8, later standards relaxed restrictions. Tilde (~) exemplifies evolution: RFC 1738 required %7E, RFC 2396 (1998) moved tilde to unreserved, RFC 3986 confirmed unreserved status. Evolution reflects deployment lessons. Standards preserve backward compatibility.

RFC normalization permits decoding unnecessarily percent-encoded unreserved characters. %41 safely normalizes to A. Encoded reserved characters like %2F never decode; changing meaning breaks structure. Modern consensus uses RFC 3986 as reference baseline. URL encoder & decoder follows RFC 3986 throughout, offering fixed reference separate from browser behavior. WHATWG URL Standard adds component-specific encoding sets beyond RFC. Standards coexist: RFC 3986 for general URL parsing, WHATWG for web browsers. Libraries differ; check documentation.

Normalisation guidance in section 6 — hex case, unreserved decoding and path segment rules

Testing against RFC 3986 ensures URLs work across software spanning decades. URL encoder & decoder gives RFC 3986 encoding baseline to apply to components constructed. Read standard documentation explaining every encoding decision in URL libraries. WHATWG URL builds on RFC 3986 rather than replacing it entirely. Building URLs for general browsers? Follow RFC 3986; browsers apply WHATWG rules on top. Older systems? Test actual implementations. Normalizing for storage? Apply RFC 3986 consistently. Understanding reserved/unreserved distinction tells you safe characters.

URL encoding is not security sanitization. Each context—SQL, HTML, JavaScript, URI—needs its own output encoding. Percent-encoding protects URL structure only. Apply right defense at right layer.

What this does not cover — the WHATWG URL Standard's different encode sets and IRI handling

RFC 2396 (1998) clarified character sets more rigorously than RFC 1738. It formalized reserved characters serving URI structure and unreserved as literal data. Unreserved expanded including hyphen, period, underscore, tilde over original definitions. RFC 2396 introduced distinction between gen-delims (:, /, ?, #, [, ], @) and sub-delims (!, $, &, ', (, ), *, +, ,, ;, =). Each group has different structural roles in URLs. Naming clarifies reserved characters divide into two groups. Knowing names helps technical discussions.

RFC 3986 (2005) is modern reference. It kept reserved/unreserved distinction but simplified notation. Standards bodies do not retroactively break web. Encode deliberately knowing your standard. URL encoder & decoder provides RFC 3986 reference.

Takeaway: the standard is short and precise — how the URL encoder & decoder's two modes correspond to encoding data versus preserving delimiters

Choosing standards depends on context. Building URLs for general browsers? Follow RFC 3986; browsers apply WHATWG rules. Older systems? Test actual implementations. Normalizing for storage? Apply RFC 3986 consistently. Percent-encoding rules evolved from conservative RFC 1738 through clarified RFC 2396 and RFC 3986 to layered WHATWG URL Standard. Each generation reflected experience. Modern builders follow RFC 3986 or WHATWG contextually. Old and new URLs coexist requiring compatibility thinking. Understanding categories prevents encoding mistakes.

Verify URLs encode correctly before deployment. URL encoder & decoder demonstrates RFC 3986 rules end-to-end. See exact hex values and understand which characters encode. Use this tool when building URLs by concatenating pieces. RFC 3986 reserved and unreserved categories partition character sets for consistent URI parsing.