English

Developer tools · URL encoder & decoder

From RFC 1738 to the URL Standard: how percent-encoding rules have evolved

· Background

url-encoding rfc-history web-standards

Evolution of URL percent-encoding standards from RFC 1738 through RFC 3986 to the WHATWG URL Standard
Original ToolAcre vector illustration

The rules for escaping characters in URLs have been rewritten several times since 1994. This post follows RFC 1738, RFC 2396, RFC 3986 and the WHATWG URL Standard and explains what changed each time.

From RFC 1738 to the URL Standard—how percent-encoding rules evolved

Tilde (~) exemplifies how encoding rules change across standards generations and deployments. RFC 1738 (1994) required %7E everywhere; RFC 2396 (1998) moved tilde to unreserved allowing unencoded. RFC 3986 (2005) confirmed unreserved status. Old URLs with %7E remain valid; new builders output ~. Evolution reflects deployment lessons as web matured and infrastructure standardized. RFC 1738 was conservative because early infrastructure was heterogeneous and varied.

RFC 1738 codified 1994 browser behavior. Deployments standardized on UTF-8; restrictions proved unnecessary. Later standards relaxed character restrictions. RFC 3986 permits safe decoding of unreserved characters.

RFC 1738 (1994): 'unsafe' characters and the first escaping rules — what was considered dangerous and why

RFC 1738 defined "unsafe" characters as those conflicting with URI syntax (space, slash), historically used in protocols (control characters), or that systems could not safely transmit. Conservative list percent-encoded far more than necessary for modern internet. Many early systems predated RFC; it codified their behavior. Control characters were genuinely dangerous in protocols; spaces were transmission problems for HTTP clients reading from command lines. Modern systems handle these cases more gracefully through explicit encoding.

Testing against RFC 1738 reveals what old systems expected. Encode a character from 1990s URL specification and compare with modern RFC 3986. Differences show what got relaxed. Unreserved set expanded over time. Hyphen, period, underscore were always safe. Tilde needed RFC 2396 to become safe. Conservative approach meant backwards compatibility. Old URLs built under RFC 1738 rules remain valid today. Normalization in RFC 3986 section 6 allows decoding unnecessarily percent-encoded unreserved characters safely.

RFC 2396 (1998) — reserved versus unreserved, the generic syntax, and the tilde rehabilitated

RFC 2396 (1998) clarified character sets more rigorously than RFC 1738. It formalized reserved characters serving URI structure versus unreserved as literal data. Unreserved expanded including hyphen, period, underscore, tilde. It acknowledged generic URI syntax separate from scheme-specific rules. RFC 2396 introduced distinction between gen-delims (:, /, ?, #, [, ], @) and sub-delims (!, $, &, ', (, ), *, +, ,, ;, =). Naming clarifies reserved characters divide into two groups with different structural roles.

RFC 2396 introduced normalization guidance specifying which percent-encoded characters could decode without meaning change. Unreserved character decoding normalizes. Reserved character encoding sticks. RFC 3986 simplified notation further. Standards maintain backwards compatibility fiercely.

RFC 3986 (2005) — ! * ' ( ) move to sub-delims, gen-delims are named, and normalisation guidance arrives

RFC 3986 (2005) is modern reference standard for percent-encoding. It kept reserved/unreserved distinction but simplified notation and added normalization guidance. Tilde moved unambiguously unreserved. Standard clarified percent-encoded unreserved characters can decode without meaning change. RFC 3986 section 3 describes URI syntax with precision. Section 2 defines character categories. Section 6 dedicates formal rules to syntactic normalization. Comparison-based normalization considers URIs identical if normalized forms match. Dot-segment removal from paths normalizes without meaning change.

Normalization matters for analytics, caching, link following. URLs differing only in hex digit case (RFC 3986 prefers uppercase) should be identical in practice. Normalizing prevents duplicate log entries and cache misses. Caches keyed on normalized URLs serve content regardless of requestor encoding preference. RFC 3986 normalization guidance lets systems make consistent decisions. But strict enforcement breaks URLs working fine in current internet.

The WHATWG URL Standard: parsing what browsers actually receive — encode sets, special schemes and error tolerance

WHATWG URL Standard (2016–present) emerged from browser experience with URLs not following RFC 3986 perfectly. Browsers faced unencoded spaces, mixed encodings, quirks. WHATWG describes real browser parsing, not theoretical grammar. Real-world browsers had developed practical rules for tolerating spaces, handling escaped characters, recovering from malformed input. RFC 3986 arrived in 2005 and defined formal grammar, but browsers in practice had already diverged slightly.

WHATWG defines nine encode sets with context-specific rules. Space in path becomes %20; slash in userinfo becomes %2F. Browser applies narrower standard for web. RFC 3986 provides baseline; WHATWG builds on it.

Worked example: one URL with a tilde, a space and a non-ASCII character — how each generation of rules encodes it

Internationalized domains use punycode encoding (München becomes xn--mnchen-3ya). Paths and queries still use percent-encoding. The domain portion uses punycode; path and query portions use percent-encoding. Layers do not mix or interfere.

IDNA (Internationalized Domain Names in Applications) solves hostname problem. Punycode encodes non-ASCII into ASCII for DNS compatibility. The xn-- prefix signals punycode encoding. Algorithm is deterministic: münchen always becomes xn--mnchen-3ya. Non-ASCII characters must convert before DNS resolution. Percent-encoding does not work for hostnames due to DNS constraints and label limits. Each approach solves different problem correctly. Standards evolved separately for good reasons.

What this does not cover — IRIs and internationalised domain names, which have their own history

Choosing standards depends on context. Building URLs for browsers? Follow RFC 3986; browsers apply WHATWG. Older systems? Test implementations. Understanding evolution prevents confusion.

Modern builders should follow RFC 3986 or WHATWG contextually. Old and new URLs coexist requiring careful compatibility thinking. URL encoder & decoder follows RFC 3986 throughout, offering fixed reference separate from browser behavior. WHATWG adds component-specific encoding sets beyond RFC basics. Knowing what changed when helps understand why systems disagree. Testing your URL with both standards reveals which standard controls in your environment. Both standards are correct.

Takeaway: know which rulebook your code follows — how the URL encoder & decoder gives you the RFC 3986 behaviour as a fixed reference point

RFC 3986 normalization permits decoding unnecessarily percent-encoded unreserved characters safely. %41 normalizes to A. Encoded reserved characters like %2F never decode; changing meaning breaks structure. Standards bodies maintain backwards compatibility fiercely. Fixing would require worldwide coordination impossible after decades. Standards books do not retroactively break the web. Changing encoding decisions breaks billions of existing systems simultaneously.

Percent-encoding spans three decades of careful evolution: from RFC 1738 through RFC 2396 and RFC 3986 to the modern WHATWG URL Standard. Modern code should follow RFC 3986 baseline. Old URLs with earlier encoding remain valid. Testing round-trips ensures correctness and compatibility.