English

Developer tools · URL encoder & decoder

escape() vs encodeURIComponent: how JavaScript's URL encoding evolved

· Background

javascript url-encoding history

Three generations of JavaScript URL encoding functions
Original ToolAcre vector illustration

JavaScript has had three generations of URL-encoding functions, and the oldest still lurks in production code. This post explains what escape() does wrong, why ES3 added the URI functions and why they preserve ! * ' ( ).

The %u20AC in a legacy log — the unmistakable fingerprint of escape() and the decoding failure it causes

A legacy JavaScript file contains a URL-encoding call using the deprecated escape() function. The output in a log file or error message includes the sequence %u20AC—an unmistakable fingerprint of the deprecated escape() function that nobody else uses. This sequence does not match any standard URL encoding, and a decoder built on RFC 3986 or WHATWG rules will not recognize it. The data cannot round-trip through modern tools. It is a common sign of code that predates ES3 and has not been updated since the 1990s.

The escape() function was designed in the Netscape era, before JavaScript had standards or formal URL encoding rules. It encodes most non-ASCII characters using %uXXXX notation, a four-digit hexadecimal code that nobody else uses and no standard defines anywhere. This made sense for one-off uses within a browser, but it broke compatibility with URL standards and made data impossible to decode elsewhere.

escape() and unescape(): a Netscape-era design — Latin-1 assumptions, the %uXXXX invention and why it never matched any standard

escape() and unescape() assume the input is Latin-1 (ISO 8859-1), the character encoding that predated UTF-8 and Unicode. They convert each character to a hexadecimal code, using %XX for high-bit Latin-1 characters and %uXXXX for everything outside Latin-1. A non-Latin-1 character like an emoji cannot be represented at all. The functions are simple and fast, but they are also completely wrong for any modern use case.

Both functions were added to JavaScript before standards existed. They were deprecated immediately after ES3 introduced proper URL encoding in 1999. They remain in JavaScript for backward compatibility—removing them would break ancient code. But any new code should never use them. They are a legacy relic.

ES3 (1999) adds encodeURI and encodeURIComponent — UTF-8 percent-encoding aligned with RFC 2396

ES3 introduced two functions: encodeURI and encodeURIComponent. Both perform UTF-8 percent-encoding: convert non-ASCII characters to UTF-8 bytes, then write each byte as %HH. Both align with RFC 2396, which was current at the time. RFC 3986 came later and did not change encoding behavior. These functions are still the standard today and should be used.

encodeURI is meant for encoding a complete URI; encodeURIComponent is meant for encoding a component inside a URI, like a query value or path segment. The difference is absolutely critical and easy to misunderstand. encodeURI preserves structural characters like : / ? # @ = & and ;. encodeURIComponent encodes all of those, leaving them safe to embed inside a larger URI.

Why ! * ' ( ) are still left unencoded — the RFC 2396 'mark' characters frozen into the language after RFC 3986 moved them

Both functions leave these characters unencoded: letters, digits, hyphen (-), underscore (_), period (.), tilde (~), and the five punctuation marks ! * ' ( ). The marks came from RFC 2396, which listed them as unreserved "mark" characters. RFC 3986 came out in 2005 and moved those five into a different category, but JavaScript had already frozen encodeURI and encodeURIComponent in 1999. Changing which characters they leave alone would break existing code, so they stayed.

The decision to keep those five marks unencoded for backward compatibility means JavaScript's encoding does not perfectly match either RFC 3986 or the WHATWG standard. It is close enough for practical use, and changing it now is completely impossible. This is a lesson in API stability: once you freeze behavior, you cannot change it even if the standard evolves.

Worked example: the same string through escape, encodeURI and encodeURIComponent — three outputs compared

Take the string "R&D (research) = café's". Run it through escape(), encodeURI and encodeURIComponent. escape() produces "R%26D%20(research)%20%3D%20caf%E9's", mixing unencoded parentheses and apostrophe with percent-encoded ampersand and equals. encodeURI produces "R&D%20(research)%20=%20caf%C3%A9's", leaving the ampersand and equals signs alone because they are structural. encodeURIComponent produces "R%26D%20%28research%29%20%3D%20caf%C3%A9%27s", encoding everything including the parentheses and apostrophe.

Paste the same string into the URL encoder & decoder and switch between encodeURI and encodeURIComponent to see the difference. Then inspect what escape() would produce (you can call it in the browser console, though it will warn you). You see immediately that the three functions produce three completely different results.

Migrating away from escape() — mapping old calls to the right modern function and handling stored %uXXXX data

Old code using escape() must be updated. If escape() was used to encode a URI component, replace it with encodeURIComponent. If it was used to encode a complete URI, use encodeURI. For stored data that contains %uXXXX sequences, you need a custom decoder: convert each %uXXXX to a Unicode code point, then collect the code points into a string. JavaScript's built-in unescape() will read the %uXXXX, but the result might not be correct UTF-8.

After replacing escape(), test the code with strings containing non-ASCII characters, punctuation and special characters. The output should now match what modern tools and standards expect. If your code predates ES3 significantly, it might also use other outdated patterns; a comprehensive audit is worth the effort.

What this does not cover — the URL and URLSearchParams APIs, which are covered separately

The URL and URLSearchParams APIs, added much later, provide higher-level interfaces for URL construction and component encoding. They handle all the escaping automatically and match the WHATWG URL Standard exactly. They are the preferred way to build URLs programmatically in modern JavaScript.

This post covers only the encoding functions, not those higher-level APIs. URL and URLSearchParams parse structure, select component rules and serialize the result, whereas encodeURIComponent transforms one supplied string without knowing where it will be placed. That distinction is the boundary: migrate an old escape() call according to whether it handled a value or an address, then consider replacing surrounding manual concatenation with the structured APIs as a separate refactor.

Takeaway: three functions, one surviving pair — how the URL encoder & decoder shows the modern encodeURI and encodeURIComponent behaviour side by side

Modern JavaScript development should use encodeURI or encodeURIComponent, never escape(). The functions were standardized in 1999 and have not changed since then. They encode non-ASCII characters as UTF-8 bytes and handle standard reserved characters correctly. The URL encoder & decoder tool implements both functions and lets you see their behavior side by side, making it easy to choose the right one for your component.

If you encounter %u sequences in old logs or stored data, they are escape() output and should be migrated. The migration is straightforward once you identify the pattern. Modern code should never produce them.