English

Developer tools · URL encoder & decoder

How percent-encoding works: from characters to UTF-8 bytes to %XX sequences

· How it works

url-encoding utf-8 percent-encoding developer

Character codes mapped through UTF-8 encoding steps into percent-encoded hex sequences
Original ToolAcre vector illustration

Percent-encoding does not encode characters; it encodes bytes. This post shows how a character becomes UTF-8 bytes and then hex pairs, and why an accented letter takes two %XX groups while an emoji takes four.

Why 'é' turns into %C3%A9 rather than %E9 — the observation that reveals the byte layer underneath

When a junior developer sees %C3%A9 in a URL, percent-encoding operates on bytes, not characters. The character é is not one byte; UTF-8 encodes it as two: C3 A9. The percent-encoding rule from RFC 3986 is simple: encode each byte as a percent sign followed by two hex digits. That distinction transforms the explanation from mysterious to logical.

Understanding percent-encoding requires understanding UTF-8. Text must be converted to bytes using character encoding. UTF-8 is the standard for URLs and web. It expresses characters as variable-length byte sequences: ASCII uses one byte, accented letters use two, emoji use four. Each stage is distinct: character, Unicode code point, UTF-8 bytes, then %XX pairs. Skipping to hex without understanding bytes misses the point.

The percent-encoding rule from RFC 3986 — one % followed by two hex digits per byte, uppercase preferred

RFC 3986 defines one rule: encode each byte as percent followed by two uppercase hexadecimal digits. Unreserved characters needing no encoding are letters, digits, hyphen, underscore, period and tilde. Everything else must be encoded. Spaces become %20, slashes become %2F, and the percent sign becomes %25. This prevents special characters in query values from breaking URL structure.

A space encodes as byte 0x20, becoming %20. A forward slash is 0x2F, becoming %2F. These are ASCII characters needing one byte. Accented letters and emoji differ. The percent sign becomes %25. Reserved delimiters like colons get encoded to preserve structure. This prevents an embedded ampersand or equals in a query parameter from breaking parsing. Each byte becomes %HH.

UTF-8 as the assumed charset — why modern URLs are UTF-8 and where the legacy exceptions are

UTF-8 uses variable-length encoding. ASCII from code points 0 to 127 is one byte. Characters from 128 to 2047, including accented Latin letters, are two bytes. Characters from 2048 to 65535, common in East Asian scripts, are three bytes. Characters above 65535, including most emoji, are four bytes. Each byte is prefixed with bits signaling how many bytes follow.

The accented letter é is Unicode code point U+00E9. UTF-8 encodes it as two bytes: 0xC3 and 0xA9. Percent-encoding produces %C3%A9. German ü (U+00FC) encodes as 0xC3 0xBC, becoming %C3%BC. Spanish ñ (U+00F1) encodes as 0xC3 0xB1, becoming %C3%B1. The pattern is consistent: the first byte signals a two-byte sequence. One accented letter expands by six characters in encoding.

Worked example: encoding 'café 😀' byte by byte — the code points, the UTF-8 bytes and the resulting string

Emoji makes the byte layer obvious. The thumbs-up emoji 👍 is code point U+1F44D. UTF-8 encodes it as four bytes: F0 9F 91 8D. Percent-encoding produces %F0%9F%918D: twelve characters for one symbol. The smiley 😀 (U+1F600) encodes as F0 9F 98 80, becoming %F0%9F%9880. Four-byte sequences become twelve percent-encoded characters.

Mixed text shows why understanding bytes matters. The phrase "café 😀" contains plain ASCII, an accent, and an emoji. Letters c, a, f encode as 63, 61, 66. The é encodes as C3 A9. Space encodes as 20. The emoji encodes as F0 9F 98 80. The result is "caf%C3%A9%20%F0%9F%9880". Understanding which bytes need encoding makes output predictable.

Decoding in reverse — collecting %XX groups into bytes and only then interpreting them as UTF-8

Decoding reverses the process. A decoder scans for %XX pairs and collects them into byte values. Seeing %C3%A9, it extracts bytes C3 and A9. UTF-8 decoding interprets those as the character é. If a sequence is incomplete, like %C3 alone, the result is an error. The decoder knows from UTF-8 prefix bits that C3 requires a second byte.

Case does not matter in hex digits; %C3%A9 and %c3%a9 decode identically. RFC allows uppercase or lowercase, though uppercase is preferred. But case matters for characters: é (as %C3%A9) is not the same as É (as %C3%89). URL comparison must normalize percent-encoding or risk treating identical resources as different. Frameworks normalize before caching.

Why case does not matter in hex digits but does elsewhere — normalisation rules and URL comparison

RFC 3986 mentions punycode for domain names and form encoding for submissions as separate rules. Punycode encodes non-ASCII domain names without percent signs for DNS compatibility. The domain 😀.example becomes "xn--js8h.example". Form encoding modifies percent-encoding with one exception: spaces become plus signs instead of %20. Submitted forms as application/x-www-form-urlencoded use plus for spaces.

The URL encoder tool shows all three modes: component encoding, whole-URL encoding, and form encoding. Component encoding with encodeURIComponent encodes every special character including delimiters, suitable for query values. Whole-URL encoding with encodeURI preserves structural characters for complete URLs. Form encoding is for POST bodies. Each uses UTF-8; they differ only in which bytes get left unencoded.

Punycode and form encoding: sibling standards, not percent-encoding extensions

The byte perspective resolves URL mysteries. Why does one emoji need twelve characters? Because UTF-8 uses four bytes, each becoming %HH. Why do some URLs have %2F for slashes while others have plain slashes? Because the encoding mode decides: a slash in a path segment stays unencoded, but inside a query value it must be %2F to avoid misreading.

Think in bytes for predictable percent-encoding. A character is a Unicode code point. UTF-8 is its byte representation. Percent-encoding is the transmission format. Character expansion happens at the UTF-8 layer. Hex case does not affect decoding but character case does. Invalid byte sequences fail at UTF-8 due to strict prefix rules. The URL encoder tool shows this progression.

Takeaway: think in bytes — how the URL encoder & decoder shows the exact %XX output for any text you paste, in the browser

Worked example: encoding "café 😀". The word café has letters c, a, f as ASCII single bytes: 63, 61, 66. The é is UTF-8 two bytes: C3 A9. Space is 20. Emoji 😀 is four bytes: F0 9F 98 80. Unreserved ASCII letters stay visible. Result: "caf%C3%A9%20%F0%9F%9880". This shows why one emoji expands to twelve characters.

Takeaway: think in bytes, not characters. Percent-encoding is applied after UTF-8 encoding. Each byte becomes %HH. Variable-length UTF-8 means characters expand differently: ASCII becomes %XX (two chars), two-byte accents become %XX%XX (six chars), four-byte emoji become %XX%XX%XX%XX (twelve chars). Paste text into the URL encoder tool and observe the progression.