English

Developer tools · URL encoder & decoder

Punycode vs percent-encoding: how non-ASCII domains and paths are handled

· Background

internationalization punycode url-encoding

Domain names converted to punycode while paths stay percent-encoded
Original ToolAcre vector illustration

A URL with a non-ASCII hostname and a non-ASCII path uses two completely different encodings. This post explains IDNA and punycode for the host, percent-encoding for everything else, and why the split exists.

The address that shows 'münchen.example' in one browser and 'xn--mnchen-3ya.example' in another — one host, two spellings

The city München appears in a German domain name. In your browser's address bar, you might see münchen.example displayed normally. Copy the address from a different application, and it appears as xn--mnchen-3ya.example, an ASCII-only string that looks nothing like German text. One URL, two spellings, both absolutely valid. Neither is wrong; they represent the same domain using completely different character sets. The difference reflects a fundamental constraint on how DNS works and how Internet infrastructure expects hostnames to be transmitted.

Path segments like /café/ need encoding but use a different system. Non-ASCII é becomes %C3%A9 in paths. Why the difference? DNS constraints require punycode for hostnames.

Why hostnames cannot use percent-encoding — DNS labels, allowed characters and length limits

DNS labels, the individual segments of a hostname separated by dots, have very strict rules. They can only contain ASCII letters, digits, hyphens and underscores. They have length limits: each label can be at most 63 octets, and the full hostname cannot exceed 255 octets. These are hard constraints from the DNS protocol itself, defined decades ago before international domain names were even a concept. Percent-encoding cannot work for hostnames because the resulting string would likely exceed label limits on longer words.

More importantly, DNS is a global system operated by routers and servers worldwide. Not all of them understand UTF-8 or Unicode. A percent-encoded character like %C3%A9 is still three ASCII characters, so it fits the DNS constraints. But that approach means every lookup has to percent-encode on the way in and decode on the way out, adding complexity to the protocol layer itself. A better solution was needed for hostnames specifically.

IDNA and punycode in outline — the xn-- prefix and the bootstring algorithm, described qualitatively

IDNA, the Internationalized Domain Names in Applications specification, solves the hostname problem by encoding non-ASCII domain names into ASCII that DNS can handle. The encoding used is called punycode, a compression algorithm that turns Unicode text into ASCII using the prefix xn-- followed by a bootstring-encoded representation. The algorithm is deterministic: münchen always becomes xn--mnchen-3ya every single time. Any non-ASCII hostname must be converted this way before DNS resolution can happen.

The xn-- prefix signals to DNS and to IDNA-aware software that the following characters are punycode, not literal ASCII letters. A domain like example.xn--mnchen-3ya.com is understood to mean example.münchen.com by IDNA-aware software. Punycode uses only ASCII letters, digits and hyphens, so it fits cleanly into DNS labels without problems. The algorithm compresses the non-ASCII information into this ASCII representation.

Paths, queries and fragments stay percent-encoded — UTF-8 bytes to %XX, as elsewhere

Everything else in a URL—the path, query string, fragment—uses percent-encoding instead. A non-ASCII character is first converted to UTF-8 bytes, then each byte is written as %HH where HH is hexadecimal. The path /café/ becomes /caf%C3%A9/. The query string ?name=josé becomes ?name=jos%C3%A9. Percent-encoding is standard everywhere on the web: in HTTP request URLs, in HTML forms, in APIs. It does not need special handling by DNS or routers.

Percent-encoding also allows other special characters to be represented safely. A space becomes %20, a slash (if it must appear inside a value) becomes %2F, and so on. The scheme is consistent and universal. It is not used for hostnames because DNS does not understand URLs or percent-encoding; it only understands ASCII labels.

Worked example: one URL with both — the host converted to punycode, the path percent-encoded, side by side

Take the URL "https://münchen.example/café?city=münchen". The hostname münchen must be converted to punycode before DNS lookup: https://xn--mnchen-3ya.example/café?city=münchen. But wait—the path and query also have non-ASCII. Convert those too: https://xn--mnchen-3ya.example/caf%C3%A9?city=m%C3%BCnchen. Now the hostname is punycode, the path and query are percent-encoded. A browser displays the original Unicode version for readability; the HTTP request carries the encoded version.

In the URL encoder & decoder tool, paste a path containing non-ASCII text and compare the single-value mode (just the path) with the whole-address mode (full URL). The tool shows you the percent-encoded result for the path. The hostname, however, requires separate punycode conversion; most encoding tools do not handle that inline, so read it from the tool documentation.

Homograph attacks and why browsers sometimes show punycode — the security reasoning behind the display rules

A malicious actor could register a domain using Cyrillic letters that look visually identical to Latin letters, like "https://xn--80akhbyknj4f.example" (a Cyrillic version of "example.example" in punycode). If a browser displays it decoded as Cyrillic text, users might not notice the difference. To prevent homograph attacks, browsers sometimes show the punycode version instead of decoding it. A warning appears: this domain is all or mostly non-ASCII, and you may not recognize the characters.

The URL encoder & decoder is a tool for encoding and decoding, not for security assessment. If you are working with international domain names, be aware that the punycode representation is what the network sees.

What this does not cover — running the punycode algorithm by hand or the IDNA 2003 versus 2008 differences

IDNA has gone through multiple versions over time: IDNA 2003 and IDNA 2008 handle certain edge cases differently, particularly around normalization and which Unicode characters are allowed by specification. Some older systems still use IDNA 2003 while others have migrated to IDNA 2008 for better compliance. The differences matter significantly if you are building systems that must be compatible across multiple versions. Check your system requirements carefully always.

Punycode uses bootstring compression. Implementations exist in common languages, but verify IDNA policy with your hostname system. Test resolution and display behavior rather than assuming.

Takeaway: two encodings for two jobs — how the URL encoder & decoder handles the percent-encoded parts, and why a percent-encoder is the wrong tool for the hostname

Hostnames need punycode because DNS is an old protocol that only understands ASCII labels and has strict length and character constraints. Paths, queries and fragments use percent-encoding because it is universal on the web and does not have those constraints. They are two separate solutions for two completely different problems. When you encounter a non-ASCII URL, the hostname gets punycode conversion first, then the rest uses percent-encoding.

For most development work, your framework or library handles this conversion behind the scenes automatically. But understanding why two different encodings exist prevents confusion when debugging international URLs or implementing your own URL handling code successfully.