English

Developer tools · URL encoder & decoder

Anatomy of a URL: scheme, authority, path, query and fragment explained

· Background

url-structure web-standards developer-tools

Five URL components labeled with their delimiters
Original ToolAcre vector illustration

Every percent-encoding decision depends on which part of the URL a character sits in. This post names the five components from RFC 3986, shows what each delimiter means and why the fragment never reaches the server.

Why the same # is fine in one place and destroys a link in another — a component question, not a character question

A hash character (#) in a URL means completely different things depending on where it appears. Inside a query string value like ?search=C%23sharp, it must be encoded to %23 to be safe. At the end of a URL like https://example.com/page#section, it marks the fragment delimiter and everything after it is a fragment. One character, two contexts, two different meanings. This is why encoding decisions depend on knowing which part of the URL you are working with.

The question mark (?) also has dual nature. Inside a path or query value, it must be encoded as %3F to appear as data. As a literal character between path and query string, it is structural syntax. Understanding these five components—scheme, authority, path, query and fragment—is the foundation of correct URL handling.

The five components: scheme, authority, path, query, fragment — and the delimiters that separate them

RFC 3986 formally defines URLs as having five components: scheme, authority, path, query and fragment, separated by specific delimiters. The scheme comes first, followed by ://, then the authority, then /, then the path, then ?, then the query, then #, then the fragment. Not all components appear in every URL. A minimal URL might have only scheme and path, like "mailto:user@example.com". A full URL includes all five.

Each component has its own syntax rules. A colon is reserved in scheme, slash in path, ampersand in query. Reserved characters need encoding when they appear as data: slash in path values becomes %2F.

Inside the authority — userinfo, host and port, and why @ and : are reserved there

The authority component contains the network address of the resource: username, password, hostname and port. The format is [userinfo@]host[:port]. Userinfo and host are separated by @; host and port are separated by :. These @ and : characters are reserved within the authority to delimit these subcomponents. If you have a username containing an @ symbol, it must be percent-encoded before you concatenate it. For example, "user@email.com:password" as a username would become "user%40email.com:password" before the final @.

The hostname can be a registered domain like example.com, an IP address in dotted decimal like 192.0.2.1, or an IPv6 address wrapped in brackets like [::1]. The port is optional; if omitted, the scheme determines the default (80 for http, 443 for https, etc.). The userinfo part is rarely used in modern URLs but remains part of the syntax.

Path segments and the meaning of / — hierarchical structure and dot-segment resolution

The path is a sequence of segments separated by forward slashes. The path /a/b/c has three segments: a, b and c. Each segment can contain unreserved characters, percent-encoded characters, or certain reserved characters that are safe in this context. A slash within a segment must be encoded as %2F to avoid confusion with segment separators. The path is hierarchical; it implies that a is a location, then a/b is more specific.

Paths also support special dot segments: a single dot (.) means "current directory" and two dots (..) mean "parent directory". A path like ../../etc/passwd resolves upward. Modern URLs and HTTP avoid using these, but they exist in the syntax. A path segment containing a literal dot or double-dot must be percent-encoded if the literal dot-meaning is not intended.

Query and fragment — key-value conventions, and why the fragment stays in the browser

The query string follows the path and starts with ?. It is traditionally a series of key=value pairs separated by &, although the syntax is actually unstructured—anything can go in a query. If you have a value containing & or =, those characters must be percent-encoded so they are not mistaken for delimiters. The query is sent to the server; the server decides what to do with it.

The fragment follows the query and starts with #. Everything after # is a fragment, and it never reaches the server. The browser handles the fragment locally, usually to jump to a named anchor or to indicate state within a single-page application. Because the fragment never reaches the server, a URL with a different fragment is considered to point to the same resource.

Worked example: dissecting a long real-world URL — labelling every component and every delimiter

Take the URL "https://user:pass@example.com:8080/path/to/page?search=hello&sort=date#results". The scheme is https. The authority is user:pass@example.com:8080, split into userinfo (user:pass), host (example.com) and port (8080). The path is /path/to/page, with segments path, to and page. The query is search=hello&sort=date, containing two parameters. The fragment is results. Enter this URL into the URL encoder & decoder to see how the tool labels and encodes each component.

If search contained &, like ?search=R&D, it becomes ?search=R%26D when encoded properly. Percent-encoded characters do not create visual boundaries in text, so careful encoding and decoding is absolutely essential for correct parsing.

What this does not cover — relative reference resolution and special schemes such as mailto: and data:

Relative references like "../page" or "?query=value" are valid inside HTML and interpreted relative to the current document, but they have their own separate resolution rules. Special schemes like mailto:, data:, and file: follow completely different rules and are not standard absolute URLs.

This post focuses only on the standard absolute URL structure demonstrated by the inspector. Relative references need a base URL before their components can be interpreted, while schemes such as mailto and data do not share the same authority-and-path shape. Keeping those cases separate prevents a rule learned from an HTTPS address from being applied blindly to syntax with different delimiters, resolution steps or transport behavior.

Takeaway: know the component before you encode — how the URL encoder & decoder's single-value and whole-address modes map onto this structure

The component you are encoding determines which characters need escaping and which are safe. A slash is literal syntax in the path, so a slash in a value must be %2F. In queries, & and = must be encoded if they appear in values. The URL encoder & decoder has two modes: "component" for encoding a single value, and "whole address" for a complete URL. Use component mode when building URLs by concatenating parts; use whole-address mode to verify an existing URL.

Knowing the five components and their delimiters lets you choose correctly every time you encounter an encoding task.