English

Developer tools · URL encoder & decoder

Spaces in file download links: why %20, + and a raw space are not the same

· Why it matters

url-encoding http developer-workflow

Three encoding methods for spaces in download filenames compared
Original ToolAcre vector illustration

A file called 'Q3 report (final).pdf' can be linked three different ways, and only one is reliably correct. This post explains why filenames break download links and how to encode them so every client agrees.

The download that fails for some users and works for others — a filename with spaces and a plus sign

Files like "Q3 report.pdf" work fine when stored locally but fail via download links for some users while others succeed without issue. Raw spaces are invalid in URLs per RFC 3986 specification. Browsers tolerate them in address bars but HTTP clients strictly reject them. Understanding %20, plus signs, and raw spaces is absolutely essential for reliable distribution. The distinction between encoding methods directly affects download success rates across different platforms, various automation tooling, and HTTP client implementations worldwide. Developers must understand this distinction when building download systems. Context matters for encoding choices and system compatibility.

Developers must choose between raw spaces, %20, or plus signs when creating download links. Testing with curl, wget, and Python reveals which clients enforce RFC compliance. Browser downloads succeed due to error recovery, but API integrations fail when encountering unencoded spaces.

Why a raw space is invalid in a URL — and why browsers tolerate it in the address bar but HTTP clients do not

Raw spaces in URLs have historical roots in protocol design. URLs traverse systems treating spaces as delimiters between tokens. A space in a URL could be misinterpreted as its terminator. HTTP clients reading from command lines truncate at first spaces. This foundational design remains in protocol implementations and is unlikely to change.

Browsers tolerate raw spaces through silent conversion to %20 before sending HTTP requests. This user-friendly behavior hides protocol requirements from end users pasting URLs into address bars. Automated systems lack this recovery layer. Scripts fail on URLs with raw spaces. Email clients encounter failures opening such links.

%20 versus + in a path segment — the form-encoding convention that does not apply to paths

%20 versus plus signs represents a fundamental distinction in URL encoding contexts. In path segments, spaces must encode as %20 per RFC 3986. The plus sign is not a space encoding in paths. This convention originated in HTML form encoding where it serves as space encoding in query strings. Developers often incorrectly apply form rules to paths.

Form-encoding conventions allowing plus signs do not apply to paths with different structural requirements. In query strings, ampersands and equals delimit parameters. Using plus for spaces in query values creates no ambiguity since plus is not a delimiter. In paths, plus has no special meaning. Mixing conventions creates broken download links.

Non-ASCII filenames — UTF-8 percent-encoding and object storage keys that store the raw name

Non-ASCII filenames require UTF-8 percent-encoding before safe transmission in URLs. A filename like "Über report.pdf" contains "Ü" (U+00DC) outside ASCII range. UTF-8 encoding converts this to bytes C3 9C. These bytes percent-encode as %C3%9C in URLs. Each UTF-8 byte gets its own triplet, producing longer encoded filenames.

Object storage services like Amazon S3 present interesting cases for non-ASCII filenames. Some systems allow raw UTF-8 bytes in keys while others require percent-encoding. The encoding strategy depends on storage providers and URL usage. URL-based access requires percent-encoded UTF-8. Developers must coordinate storage and URL generation layers.

Worked example: encoding 'Über Q3 report (final)+notes.pdf' for a path — the exact output and why the + must become %2B

Worked example: encoding "Über report (final)+notes.pdf" demonstrates complete encoding. The filename contains spaces, non-ASCII characters, and a literal plus. UTF-8 encoding of "Ü" produces %C3%9C. In path encoding, spaces become %20 (unlike form encoding using plus). The literal plus becomes %2B. Parentheses encode as %28 and %29. Result: %C3%9CberQ3%20report%20%28final%29%2Bnotes.pdf.

Testing with URL encoder & decoder shows exact transformation. Pasting the filename into single-value mode produces correct percent-encoded segments using path rules. The tool preserves path delimiters while encoding only filename components. Visual comparison of input and output makes rules clear and verifiable before production. Compare this with form mode to see context differences.

Content-Disposition and the filename* parameter — a separate encoding for the download prompt, mentioned for completeness

Content-Disposition and filename* parameters represent alternative encoding layers for download prompts. Servers include Content-Disposition headers specifying filenames for download dialogs. The filename parameter uses RFC 2183 encoding while filename* uses RFC 5987 with percent-encoding. Browsers interpret these headers to decide save-file names. The same filename encodes twice with different schemes.

Two encoding layers create opportunities for transcoding errors. URL-encoded and header-encoded filenames might not round-trip correctly if servers and clients disagree. For maximum compatibility, developers should encode filenames in URL paths using %20 and UTF-8 percent-encoding, and set Content-Disposition headers with decoded filenames. This ensures all HTTP clients and browsers work correctly.

What this does not cover — reserved filenames on specific operating systems and storage-provider quirks

Reserved filenames on specific operating systems add complexity to URL encoding. Windows reserves names like CON, PRN, and AUX for devices. Files literally named "CON.pdf" cannot exist on NTFS. macOS has naming conventions and extended attribute rules. Linux is case-sensitive. Valid URL-encoded filenames might not be valid for storage on certain systems.

Storage-provider quirks add complexity to cross-platform distribution. Amazon S3 accepts UTF-8 keys and is case-sensitive. Google Cloud Storage behaves similarly with additional restrictions. Azure Blob Storage has different character rules. Filenames working on S3 might fail on Azure. Architects must check provider documentation and test with real non-ASCII filenames.

Takeaway: encode the segment, not the URL — how the URL encoder & decoder's single-value mode produces a path-safe filename

Takeaway: encode the segment, not the URL—the URL encoder & decoder single-value mode produces path-safe filenames. The tool accepts raw filenames and produces percent-encoded segments. This prevents double-encoding and mixing contexts. Using single-value mode avoids balancing path, query, and fragment encoding rules. Generated segments are safe to insert into URLs.

Best practice encodes filenames where they enter URL construction. Do not assume browsers fix encoding issues. Test with actual HTTP clients used by target users: curl, wget, Python, Java httplib, and browser fetch APIs. Verify filenames survive round-tripping through entire systems. URL encoder & decoder is the starting point ensuring correctness.