English

Developer tools · URL encoder & decoder

Plus versus %20: the history of application/x-www-form-urlencoded

· Background

url-encoding html-forms http-standards

Form submission showing plus sign encoding space in query string versus percent-twenty in URI syntax
Original ToolAcre vector illustration

Forms encode a space as + while the URI standard says %20, and the reason is historical. This post traces the convention from early HTML forms to today's WHATWG definition and explains why it never went away.

Plus versus %20—why forms and URIs encode spaces differently

HTML forms submitted as GET encode spaces as plus signs in the query string. The same space becomes %20 in URLs following RFC 3986. Both are correct because they follow different standards. Form fields containing spaces become name=value+with+spaces in form encoding but %20 in RFC 3986. The plus-percent-twenty distinction marks which standard applies to your data.

Form encoding uses plus for spaces as historical convention from RFC 1866 (1995), HTML 2.0's original form submission definition. GET requests encoded spaces as plus signs, reserving plus for literal + encoded as %2B. This rule applied only to application/x-www-form-urlencoded, not general URI syntax. Billions of server frameworks became dependent on this convention. Testing both shows clear differences: form mode produces plus; URI mode produces %20. URL encoder & decoder offers both modes to compare directly.

Early HTML forms and GET submission — how the form encoding was defined and why + was chosen

RFC 1866 (1995) defined form submission where spaces become plus and literal plus becomes %2B. This applied only to application/x-www-form-urlencoded. RFC 3986 specified %20 for general URI syntax. Two standards coexisted deliberately.

RFC 2396 clarified character sets more rigorously than earlier standards. It formalized reserved characters serving URI structure versus unreserved as literal data. Standards bodies codify browser and proxy behavior as it evolves. RFC 3986 came later without changing encoding behavior, just clarifying notation. All current browsers standardize on UTF-8 encoding. HTML forms via submit buttons send application/x-www-form-urlencoded format with plus for spaces. Manual URI construction uses %20. Understanding both standards prevents integration surprises.

RFC 1866 and later HTML specifications — where the rule was written down and how it diverged from URI syntax

WHATWG URL Standard specifies URLSearchParams.toString() produces application/x-www-form-urlencoded output with plus signs for spaces. URL constructor percent-encodes following RFC 3986. Browser navigating to URL with space encodes %20; form submitted as GET encodes plus. These fundamentally different tools serve different purposes. Manual encodeURIComponent gives %20 for spaces—RFC 3986 style. Forms submitted to same URL send plus. Servers parsing form submissions expect plus; receiving %20 causes silent parameter failures.

Testing both reveals server-side assumptions you depend on. JavaScript URLSearchParams provides safe form-style encoding hiding plus complexity. Construct URLSearchParams, append entries, call toString() to get application/x-www-form-urlencoded with proper plus. Alternatively build query strings with encodeURIComponent; you get RFC 3986 %20. Never mix approaches. Query strings with manual plus and encodeURIComponent create ambiguity. Receivers cannot distinguish if plus means space or literal plus. Standard approaches handle consistently.

The URL Standard today — application/x-www-form-urlencoded as a separate serialiser with its own rules

JSON APIs typically reject plus as space, expecting %20 per RFC 3986. Clients sending plus fail silently: parameters vanish. Testing APIs with both encodings reveals which standard they accept. URLSearchParams in JavaScript handles form encoding. URL encoder & decoder produces RFC 3986 %20.

HTML form submission automatically handles encoding. Your server framework decides which rules apply. Rails, Django, PHP all treat plus as space in received form data automatically. But manually building query strings for same endpoints matters hugely. A plus uploaded creates ambiguity. Specification compliance and real-world server behavior diverge slightly. Document which standard your endpoints expect. Test both encoding styles. Defensive code handles both gracefully.

Worked example: the same form field seen as a query string and as a request body — with + in one place and %20 in another

JavaScript URLSearchParams applies form-encoding: space becomes plus, not %20. new URLSearchParams({q: "hello world"}) produces "q=hello+world", not "q=hello%20world". This is historical application/x-www-form-urlencoded rule built into JavaScript specifically. But passing this string as raw query to new URL keeps plus as plus; only URLSearchParams decodes it as space. Constructor is faithful to what it sees. Plus-sign difference causes common bugs when mixing functions incorrectly.

URL constructor and encodeURIComponent are different tools. encodeURIComponent encodes almost everything except unreserved letters, digits, and - _ . ! ~ * ' ( ). It assumes no context. URL constructor parses actual URL and applies WHATWG rules per component. encodeURIComponent turns "hello/world" into "hello%2Fworld"; new URL sees slashes as path separators. Same input, different output. Use encodeURIComponent when building URLs by concatenating pieces. Use URLSearchParams or URL constructor for complete or partial URLs.

Why it cannot be fixed — decades of servers and clients that depend on the current behaviour

Percent-encoding rules evolved from RFC 1738 (1994) through RFC 2396 (1998) to RFC 3986 (2005). Each generation clarified ambiguities. RFC 1738 was conservative, treating characters unsafe because early web had limited character support. Deployments standardized on UTF-8, implementations became consistent. Later standards relaxed restrictions on characters proving safe across systems. Modern consensus: UTF-8 everywhere. Standards bodies maintain backward compatibility fiercely. Fixing would require worldwide coordination—impossible after three decades. Two standards coexist deliberately.

Testing with both plus and %20 reveals server assumptions. Server logs show what clients send. Forms use plus; manual URLs use %20. Choose by context and follow API documentation.

What this does not cover — multipart/form-data and JSON bodies

Testing both encodings reveals server behavior. Send a+b both ways. Many production servers expect form encoding; newer APIs expect %20. Your choice depends on receiver expectations. URLSearchParams handles form encoding; encodeURIComponent handles RFC encoding.

Never combine encoding methods. Value encoded with encodeURIComponent %2B then passed to URLSearchParams gets double-encoded as %252B. Decoding once yields %2B instead of plus. Character becomes literal percent-two-six string rather than plus sign. Check intermediate steps in your build process. Encoding happens exactly once per value only. Document which encoding standard your pipeline uses. Test with special characters including plus, space, ampersand.

Takeaway: two standards, both correct in their context — how the URL encoder & decoder gives you the RFC 3986 form, with %20 for spaces, so you know which one you are looking at

Plus-versus-twenty split is not a bug to fix. It is historical artifact of standards solving distinct problems differently. Fixing would require worldwide coordination—impossible after thirty years. Standards bodies do not retroactively break web. RFC 3986, form rules, browser URL construction each have standards and reasons. RFC 1866 form encoding and RFC 3986 URI encoding serve different layers. Encode deliberately knowing your standard. Test against realistic payloads.

Choose encoding by context. Forms use plus per HTML standards. Manual URIs use %20 per RFC 3986. APIs specify which to expect; follow documentation or test both. URL encoder & decoder shows RFC 3986. Need form encoding? URLSearchParams does that. Tool does not mix encodings; understanding standards prevents surprises. Encoding inconsistency between layers causes subtle parameter loss, truncation, data corruption. Both standards are correct in their domain. Apply deliberately and document.