Developer tools · URL encoder & decoder
Why inconsistent URL encoding splits one page into many rows in analytics
· Why it matters
analytics url-encoding normalization
%20 and +, %2F and /, %c3 and %C3 can all describe the same URL, but reports treat them as different pages. This post explains where the variants come from and how to normalise them before counting.
The landing page with six URLs in the report — the variants side by side and the traffic they split
Data analysts notice one landing page appearing as six different URLs in analytics dashboards. Same pages can be: /landing?utm_source=email, /landing?utm_source=%65mail, /landing?utm_source=email%20campaign, /landing?utm_source=email+campaign, /landing?utm_source=email%20Campaign, /landing?utm_source=email%2bcampaign. Each variant counts as separate page views, fragmenting traffic. Data from spreadsheets, email, and forms introduce encoding variations.
Inconsistent encoding stems from multiple data sources and transformations. Hand-written links use raw spaces or no encoding. Spreadsheet exports produce percent-encoded URLs. Email clients mangle or re-encode URLs. Redirect chains normalize inconsistently. API integrations, JavaScript frameworks, and analytics code apply different rules. The same URL concept passes through layers, getting encoded and re-encoded differently.
The sources of variation — hand-written links, spreadsheet exports, mail clients and redirect chains
Case in hex digits presents the first normalization issue. RFC 3986 specifies hex digits should be uppercase: %2F, not %2f. Uppercase and lowercase hex encode identical bytes. Strict comparison treats %2F and %2f differently. The character "e" as %65 should normalize to unencoded "e" because RFC 3986 classifies letters as unreserved. Over-encoding entire URLs produces different analytics records.
The unreserved set in RFC 3986 includes: A-Z, a-z, 0-9, hyphen, period, underscore, and tilde. These should never be percent-encoded in normalized URLs. RFC normalization specifies that %41 decoding to "A" should normalize to unencoded "A". Applying this across URLs removes redundant encoding. URLs like %2f%6c%61%6e%64%69%6e%67 become /landing after decoding.
Case in hex digits and the unreserved set — what RFC 3986 says is equivalent and what is not
Reserved characters are not interchangeable and must remain distinct during normalization. RFC 3986 reserves gen-delims (:, /, ?, #, [, ], @) and sub-delims (!, $, &, ', (, ), *, +, ,, ;, =). These have structural meaning. Forward slashes in paths function as separators and should not encode. When the same character appears as data in query values, it should encode as %2F. Blindly decoding breaks URL structure.
Normalization nuances create challenges requiring contextual understanding. Only decode unreserved characters, leaving reserved characters encoded. URLs like /landing?data=%2F%20%2f remain ambiguous. Query strings begin with ? (reserved, structural). Inside query values, anything can appear—question marks require %3F encoding. URLs encoded as %2f%6c%61%6e%64%69%6e%67%3fkey%3dvalue normalize to /landing?key=value.
Reserved characters are not interchangeable — why %2F and / can legitimately mean different things
Worked example: normalising six URL variants demonstrates complete normalization. Base URL represents /page?utm_source=email&campaign=test. Six variants: 1) /page?utm_source=email&campaign=test (canonical), 2) /page?utm_source=%65%6d%61%69%6c&campaign=test (lowercase hex), 3) /page?utm_source=email%20&campaign=test (space in value), 4) /page?utm_source=email+&campaign=test (plus as space), 5) /page?utm_source=EMAIL&campaign=test (different case), 6) /page?utm_source=email&%63ampaign=test (hex in name).
Normalizing variant 2 requires fixing hex cases and decoding unreserved letters: %65%6d%61%69%6c becomes email. Variant 4 with plus signs requires context awareness—if sources are HTML forms, plus means space; otherwise plus is literal. Variant 5 has uppercase "EMAIL"; lowercase "email" is canonical since emails are case-insensitive. Variant 6 has %63 (hex for "c"); unreserved decoding produces "campaign" matching canonical.
Worked example: normalising six variants of one URL — decoding safe characters, fixing hex case, and what remains distinct
Implementing normalization in pipelines—normalise on ingestion and keep raw values—is recommended architecture for analytics. At ingestion points where URLs enter databases (logging endpoints), apply normalization before storing or deriving page-view keys. Normalization: 1) Parse URLs into components, 2) Decode unreserved sequences (fix hex case), 3) Normalize parameter order, 4) Produce canonical forms for grouping, 5) Store normalized forms and raw values. This ensures all six variants hash to same group key.
Hash functions based on normalized URLs ensure all variants map to identical pages in reports. If analytics systems lack built-in normalization, data engineering layers (ETL pipelines) normalize before database writes. For tools like Google Analytics, configurable filters allow regex grouping or sending titles separate from URLs. Most robust approaches normalize at sources: when tracking code sends URLs to analytics, ensure canonicalized forms.
Doing it in a pipeline — normalise on ingestion and keep the raw value, described as a pattern
What this does not cover includes tracking-parameter stripping and canonical tags for SEO, which are related but different. Tracking parameters like utm_source and utm_campaign might be stripped from analytics to group by organic content. This is separate business logic. HTML canonical tags consolidate page views across variants for SEO but do not affect internal analytics. Comprehensive strategies employ multiple deduplication layers combining both approaches.
Analytics space normalization support varies widely. Google Analytics handles some normalization automatically but might miss variants. Other tools require manual configuration. Paid search platforms apply different normalization to campaign URLs. Server logs record URLs as received without normalization. Comprehensive strategies document normalization applied at each layer and raw data preserved for auditing. URL encoder & decoder helps inspect variants.
What this does not cover — tracking-parameter stripping policies and canonical tags for SEO
Takeaway: normalise before counting—the URL encoder & decoder helps inspect any variant showing what it encodes and whether it matches canonical forms. For suspicious analytics variants, paste into decoders examining decoded outputs. If two URLs decode to identical forms, they represent identical pages and should consolidate. The tool shows exactly which characters encode, their hex values, and results. This inspection is the first troubleshooting step.
When troubleshooting analytics discrepancies, create lists of all observed URL variants and decode each with URL encoder & decoder. Compare decoded forms. If forms differ in data content (like different utm_source values), they are legitimately different pages. If they differ only in encoding (like %65mail vs email), they are duplicates needing normalization. Document canonical forms and implement normalization. URL encoder & decoder provides diagnosis; analytics pipeline provides solution.
Takeaway: normalise before you count — how the URL encoder & decoder helps you inspect any variant to see what it actually encodes
Takeaway: normalise before counting—the URL encoder & decoder helps inspect any variant showing what it encodes and whether it matches canonical forms. For suspicious analytics variants, paste into decoders examining decoded outputs. If two URLs decode to identical forms, they represent identical pages and should consolidate. The tool shows exactly which characters encode, their hex values, and results.
When troubleshooting analytics discrepancies, create lists of all observed URL variants and decode each with URL encoder & decoder. Compare decoded forms. If forms differ in data content (like different utm_source values), they are legitimately different pages. If they differ only in encoding (like %65mail vs email), they are duplicates needing normalization. Document canonical forms and implement normalization.