Developer tools · HTML entity escaper
One character, four escapes: é, \u00e9, %C3%A9 and =C3=A9 compared
· Background
html unicode url-encoding
The same é can appear as an HTML entity, a JavaScript escape, a percent-encoded byte pair or a quoted-printable sequence. This post lines up the four notations, explains what each layer needs and shows how to move between them.
The é that looked different in the page source, the console, the address bar and the raw email — one character, four costumes
The é that looked different in the page source, the console, the address bar and the raw email — one character, four costumes. The same é appears differently because each surrounding protocol represents a different unit. Page source, JavaScript source, a URL and an email transport are not interchangeable escape contexts.
To verify unicode escape formats compared, construct the that looked different for a developer chasing a character through HTML, JSON, URLs and email. Preserve in the page source while multi-format comparison produces the console the address; identify where bar and the raw is consumed. The observation about email one character four belongs to HTML text only.
HTML: code-point references — é and é name the Unicode code point
HTML: code-point references — é and é name the Unicode code point. HTML decimal é and hexadecimal é identify Unicode code point U+00E9. ToolAcre decodes both forms when they end with a semicolon.
A developer chasing a character through HTML, JSON, URLs and email can test html code point references by recording 233 and xe9 name before the multi-format comparison pass. Compare the unicode code point afterward and locate the parser responsible for multi-format comparison evidence. This unicode escape formats compared result explains multi-format comparison evidence, not executable contexts.
JavaScript and JSON: \u00e9 — UTF-16 code units, and why an emoji needs a surrogate pair here
JavaScript and JSON: \u00e9 — UTF-16 code units, and why an emoji needs a surrogate pair here. JavaScript and JSON é describe a UTF-16 code unit. An astral emoji needs a surrogate pair in that notation, whereas an HTML numeric reference names its single Unicode code point.
Isolate javascript and json u00e9 in a short multi-format comparison sample. Show utf 16 code units as literal source, follow and why an emoji to its destination, and name the API reading needs a surrogate pair. For unicode escape formats compared, here remains parser-bound evidence.
URLs: %C3%A9 — UTF-8 bytes, not code points, so the same character needs two groups
URLs: %C3%A9 — UTF-8 bytes, not code points, so the same character needs two groups. URL percent encoding represents UTF-8 bytes. É’s lowercase counterpart é becomes bytes C3 A9 and therefore %C3%A9, not %E9 in a UTF-8 URL component.
Treat urls c3 a9 utf as a boundary experiment. A developer chasing a character through HTML, JSON, URLs and email should retain 8 bytes not code, perform one multi-format comparison operation, and inspect points so the same character by character before changing character needs two groups. The claim about multi-format comparison evidence stops at this HTML layer.
Email: =C3=A9 in quoted-printable and w6k= in Base64 — MIME's two ways to carry the same bytes
Email: =C3=A9 in quoted-printable and w6k= in Base64 — MIME's two ways to carry the same bytes. Quoted-printable email can render those same UTF-8 bytes as =C3=A9, while Base64 encodes the byte sequence into different ASCII characters. MIME headers determine how recipients interpret them.
Reproduce email c3 a9 in with harmless input instead of customer material. Record quoted printable and w6k, observe in base64 mime s, and count every intentional multi-format comparison pass. That unicode escape formats compared trail lets a developer chasing a character through HTML, JSON, URLs and email evaluate two ways to carry and the same bytes without guessing.
Worked example: taking 'café' through all four notations — the exact strings side by side, noting which encode bytes and which encode code points
Worked example: taking 'café' through all four notations — the exact strings side by side, noting which encode bytes and which encode code points. For café, the forms are café or café in HTML, café in escaped JavaScript notation, caf%C3%A9 in a URL component, and caf=C3=A9 in UTF-8 quoted-printable.
Place worked example taking caf, through all four notations, and the exact strings side side by side during the multi-format comparison review. A developer chasing a character through HTML, JSON, URLs and email can then decide whether by side noting which changed at conversion or downstream. Keep the unicode escape formats compared conclusion about encode bytes and which out of generic security claims.
What this does not cover — CSS escapes and punycode for hostnames
What this does not cover — CSS escapes and punycode for hostnames. CSS escapes and internationalized hostname encoding use additional grammars and are deliberately omitted. Choosing an encoder starts by identifying which parser consumes the output.
Define what this does not before running multi-format comparison. Save cover css escapes and as a control, inspect the code points behind punycode for hostnames, and map multi-format comparison evidence to the next interpreter. This makes multi-format comparison evidence auditable for a developer chasing a character through HTML, JSON, URLs and email investigating unicode escape formats compared.
Takeaway: know whether the layer wants bytes or code points — how the HTML entity escaper, URL encoder & decoder and Base64 encoder & decoder are panels in one product, so you can check each form without a page load
Takeaway: know whether the layer wants bytes or code points — how the HTML entity escaper, URL encoder & decoder and Base64 encoder & decoder are panels in one product, so you can check each form without a page load. Use the entity, URL and Base64 panels for their own layers and compare intermediate text deliberately. None substitutes for the others, and none turns untrusted content into universally safe data.
Connect takeaway know whether the to an observable multi-format comparison output. Keep layer wants bytes or beside the one-pass result, then verify where code points how the enters html entity escaper url. A developer chasing a character through HTML, JSON, URLs and email can now review encoder decoder and base64 as a narrow unicode escape formats compared finding. The practical decision behind this article is specific: The same é can appear as an HTML entity, a JavaScript escape, a percent-encoded byte pair or a quoted-printable sequence. This post lines up the four notations, explains what each layer needs and shows how to move between them. The reader action is equally concrete: Links to the HTML entity escaper for the entity form and demonstrates switching to the URL encoder & decoder and Base64 encoder & decoder panels to see the percent-encoded and Base64 forms of the same text.