Developer tools · Base64 encoder & decoder
RFC 4648 explained: the standard that defines Base64, base32 and base16
· Background
base64 encoding
RFC 4648 is the short, readable document behind every Base64 implementation. This post walks through what it specifies, what it deliberately leaves open, and why implementations still differ.
Two libraries, two answers for the same string — a real interoperability puzzle that only the standard settles
Two JavaScript libraries can return different Base64 for the same string, each claiming correctness. RFC 4648 is the readable twelve-page document that should settle such disagreements, yet implementations still differ because the RFC deliberately leaves certain decisions to applications. This article walks through what RFC 4648 specifies, what it intentionally delegates to callers, and why reading the standard once resolves most real interoperability puzzles. The Base64 encoder & decoder tool includes RFC 4648 test vectors so you can verify an implementation against the authoritative examples.
RFC 4648 replaced and consolidated several earlier documents: Base64 from MIME (RFC 2045), Base64 from Privacy-Enhanced Mail (RFC 1421), base32 from S/MIME (RFC 2630) and base16 from various sources. The consolidation was necessary because MIME and PEM each had their own alphabet and rules, and MIME line wrapping conflicted with PEM's 64-column blocks. RFC 4648 defines five encoding families in one place: base64, base64url, base32, base32hex and base16, each with its own alphabet, padding rules and example test vectors. The base64 alphabet is A-Z, a-z, 0-9, plus and slash, in that order.
What the implementation and test vectors establish — standard and URL-safe alphabets, padding and whitespace handling
Each character represents 6 bits; three input bytes (24 bits) map to four output characters. The alphabet is not arbitrary: it avoids characters that differ between EBCDIC and ASCII, avoiding control characters, quotes, and the backslash that would need escaping in C string literals. The base64url variant replaces plus with dash and slash with underscore to avoid reserved characters in URLs and filenames. Both variants are equally valid; RFC 4648 section 2 specifies base64, section 5 specifies base64url, and an application must state which one it uses.
Padding with equals characters brings the output to a multiple of four characters. If the input is 1 byte (8 bits), the output is two characters plus two equals signs. If the input is 2 bytes (16 bits), the output is three characters plus one equals sign. If the input is a multiple of 3 bytes, no padding is needed. Some applications omit padding or allow missing padding on decode; RFC 4648 section 3.2 defines canonical encoding as always padded, but section 3.3 notes that decoders may accept missing padding for compatibility.
The alphabets this tool implements — standard Base64 and Base64url; other bases remain outside its scope
The padding distinction is why implementations disagree: a strict decoder rejects missing equals, while a lenient one accepts it. RFC 4648 explicitly says: the pad character equals is typically percent-encoded when used in URLs, so if base64url output is used directly in a URL parameter, the padding is not required and should be omitted. This sentence is one reason URL-safe mode and padding omission are often paired, though they are independent choices. Section 5 (base64url) does not forbid padding; it merely notes the common practice.
A caller that chooses base64url must decide whether padding is required for the receiving system. Non-alphabet characters in the input are handled differently by different decoders. RFC 4648 section 3.1 states: Implementations MUST reject the encoding if it contains characters outside the base alphabet. However, section 3.3 notes that MIME Base64 (RFC 2045) allows line breaks for 76-character wrapping, and decoders for MIME must skip whitespace. The RFC distinguishes between strict decoding (reject all non-alphabet) and MIME-compatible decoding (skip whitespace, reject other characters).
Padding, non-alphabet characters and canonical encoding — the sections that explain most decoder disagreements
An application must choose which rule to follow; the standard defines both. Base32 uses A-Z and 2-7 (32 characters total), encoding five input bytes (40 bits) to eight output characters. Base32hex substitutes 0-9 and a-v for the alphabetic characters, useful in contexts where lowercase letters are preferred.
Base16 is hexadecimal: 0-9 and a-f. Base32 and base32hex have their own padding rules in sections 6 and 7, and the RFC provides separate test vectors for each alphabet. Most developers need only base64 and base64url; base32, base32hex and base16 are included in the RFC for completeness and for applications like TOTP secrets (RFC 4226) and DNS encoding.
Application choices visible in this implementation — line wrapping, strict text decoding and error handling
The test vectors in RFC 4648 are the ground truth for checking an implementation. Encoding the strings f, fo, foo, foob, fooba and foobar produces specific base64 output: Zg==, Zm8=, Zm9v, Zm9vYg==, Zm9vYmE= and Zm9vYmFy. An implementation that produces different output for these strings is incorrect. The RFC provides equivalent test vectors for base32, base32hex and base16. The Base64 encoder & decoder tool includes these vectors so you can verify its output against the standard. Line wrapping is a MIME concern, not a base64 concern.
RFC 2045 specifies 76-character lines; RFC 4648 section 3.1 notes this in the context of MIME but does not make it a requirement of base64 itself. Some applications wrap at 64 characters (the original PEM standard); others do not wrap at all. A strict RFC 4648 base64 decoder operates on the alphabet and padding only. A MIME-compatible decoder must skip line breaks (CR, LF, CRLF). An application using base64 outside of MIME should not add line breaks unless the receiving system requires them; the RFC does not define line wrapping as part of base64.
Worked example: the RFC's own test vectors — encoding the 'foobar' prefixes and checking them in the browser
Whitespace handling is another point of implementation variance. RFC 4648 says strict decoders must reject non-alphabet characters. MIME-wrapped base64 (RFC 2045 base64) allows whitespace for formatting. The two standards agree on what the output bytes should be but differ on what input is valid. Most JavaScript implementations choose MIME compatibility and skip whitespace; the strict rule is rarely used in browsers. The Base64 encoder & decoder accepts both whitespace-containing (MIME) and strict input, making the distinction explicit. Canonical vs. forgiving decoding is the final major variance.
Canonical decoding follows RFC 4648 section 3.2: reject malformed padding, reject missing padding, reject non-alphabet characters. Forgiving decoding, used in web standards (the HTML spec calls it forgiving-base64), adds rules: ignore whitespace, accept missing padding, allow dash and underscore as plus slash equivalents even in standard base64 mode. JavaScript's atob() is forgiving; a strict RFC 4648 decoder is stricter. Neither is wrong; they serve different contexts. An application reading data from a user or from the network should know which rule the other side expects.
What this does not cover — the MIME and PEM documents themselves, and language-specific APIs
The RFC leaves nine choices to the application: which of the five alphabets, whether to require or allow padding, whether to require or allow whitespace, whether to treat dash underscore as plus slash equivalents, how to report errors, how to handle end-of-input, whether to accept missing padding, how many output bytes to allocate, and how to signal a size limit. These choices explain why two RFC 4648 implementations can disagree on the same input. Read the RFC once; check your implementation against its test vectors; state which options your application uses; test interoperability with the actual peer, not assumptions.
Understanding RFC 4648 settles most Base64 disputes because the disagreement is usually not about the RFC itself but about which options each side chose. The RFC is brief enough to read end-to-end in an hour. The standard defines the alphabets, provides test vectors, and warns where implementations must decide. The Base64 encoder & decoder tool lets you experiment with the test vectors and see the standard alphabet in action. Most everyday base64 use does not require deep RFC knowledge; but when debugging encoding mismatches or integrating with an unfamiliar API, reading the standard once removes guesswork.
Takeaway: read the standard once — how the Base64 encoder & decoder gives you a quick way to check the standard alphabet's test vectors
RFC 4648 is the consolidation of decades of ad-hoc base encoding practice into one readable specification. It does not define when to use base64 (MIME, PEM, JWT, data URIs, etc. each have their own specs); it defines what base64 is. By defining five encoding families and noting which options are canonical, the RFC makes it possible to check whether an implementation is correct. The authoritative test vectors are the starting point: if your implementation encodes foobar and produces anything other than Zm9vYmFy, the RFC says the implementation is wrong.
Use that authority as a verification checkpoint: encode each RFC test vector, compare the exact characters, and then decode the result to confirm that the original bytes return unchanged. This browser-based check separates an alphabet or padding mistake from a problem elsewhere in an integration, while keeping the standard itself as the reference rather than relying on a library label.