Developer tools · Base64 encoder & decoder
How Base64 turns three bytes into four characters, step by step
· How it works
base64 encoding unicode
Base64 is nothing more than regrouping bits: 24 bits in, four 6-bit indices out. This post walks through the table lookup, the bit shifting and the reverse trip so the format stops being a black box.
The string 'TWFu' and the word it hides — starting from a real four-character block and asking where each letter came from
The four-character Base64 string TWFu decodes to the three-byte sequence Man. How three bytes become four characters reveals that Base64 is not encryption or compression but pure bit regrouping. Once you see bit layout, Base64 output stops being opaque and becomes predictable. You can encode Man by hand, verify it against TWFu, and understand why Base64 always outputs four characters per three bytes of input.
The magic of Base64 is that three bytes (24 bits) regroup perfectly into four chunks of six bits. Six bits represent 0 to 63, which is why alphabet contains exactly 64 symbols: A–Z (26), a–z (26), 0–9 (10), and + and / (2). Each six-bit chunk indexes into alphabet to produce one output character. The reverse is equally clean: four characters index into alphabet to recover four six-bit chunks, which regroup into three bytes.
From bytes to 6-bit indices — how 24 bits are split into four groups and why 64 symbols are exactly enough
This is why Base64 feels natural everywhere. Take three bytes M, a, n in ASCII: 0x4D, 0x61, 0x6E. Write in binary: 01001101, 01100001, 01101110. Concatenate all 24 bits: 010011010110000101101110. Regroup into four chunks of six bits: 010011 010110 000101 101110. Interpret as binary numbers: 19, 22, 5, 46. Index into Base64 alphabet (A=0, B=1, ...Z=25, a=26, ...z=51, 0=52, ...9=61, +=62, /=63). Index 19 is T, index 22 is W, index 5 is F, index 46 is u.
Output: TWFu. The index lookup is mechanical. Base64 alphabet is sequence where position matters: every implementation uses same order A–Z, a–z, 0–9, +, /. Different order produces different output; changing order is exactly how base64url works. In standard alphabet, uppercase letters occupy indices 0–25, lowercase 26–51, digits 52–61, special characters 62–63. This ordering is arbitrary but fixed by RFC; every decoder expects same mapping.
The alphabet table and index lookup — A–Z, a–z, 0–9, + and / in order, and why the order matters for comparison
If you write alphabet on paper and count carefully, you can manually encode without computer: look up 19, count A B C...T, write T, repeat. Reversing is equally straightforward. Given TWFu, look up each character in alphabet: T is 19, W is 22, F is 5, u is 46. Convert to binary (leading zeros for six bits): 010011, 010110, 000101, 101110. Concatenate: 010011010110000101101110.
Group into three bytes of eight bits: 01001101, 01100001, 01101110. Interpret as decimal or hex: 77, 97, 110 or 0x4D, 0x61, 0x6E. Convert to ASCII: M, a, n. You recovered original three bytes. This is why Base64 is reversible and why padding becomes necessary only for inputs not divisible by three. Base64 encodes exact bytes and nothing more. Encoding Man and encoding bytes (77, 97, 110) are identical operations; Base64 does not know or care about characters, language or encoding.
Worked example: encoding 'Man' by hand — the binary of M, a and n, the four indices and the four output characters
It sees bytes. The tool's encoder and decoder separate concern: textual input like Man goes through TextEncoder first, turning into UTF-8 bytes. Those bytes are Base64 input. The output TWFu is text (ASCII characters), but stands for bytes, not for word. Different tool reading TWFu recovers bytes (77, 97, 110) and must independently decide whether they represent word, image, message in another encoding or something else.
Large inputs are many repetitions of this pattern. A 300-byte file uses 300/3 = 100 blocks of three bytes, each becoming four characters, producing 400 output characters. When the last block is padded, a decoder discards the zero fill rather than manufacturing another byte. That boundary is visible with a two-byte input: three useful indices survive, the fourth position is an equals sign, and only sixteen reconstructed bits belong to the result.
Reversing the process — index lookup, bit packing and where the padding bits go when decoding four characters to three bytes
Because pattern is regular, operation is fast: bit shift, look up, write. The only irregularity is final block when input length is not a multiple of three, handled by padding. Because every block is independent—bits of one block do not affect next—Base64 can encode incrementally: feed bytes in, get characters out, without waiting for entire input.
Base64url differs only in alphabet substitution. Indices 62 and 63 become - and _ instead of + and /. Bit regrouping is identical; byte-to-character mapping is identical; only the lookup table changes. A hand decoder can therefore reuse every shift and mask from standard Base64, replacing only those two terminal symbols.
Why the result is a sequence of bytes, not text — the separate step that turns bytes into UTF-8 characters
This is why RFC 4648 section 5 describes it as distinct alphabet, not different encoding. String TWFu in standard Base64 is unambiguous: it can only mean indices (19, 22, 5, 46). In base64url, string would need to contain - or _ to differ, and without those present, same indices apply.
Errors in implementation usually involve off-by-one mistakes in bit shifts or incorrect alphabet mapping. An encoder using wrong alphabet order produces different output if a and A were swapped. A decoder mishandling last partial block (when padding present) might recover wrong number of bytes. The Base64 encoder & decoder uses standard alphabet and handles padding by RFC 4648, so you can paste any hand-computed example and check work.
What this does not cover — base64url, MIME line wrapping and the performance of large buffers
Because bit math is deterministic, any error in manual encoding will produce different output when decoded, making mistake immediate. Base32 (RFC 4648 section 6) extends principle to five-bit chunks: 32 symbols (A–Z and 2–7), so five bits fit exactly into one character, and 40 bits (five bytes) regroup into eight characters. Same regrouping logic applies; difference is alphabet size and consequently ratio of input bytes to output characters.
Hexadecimal (base16) uses eight of 256 possible symbol combinations and maps one byte to two characters with no regrouping. Understanding Base64 as bit regrouping makes variants conceptually simple: pick bits per character, group input accordingly, look up each group in alphabet. When debugging Base64, bit picture is your tool. If bytes were corrupted, encode them again and compare output character by character. If unsure what bytes TWFu contains, decode it and examine output in hex.
Takeaway: Base64 is a reversible regrouping of bits — how the Base64 encoder & decoder lets you check any hand-computed block instantly in the browser
The Base64 encoder & decoder shows both characters and hex view, making simple to verify whether looking at text bytes (will decode to legible text) or binary data (shows as hex and best kept as bytes, not text). The step-by-step process—bytes to bits, bits to indices, indices to characters—is deterministic, fast, same in every compliant implementation. RFC 4648 defines Base64 formally so implementations can be compared.
Standard specifies alphabet, bit layout, padding rules and how line wrapping is handled in MIME. Knowing standard makes easy to verify whether decoder follows it strictly (canonical Base64) or accepts variants (missing padding or URL-safe characters). Many real-world applications use Base64 slightly differently: some omit padding, some use URL-safe characters, some wrap at different line lengths. The Base64 encoder & decoder handles variations automatically, but understanding standard makes debugging integration issues much simpler.