Developer tools · Base64 encoder & decoder
What atob and btoa stand for, and why they only understand Latin-1
· Background
base64 javascript unicode
atob and btoa date from Netscape and the names mean 'ASCII to binary' and 'binary to ASCII'. This post covers where they came from, how the WHATWG standards define them, and why they never learned Unicode.
A function name that reads like a typo — the confusion the names cause and the one-line answer
atob and btoa are JavaScript built-in functions that were introduced in Netscape in the 1990s. The names are abbreviations: btoa stands for binary to ASCII and atob stands for ASCII to binary. The names reflect their age and design: they were built when binary meant a string of byte values (0-255) rather than the more modern Uint8Array or Buffer. The mnemonic often given for the names is less important than the observable contract: one function maps a binary string to Base64 and the other reverses it. This repository does not document the original naming decision, so the article avoids presenting folklore as sourced browser history.
The functions expect a binary string: each character's code unit must be in the range 0-255, representing one byte. If you pass a character with a code unit above 255 (like an emoji or an accented letter from outside Latin-1), the function throws InvalidCharacterError or silently produces incorrect output. btoa (binary to ASCII) encodes a binary string to base64.
What the names suggest and what the byte-string contract actually proves
The input must be a string where each character is a byte (code unit 0-255). btoa(hello) encodes the ASCII bytes as base64 and returns aGVsbG8=. btoa with the letter e-acute appears to work because the precomposed Latin-1 letter e-acute (U+00E9) has a code unit of 233, which is within 0-255. However, btoa encodes it as a single byte, 0xE9, not the UTF-8 bytes 0xC3 0xA9 that e-acute should produce. Before typed arrays became the normal byte container, JavaScript APIs used strings whose code units stood for bytes. That model remains visible because btoa rejects code units above 255. The precise product chronology is not established by these files; the failure boundary is established by executable tests.
This silent corruption is more dangerous than an error: the result looks fine but is wrong. atob (ASCII to binary) decodes base64 back to a binary string. atob(aGVsbG8=) returns hello. The output is a binary string where each character's code unit is 0-255, representing one byte. If you want to convert this to proper Unicode text, you need to interpret the bytes as UTF-8 and decode them with TextDecoder.
The legacy binary-string model — observable behavior without an unverified browser-history claim
For ASCII, this extra step is unnecessary (ASCII is a subset of UTF-8), but for any non-ASCII bytes, it is essential. atob does not do that interpretation; it returns the raw bytes as a binary string.
The WHATWG standards (the living standards for web APIs) define atob and btoa in the HTML specification. The definition includes a forgiving-base64 decode algorithm for atob: it skips whitespace and accepts missing padding, making real-world base64 (including MIME-wrapped base64 with line breaks) decodable. ToolAcre normalizes whitespace, URL-safe punctuation and missing padding before calling the browser decoder. It then copies returned code units into Uint8Array and applies a fatal UTF-8 decoder. That combination separates forgiving Base64 syntax from strict text interpretation.
Current behavior in this implementation — forgiving alphabet normalization and strict UTF-8 text decoding
The function signature has not changed, but the standards definition is the authority for what the function does. Why do atob and btoa only accept Latin-1? Because when they were designed in the 1990s, JavaScript did not have a way to represent bytes directly (no Uint8Array or ArrayBuffer). The only way to pass bytes to a function was as a string where each character represents one byte.
This is called a binary string and is confusing by modern standards. A JavaScript string is Unicode text, not a sequence of bytes. The design conflated the two: a string where each code unit is 0-255 is a binary string. The naming reflects the era: ASCII in btoa literally meant seven bits for ASCII text, but the implementation accepts any byte (0-255). Adding a Unicode mode directly to btoa would change its long-standing byte-string contract and risk compatibility. The reviewed source instead composes TextEncoder before encoding. This article can verify that composition; it omits claims about standards-committee motives not recorded in the repository.
Worked example: tracing forgiving-base64 on a string with spaces and missing padding — what atob accepts that a strict decoder rejects
Modern alternatives avoid the binary string model. The Encoding API provides TextEncoder to convert text to UTF-8 bytes, and TextDecoder to convert UTF-8 bytes back to text.
Base64 encoding and decoding is now specified in the HTML spec for both strings (atob and btoa) and typed arrays. The Base64 encoder & decoder tool uses TextEncoder and TextDecoder around atob and btoa, so you can safely encode and decode Unicode text without the Latin-1 limitations. A spaced or unpadded value succeeds because normalization removes whitespace and restores the required block length. A value whose cleaned length leaves remainder one is rejected before atob. This distinction shows what “forgiving” means here: recoverable formatting is accepted, structurally impossible input is not.
Newer standards work on Base64 for typed arrays — described qualitatively, with a note to check current browser support
Handling Unicode with btoa requires encoding text to UTF-8 bytes first. The old workaround was btoa(unescape(encodeURIComponent(text))), which is confusing but works: encodeURIComponent percent-encodes UTF-8 bytes, unescape converts triplets back to characters, and btoa encodes the resulting binary string. This works but relies on deprecated functions and is hard to read. Modern code should use TextEncoder(text).map(byte => String.fromCharCode(byte)) followed by btoa, or better, directly convert to Uint8Array and use the Encoding API.
Atob does not automatically give you text; it gives you binary. atob(Y2Fmw6kg8J+YgA==) returns a binary string containing the bytes of UTF-8-encoded text with caf-accent and emoji. To recover the text, convert the binary string to a Uint8Array and pass it to TextDecoder(utf-8). The Base64 encoder & decoder tool does this automatically: you paste text, it encodes it to UTF-8 bytes, then to base64. Typed-array Base64 APIs are evolving across browsers, but this source does not use them. Depending on one requires a current compatibility check and fallback plan. ToolAcre’s explicit byte-array conversion remains inspectable and covered by its present test suite.
Typed-array alternatives are evolving — verify current browser support before depending on them
You paste base64, it decodes to UTF-8 bytes, then to text. The intermediate binary-string step is hidden because it is an implementation detail of the 1990s API. Understanding atob and btoa is useful for debugging legacy code or working with old APIs that hand you binary strings. Most new code should avoid the binary string model entirely.
If you need to encode or decode base64, the Base64 encoder & decoder tool handles Unicode correctly. If you are building an API, accept Uint8Array or a typed-array view, or clearly document whether your base64 is UTF-8 or Latin-1. When reviewing code that uses btoa with non-ASCII text without TextEncoder, it is a bug: the output encodes the wrong bytes. Node Buffer and non-browser runtimes define different APIs and acceptance rules. They are intentionally excluded. The claims in this article concern the browser primitives and the wrapper implemented in apps/dev, not every function named atob or btoa in every environment.
Takeaway: two 1990s functions with a byte-string contract — how the Base64 encoder & decoder does the UTF-8 step around them so accents, CJK and emoji round-trip
The names atob and btoa are peculiar artifacts of 1990s computing. Modern naming would be base64Encode and base64Decode, and the APIs would accept Uint8Array or strings with explicit encoding declarations. But atob and btoa persist in browsers for backward compatibility. Understanding what they mean (and what they cannot do) helps you avoid silent corruption when encoding Unicode text.
The Base64 encoder & decoder tool bridges the gap: it speaks the UTF-8 and base64 languages that modern code needs. The robust pattern is compositional: encode text to UTF-8 bytes, convert bytes to the binary-string contract, then call btoa; reverse those steps around atob. Try an accent, CJK characters and emoji, then require the decoded text to match every original code point.