English

Developer tools · Base64 encoder & decoder

Why email attachments are Base64: MIME, 7-bit transport and 76-column lines

· Background

base64 encoding

MIME Content-Transfer-Encoding header and 76-character Base64 line wrapping for email transport
Original ToolAcre vector illustration

Email was built for 7-bit ASCII text, and attachments had to fit through it. This post traces how MIME adopted Base64, why lines are wrapped at 76 characters and what that means for size and debugging.

The attachment that arrived corrupted through an old relay — the 8-bit problem that MIME was invented to solve

Email was designed in the 1970s and 1980s for 7-bit ASCII text only. SMTP, the protocol that carries email, expects each line to be at most 998 characters of 7-bit ASCII (characters 0-127). Sending a binary file like a PDF or an image directly through SMTP would fail: bytes 128-255 would be corrupted or rejected by old mail servers and relays. Attachments need encoding. MIME (Multipurpose Internet Mail Extensions, RFC 2045) solved this by defining Content-Transfer-Encoding header values, including base64, which represents any byte sequence as 7-bit ASCII text.

MIME offers several Content-Transfer-Encoding choices: 7bit (no encoding, only for safe ASCII), 8bit (for servers that support 8-bit bytes, not universal), quoted-printable (encodes only non-safe bytes, keeping ASCII readable), and base64 (encodes everything, maximizing compatibility). Base64 was chosen for binary attachments because it is simple, standardized, and guarantees safety on any mail system, no matter how old or strictly 7-bit-only. The trade-off is size: Base64 is about one-third larger than the original bytes.

The transport problem Base64 solves — representing arbitrary bytes with printable characters

A 3 KB PDF becomes roughly 4 KB of Base64 text. The 76-character line limit comes from RFC 2045.

SMTP allows lines up to 998 characters, but older mail systems and some spam filters reject long lines. RFC 2045 specifies that MIME Base64 lines must not exceed 76 characters (plus a CRLF line ending), so a mail server will never break the transport. The limit is not magical; it is a historical compromise between readability (76 characters fits most 1980s terminals), compatibility with old systems, and avoiding detection as spam or virus patterns.

The output choices visible in this tool — canonical padding and optional 76-character wrapping

Modern mail systems usually support longer lines, but encoding to 76-character lines ensures the attachment reaches even the oldest receiver. After RFC 2045 defined MIME Base64, RFC 4288 (media types) and RFC 2183 (Content-Disposition) added standardized ways to label attachments. A message with a PDF attachment includes a Content-Transfer-Encoding: base64 header, a Content-Type: application/pdf header, and the PDF bytes encoded as Base64 with 76-character lines. A mail reader decodes the lines by removing the line breaks (CRLF characters) and then decoding the Base64 to recover the original bytes.

Decoding a MIME Base64 attachment requires ignoring whitespace. The RFC says: decoders must skip line breaks (CR and LF characters) during decoding. This is why a Base64 decoder that accepts whitespace is practical; most real MIME mail will have line breaks. Some decoders are strict and reject whitespace (suitable for contexts like JWT, where line breaks should not be present), while others are lenient and skip whitespace (suitable for MIME).

The 76-character option in practice — how the encoder inserts and decoder ignores line breaks

The Base64 encoder & decoder tool can handle both: it accepts a multi-line pasted attachment and ignores the line breaks during decode. The size impact is predictable. RFC 2045 base64 wrapping adds one CRLF (2 bytes) per 76 characters of output. For a 10 KB file, the Base64 is roughly 13.3 KB, plus CRLF every 76 characters: about 13.5 KB total. The overhead is roughly one-third more bytes.

Email size limits are usually stated for the encoded size, not the original file size; a mail server with a 25 MB limit means 25 MB of the encoded message, not 25 MB of attachments. Calculating the original file size requires dividing by 1.33 (or more precisely, by 4 divided by 3). Quoted-printable encoding is an alternative that keeps printable ASCII unchanged and encodes only bytes 128-255 and a few special characters.

Worked example: reading a raw message source — finding the Base64 part and decoding a small text attachment

A text file with mostly ASCII remains readable if you open the raw message source. Base64 obfuscates everything, even plain ASCII text. Quoted-printable is rarely used for binary files (it would be very inefficient for a PDF) but is sometimes used for text. A mail reader chooses the encoding based on the attachment type; a browser does not usually ask the user which encoding to apply.

The Base64 body of an email message is just the bytes themselves, not a separate file. When you see an attachment in a mail reader, the reader has already decoded the Base64 and is showing the original file.

The size cost in practice — roughly a third more bytes, and why mail size limits are stated for encoded size

If you view the raw message source (an option in most mail clients), you will see the MIME headers and the Base64-encoded body. The Base64 encoder & decoder tool can help you manually decode a fragment of a message source; copy the Base64 part, remove line breaks, and paste it into the tool.

Multiple attachments in a MIME message use a multipart boundary. Each part has its own headers (Content-Type, Content-Transfer-Encoding) and body. A plain-text alternative version of the message appears as one part, and each attachment appears as another part. The boundary string separates the parts; it is chosen to not appear in any part's content. A mail reader reconstructs the message by parsing the boundaries and decoding each part according to its Content-Transfer-Encoding header.

What this does not cover — encoded-word headers, S/MIME and the 8BITMIME extension in depth

RFC 2045 base64 encoding is not universal today. Some mail systems support 8-bit transport and no longer require base64. Some systems use different encoding names or add custom headers. But base64 with 76-character lines remains the most compatible choice for attachments that must reach any mail system, anywhere. When you attach a file using a mail client, the client usually chooses base64 automatically for binary files, handles the line wrapping, and adds the MIME headers.

Understanding the mechanism helps you debug when an attachment seems corrupted or when you are manually working with a message source. Building or parsing an outgoing email message requires understanding MIME structure. A library should handle the encoding, line wrapping and headers; you do not usually construct MIME manually. But if you are parsing a raw message source (debugging a delivery issue, or extracting attachments programmatically), knowing that Content-Transfer-Encoding: base64 means the following body is 76-character-wrapped base64 lets you apply the right decoder.

Takeaway: Base64 is email's compatibility layer — how the Base64 encoder & decoder lets you read a small text part from a raw message locally

The base64 itself is standard RFC 4648; the wrapping and MIME headers are specific to email. Email attachments are base64 because email was built for plain text and base64 is the simplest, most universal compatibility layer to send binary data through a text-only protocol. The 76-character line limit is a historical artifact of 1980s terminals and slow networks, but it persists as the standard for compatibility.

Understanding this history explains why MIME exists, why there are multiple encoding options, and why base64 remains the default for attachments even though modern mail systems could support binary directly. The Base64 encoder & decoder lets you manually work with MIME bodies to verify or debug the encoding.