English

Developer tools · Base64 encoder & decoder

PEM explained: why certificates and keys are Base64 between BEGIN and END

· Background

base64 encoding

PEM armor: BEGIN and END labels and 64-column Base64 wrapping of DER binary
Original ToolAcre vector illustration

A PEM file is DER binary wrapped in Base64 with labelled armor lines. This post explains the format's origins, its line rules and what you can and cannot learn by decoding one.

The certificate that 'looks like text' but fails to parse — a header typo, a stray carriage return and the format underneath

A PEM file is Base64-encoded binary wrapped in text labels. The name comes from Privacy-Enhanced Mail (RFC 1421, 1992), which used this format for encrypted messages. The format survives today in TLS certificates, SSH keys, and GPG keys. The structure is simple: a line saying -----BEGIN CERTIFICATE----- (or BEGIN PRIVATE KEY, BEGIN PUBLIC KEY, etc.), followed by 64-character lines of base64 text, followed by -----END CERTIFICATE-----.

The base64 body decodes to a binary format called DER (Distinguished Encoding Rules), which is a way to serialize structured data (specifically, ASN.1 structures). Decoding the base64 gives you binary; reading the binary requires understanding ASN.1, which is complex. The PEM armor exists because binary files are hard to email and edit. A certificate file in pure binary form would corrupt when passed through old mail systems, USENET, or web forms.

The text armor visible in a PEM block — labels around a Base64 body

By base64-encoding the binary and wrapping it in text labels, the entire certificate becomes 7-bit ASCII text that survives any transport. A text editor can open it; a mail system will not corrupt it. The -----BEGIN and -----END lines are labels for humans and automated tools; they clearly mark what kind of data is inside. A certificate is labeled CERTIFICATE; a private key is labeled PRIVATE KEY.

The label is not verified by cryptographic software; it is just a hint for humans and tools. The 64-character line limit in PEM comes from RFC 1421 and the same MIME reasoning as email base64: old mail systems had line-length limits, and 64 characters fit on a 1980s terminal. PEM wraps the base64 output at 64 characters with a line ending (CR LF on Windows, LF on Unix).

From labeled text to decoded bytes — this repository does not establish the format history

When decoding a PEM certificate, the parser must strip the armor lines (-----BEGIN..., -----END...) and the line breaks, then base64-decode the remainder. A stray carriage return or a mismatched label can break parsing. The line wrapping is not part of the base64 standard (RFC 4648 base64 is unwrapped); it is specific to PEM. Inside the base64 is DER-encoded binary. DER is ASN.1 (Abstract Syntax Notation), a complex specification for representing data structures.

A certificate is a structured record containing a subject name, a public key, a signature, and metadata. ASN.1 does not describe the bytes directly; it describes how the structure should be encoded.

Anatomy of the block as input to this tool — remove labels and pass only the Base64 body

The encoding starts with tag-length-value triplets. For example, a SEQUENCE in ASN.1 is encoded as tag 0x30, followed by the length of the contents, followed by the contents themselves. A certificate always starts with the bytes 0x30 0x82 (sequence, length encoded in two bytes), which appears as MII in base64.

Checking a certificate without parsing it: the first three characters of a PEM certificate body are almost always MII (which is 0x30 0x82 in base64, the start of a SEQUENCE). If a PEM block does not decode to 0x30, the base64 is corrupt or the label is wrong. The Base64 encoder & decoder tool can decode the body and show you the hex: paste the base64 lines (without the -----BEGIN and END armor), remove line breaks, and decode.

What decoding exposes: binary bytes, not parsed certificate fields

If the output is binary starting with 30 82, it is likely a valid certificate structure. If it is gibberish or text, the decode failed or the base64 is wrong. Common PEM errors: a label mismatch (e.g., a certificate body with a PRIVATE KEY label), a Windows line-ending issue (some parsers choke on CRLF), a typo in the armor line (extra spaces or characters), or missing line breaks.

Tools expect -----BEGIN CERTIFICATE----- not -----BEGIN CERT----- or BEGIN CERTIFICATE. Copy-pasting a PEM from a web browser or a PDF can introduce Unicode quotes or smart quotes instead of ASCII quotes, breaking the label. Pasting a private key into a certificate field is a common mistake; the parser will reject it because the label does not match. PEM supports multiple blocks in one file.

Worked example: decode a short body and inspect bytes without asserting a certificate signature

An SSH key file might contain both a private key (labeled PRIVATE KEY) and a public key (labeled PUBLIC KEY), or multiple certificate blocks. A parser reads the file from the top, looking for lines starting with -----BEGIN. When it finds one, it reads until -----END with matching label, extracts and base64-decodes the body, and processes it. Then it continues looking for the next block.

An accidentally concatenated certificate chain (multiple PEM blocks for a certificate and its intermediates) in one file is valid if all the labels are correct. The PEM format was standardized for Privacy-Enhanced Mail (RFC 1421) in the early 1990s.

What this does not cover — parsing ASN.1 structures, private key encryption and PKCS#12 bundles

RFC 7468 (2015) modernized the definition, clarifying the line-length rules, armor-line format, and edge cases. Most tools and standards reference RFC 7468 now. Other binary-to-text formats exist (like DER-to-hex for some protocols), but PEM with base64 and ASCII labels is the de facto standard for cryptography and TLS because it is human-readable, plain text, and easy to copy or send.

Building a PEM block: take the DER binary (e.g., a certificate from a cryptographic library), encode it to base64, wrap the result at 64 characters with line breaks, and surround it with -----BEGIN CERTIFICATE----- and -----END CERTIFICATE----- lines. Parsing a PEM block: find the -----BEGIN and -----END lines, extract the base64 body (removing the armor and line breaks), base64-decode to get the binary, then parse the DER and ASN.1 binary.

Takeaway: PEM is Base64 with labels — how the Base64 encoder & decoder gives you a local place to try a block's Base64 body, entirely in the browser

Most tools automate this; you rarely construct PEM by hand. But understanding the structure is useful when debugging a parsing error or when manually verifying a certificate. A PEM certificate looks like text, but the content is binary data. Reading the begin and end labels does not tell you what the certificate contains; you must decode the base64 and parse the ASN.1 to see the subject name, public key, issuer and expiration.

The Base64 encoder & decoder tool can decode the body so you can inspect the first few bytes. For full parsing, you need an ASN.1 parser (most programming languages have libraries for this). The key insight is that PEM is a container format: it holds any DER-encoded data, not just certificates. The label tells you the intended use, but the parser must handle the data type correctly.