English

Developer tools · SHA hash calculator

Hex, Base64 and Raw Bytes: Three Ways to Write the Same SHA Digest

· How it works

sha-256 base64 encoding file-formats

One 32-byte digest shown as 64 hex characters, 44 base64 characters with padding, base64url, and raw bytes with byte boundaries marked
Original ToolAcre vector illustration

sha256sum prints hex, package-lock.json stores base64 and Docker uses a sha256: prefix. They can all be the same 32 bytes. This post explains each representation and how to convert between them.

The hashes that look different but agree — a lockfile string and a terminal checksum for the same file

A SHA-256 digest is fundamentally 32 bytes. How you write those bytes determines what the digest looks like. One digest, 32 identical bytes, appears as 64 hexadecimal characters (two per byte), or 44 base64 characters (roughly four per three bytes), or different lengths and formats depending on the encoding. The confusion arises because a lockfile might show one representation and a terminal shows another, both for the same underlying 32 bytes.

Understanding the encoding is the step that turns "why do these look different?" into "I can confirm they are the same." All three representations are equivalent once you decode them back to bytes.

A digest is bytes — 20, 32, 48 or 64 of them depending on the algorithm, before any text encoding

Before any textual representation exists, the result is an ArrayBuffer of digest bytes. ToolAcre wraps that buffer with Uint8Array, then either writes each byte as two hexadecimal digits or converts each byte to a binary character before calling btoa. Neither formatter reruns the hash, and neither changes a single digest bit.

The byte width follows the selected algorithm in this tool: SHA-1 returns twenty bytes, SHA-256 thirty-two, SHA-384 forty-eight and SHA-512 sixty-four. Those are supported outputs verified by the algorithm metadata and tests. Raw bytes are appropriate for programmatic comparison; hex and base64 are transport notations for channels that expect text.

Hex — two characters per byte, why it dominates command-line tools, and the case question

Hexadecimal representation uses the digits 0-9 and the letters A-F (or a-f) to represent the 16 possible values of a 4-bit nibble. Two hex digits represent one byte. The SHA-256 digest of the input abc is 32 bytes, so it displays as 64 hex characters: ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad. This is the format most command-line tools print. Hexadecimal is human-readable and unambiguous; every byte is represented by exactly the same two characters every time.

Hexadecimal is the default format for checksums and hashes in documentation and on the command line. It is easy to read and copy, and there is no padding, no case sensitivity in interpretation (though convention dictates lowercase or uppercase consistently), and no special characters that need escaping in URLs or JSON. The downside is that it takes twice as many characters as raw bytes, which is why other formats exist.

Base64 and base64url — roughly four characters per three bytes, padding, and where each appears (SRI, npm, SSH fingerprints)

Base64 encodes three bytes as four characters drawn from a 64-character alphabet: A-Z, a-z, 0-9, +, /. The three bytes 61 62 63 (the ASCII codes for abc) encode as YWJj in base64. A full 32-byte SHA-256 digest encodes as approximately 44 base64 characters. Padding with = characters brings the output length to a multiple of 4, so 44 characters plus 0 padding (because 32 is a multiple of 3, no padding is needed). Decoding reverses the process: four base64 characters decode to three bytes.

Base64 appears in package-lock.json files, npm shrinkwrap, SRI (Subresource Integrity) attributes in HTML, and SSH key fingerprints. It is compact — roughly 33% longer than raw bytes, compared to hexadecimal's 100% longer. The trade-off is that not all text representations are equally easy to read; base64 looks more scrambled to a human eye than hex.

Prefixed forms — sha256: in container digests, sha384- in integrity attributes, SHA256: in SSH

Base64url is a variant defined in RFC 4648 that substitutes - and _ for + and /. The alphabet becomes A-Z, a-z, 0-9, -, _. JWTs use base64url because + and / have special meanings in URLs (+ can be read as a space in query strings, / is a path separator). A JWT segment is always base64url-encoded, and a decoder that insists on standard base64 will reject it. Conversely, a base64url decoder that does not accept the standard alphabet will fail on standard base64.

Padding is optional in base64url. Standard base64 pads with = to ensure output length is a multiple of 4. Base64url commonly omits padding because = is itself URL-awkward. A decoder should accept base64url with or without padding, and the encoder should be explicit about which it produces. The ToolAcre base64 tool accepts both alphabets and tolerates missing padding on input, and lets you choose the format on output.

Worked example — one digest converted from hex to base64 and back, with the byte boundaries marked

Prefixed forms add a scheme identifier to the digest. Docker image digests use sha256:ba7816bf..., where sha256: is the prefix. SSH fingerprints use SHA256:, with a colon. Some tools use sha256= or SHA256= (with an equals sign). The prefix is purely informational; it tells you which algorithm produced the digest. Removing the prefix leaves the same bytes in the same encoding.

When comparing digests, the prefix is noise. If one tool prints SHA256:ba78... and another prints ba78..., they are the same digest; the prefix is just metadata about the format. Similarly, prefixes like sha256:- (used in some container contexts) or sha384- (used in integrity attributes) are formatting conventions that do not change the bytes. Strip them for comparison.

What this does not cover — which encoding a given tool emits; check the output format before comparing

One digest, the SHA-256 of the input abc, appears in multiple forms: hex (64 characters), base64 with padding (44 characters), base64url with padding (44 characters, with - and _ instead of + and /), or with various prefixes. To confirm they are the same, decode each back to bytes and compare the bytes. The hex representation ba7816bf... decodes to the bytes 0xba 0x78 0x16 0xbf 0x8f 0x01 0xcf 0xea ... The base64 representation converts to the same byte sequence when decoded.

The ToolAcre SHA hash calculator outputs in hexadecimal by default. If you need base64, you can use a separate tool to convert the hex to base64, or use the base64 utility on the same site to encode the text directly. Tools designed for specific contexts (npm for package-lock.json, Docker for image digests) output in the format their context expects. Understanding that these are all the same 32 bytes in different clothes removes confusion when tools disagree on format.

Takeaway: compare bytes, not strings — compute the digest with the ToolAcre SHA hash calculator, then convert to the representation you are checking against

To manually convert a hex digest to base64, group the hex digits into bytes, convert each byte to decimal, then encode using the base64 alphabet. The byte 0xba (hex ba) is decimal 186; 0x78 is 120; 0x16 is 22; 0xbf is 191. Grouping these four bytes and encoding as base64 gives the characters w (0 + 22 in the alphabet), as (encoding 186), AA (encoding 120), vw (encoding 191). The full digest requires doing this 10 times and padding if needed. This manual process is instructive but tedious; a base64 converter tool makes it instant.

The key insight is that a digest is bytes first, and text representation is secondary. Every encoding of the same bytes decodes back to the same bytes and thus is interchangeable for the purposes of integrity verification. Case differences in hex, padding differences in base64, prefixes, and spacing are all formatting choices that do not affect the actual value. When you master the ability to convert between representations, digest format mismatches become debugging problems you can solve instead of mysteries.