Developer tools · SHA hash calculator
Same Text, Different SHA-256: Newlines, Encodings and Hidden Bytes
· How it works
sha-256 encoding text-processing debugging
The command line says one thing and the browser says another for what looks like the same text. Trailing newlines, UTF-16 and CRLF explain almost every case; this post shows how to find the hidden bytes.
echo says one hash, the tool says another — the everyday mismatch and why neither is wrong
A terminal reports one SHA-256 digest and a browser reports another for what appears to be identical text. The command-line tool has not broken, and the browser has not either. The bytes being hashed are not the same, even though the visible characters look identical. This post traces the most common sources of those hidden bytes and shows how to find them with a text field and a hex viewer.
The problem is almost never the SHA-256 algorithm itself. SHA-256 is deterministic: the same bytes always produce the same digest, and the digest is correct. When outputs differ, the bytes differ. The confusion arises because "the same text" is ambiguous: a person sees characters, but a hash function sees bytes, and the translation between them is where the hidden differences hide.
The trailing newline — how echo appends a byte and printf does not, and what that does to a digest
The echo command in a shell appends a newline character (U+000A, byte 0x0A) to its output. This is by design: the convention in Unix that text files end with a newline gives echo one simple job. When you type echo abc into a terminal and pipe it to sha256sum, the digested bytes are 61 62 63 0A (the ASCII codes for a, b, c, and the byte for newline), not 61 62 63. The tool ToolAcre uses to compute hashes hashes the bytes 61 62 63 and produces a different result.
The printf command does not add a newline unless you write one into the format string. printf abc | sha256sum computes the digest of bytes 61 62 63 alone, which matches the browser tool. This is why comparing hashes often means running printf instead of echo, or piping to sha256sum with the -z flag, or specifying raw input in whatever form your tool provides. The hidden newline is the single most common reason a browser tool and a command-line tool disagree.
UTF-8 versus UTF-16 — why the same characters are different bytes in some shells and editors
UTF-8 and UTF-16 encode the same characters as different byte sequences. The character é (U+00E9, an e with an acute accent) encodes as two UTF-8 bytes: 0xC3 0xA9. In UTF-16, which is how JavaScript internally represents strings, the same character occupies two bytes in a different order (depending on endianness) or a different form entirely if composed from a base character and a combining mark. When you copy café from a Windows application and paste it into a browser hash tool, the bytes the tool hashes might not match what a Mac terminal hashes, because the systems defaulted to different encodings or normalisation forms.
The ToolAcre hash tool explicitly converts text to UTF-8 before hashing via TextEncoder. This is the same encoding the Unix command-line uses by default. The source code in apps/dev/src/lib/base64.js shows the function textToBytes calling new TextEncoder().encode(), which guarantees UTF-8. If another system is using UTF-16 or Latin-1 or any other encoding, the bytes produced will differ. The tool displays the byte count alongside the digest, which is why pasting café and comparing with a command-line hash will show different byte counts if the encodings diverge.
CRLF, BOM and normalisation — line endings, byte-order marks and composed versus decomposed accents as invisible input differences
CRLF (carriage return + line feed, bytes 0x0D 0x0A) is the line ending convention on Windows; LF (line feed alone, byte 0x0A) is the Unix convention. A text file that looks identical when opened in a text editor can carry different line endings, and those bytes are part of the input to the hash. A file edited on Windows and checked against a SHA-256 computed on a Unix system will not match if one system has converted the line endings and the other has not.
The byte-order mark (BOM, bytes 0xEF 0xBB 0xBF for UTF-8) is an optional sequence at the start of a file that signals the encoding. Some editors add it; some tools strip it; some ignore it. If a file carries a BOM and you hash it byte-for-byte, the BOM bytes are part of the digest. If you then copy the visible text (which a viewer hides the BOM from) into a tool that does not add a BOM, the digests will not match. Text normalization forms (NFD versus NFC for composed and decomposed accents) add another layer: the same accented letter can be represented as a single precomposed character or as a base character followed by a combining accent, and the byte sequences are different.
Worked example — one string hashed with and without a newline, then in two encodings, with each byte difference shown
Diagnostic method one: use a hex dump tool or online converter to see exactly what bytes your tools are operating on. Paste the text into a base64 encoder, encode it, and you have a textual record of the bytes. Then decode the base64 on the command line with base64 -d and pipe it to od -A x -t x1z to see the hex byte sequence. If the bytes match, the algorithm is correct; if they do not, the difference will be visible.
Diagnostic method two: use the ToolAcre SHA hash calculator to hash progressively longer inputs, starting with a single character. Add a newline (which means typing Enter inside the text box), add spaces, add the same text with UTF-16 escape sequences if the input came from a non-ASCII source. Watch the digest change with each addition. The byte count displayed alongside the digest tells you how many bytes the tool is hashing, which narrows the search dramatically.
Hex case and whitespace in the output — the differences that are purely cosmetic
The hexadecimal representation of a digest is case-insensitive. Uppercase and lowercase letters both represent the same bytes: A = 10, a = 10. Some tools emit uppercase, some lowercase, some allow either. If one digest is lowercase and another is uppercase, they are the same digest. Whitespace in the digest display is purely cosmetic. A digest displayed as ba78 16bf versus ba7816bf is the same; the space is just a formatting choice. Mismatches caused by case or whitespace are not real mismatches.
Fixed-width formatting differences are also invisible at the byte level. A digest shown with hyphens, spaces or colons (like ba-78-16-bf) is a formatting convention that makes it easier for humans to read, not a change to the actual bytes. The ToolAcre tool always emits lowercase with no separators, which is the format most command-line tools print. If you are comparing with a tool that emits differently, convert to the same representation first.
What this does not cover — hashing files, where the same principles apply but the bytes come from disk rather than a text field
The ToolAcre SHA hash calculator performs UTF-8 conversion before hashing, displays the input byte count, and offers base64 and hexadecimal output forms. The source file apps/dev/src/lib/hash.js shows the hashText function calling digestBytes, which passes bytes.slice().buffer to crypto.subtle.digest. The comments in that file explicitly document that the UTF-8 step is intentional and note the difference between different encodings. Testing your input against the known vector (abc hashes to ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad in hex) establishes that the browser tool is working correctly; any deviation points to a byte difference in the input.
File hashing follows the same principle. The bytes in the file are what matter: a line-ending difference when you export from one system and import to another can change every digest. Some tools offer options to handle line endings during comparison; others hash the file as-is. Knowing whether your tool hashes the file as binary or performs text normalization first is essential for reproducibility.
Takeaway: hash bytes, not text — the ToolAcre SHA hash calculator hashes the bytes you paste, so check what you pasted first
Comparison and verification work only when you hash the same bytes. Start by confirming you are hashing the exact same input: run echo -n (or printf) instead of echo to avoid the newline, specify UTF-8 encoding explicitly if your tool allows it, check that CRLF has not been inserted by an editor or system utility. Then hash with the ToolAcre calculator and the command-line tool side by side. If the digests match, the bytes were identical. If they do not, use the byte-count display and the hex dump method to find the hidden difference.
Once you understand where the bytes diverged, you can choose whether to normalize them for comparison. Some checksums are meant to verify file integrity exactly as it exists on disk, in which case byte-for-byte hashing is the goal. Others are meant to verify that the visible content is the same, in which case normalizing line endings and encoding is correct. Neither is wrong; they answer different questions. The SHA-256 algorithm is always right; the question is only whether you are asking it to hash the same input in both cases.