Developer tools · SHA hash calculator
Content Addressing: How Git, Docker and npm Use SHA Digests as Names
· Background
sha-256 docker cryptography developer-workflow
Git commits, container image digests and lockfile integrity strings are all the same idea: naming data by its hash. This post explains content addressing and what each ecosystem gains from it.
The sha256: in your docker pull — what that string is and why it never changes for the same image
Version control systems, container runtimes and package managers all use the same naming idea: a file or a collection of bytes is named by its SHA digest. In Git, a commit's 40-character SHA-1 identifier (or 64-character SHA-256 in modern repositories) is computed from the commit's content—the tree, author, timestamp and message. Change a single byte and the SHA changes. In Docker, each layer digest is a SHA-256 hash of the layer's contents, and the image digest is computed from the manifest. In npm and other package managers, integrity fields store SHA-512 digests of tarballs to verify downloads. Content addressing means the name depends only on the bytes, not on a central database or a timestamp.
The benefit is immutability within each system. A git commit SHA-256:abc... will always refer to the same tree and message because the hash determines the identity. If someone claims to have a different commit with the same SHA, they are claiming the same bytes produce two different hashes, which breaks cryptography. Deduplication becomes automatic: two files with identical bytes produce the same digest, so storage systems can store the bytes once and reference them twice. Integrity checking becomes as simple as recomputing the digest and comparing: if the bytes were modified in transit or at rest, the digest no longer matches.
Naming data by its hash — the content-addressing idea and why it makes deduplication and integrity free
Git stores objects—commits, trees, blobs and tags—keyed by their SHA digest. The `git cat-file` command takes an object ID and retrieves the bytes. The object store is content-addressed: you request by digest, not by location or by name. When you clone a repository, git verifies each object by recomputing its digest and checking against the digest packed in the transfer. The transition from SHA-1 to SHA-256 is gradual; repositories can support both for compatibility. The on-disk format stores object type, size and compressed bytes. The digest is computed over the canonical uncompressed form.
Modern Git repositories can use SHA-256, and the transition is ongoing because SHA-1 collisions are now practical (demonstrated in 2017 and refined in 2020). A commit in a repository using SHA-256 has a 64-character hex identifier instead of 40. The `git hash-object` command computes the SHA of a blob (file contents) without storing it; `git commit-tree` computes the SHA of a tree structure and message. Both operations are deterministic: the same bytes always produce the same digest. This is how GitHub and other forges can display commit SHAs consistently—they compute the same digest the author's clone computed.
Git uses content-addressed objects and this tool can reproduce text digests; migration details require Git-specific sources
Docker images are built in layers, where each layer is a filesystem delta (changes from the previous layer). The OCI Image Spec defines how to compute the digest of a layer and the digest of the image manifest. The layer digest is the SHA-256 of the compressed tar file containing the layer files. The manifest is a JSON document listing layers, their digests and metadata. The image digest is the SHA-256 of the manifest JSON itself. When you pull an image using a tag like `latest`, the registry looks up the tag and returns the manifest digest. You can then pull by digest directly, ensuring you get exactly the same bytes—all layers and metadata—every time.
The `docker inspect` command on a local image shows its digest. Running the same image from the same tag on two machines produces the same digest if the registry still holds that tag pointing to the same manifest. Content addressing makes image supply chains auditable: a CI/CD pipeline can verify that the image it deployed matches the digest in the build log, and a security scanner can report on all images known to have a specific vulnerability by their digest rather than by tag, which can move.
Containers — OCI manifests and layer digests, and why a tag can move but a digest cannot
Package managers use digests to verify downloads against tampering or corruption. In npm, the `package-lock.json` file includes an `integrity` field for each dependency, containing a hash (usually SHA-512) and the encoding (usually base64). When npm downloads a tarball, it recomputes the hash and compares. If the hashes do not match, the install fails. Go uses a `go.sum` file with similar structure: module path, version and SHA-256 of the module source. Cargo uses checksums in `Cargo.lock`. The principle is identical: the digest is computed once when the dependency is first resolved, and checked on every subsequent install.
Integrity checking does not require uploading the package to a signing authority or storing signatures separately. The digest IS the integrity check. For maximum assurance, projects use `go.sum` which is signed by the Go project's transparency system, or npm integrity combined with other verification, but the base case is simple: the publisher computes the digest once, records it in the lockfile, and tools on the consumer side verify that the downloaded bytes match.
Package integrity encodings vary; only this calculator’s supported SHA outputs are asserted here
The same bytes through the same algorithm always produce the same digest, regardless of where the bytes come from. A developer's local build of a commit produces the same SHA-256 as a CI/CD system checking out the same revision from the same repository. This reproducibility is why content addressing works: you can verify an artifact without trusting the delivery mechanism. The digest becomes a cryptographic commitment: changing even one byte invalidates it.
Distributing the digest separately (before distributing the artifact) protects against in-flight modification. A web page published before release can display "expect SHA-256:abc..." and then users can verify downloads against it. A git commit published in a public repository is a commitment to the bytes; the digest proves it.
Worked example — following one blob from bytes to digest to the name a tool uses for it
Different systems encode their digests differently. Git uses lowercase hexadecimal by default (40 or 64 hex characters). Docker uses the format `sha256:` followed by hex. npm and Go use base64 in integrity fields. The bytes are the same; only the representation differs. A SHA-256 digest over "abc" is always the same 256 bits, but you may see it as a 64-character hex string, a 44-character base64 string or a label like `sha256:` followed by either. Converting between encodings is lossless; the digest is the same value in every representation.
Understanding the encoding matters when comparing digests across tools. If Git prints a hex digest and a tool shows base64, you must convert one representation to the other to verify they match. ToolAcre's SHA hash calculator displays both hex and base64 for every digest, making it easy to convert or cross-reference with other systems.
What this does not cover — the specific encodings each tool uses (hex versus base64), covered in a separate post
Content addressing is not specific to cryptography, though cryptographic hashes make it secure. A CRC32 checksum also content-addresses data, but CRC32 collisions are common and collisions can be manufactured; this repository does not mark SHA-256 as broken, while CRC32 is not offered as an adversarial integrity primitive. The choice of hash algorithm matters for security: SHA-256 is the modern standard for systems that need integrity protection against adversaries. SHA-1 is legacy only (Git and others are migrating away). Choosing the right algorithm is a separate decision from choosing content addressing as the naming scheme.
Content addressing combined with cryptographic hashes is the foundation of supply-chain integrity in modern software. Every package you install, every container you run and every commit you check out can be verified to be the bytes the original publisher intended, without relying on secure transfer (though secure transfer is still good practice).
Takeaway: the hash is the identity — the ToolAcre SHA hash calculator lets you compute the same digests these systems rely on
Content addressing is independent of the encoding, the storage location or the transfer mechanism. The same bytes produce the same digest whether stored locally, in a CDN, in a registry or transmitted over HTTP or secure HTTPS. The digest is a cryptographic commitment to the bytes, and verifying it requires only the bytes and the algorithm, not any external service. This is why content addressing enables offline verification: you can download a file over an untrusted channel, check the digest and know whether the bytes are authentic.
The ToolAcre SHA hash calculator lets you compute the same digests these systems rely on. Paste a string or watch a file, run the calculator and see the SHA-256, SHA-384 and SHA-512 digests that Docker, Git, npm and other tools use internally. Compare your computed digest with the one from the original source to verify the bytes have not been modified. The calculator hashes UTF-8 text that you paste; it does not hash files or key material, so the boundary between what it can hash (text input) and what it cannot (binary files, cryptographic keys in their encoded forms) is clear and documented.