English

Developer tools · UUID generator

Name-Based UUIDs (v3 and v5): Deterministic IDs From a Namespace

· Background

uuid cryptography browser-apis

Namespace and name concatenated, hashed with SHA-1, and resulting bytes formatted as a UUIDv5
Original ToolAcre vector illustration

When the same external record must always get the same identifier, random UUIDs will not do. Version 3 and 5 UUIDs hash a namespace and a name into a stable ID; this post explains how and when to use them.

Re-importing the same customer twice — the duplication problem that deterministic IDs solve

A data import pipeline receives customer records from an external system with external IDs stable within that system. If you generate a new random UUID for each import run, importing the same customer twice produces two different identifiers and duplicate records. This duplication flows downstream into reporting, billing, and support systems. If you derive a UUID from the customer's external ID and a stable namespace representing your import source, every import produces the same UUID for the same customer, enabling you to identify and update existing records. This determinism is the defining characteristic of v3 and v5 UUIDs: they are not generated independently but derived from inputs, and the same input always produces the identical UUID.

Namespace plus name — how the input is concatenated and hashed, and why the namespace prevents collisions between different sources

A v3 or v5 UUID is derived from three components: a namespace UUID (typically predefined), a name (any byte string), and a hash algorithm (MD5 for v3, SHA-1 for v5). Concatenate the 16 bytes of the namespace UUID with the UTF-8 bytes of the name, hash the concatenation, take the first 16 bytes of the hash output, and interpret those bytes as a UUID with the version nibble set to 3 or 5. The namespace partitions the ID space: v5 UUIDs from the DNS namespace never collide with v5 UUIDs from the URL namespace. RFC 9562 defines four predefined namespaces: by DNS name, by URL, by OID, and by X.500 distinguished name. Organizations can mint their own namespace by generating a v4 UUID.

MD5 in v3 and SHA-1 in v5 — why a weakened hash is acceptable here, since the ID is not a security control

Version 3 uses MD5 and version 5 uses SHA-1, choices dating to their specification dates and available implementations. For name-based UUIDs, this distinction is immaterial because the hash function is not a security boundary or cryptographic control. The UUID is not proving authenticity or integrity; it is simply converting a variable-length string into a fixed 128-bit value. The attack model is irrelevant because UUIDs are stored and compared as opaque values, not as proofs or security controls. New implementations should use v5 (SHA-1) rather than v3 (MD5), not for compelling security reasons but because v5 is the modern standard and widely available.

The predefined namespaces — DNS, URL, OID and X.500, and when to mint your own

RFC 9562 specifies exactly four predefined namespace UUIDs with specific byte representations: 6ba7b810-9dad-11d1-80b4-00c04fd430c8 for DNS, 6ba7b811-9dad-11d1-80b4-00c04fd430c8 for URLs, 6ba7b812-9dad-11d1-80b4-00c04fd430c8 for OIDs, and 6ba7b814-9dad-11d1-80b4-00c04fd430c8 for X.500 distinguished names. A v5 UUID derived from the DNS namespace and the name www.example.com will always be identical and will never collide with a v5 UUID from the URL namespace. Using a predefined namespace ensures interoperability: if multiple teams independently use v5 with the DNS namespace, they generate identical UUIDs for the same DNS names. Choosing or minting a namespace is part of schema design.

Worked example — deriving a v5 UUID conceptually from the URL namespace and a record URL, step by step

Derive a v5 UUID conceptually from the URL namespace and the name https://example.com/api/users/42. The namespace UUID as 16 bytes is 6b a7 b8 11 9d ad 11 d1 80 b4 00 c0 4f d4 30 c8. The name is the UTF-8 string https://example.com/api/users/42, which is 30 bytes. Concatenate the namespace bytes (16) and name bytes (30) to get 46 bytes total. Compute the SHA-1 hash, producing a 20-byte hash. Take the first 16 bytes and interpret them as a UUID with the version nibble set to 5 and variant bits set to RFC standard. Computing it again with identical inputs produces the identical result. Most developers use their language UUID library to compute v5.

Where the pattern breaks — when names change, when the namespace is inconsistent across teams, and when inputs are secret

Name-based UUIDs assume that the name is stable and consistent across systems and import runs. If the same external record has different names in different systems, generating v5 from each name produces different UUIDs and fails to identify the same person. If a namespace is not agreed upon across teams (each team minting its own namespace for what is actually the same source), they generate different UUIDs and fail to match records. If the input is sensitive data, generating a v5 UUID means the UUID is a public, deterministic value that anyone can look up if they know the inputs. Determinism breaks when inputs change or namespace definitions are inconsistent.

What this does not cover — the ToolAcre generator draws from the CSPRNG, so name-based IDs need your language's UUID library

ToolAcre generates v4 UUIDs only, drawn from the browser cryptographically secure generator for independence. Name-based UUID derivation requires your language UUID library or an implementation that computes SHA-1 and formats the result correctly. This post explains the concept and use cases; implementing v5 generation is straightforward in any language with access to standard cryptographic libraries. The mechanics of v5 derivation are simple; the challenge is integrating it into a system schema where the namespace is stable, the name is consistent, and the approach is well-documented for your team. Development teams should document namespace choices.

Takeaway: deterministic when you need it, random otherwise — use v5 for stable mappings and the ToolAcre generator for everything that should be unguessable

Use v5 for stable mappings between external identifiers and your internal records. The determinism prevents duplicate imports and makes matching records across systems straightforward and reliable. Do not use name-based UUIDs for identifiers that need to be unguessable or for scenarios requiring strong confidentiality and secrets. ToolAcre generates random v4 UUIDs for identifiers that must be independent and distinct without predictability. When your systems need deterministic IDs that map inputs to fixed identifiers, your language UUID library can compute them. Determinism is a powerful feature when you control the input..