English

Developer tools · UUID generator

How Many Random Bits Does a v4 UUID Have? 122, Not 128

· How it works

uuid cryptography browser-apis

A 128-bit field with 6 fixed bits (version and variant) highlighted and 122 random bits shown, illustrating why collision probability is calculated from 122 bits
Original ToolAcre vector illustration

Six of the 128 bits in a random UUID are fixed by the standard, leaving 122 for randomness. This post shows how to estimate collision odds honestly and why real collisions come from broken generators, not from the mathematics.

The architect's question: will we ever get a duplicate? — where the worry comes from and why the answer depends on the generator

The architect asks: if we generate 10 million UUIDs per day for ten years, will we ever get a duplicate? The honest answer is: almost certainly not, if the generator is cryptographically secure; almost certainly yes, if the generator is broken. The RFC 9562 math is sound: a v4 UUID with 122 random bits has collision probability of roughly n² / 2 to the 123rd power, where n is the number of identifiers generated. For most real systems, this probability is negligible. The catch is that this formula assumes every bit is truly random. If the generator leaks or repeats or was seeded predictably, the formula is wrong, and duplicates become inevitable. The 128-bit layout includes four version bits (0100 for v4) and two variant bits (10 for RFC 9562), which are fixed and set by the standard.

Which six bits are spoken for — the four version bits and two variant bits, and why they are set rather than random

That leaves 122 bits for randomness. The formulation is sometimes called 2 to the 122 random bits, producing unique values. Using the birthday paradox approximation, the probability of at least one collision among n randomly generated values is roughly n² / 2 to the 123rd. For n = 1 million, this is (10^6)² / 2^123 = 10^12 / 9. 3 × 10^36, which is about 10^-25. For n = 10 billion, it is still about 10^-16. These are not "effectively zero"; they are "you will never observe this. " The birthday approximation gives a concrete way to calculate the risk: count the UUIDs you plan to generate, square that number, divide by 2 to the 123rd power. If the generator is the browser's crypto.getRandomValues, every bit is backed by operating-system entropy. If it is a Math.

The birthday approximation in plain terms — the probability of at least one collision among n identifiers is roughly n squared divided by 2 to the 123

For a UUID produced by an unsuitable or repeated-state generator, the mathematical model breaks down because its independence assumption is false. A fixed test seed, copied fixture or process snapshot can replay values even though the text still carries the version-4 nibble. Those are implementation defects, not evidence that the 122-field calculation was wrong. Another mundane source is copying the same literal identifier into several fixtures and later merging their data. When investigating a duplicate, preserve the generator, seed policy, process lifecycle and import history. Do not jump from one repeated value to a claim that independent CSPRNG output exhausted the UUID space.

Worked example — plugging a stated generation rate and time span into the approximation, with every step shown so you can substitute your own figures

Process forking without random-state resynchronization. A bug where Math.random was used instead of crypto.getRandomValues. A test fixture that was hand-crafted with the same UUID in multiple rows and was accidentally used in production. An old version of a UUID library that had a range limitation or state bug. None of these scenarios involve the birthday approximation mathematics; they involve broken implementation or operational mistakes. Calculate collision risk honestly for your system: count the rate of UUID generation (per second, per day, per year), project it over the time the system will run, and plug the total into the birthday formula. If your system generates 100,000 UUIDs per day for five years (182 million total), the collision probability is (1. 82 × 10^8)² / 2^123 ≈ 3. 3 × 10^-22, which is negligible.

Where duplicates actually come from — Math.random seeds, cloned virtual machines, forked processes with copied state, and copy-paste in fixtures

If you generate 10 million per second for a year (315 trillion total), the probability is (3. 15 × 10^14)² / 2^123 ≈ 10^-10, which is still vanishingly small. These estimates assume that every bit is independent and random. The ToolAcre generator uses crypto.getRandomValues, which gives you CSPRNG-backed randomness; the operation is the only part you need to trust. Never rely on the collision probability as an excuse to skip proper authorization checks. A UUID is not a password, not an access token, and not a secret even if it is 122 random bits. The uniqueness is the benefit; the unpredictability is a separate (and more important) property that prevents guessing. The birthday mathematics handles uniqueness; it does not address lifetime (should this UUID expire? ), secrecy (does it need to be hashed before storage?), or authorization (does possessing this UUID prove anything about the caller?). The ToolAcre generator gives you CSPRNG-backed UUIDs with 122 random bits, which means the uniqueness mathematics holds and the unpredictability is sound. Everything else—token validation, expiry, access control—is your application's responsibility. The probability formula assumes independence of each generated UUID from previous generations. If your system generates identifiers from a single CSPRNG instance, and each call draws new randomness from the OS, the independence assumption holds. If your system uses a cached CSPRNG state or a seeded generator without OS reseeding, the assumption breaks down. Collision risk increases dramatically if the entropy source is exhausted (occurs on some embedded systems or virtual machines under load) or if the random state is never reset between processes (process forking without re-seeding the CSPRNG).

What this does not cover — v1 and v7 uniqueness, which rests on timestamps and clock sequences rather than randomness alone

The ToolAcre fallback calls crypto.getRandomValues for each byte array and holds no application-level PRNG state. Platform internals remain the browser and operating system's responsibility. Simulating collisions with real generators shows the difference between theory and broken practice. A generator built with Math.random starting from the same seed will produce identical sequences; you will see the first UUID collision within a few hundred to a few thousand generated values, not after 2^60 values (the square root of 2^122) as the birthday approximation predicts. A generator using crypto.getRandomValues from sound OS entropy will produce collisions only when the theoretical probability becomes unavoidable (around 2^60 UUIDs), a number so large you will never reach it. A generator using a weak or re-used entropy source (common in poorly-implemented UUID libraries or test frameworks) will produce collisions somewhere in between.

Takeaway: trust the maths, audit the generator — the ToolAcre generator uses the browser's CSPRNG, which is the part that must not be faked

A batch test can catch an implementation that returns a constant or replays an obvious sequence, but a passing sample cannot prove future uniqueness. Production systems should still enforce a unique constraint wherever duplicate identifiers would corrupt data. Imports deserve special attention because two valid source systems may already contain the same literal identifier, and fixtures can be copied across environments. ToolAcre's tests check that a batch of 500 contains 500 distinct values; that is a regression check for this implementation, not a statistical guarantee. If a duplicate appears, preserve evidence and inspect generation, import, fixture and storage paths before attributing a cause.