English

Text & everyday tools · Password Generator

Modulo bias explained: picking a random word from a list without skew

· How it works

passwords randomness cryptography

Uneven numbered buckets beside an even set produced after rejecting a tail
Original ToolAcre vector illustration

Shows why 'random number modulo list length' favours some words over others when the range does not divide evenly, how large the skew is, and how rejection sampling removes it.

A fair die and an unfair shortcut — why taking a random number modulo 7,776 is not the same as rolling dice

A fair source can still feed an unfair selection routine. If a program reads an integer and immediately takes the remainder after division by a list length, some indices receive more source values whenever the source range is not exactly divisible by that length. The flaw is in the reduction step, not necessarily in the bytes. ToolAcre avoids that shortcut before choosing any EFF word.

Physical dice provide an intuitive comparison only when complete outcome groups map evenly. The browser implementation has a different raw range, so it computes an acceptance window for the requested bound. This article discusses that shipped mechanism rather than claiming that browser bytes literally reproduce five dice. Both can select indices, but their procedures and audit evidence are distinct.

Where the bias comes from — a 32-bit value does not divide evenly into 7,776 buckets, so the first few words get one extra chance

For one byte and a bound of 100, the raw range has 256 possible values. Two complete groups of 100 fit, leaving 56 values. Reducing every byte modulo 100 gives indices zero through 55 three preimages each, while indices 56 through 99 receive only two. The input bytes may be uniform, yet the selected bucket distribution is not.

The implementation comment also derives the corresponding finite-range issue for a 7,776-entry list when two bytes are reduced directly. Those figures come from the actual ranges named in source, not from an assumed attack rate. The important review question is whether the leftover values are used, not whether the skew looks small in a handful of generated phrases.

How big the effect is — small for a large random range, but nonzero, and why cryptographic code refuses to accept it

The deterministic tests make the uneven tail observable. Bounds 3, 5, 7, 100 and 7,776 are deliberately awkward, while a power-of-two bound demonstrates the no-rejection case. Another statistical control uses the same Web Crypto source for the correct function and a naive modulo helper, so the test isolates the reduction strategy rather than blaming the entropy source.

ToolAcre does not convert those distribution checks into a crack-time prediction. Bias reduces uniformity, but translating that reduction into a particular attacker’s cost requires a complete threat model. The engineering requirement is cleaner: every eligible index should have the same number of accepted raw values, and the code can enforce that exactly.

The verified skew examples in the implementation comments and deterministic tests

`secureRandomInt` finds the smallest whole byte count capable of representing the highest valid index. It computes the byte range, subtracts the remainder after division by the bound, and calls the result `limit`. Values below that limit belong to complete groups; values at or above it are discarded before the modulo operation is allowed to run.

A new draw follows each rejection. The loop has a high fixed ceiling so a broken injected source cannot hang the tab forever; after repeated out-of-range values it throws instead of returning a biased answer. This failure path is part of correctness: refusing a suspicious source preserves the promise that a returned index came from the uniform acceptance window.

Alternatives — drawing exactly enough bits and discarding out-of-range values, or using a library's uniform integer function

Other uniform-integer designs are possible, but they are not the behavior of this package and are therefore not presented as interchangeable ToolAcre options. The engine exposes one audited path. `secureRandomChoice` validates that its input is a nonempty array and delegates to `secureRandomInt(items.length)`, making word selection a direct consumer of the bounded integer contract.

The random-password generator uses the same path for characters, then invokes a Fisher–Yates shuffle driven by secure bounded swaps. It does not use `sort` with a random comparator. Keeping one primitive underneath several features makes a source review tractable: fix or test uniform reduction once, then follow its callers.

Alternative designs are outside this module; the shipped path uses rejection sampling

The bound-100 regression test supplies every tail byte from 200 through 255 and then a final 42. Correct code consumes all 56 rejected bytes and answers from 42. A second focused case supplies 200 and 7; because 200 modulo 100 would be zero, returning 7 proves that the first byte was rejected rather than quietly reduced.

This worked example is deterministic by design and contains no credential. Production bytes remain private to the browser call and are not logged. The test source is injectable only so behavior can be forced at the boundary; the public package offers no production seeded mode that could reproduce a generated passphrase.

Worked example: bound 100 rejects byte 200 and accepts the following 7

Floating-point random reduction, library-specific uniform APIs and unrelated shuffling algorithms fall outside this implementation. Evaluating them would require their source and contracts. ToolAcre’s code is integer-based and explicitly bounded, so adding a survey of alternatives would blur the narrow claim the tests actually prove.

The article also does not infer that a uniformly generated password is suitable for every policy or device. Uniform index selection addresses one mechanism. Storage, reuse, clipboard handling, malware and destination constraints remain separate questions even when every list element had an equal chance of selection.

The takeaway — a passphrase generator must choose words uniformly; check the tool's Technical notes for how the Password Generator does it

A passphrase generator must not favour early entries merely because its raw range leaves a remainder. ToolAcre takes bytes from Web Crypto, rejects the incomplete tail and reduces only values inside complete equal-sized groups. The tests force both rejection and acceptance boundaries so the claim is not based on visual inspection of outputs.

When reviewing similar code, calculate the raw range for the chosen byte width, divide it by the requested bound and look for a discarded remainder. If no tail is rejected, demand another proof of uniformity. The wordlist can be public and the source cryptographic while a careless reduction still introduces avoidable skew.