Text & everyday tools · Password Generator
The EFF wordlists explained: long, short and unique-prefix lists
· Background
passwords passphrases wordlists
Describes the three wordlists the Electronic Frontier Foundation published in 2016, the criteria behind them — memorability, distinctness, no confusing or offensive words — and how they differ from the original diceware list.
Why a wordlist needs designing — what went wrong with lists containing obscure words, near-duplicates and awkward strings
A wordlist is part of the generator’s choice space, so its exact contents and size matter. ToolAcre ships three lowercase English text files under the password app and exposes matching option records. Build-time tests count nonblank lines and fail if a file no longer matches its advertised size.
The workbook adds editorial history about obscure words, near-duplicates and offensive terms. Those claims require EFF primary material not included in the cited implementation set, so this article does not repeat them from memory. It focuses on committed assets, labels, provenance and behaviors the repository can verify.
What this repository verifies about the three shipped wordlists
`eff-long.txt` contains exactly 7,776 nonblank entries and is selected by default. The option table describes it as the EFF long list, associates one five-dice-sized index range with each word and derives about 12.9 bits per unfiltered draw as log₂(7,776). The test recomputes that logarithm rather than trusting folklore.
The browser fetches this file from the page’s own origin on first load. Filtering by minimum and maximum word length removes ineligible entries and deduplicates in first-seen order before generation. Consequently, the full file size is not always the current pool size.
The long list ships with exactly 7,776 lowercase entries
`eff-short.txt` contains exactly 1,296 entries. Its interface description says the words are shorter and quicker to type, and calculates about 10.3 bits per unfiltered draw from log₂(1,296). A smaller pool means more word draws are required to reach the same calculated choice count, but this article sets no universal target.
The list is loaded on demand when selected and cached as a promise for the current page, preventing duplicate simultaneous downloads. A failed load is evicted so another selection can retry. Missing, empty and network-failure cases produce distinct human-readable errors.
The first short list ships with exactly 1,296 entries
`eff-special.txt` also contains exactly 1,296 entries. ToolAcre labels it “EFF Short, distinct prefixes” and states that the first three letters identify a word and no word is a prefix of another. The repository does not independently establish broader typo-forgiveness or minimum-edit-distance history, so those outline claims are not expanded here.
Its equal size gives the same unfiltered log₂(1,296) contribution per uniform draw as the other short list. The vocabulary differs, and browser tests confirm switching lists changes generated vocabulary rather than merely changing a label.
The distinct-prefix short list also ships with exactly 1,296 entries
The credits and third-party records say the redistributed EFF lists had their dice-roll columns removed while words and order remained unchanged. They record attribution and licence provenance. They do not contain the full human review process, frequency datasets or rejection criteria implied by the workbook outline.
That evidence boundary matters because list design history is not needed to verify current generator behavior. Reviewers can inspect every committed line, count entries and follow selection. Historical claims should wait for an EFF source rather than being inferred from a filename or product description.
Selection-history claims are omitted; committed files and provenance are the evidence
Publication does not weaken uniform index selection because list secrecy is not assumed. ToolAcre’s entropy model treats settings and vocabulary as known. The unknown value is the fresh sequence of indices selected through Web Crypto and rejection sampling. An attacker learning the file does not retroactively change how many eligible indices existed.
Publishing a particular generated phrase is different: that reveals the selected sequence itself. Therefore examples should discuss settings and pool sizes without printing credentials readers might copy. Never use a phrase from documentation as an account secret.
What this does not cover — wordlists in other languages and community-maintained variants
Only these three English lists ship. The limitations page says other languages were not added because unverified or unsuitable assets would undermine the calculation. This article does not survey community variants, licences or linguistic quality without those sources.
Users needing another language should choose tooling whose list, provenance, licence and size they can verify. Translating words ad hoc changes the pool and may introduce duplicates or associations, so the original estimate would no longer describe the new process.
The takeaway — the Password Generator draws from the EFF wordlists, so every word is familiar, distinct and typeable
Choose the long or either short list based on the committed properties and your destination’s accepted format, then inspect the eligible count after filtering. The generator calculates from the words actually loaded, so a narrower range is reflected rather than hidden behind the advertised file size.
The durable facts are three assets, exact tested counts, same-origin loading and uniform selection. Claims about their cultural history, universal memorability or suitability for every account are intentionally absent. The list is one input to generation, not a manager or security guarantee.