English

Developer tools · UUID generator

What Makes a UUID String Well Formed, and What a Checker Can't Know

· How it works

uuid cryptography browser-apis

Multiple UUID representations shown side-by-side: canonical 8-4-4-4-12, uppercase, with braces, hyphenless, and with urn: prefix
Original ToolAcre vector illustration

Uppercase, braces, urn: prefixes and missing hyphens all turn up in real input. This post defines the canonical form, shows what a lenient validator should accept, and separates well-formedness from existence.

The 400 that should have been a 404 — how sloppy UUID validation produces confusing API errors

An API endpoint receives an identifier from a client: {12345678-90AB-CDEF-1234-567890ABCDEF}. The validation code checks if it matches /[0-9a-f]{32}/ and rejects it as invalid. The client gets a 400 Bad Request where 404 Not Found was meant. The identifier is well-formed—it is a valid UUID in braced format—but the validator is too strict. Conversely, an endpoint that accepts any 32-character hex string (without dashes) will accept 123456789012345678901234567890123456, parse it as valid, and miss the typo. RFC 9562 defines the canonical textual representation, but real-world input arrives in five different formats, and a validator that only accepts the canonical form will reject 1 to 5 percent of well-intended input.

The canonical textual form — 32 lowercase hex digits in 8-4-4-4-12, exactly 36 characters, as the standard specifies for output

The canonical textual form is 32 lowercase hexadecimal digits in five groups separated by hyphens: 8-4-4-4-12. Represented as 550e8400-e29b-41d4-a716-446655440000. The standard mandates lowercase for output; on input, case-insensitive matching is recommended. This form is unambiguous, parses into bytes the same way on every platform, and is what every UUID library outputs by default. If you are generating a new UUID from a CSPRNG, the canonical form is what you should produce and what ToolAcre produces. Real-world input deviates in predictable ways. Uppercase identifiers (550E8400-E29B-41D4-A716-446655440000) are common from systems that default to uppercase; they represent the same bytes and should be accepted after normalization to lowercase.

Variants you will meet in the wild — uppercase hex, {braces}, the urn:uuid: prefix and 32-character hyphenless forms, and which the standard says to accept

Braced form ({550e8400-e29b-41d4-a716-446655440000}) is standard output from Python's uuid module and Microsoft's systems; stripping the braces gives valid canonical form. The URN prefix (urn:uuid:550e8400-e29b-41d4-a716-446655440000) is defined by RFC 8141 for uniform resource names; removing the scheme and stripping-identifier prefix leaves canonical form. The hyphenless form (550e8400e29b41d4a716446655440000) is 32 hex digits without structure; it is valid bytes but loses the 8-4-4-4-12 grouping that makes versions and variants readable. These variants all map to the same 128-bit value. RFC 9562 section 3 states that on input, uppercase variants SHOULD be accepted. It does not forbid other variants; it says that on output, the canonical lowercase form MUST be used.

Version and variant sanity — whether to reject a UUID whose third group starts with 0 or whose fourth group starts with f

A well-formed validator should: accept the canonical 8-4-4-4-12 form in lowercase or uppercase; accept braced and urn: variants by stripping them and validating the core form; accept hyphenless 32-digit hex strings and format them as canonical for comparison; reject strings with the wrong number of hex digits or non-hexadecimal characters. The most common mistake is rejecting uppercase or braced input because the validator was hand-written to match the canonical form only. A sanity check on the version and variant can catch typos. If the third group starts with 0 or 9, the UUID is invalid or reserved; if the fourth group starts with e or f, the variant is not RFC 9562.

Worked example — six candidate strings run through a strict check and a lenient one, with the reasons each passes or fails

A lenient validator accepts these values; a strict validator can reject them. The ToolAcre well-formed check performs strict validation: it confirms the canonical 36-character form with dashes in the right places, verifies hex digits in every position, and checks that the version and variant bits are in range. It does not check that the UUID exists in your database or that it was generated from a cryptographically secure source; those are separate checks done by your application logic. Well-formed is not the same as real. A UUID string that parses correctly according to its shape might not identify any row in your database.

Well formed is not real — why a syntactically perfect UUID may not exist in your data, and why the checker should never be your authorization layer

A UUID that is perfectly formed might have been guessed or copy-pasted incorrectly. Format validation is the first gate; existence checks and authorization checks are the second and third. Running a lookup against the database for every format-invalid input is wasteful; rejecting format-invalid input before database queries saves time. The ToolAcre generator outputs standard 36-character UUIDs; if you are building your own validator, accept braced and urn: variants to match real-world input, and reject strings that fail the basic shape rules before asking your database. Implementing a strict validator requires regular expressions and edge-case handling. The canonical form is straightforward: /^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/i (case-insensitive). The braced form adds curly braces: /^\{[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\}$/i. The urn: variant adds a scheme: /^urn:uuid:[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/i.

What this does not cover — normalising IDs for storage and choosing a column type, which are separate decisions

A single regex that handles all variants is less readable but possible. Most validators normalize first: strip braces and urn: prefix, convert to lowercase, then match the canonical pattern. The version and variant bits can be checked after pattern matching by examining position 14 and position 19 as described in the 403 article. Handling invalid input gracefully is part of validation design. When a client submits a malformed UUID, do not expose the regex pattern or the internal validation rules in the error message. Return a clear error: "Invalid UUID format. Expected 8-4-4-4-12 format, such as 550e8400-e29b-41d4-a716-446655440000. " Do not attempt to correct the input; ask the client to resubmit.

Takeaway: validate shape early, look up existence separately — the ToolAcre check confirms shape in the browser before you touch a database

Some systems log invalid input for security auditing (detecting attempted injection or format confusion attacks). The ToolAcre validator rejects non-canonical forms with a clear error message and does not attempt to auto-correct. Why canonical form matters for interoperability: if one system stores UUIDs as hyphenless hex and another stores them as canonical 8-4-4-4-12, comparing them for equality requires normalization. Uppercase vs lowercase requires case-insensitive comparison. Braced vs bare requires stripping. These variations make bulk operations (imports, migrations, comparisons) harder. Standard tools that output canonical form reduce friction. The ToolAcre generator always outputs 36-character lowercase canonical form; when you import UUIDs from other systems, normalize them to this form in your ETL process to ensure consistency.