English

Developer tools · UUID generator

From 128 Bits to 36 Characters: How UUID Text Encoding Works

· How it works

uuid cryptography browser-apis

A 128-bit value shown as bytes, then as 36-character hex with hyphens, then as 22-character base64url, illustrating the space trade-off
Original ToolAcre vector illustration

A UUID is 16 bytes, yet its familiar form is 36 characters. This post explains the hex doubling, the hyphens, the case rules, and the shorter encodings people use when the standard form is too long.

Why the column is wider than the value — the 16-byte value that costs 36 characters in text and what that means for URLs and storage

The storage column width explodes when you choose a UUID format. A 128-bit value is 16 bytes, but its text representation depends on the encoding: hexadecimal (36 characters with hyphens, 32 without), base64url (22 characters), base58 (22–23 characters), Crockford base32 (26 characters). If your schema stores UUIDs as VARCHAR(36), you are spending 36 characters in every row. In a table with 1 billion rows and no other columns, that is 36 gigabytes of text overhead compared to 16 gigabytes of binary. The choice is not just cosmetic; it affects query size, network round-trips, and cache pressure. The canonical format is 36 characters: eight hex digits, hyphen, four hex digits, hyphen, four hex digits, hyphen, four hex digits, hyphen, twelve hex digits.

Hex doubles everything — each byte becomes two characters, and four hyphens complete the 36

Each byte becomes exactly two hexadecimal characters (0–9, a–f). The hyphens are there for readability and legacy reasons from when UUIDs were first specified. Hexadecimal encoding doubles the byte count: 16 bytes become 32 hex digits plus 4 hyphens. It is the slowest encoding and the longest, but human-readable and supported everywhere. Case rules: RFC 9562 mandates lowercase for canonical output, but input is case-insensitive. Storing uppercase wastes an opportunity to normalize, so store lowercase and compare case-insensitively on input. Base64url encoding represents three bytes as four characters using a 64-character alphabet (A–Z, a–z, 0–9, minus, underscore). Sixteen bytes become 21 characters plus one padding character, totaling 22 characters. Base64url removes the padding and the standard characters (plus and slash) that are reserved in URLs.

Case rules — lowercase on output, case-insensitive on input, and why mixed-case comparisons cause silent mismatches

A UUID in base64url format saves 14 characters compared to hex and is valid in URLs without percent-encoding. The trade-off: it is less readable (lower-case letters look like digits; b, 8, B, and 8 are easy to confuse). Base58 is used by Bitcoin and other blockchains and removes ambiguous characters (0, O, I, l), making the result 22–23 characters while remaining readable. Crockford base32 (designed for checksummed ISBN-like formats) uses 26 characters and prioritizes correctness over brevity. The Microsoft GUID byte-order trap applies to UUID storage in some databases. RFC 9562 specifies network byte order (big-endian) for all bytes. Some Microsoft SQL Server configurations store GUIDs with little-endian byte order in the first three fields.

Shorter encodings — base64url at 22 characters, base58 and Crockford base32, with their trade-offs in readability and copy-paste safety

The same 128-bit value stored in big-endian and little-endian produces different hex strings. A UUID 550e8400-e29b-41d4-a716-446655440000 stored as a Microsoft GUID might be retrieved as 00840e55-9be2-d441-a716-446655440000 (bytes 0–3 and 4–5 and 6–7 reversed). If your system bridges between RFC-compliant and Microsoft systems, you must be aware of this and either normalize at the boundary or document which format you are using in each column. Worked example: the v4 UUID 9b2e4f1a-4f3e-4c1a-8a7d-1b2c3d4e5f60 in hexadecimal occupies 36 characters. As 16 bytes it is 9b 2e 4f 1a 4f 3e 4c 1a 8a 7d 1b 2c 3d 4e 5f 60. In base64url: split into three-byte chunks, convert to base64, strip padding: my5PGk8-TBqKfRssPTRPX2A. In hexadecimal without hyphens: 9b2e4f1a4f3e4c1a8a7d1b2c3d4e5f60 (32 characters).

The Microsoft byte-order trap — how a GUID's first three fields are stored little-endian, so the same bytes can print as two different strings

Base64url saves 14 characters; base58 would save roughly the same; hexadecimal is the standard. Choose based on your use case: if the identifier appears in URLs and every character matters, use base64url; if it appears in logs and UIs where humans read it, use hexadecimal canonical form; if you are building a blockchain system or distributed system where checksumming matters, use base58 or Crockford base32. When choosing a column type, store the value that optimizes for your actual access pattern. If you query UUIDs frequently and need case-insensitive matching, store binary(16) and let the database handle the representation. If you query by substring (searching for UUIDs that start with a prefix), hexadecimal is more readable in debug output.

Worked example — one identifier written as bytes, canonical hex and a shortened form, showing each conversion step

If you export to CSV and email to non-technical users, hexadecimal is more recognizable. If you are space-constrained (mobile app with local cache), base64url or base58 saves bandwidth. The ToolAcre generator outputs canonical 36-character hexadecimal format; if you need a different encoding, the well-formed check still works because it normalizes any valid representation before checking format. Performance considerations matter when encoding or decoding millions of UUIDs. Hexadecimal encoding is simple: convert each byte to two characters in O(1) time per byte. Decoding is equally simple. Base64 encoding and decoding use lookup tables and are slightly slower (roughly 2–3x slower than hex per byte, depending on hardware and implementation). Base58 is significantly slower because it is essentially base conversion and requires modular arithmetic.

What this does not cover — database column choices such as native uuid types versus binary(16), covered separately

If your system encodes or decodes UUIDs in a hot loop (high-frequency identifier generation, bulk export), hexadecimal is faster. If encoding happens infrequently and the 14-character savings matter, base64url is a reasonable trade-off. The ToolAcre generator outputs hex, so you are getting the performance advantage without sacrificing compatibility. String comparison semantics differ by encoding. Hexadecimal UUIDs can be compared as strings: 550e8400-e29b-41d4-a716-446655440000 < 550e8400-e29b-41d4-a716-446655440001 (lexicographic comparison works). Binary UUIDs can be compared as bytes: byte-by-byte comparison is the same as numeric comparison. Base64url and base58 encoded UUIDs, however, do not preserve the numeric order in lexicographic string comparison. If your system relies on lexicographic sorting of UUIDs (a surprisingly common pattern for building indexes or database keys), you must either use hexadecimal, binary, or a sortable UUID variant (v6 or v7).

Takeaway: keep the canonical form at boundaries — the ToolAcre generator outputs standard 36-character UUIDs and its check accepts strings in that form

The ToolAcre generator currently produces v4 UUIDs, which are not sortable by encoding order. Interoperability requires standardization on a single encoding. A system that accepts UUIDs in hex, base64, and base58 simultaneously must normalize all input to a canonical form before processing. This is possible but adds complexity. External APIs or databases might require a specific encoding: some APIs expect urn:uuid: prefixed hex, others expect hyphenless hex, still others expect base64url. Document your system's UUID encoding expectation clearly in API contracts. The ToolAcre generator always outputs canonical hex; if you need other encodings, perform the conversion explicitly and document the trade-offs (space, performance, readability, sortability) to the team.