Text & everyday tools · Text Toolkit
How slug generators fold accents: NFD decomposition explained
· How it works
url-slugs text-conversion javascript
Explains how Unicode canonical decomposition separates a base letter from its accent so 'Café Crème' becomes cafe-creme rather than caf-cr-me, and where decomposition alone falls short.
caf-cr-me — the common slug bug that mangles names, and why it happens
A weak slug routine can turn `Café Crème` into `caf-cr-me` when it deletes every character outside a narrow ASCII range. The visible accents disappear, but the underlying base letters disappear with them, leaving a URL that no longer resembles the article title. That damage is especially obvious in names, places and repeated editorial categories.
ToolAcre takes a different path in `slugify`. It first normalizes the input, then removes a specific range of combining marks while retaining the resulting letters. Only afterward does it lowercase the text and shape separators. The order is the reason `Café Crème` becomes `cafe-creme` instead of a fragment with missing vowels.
One character, two representations — how é can be a single code point or an e followed by a combining acute accent
Text that looks identical can have different internal sequences. An `é` may arrive as one precomposed character, or as an ordinary `e` followed by a combining acute mark. A content editor usually cannot see which representation came from a CMS, document or clipboard, yet a character-by-character filter can treat the two inputs differently.
That hidden difference matters when a replacement rule recognizes one form but not the other. ToolAcre avoids writing a separate replacement for each precomposed spelling. Normalization gives the slug pipeline a more consistent intermediate form, so supported accent marks can be removed by one later step while the base letters remain available for the URL.
Normalisation Form D — how canonical decomposition rewrites every accented letter into base letter plus combining marks
The implementation calls `.normalize("NFD")` before any lowercasing or separator work. For characters that have a canonical decomposition handled by the JavaScript runtime, this produces a base character followed by one or more combining marks. The function does not maintain its own catalog of French or Spanish spellings and does not inspect words semantically.
The outline says NFD rewrites every accented letter, but the source supports a narrower statement. Decomposition depends on the character, and the following removal expression covers code points from `U+0300` through `U+036F`. The article should therefore describe the behavior demonstrated by the code rather than promise universal accent removal for every script or mark.
NFD decomposes supported characters; the implementation does not promise that every accented letter separates
After normalization, `slugify` applies `/[̀-ͯ]/g` and replaces each matching mark with an empty string. In the decomposed form of `é`, the `e` does not match that range, while the acute mark does. Removing only the mark leaves the readable base letter that the earlier ASCII-only approach would have discarded.
This is accent folding, not a general text cleanup pass. The regular expression is deliberately placed before the rule for separators, allowing the base letter to participate as a letter later. If mark removal happened after unsupported runs had already been collapsed, a decomposed mark could influence separator placement and produce a less faithful slug.
The rest of the slug pipeline — lowercasing, collapsing non-alphanumeric runs to single hyphens, trimming leading and trailing separators, dropping emoji
The remaining pipeline lowercases the normalized text and replaces each run that is not a Unicode letter or number with the configured separator, which defaults to a hyphen. A second expression trims repeated separators from both ends. Emoji and punctuation therefore disappear as content, while adjacent unsupported characters become one boundary rather than several hyphens.
The outline describes a non-alphanumeric collapse, but the actual pattern uses Unicode property escapes, not an ASCII-only alphabet. Letters from non-Latin scripts can remain in the slug after lowercasing. The configuration confirms there is no transliteration step: symbols are removed, but retained letters are not automatically rewritten as approximate Latin spellings.
The rest of this slug pipeline keeps letters and digits from any script while replacing other runs with the chosen separator
Follow `Café Crème & Co. — Été 2024!` through the implementation. NFD separates the supported accented characters into base letters and marks. The mark-removal expression leaves `Cafe Creme & Co. — Ete 2024!`, and lowercasing produces `cafe creme & co. — ete 2024!` before punctuation has been processed.
The non-letter and non-number runs then become hyphens, yielding the meaningful sequence `cafe-creme-co-ete-2024` after leading and trailing separators are trimmed. The ampersand, full stop, dash and exclamation mark do not receive spoken names or custom substitutions. They serve only as boundaries between letter-and-number runs in this conversion.
What decomposition cannot do — letters like ø, ł, ß and æ have no accent to strip and need a transliteration table
Decomposition is not transliteration. Characters such as `ø`, `ł`, `ß` and `æ` are still Unicode letters after this pipeline, so the property-based filter keeps them rather than consulting a table for `o`, `l`, `ss` or `ae`. Claiming that users need a transliteration table may be useful design advice elsewhere, but no such table exists in this tool.
That distinction also explains why the result may be a valid project slug without being ASCII-only. Editors whose publishing system requires ASCII should check that separate system constraint before using the output. ToolAcre promises accent folding for the implemented decomposition-and-mark range; it does not promise language-aware spelling, reversible conversion or Latin output for every title.
Characters without removable marks remain letters; this tool has no transliteration table
The reliable takeaway is procedural: normalize first, remove the supported combining marks, lowercase, collapse unsupported runs and trim separators. Each stage has one visible responsibility, and their sequence preserves base letters before punctuation is discarded. That is enough to prevent the common `caf-cr-me` failure without inventing language rules that the source does not contain.
Paste the worked title into the Text case converter and select the slug option to inspect the final result in the same text box. If a title includes letters outside the demonstrated accent cases, review the output against the destination platform. The converter supplies a predictable browser-side transformation, while the editor remains responsible for route conventions.