Text & everyday tools · Text Toolkit
Full-width punctuation: why 。 and ! matter for sentence case and counting
· Background
text-tools unicode cjk-punctuation
Explains what full-width and ideographic punctuation are, why CJK text uses them, and why a sentence-case converter or sentence counter that only knows ASCII full stops gets bilingual text wrong.
The Japanese paragraph counted as one sentence — how ASCII-only tools miss 。 and ! entirely
Paste `Release ready. 次の版です。確認してください!` into a counter that recognises only the ASCII full stop and the Japanese portion may be absorbed into one long remainder. The visible marks are not decorative variants: they are the boundaries readers use, so ignoring them produces a misleading sentence total.
ToolAcre includes `.`, `!`, `?`, `。`, `!` and `?` in its sentence-matching set. That fixes the specific ASCII-only failure, but it does not make the counter a Japanese parser. The result still comes from runs of text divided by recognised punctuation, with no grammatical model deciding where a thought ends.
Half-width and full-width — the legacy of fixed-pitch CJK typesetting and the Unicode blocks that carry it
“Full-width” describes characters designed to occupy the width associated with an ideographic cell in fixed-width East Asian layout. Unicode preserves compatibility forms such as `!` and `?`, while Japanese also uses ideographic punctuation including `。` and `、`. Similar appearance does not mean identical code points or interchangeable semantics.
The Unicode Standard places fullwidth ASCII variants in the Halfwidth and Fullwidth Forms block, while U+3002 IDEOGRAPHIC FULL STOP and U+3001 IDEOGRAPHIC COMMA belong to CJK Symbols and Punctuation. That distinction matters to software: a regular expression must name or classify the actual characters it intends to recognise.
The ideographic full stop and its relatives — 。、!? and the full-width forms of ASCII punctuation
`。` normally closes a Japanese sentence, `、` separates material within one, and `!?` provide question and exclamation forms familiar in modern copy. Fullwidth `!` and `?` correspond visually to wide versions of ASCII marks; they are not converted automatically merely because an editor displays them at a similar size.
ToolAcre treats `。!?` as sentence terminators but not `、`, which is appropriate for its narrow counting rule. Repeated terminators are consumed with the preceding text as one matched run, so `本当!?` contributes one sentence rather than two. Quotes, brackets and editorial conventions receive no separate linguistic interpretation.
What a sentence-aware transform needs — recognising sentence ends across scripts before capitalising or counting
The Sentence case transform first lower-cases the entire input. It then capitalises a lower-case letter at the beginning, or after one of the six recognised terminators only when that terminator is followed by whitespace. Therefore `hello。 world` becomes `Hello。 World`, while `hello。world` leaves the second English word lower-case.
That whitespace condition corrects the outline’s broader claim that recognising an ending is sufficient. Japanese commonly begins the next sentence immediately after `。`, with no space, and Japanese characters generally have no upper-case form anyway. In mixed copy, add whitespace only when editorial style calls for it; do not insert it merely to drive conversion.
What this sentence-case transform actually recognises: a terminator followed by whitespace
The word figure uses whitespace-separated runs. A Japanese passage without spaces can therefore count as one “word” even when a reader identifies many lexical units. Conversely, punctuation without surrounding whitespace does not split a run: `end,start` is one word under this definition. The number is mechanical, not a language-aware token count.
Sentence counting is also punctuation-based. A full stop in `Dr. Smith` can create an extra sentence, while an unpunctuated line break creates no new sentence boundary. The counter recognises CJK terminators and repeated marks, but quotations, abbreviations, ellipses and malformed punctuation can still make its output differ from editorial judgement.
Spaces, words and punctuation-based counts: useful figures with explicit limits
Try `LAUNCH READY. 次の版です。 check names! FINAL PASS?` in the Text Toolkit. Sentence case produces `Launch ready. 次の版です。 Check names! Final pass?`: all cased text is lowered first, then the opening letters after terminator-plus-space boundaries are raised. The Japanese text remains visually unchanged because it has no case distinction.
Read the statistics beside that result as definitions, not verdicts. The sample has four punctuation-delimited sentence runs, while its word total follows the five whitespace-separated chunks rather than Japanese morphology. Character totals use grapheme clusters where `Intl.Segmenter` is available, and the no-spaces total removes Unicode whitespace before recounting.
Worked example: inspect the transform and every count instead of assuming linguistic analysis
This behaviour does not implement Japanese line-breaking rules, vertical composition, ruby annotations or kinsoku shori restrictions on where punctuation may appear. It also does not normalise halfwidth and fullwidth characters. If publication requires typography checks, use an editor or layout system with explicit Japanese-language support after the text transformation.
Locale-specific casing is outside the promise too. The converter uses default JavaScript Unicode mappings, lower-cases proper nouns and acronyms, and can mistake an abbreviation boundary for a sentence boundary when whitespace follows it. Undo is available, but editorial review remains necessary for names, branded capitalisation and bilingual style decisions.
The takeaway — the Text Toolkit's Sentence case recognises full-width CJK sentence ends, so bilingual copy is handled rather than half-handled
Full-width punctuation matters because code sees characters, not visual intentions. ToolAcre explicitly includes `。!?` in its punctuation-based sentence matcher, so mixed English-Japanese copy is not limited to ASCII endings. Its Sentence case rule is narrower: capitalisation occurs only at the start or after a recognised terminator followed by whitespace.
Use the Text Toolkit to expose those rules quickly, then interpret each figure in context. Sentence totals are delimiter estimates, words are whitespace-separated runs, and characters are reader-perceived graphemes when browser support permits. Those transparent limitations make the output useful without pretending that a compact browser function performs full linguistic analysis.