Text & everyday tools · Text Toolkit
How case converters split camelCase and acronyms like parseHTTPResponse
· How it works
text-conversion developer-workflow unicode
Explains the three boundary rules a good case converter applies — separator runs, lower-to-upper transitions and acronym edges — and why parseHTTPResponse2Json is the test that exposes weak ones.
parse_h_t_t_p_response and other failures — why converting identifiers is harder than converting 'hello world'
Turning “hello world” into hello_world is easy; turning parseHTTPResponse into parse_h_t_t_p_response is a sign that an algorithm treated each capital as a word. A developer renaming JavaScript identifiers for a Python API needs the opposite: identify tokens first, then apply the destination’s joining convention. The same first step should power snake, kebab, camel and Pascal output, otherwise four buttons disagree about where the words begin.
Rule one: separator runs — spaces, hyphens, underscores and punctuation all mark a boundary, however many in a row
Runs of spaces, hyphens, underscores and other non-letter/non-number characters already mark boundaries. Treat a run as one separator: button--primary should become two tokens, not an empty token between the hyphens. ToolAcre’s case converter splits with a Unicode-aware expression for letters and numbers, so common non-ASCII letters are not thrown away simply for being outside A–Z. Whitespace normalisation is about naming conventions, not about changing the original source file automatically.
Rule two: lower-to-upper transitions — a lowercase letter or digit followed by a capital starts a new word
The pattern for a lowercase letter or digit followed by an uppercase letter inserts a boundary before the uppercase. That turns parseResponse into parse + Response and allows 2Json to become 2 + Json. This rule does not by itself separate the preceding word from its final digit: ToolAcre treats Response2 as one token before splitting the capital J. Whether you want Response + 2 is a style-guide decision, so inspect numeric identifiers instead of assuming every converter makes the same choice.
Rule three: acronym edges — a run of capitals followed by a capital-then-lowercase pair ends the acronym before the last capital
An acronym needs one more lookahead. In HTTPResponse, the capital run HTTP ends before R because R is followed by lowercase esponse. A rule matching the run and the capital-lowercase pair inserts a boundary there: HTTP + Response, not H + T + T + P + Response. When a name ends with all capitals, there is no lowercase suffix to mark a new word and the acronym stays together. That is the difference between tokenising identifiers and inserting a separator before every capital.
Digits and the edge cases — where '2Json' splits, and why no rule set satisfies every style guide
Digit conventions remain ambiguous: version2Parser, HTTP2Json and IP6Address do not all imply the same word grouping. ToolAcre groups digits with the preceding token until a subsequent capital triggers the next boundary. Likewise, a letter with locale-sensitive uppercasing may expand or behave differently from ASCII: Turkish dotted/dotless i and German ß deserve manual review. The tool makes a deterministic transformation, not a claim to understand the semantic naming of a variable.
Worked example — running parseHTTPResponse2Json and a hyphenated CSS class through snake, kebab, camel and Pascal case
Run parseHTTPResponse2Json through the actual case converter. It produces parse_http_response2_json in snake case, parse-http-response2-json in kebab case, parseHttpResponse2Json in camel case and ParseHttpResponse2Json in Pascal case. For button--primary, consecutive separators collapse into button_primary and button-primary. These outputs expose both the successful acronym rule and the digit grouping choice. Test against your destination API before using search-and-replace across a codebase.
What this does not cover — locale-specific casing such as the Turkish dotted i, which has its own post
This mechanism does not implement every language’s case folding, infer that an acronym stands for a particular domain term, or rename references throughout a program. String and identifier conversion are different from refactoring source code with symbol awareness. The tool cannot promise that a Python service accepts a renamed JSON field; callers may still depend on the old spelling. Run your schema tests after changing public field names.
The takeaway — the Text Toolkit's case converter applies these three rules and shows the result instantly so you can check it before pasting into code
First identify tokens using separators, transitions and acronym edges; only then join and case them. The Text Toolkit makes those rules visible with immediate output and lets you undo a conversion. For an acronym-heavy field name, compare the real snake/kebab results rather than blindly applying a simple capital-letter regular expression to an entire repository.