Text & everyday tools · Text Toolkit
How URL and email extraction know where an address ends
· How it works
text-extraction urls email-addresses
Looks at the two hard decisions in extracting links and addresses from prose — where a URL stops and how strict an email pattern should be — and why a pragmatic pattern beats the full RFC grammar on real text.
The link with a full stop glued on — why naive extraction returns https://example.com. and breaks the link
Meeting notes often put a useful link at the end of a sentence: `Agenda: https://example.com/roadmap.` A search that accepts every non-space character would include the final full stop. The visible text looks reasonable, but the extracted value no longer matches the address the writer intended to share or revisit.
ToolAcre first finds an HTTP or HTTPS sequence, then removes trailing full stops, commas, semicolons, colons, exclamation marks and question marks. The cleanup is deliberately attached to the match boundary rather than applied across the document. Punctuation elsewhere in a URL remains available when the initial pattern permits it.
The link with a full stop attached: why extraction must exclude trailing punctuation
The URL pattern begins only at `http://` or `https://` and continues until whitespace or one of several closing delimiters appears. Quotes, angle brackets, a closing parenthesis and a closing square bracket all stop the match. That rule lets prose such as `(https://example.com/notes)` return the address without its surrounding bracket.
This is a practical boundary rule, not a general parser for every possible URI. An opening parenthesis can remain inside a match while a closing one terminates it, so an address whose own path genuinely contains that character may be shortened. The extractor favors common prose boundaries and leaves unusual cases for manual review.
How strict should email matching be — what RFC 5322 technically permits, and why matching all of it produces worse results on real text
Email extraction makes a similar trade-off. It looks for letters, numbers and familiar mailbox punctuation before `@`, then a domain-like sequence followed by a dot and at least two letters. That pattern catches ordinary addresses embedded in notes without attempting to reproduce every construction admitted by the full email grammar.
The narrower pattern is useful because extraction and validation answer different questions. Extraction identifies likely contact details in unstructured text; it does not establish that a mailbox exists, accepts mail or belongs to the named person. Treat the result as a reviewable list, especially when punctuation or uncommon mailbox syntax appears nearby.
De-duplication and order — one copy of each address, in order of first appearance, so the list mirrors the source
For URLs and email addresses, ToolAcre removes duplicates with a `Set`. JavaScript sets retain insertion order, so the first occurrence determines where an item appears in the result and later identical matches disappear. The output therefore follows the source from top to bottom instead of alphabetizing contacts or links without permission.
Equality is exact. `https://example.com` and `https://example.com/` remain separate, as do email addresses whose letter case differs. The extractor does not normalize domains, remove URL fragments or decide that two destinations are equivalent. Those changes could merge intentionally distinct text, so any further normalization belongs in a separate checked step.
Numbers too — pulling numeric values out of prose, and why decimal and thousands separators make 'a number' less obvious than it sounds
Number extraction accepts an optional minus sign, one or more digits and an optional decimal part introduced by a full stop. It can pull `-12`, `48` and `3.75` from prose. Unlike URL and email extraction, it returns every match directly, so repeated values remain repeated and preserve their occurrence count.
Grouped and localized forms expose the limit of that compact rule. `1,250` is returned as `1` and `250`, while `3,5` is also split rather than interpreted as a decimal comma. Currency signs, exponents and leading plus signs are not part of a match. Read the nearby text before assigning meaning to any extracted number.
Numbers too: signed integers and decimals, without treating grouped values as one number
Try these notes: `Email sam@example.com, then review https://example.com/plan. Backup: sam@example.com. Budget -12.5; seats 24. Reference (https://example.com/help).` Extract emails returns one `sam@example.com`. Extract URLs returns the plan and help addresses without the sentence full stop or closing parenthesis, in that order.
A manual scan should find two email appearances but one unique email result, two distinct URL appearances and two number matches. Extract numbers returns `-12.5` and `24`; digits inside the example domains do not add matches because none are present. This check separates occurrence counts from the de-duplicated contact and link lists.
What this does not cover — validating that an address actually exists, and exotic but valid addresses the pragmatic pattern misses
A matched email is only text shaped like the implemented pattern. ToolAcre does not query mail servers, send a test message or inspect domain records. A plausible-looking typo can therefore pass extraction, while an uncommon but usable address can be missed. Confirm important contacts through an appropriate trusted channel before relying on the list.
URL extraction likewise does not request a page, follow redirects or test whether a destination is safe or available. It also requires an explicit HTTP or HTTPS prefix, so a bare `example.com` is not returned. These limits keep extraction local and predictable, but they make human review part of any consequential workflow.
The takeaway — the Text Toolkit's extraction buttons make these trade-offs explicitly and give you a clean, de-duplicated list in your browser
The extraction buttons turn a mixed block of prose into newline-separated matches in the browser. URLs and emails use pragmatic patterns, trim or stop at common prose boundaries and return unique values in first-seen order. Numbers use a smaller signed-decimal pattern and retain duplicates, reflecting a different function contract rather than one universal extraction policy.
Paste only the notes you need, run Extract URLs and Extract emails separately, and compare each list with the source before copying it onward. The clean output removes repetitive scanning, while the documented boundaries show where judgment remains necessary. For unusual syntax, keep the original passage beside the extracted result during review.