Developer Extract

Extract emails from text

What to do

Transformed in your browser · nothing is uploaded

Local · finds addresses, does not validate them

Finds everything shaped like an email address, removes duplicates and lists them one per line in the order they first appeared. Two more tabs do the same for URLs and for phone numbers, reading the same box, so one paste answers all three questions.

How to use the extract emails from text

1 Paste your text into the first box.
2 Choose what you want done to it.
3 Copy the result, or save it as a text file.

The pattern is deliberately permissive rather than validating. A genuinely RFC-compliant email pattern runs to thousands of characters and still rejects addresses that work in practice, so validating here would quietly drop real addresses and you would never know which. Finding things shaped like an address and putting the list in front of you is the safer trade: a false positive is obvious to a human reading the list, and a missing address is invisible.

What that means concretely. Anything of the form [email protected] with a top-level domain of two letters or more is found. Case is preserved and compared exactly, so [email protected] and [email protected] both appear in the list as separate entries even though they almost certainly reach the same mailbox; addresses are case-insensitive in the domain and, by the letter of the specification, case-sensitive in the local part, so lowercasing them for you would be a guess. Internationalised addresses with non-Latin characters are not found at all. Obfuscated ones written as "name at example dot com" are not addresses in shape and are missed by design.

The URL tab wants a scheme, so it finds http:// and https:// links and ignores a bare example.com. Trailing sentence punctuation is stripped, which is right for a full stop at the end of a sentence and wrong for the rare URL that genuinely ends in a closing bracket; a Wikipedia link with a parenthesised title is the case that loses a character.

The phone tab is the one to check by eye. Phone numbers cannot be matched reliably, and this does not pretend otherwise: it finds runs of seven digits or more with the usual separators and an optional country code. That is enough to pull real numbers out of a document and loose enough that an invoice number such as 2024-118 comes back as a phone number, because it has seven digits and a separator. Read the list rather than trusting it.

One legal note that matters more than any of the above. A list of addresses lifted out of a page or a thread is personal data under the GDPR and comparable laws elsewhere. Extracting it does not make sending to it lawful, and unsolicited mail to harvested addresses is the specific practice those rules exist to restrict.

What people use it for

  • Pulling the addresses out of a forwarded thread with fifty people on it
  • Getting one address per line from a contact list someone pasted into a message
  • Collecting the links out of a block of text so they can be checked
  • Finding the phone numbers in a document without reading it line by line
  • Auditing what a page links to before publishing it

Questions

No, deliberately. A validating pattern is enormous and still rejects working addresses, which would lose real ones without telling you.

RFC 5321 §4.5.3.1, size limits for the local part and the domainMDN, working with strings
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding