Extract emails from text
Transformed in your browser · nothing is uploaded
Finds everything shaped like an email address, removes duplicates and lists them one per line in the order they first appeared. Two more tabs do the same for URLs and for phone numbers, reading the same box, so one paste answers all three questions.
How to use the extract emails from text
The pattern is deliberately permissive rather than validating. A genuinely RFC-compliant email pattern runs to thousands of characters and still rejects addresses that work in practice, so validating here would quietly drop real addresses and you would never know which. Finding things shaped like an address and putting the list in front of you is the safer trade: a false positive is obvious to a human reading the list, and a missing address is invisible.
What that means concretely. Anything of the form [email protected] with a top-level domain of two letters or more is found. Case is preserved and compared exactly, so [email protected] and [email protected] both appear in the list as separate entries even though they almost certainly reach the same mailbox; addresses are case-insensitive in the domain and, by the letter of the specification, case-sensitive in the local part, so lowercasing them for you would be a guess. Internationalised addresses with non-Latin characters are not found at all. Obfuscated ones written as "name at example dot com" are not addresses in shape and are missed by design.
The URL tab wants a scheme, so it finds http:// and https:// links and ignores a bare example.com. Trailing sentence punctuation is stripped, which is right for a full stop at the end of a sentence and wrong for the rare URL that genuinely ends in a closing bracket; a Wikipedia link with a parenthesised title is the case that loses a character.
The phone tab is the one to check by eye. Phone numbers cannot be matched reliably, and this does not pretend otherwise: it finds runs of seven digits or more with the usual separators and an optional country code. That is enough to pull real numbers out of a document and loose enough that an invoice number such as 2024-118 comes back as a phone number, because it has seven digits and a separator. Read the list rather than trusting it.
One legal note that matters more than any of the above. A list of addresses lifted out of a page or a thread is personal data under the GDPR and comparable laws elsewhere. Extracting it does not make sending to it lawful, and unsolicited mail to harvested addresses is the specific practice those rules exist to restrict.
What people use it for
- Pulling the addresses out of a forwarded thread with fifty people on it
- Getting one address per line from a contact list someone pasted into a message
- Collecting the links out of a block of text so they can be checked
- Finding the phone numbers in a document without reading it line by line
- Auditing what a page links to before publishing it
Questions
No, deliberately. A validating pattern is enormous and still rejects working addresses, which would lose real ones without telling you.
Yes, keeping the first occurrence so the order matches the source document.
Because they are compared exactly. Domains are case-insensitive but local parts are not, by the letter of the standard, so folding the case for you would be a guess.
No. Anything written as "name at example dot com" is not shaped like an address, which is precisely what the obfuscation is for.
No. An address with non-Latin characters in it will be missed. Those are still rare enough that supporting them would cost more false positives than it saves.
RFC 5321 sets the local part at a maximum of 64 octets and the domain at 255. Anything longer is not a valid address, whatever accepted it.
Yes. [email protected] is found whole, tag included. Stripping the tag would change which mailbox rule it matches.
Because a bare domain is indistinguishable from ordinary words in running text. Requiring a scheme is what keeps "e.g." out of the results.
Trailing punctuation is stripped, since a URL at the end of a sentence usually collects a full stop or a bracket that is not part of it. A link that genuinely ends in one loses that character.
Because 2024-118 is seven digits with a separator, which is exactly what a phone number looks like. Phone matching cannot be made reliable, so check the list.
Seven digits. Shorter internal extensions and short codes are missed, which is the trade for not returning every year and price in the document.
A leading plus and up to three digits, yes. It does not check that the code exists or that the number is the right length for that country.
Yes. Results appear in the order they were first seen in the text, which lets you match them back against the source.
Copy the text out and paste it in. This reads text, not files, so anything you can select and copy will work.
Yes, including the ones in mailto links and attributes, since it scans the raw text you paste. Strip the tags first if you want only the visible ones.
Extracting is not consent. Harvested addresses are personal data, and unsolicited mail to them is exactly what the GDPR and CAN-SPAM restrict.
Under the GDPR, generally yes if it identifies a person, so [email protected] counts. A role address such as [email protected] is a weaker case but not a safe one.
Each tab deduplicates its own results. Switching tabs re-reads the same text and gives a separate list.
As much as your browser will hold. The scan is a single pass and runs as you type, so it stays responsive on a long document.
No. It is scanned in your browser and nothing is sent anywhere.