Text cleaner
Transformed in your browser · nothing is uploaded
Twenty-four cleanup passes over one box of text: flatten curly quotes and em dashes to ASCII, collapse or remove whitespace, strip punctuation, symbols, emoji, digits, accents or HTML tags, drop blank and duplicate lines, and normalise Unicode to NFC, NFD or NFKC. Everything runs in the tab, and nothing is uploaded.
How to use the text cleaner
The invisible characters are the ones that actually cause trouble. Text copied from a website often carries zero-width spaces, non-breaking spaces and directional marks that look like nothing and break string comparisons, CSV imports and search. A field that refuses to match a value you can see is identical is almost always this, and it is worth knowing that a zero-width character is not whitespace, so removing the spaces does not remove it; Clean everything does. Smart quotes are the other frequent culprit: they are different codepoints from the straight ones, so a smart apostrophe in a config file, a CSV or a piece of code is a syntax error that looks perfectly fine on screen.
Whitespace comes in two passes because they answer different questions. Collapse reduces every run to a single space and trims each line, which is what prose wants: double spaces after a full stop go, indentation goes, line breaks stay. Remove all leaves no whitespace anywhere, breaks and tabs included, which is what a licence key or an IBAN wants after it has been copied out of a wrapped display and arrived with a line break in the middle of it. Blank lines are a third thing again: a line containing three spaces looks identical to an empty one and is not, which is why removing them with a find-and-replace on two line breaks leaves an uneven result.
What counts as a letter decides how the symbol passes behave. They match on the Unicode letter property rather than the A–Z range, so café, naïve and Müller survive while @, #, % and an em dash do not. A remover built on [^a-zA-Z0-9 ] deletes the é as well and quietly mangles every non-English name it meets, which is the single most common bug in this class of tool. Stripping the accents is therefore a separate, deliberate pass, and it decomposes the character and drops the mark, so é comes back as a plain e. A few characters look accented and are not: ø, ł and ß are distinct letters rather than a base plus a mark, so they come through unchanged. And ü becomes u here; ue is a German convention and u is the language-neutral choice.
Emoji and HTML each have a failure mode worth naming. Most emoji removers match a fixed codepoint range, which is why they leave fragments behind: a family emoji is seven codepoints joined with zero-width joiners, and a partial match deletes some of them. This matches the Extended_Pictographic property instead, so skin-tone modifiers go with the emoji they belong to and newly assigned emoji keep working. Most tag-stripping regexes have the mirror-image bug: they remove the tags and leave the contents of a script or style element behind as visible text. Stripping here removes those wholesale and turns block-level closing tags into line breaks so paragraphs survive. It is not a sanitiser. Sanitising means allowing safe markup, which needs an allowlist library; this produces plain text.
The last group is about comparison rather than appearance. An accented character has more than one valid encoding: é can be one codepoint or an e followed by a combining acute, and the two look identical and are not equal. Normalising both sides to NFC before comparing is the fix, and NFC is what to store, being the shortest and the most widely expected. NFD is the decomposed form, which is what macOS filenames historically used and where this problem is most often met. NFKC goes further and folds compatibility characters — fullwidth letters become normal ones, the fi ligature splits — which is right for building a search key and wrong for preserving what someone typed. Deduplication sits alongside because it depends on all of it: words are compared without case, lines are compared exactly after trimming, and the first occurrence is the one that survives.
The two word-level passes — de-duplicate the words, sort the words — differ from their line-level neighbours in a way that is easy to miss until it happens to you. A word pass splits on whitespace and rejoins with single spaces, so it has no line breaks left to put back: a tidy pasted column goes in and one long line comes out. That is inherent to working on words rather than lines, not a setting, so use the line passes when the shape of the list has to survive. Case is where the two levels part company as well. Words are folded, lines are not, so Apple and apple are one word and two lines — deliberately, because an identifier, a filename or the local part of an email address can differ by case and mean something else. Lowercase the text first if you want lines folded too.
What people use it for
- Stripping the accents from a name before it becomes a filename
- Taking the repeats out of a pasted keyword list
- Stripping the emojis out of a column before a CSV import
- Closing up the blank lines in text copied out of a PDF
- Stripping order numbers and amounts before sharing a support ticket
- Clearing punctuation out of text before counting the words
- Removing every space from a licence key that was copied with a line wrap in it
- Reducing a reference field to letters, digits and spaces for a legacy import
- Turning a block of HTML into readable plain text
- Working as a Unicode normaliser: NFC, NFD or NFKC so two identical-looking strings finally match
- Collapsing the double spaces and trailing whitespace in a pasted paragraph
Questions
Curly quotes and dashes become straight ASCII, runs of whitespace collapse, zero-width and non-breaking spaces go, and blank lines are dropped.
Copying from a web page or a word processor frequently brings zero-width and non-breaking spaces along. They break comparisons and imports.
They are different characters from straight quotes. In code, a CSV or a config file they are a syntax error that looks correct.
Probably a zero-width character, which is not whitespace and survives a space pass. Clean everything removes those.
Collapse leaves one space between words and trims each line. Remove all leaves no whitespace anywhere, line breaks and tabs included.
Yes. A line of three spaces looks empty and is not, which is why a find-and-replace on double line breaks leaves some behind.
No. The symbol passes match the Unicode letter property rather than A–Z, so café survives intact. Removing accents is a separate pass.
Because it is a distinct letter, not an o with a mark on it. The same goes for ł and ß.
German convention says ue; most other contexts use u. This gives u, since it is language-neutral.
They are removed, so isn’t becomes isnt. Right for text analysis, wrong for anything a person will read.
They match a fixed codepoint range. A family emoji is seven codepoints joined together, so a partial match leaves the remains. This matches the Extended_Pictographic property instead.
Yes. Removing only the tags would leave the JavaScript in the output, which is the bug in most tag-stripping regexes.
No. Stripping produces plain text. Sanitising means allowing safe markup through, which needs an allowlist library.
Because an accented character has more than one valid encoding. Normalise both to NFC before comparing.
NFC, almost always: the shortest and the most widely expected. NFKC also folds fullwidth letters and ligatures, which is right for a search key and wrong for preserving input.
Only digits are removed, so 2024-118 leaves a hyphen. Run the punctuation pass afterwards if you want both gone.
The first, so the order you had is the order you keep. Words are compared without case; lines are compared exactly, after trimming.
You used a word pass. Removing duplicate words and sorting the words both split on whitespace and rejoin with single spaces, so a pasted column comes back as one long line. Use the line passes to keep a list a list.
Yes. cat and cat, are two different tokens and both survive a de-duplication. Strip the punctuation first if that is a problem.
No. Every pass runs in this tab, and there is no practical limit on how much text you paste in.