Text Transform

Text cleaner

What to do

Transformed in your browser · nothing is uploaded

Local · zero-width characters are the usual culprit

Twenty-four cleanup passes over one box of text: flatten curly quotes and em dashes to ASCII, collapse or remove whitespace, strip punctuation, symbols, emoji, digits, accents or HTML tags, drop blank and duplicate lines, and normalise Unicode to NFC, NFD or NFKC. Everything runs in the tab, and nothing is uploaded.

How to use the text cleaner

1 Paste the text. Anything out of a PDF, a web page or a word processor is the usual case.
2 Start with Clean everything, which does the flattening, the whitespace and the blank lines in one pass.
3 Pick a narrower pass if you want only one thing gone: punctuation, symbols, digits, emoji, accents, tags, duplicates.
4 Copy the result, or save it as a text file.

The invisible characters are the ones that actually cause trouble. Text copied from a website often carries zero-width spaces, non-breaking spaces and directional marks that look like nothing and break string comparisons, CSV imports and search. A field that refuses to match a value you can see is identical is almost always this, and it is worth knowing that a zero-width character is not whitespace, so removing the spaces does not remove it; Clean everything does. Smart quotes are the other frequent culprit: they are different codepoints from the straight ones, so a smart apostrophe in a config file, a CSV or a piece of code is a syntax error that looks perfectly fine on screen.

Whitespace comes in two passes because they answer different questions. Collapse reduces every run to a single space and trims each line, which is what prose wants: double spaces after a full stop go, indentation goes, line breaks stay. Remove all leaves no whitespace anywhere, breaks and tabs included, which is what a licence key or an IBAN wants after it has been copied out of a wrapped display and arrived with a line break in the middle of it. Blank lines are a third thing again: a line containing three spaces looks identical to an empty one and is not, which is why removing them with a find-and-replace on two line breaks leaves an uneven result.

What counts as a letter decides how the symbol passes behave. They match on the Unicode letter property rather than the A–Z range, so café, naïve and Müller survive while @, #, % and an em dash do not. A remover built on [^a-zA-Z0-9 ] deletes the é as well and quietly mangles every non-English name it meets, which is the single most common bug in this class of tool. Stripping the accents is therefore a separate, deliberate pass, and it decomposes the character and drops the mark, so é comes back as a plain e. A few characters look accented and are not: ø, ł and ß are distinct letters rather than a base plus a mark, so they come through unchanged. And ü becomes u here; ue is a German convention and u is the language-neutral choice.

Emoji and HTML each have a failure mode worth naming. Most emoji removers match a fixed codepoint range, which is why they leave fragments behind: a family emoji is seven codepoints joined with zero-width joiners, and a partial match deletes some of them. This matches the Extended_Pictographic property instead, so skin-tone modifiers go with the emoji they belong to and newly assigned emoji keep working. Most tag-stripping regexes have the mirror-image bug: they remove the tags and leave the contents of a script or style element behind as visible text. Stripping here removes those wholesale and turns block-level closing tags into line breaks so paragraphs survive. It is not a sanitiser. Sanitising means allowing safe markup, which needs an allowlist library; this produces plain text.

The last group is about comparison rather than appearance. An accented character has more than one valid encoding: é can be one codepoint or an e followed by a combining acute, and the two look identical and are not equal. Normalising both sides to NFC before comparing is the fix, and NFC is what to store, being the shortest and the most widely expected. NFD is the decomposed form, which is what macOS filenames historically used and where this problem is most often met. NFKC goes further and folds compatibility characters — fullwidth letters become normal ones, the fi ligature splits — which is right for building a search key and wrong for preserving what someone typed. Deduplication sits alongside because it depends on all of it: words are compared without case, lines are compared exactly after trimming, and the first occurrence is the one that survives.

The two word-level passes — de-duplicate the words, sort the words — differ from their line-level neighbours in a way that is easy to miss until it happens to you. A word pass splits on whitespace and rejoins with single spaces, so it has no line breaks left to put back: a tidy pasted column goes in and one long line comes out. That is inherent to working on words rather than lines, not a setting, so use the line passes when the shape of the list has to survive. Case is where the two levels part company as well. Words are folded, lines are not, so Apple and apple are one word and two lines — deliberately, because an identifier, a filename or the local part of an email address can differ by case and mean something else. Lowercase the text first if you want lines folded too.

What people use it for

  • Stripping the accents from a name before it becomes a filename
  • Taking the repeats out of a pasted keyword list
  • Stripping the emojis out of a column before a CSV import
  • Closing up the blank lines in text copied out of a PDF
  • Stripping order numbers and amounts before sharing a support ticket
  • Clearing punctuation out of text before counting the words
  • Removing every space from a licence key that was copied with a line wrap in it
  • Reducing a reference field to letters, digits and spaces for a legacy import
  • Turning a block of HTML into readable plain text
  • Working as a Unicode normaliser: NFC, NFD or NFKC so two identical-looking strings finally match
  • Collapsing the double spaces and trailing whitespace in a pasted paragraph

Questions

Curly quotes and dashes become straight ASCII, runs of whitespace collapse, zero-width and non-breaking spaces go, and blank lines are dropped.

Unicode, the Unicode standardMDN, working with strings
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding