Text Lists

List sorter and line tools

What to do

Converted in your browser · nothing is uploaded

Local · sorts fold case, comparisons do not

Twenty-one operations on a list of lines: sort A–Z, Z–A or numerically, remove duplicates, reverse or shuffle the order, split a comma list into lines and join it back, number or prefix every line, chunk long text, and compare two lists for difference, intersection or union. Sorted as text, 10 comes before 9.

How to use the list sorter and line tools

1 Paste one item per line. A copied spreadsheet column already arrives that way.
2 Pick the operation. The sorts, the de-duplications and the shuffles work on what you pasted as it stands.
3 Fill in the field the operation needs: a separator for splitting and joining, the text to add for a prefix or a bullet, the characters per chunk for splitting long text.
4 For the two-list comparisons, put a blank line between the first list and the second.
5 Copy the result.

The numeric sort exists because text sorting puts 10 before 9. Sorted as strings, comparison is character by character, and 1 comes before 9, so a list of version numbers or file sizes comes out in an order that looks arbitrary. This is the single most common complaint about sorting anywhere, and the reason spreadsheets ask whether a column is text or a number. The numeric sort reads each line from the start and stops at the first character that cannot be part of a number, so 1.234,50 is read as 1.234 and €9 carries no number at all and lands wherever it happens to fall. There is no natural sort here either: item10 still comes before item2, so pad with a leading zero if that matters.

Alphabetical sorting is case-insensitive and locale-aware, which means Apple and apple sort together instead of every capital arriving first. A plain byte comparison puts every uppercase letter before every lowercase one: Zebra before apple, because that is the ASCII order, and it is almost never what anyone wants. Locale awareness also puts ä beside a, not away at the end after z, and because the sort is stable, Émile and Emile compare as equal and keep the order they arrived in. Deduplication sits alongside because sorting is when duplicates become visible: it keeps the first occurrence and trims each line before comparing it, so leading and trailing spaces are ignored and a and a count as the same line. That is what a spreadsheet export needs, since a trailing space there is both common and invisible.

But the sort and the comparison do not agree about capitals, and that is worth knowing before it surprises you. The sorts fold case, so Apple and apple sort together. Everything that compares one line against another — removing duplicate lines, and the three two-list operations — compares them exactly: Apple and apple are two different lines and both survive. That is deliberate rather than an oversight, because a line in a list is usually an identifier, a filename or an address, and those can differ by case and mean something else. For email addresses, where case genuinely does not matter, lowercase both lists first with the case converter and then compare. The word-level passes go the other way and fold case, since a word is a word. And when a duplicate survives a list that really was lowercase throughout, the line is carrying something trimming does not remove: a non-breaking space, or a zero-width character pasted out of a web page. The text cleaner strips those.

The two word-level passes have a shape of their own as well. Sorting the words and removing duplicate words both split on whitespace and rejoin with single spaces, so neither has any line breaks left to put back: a pasted column goes in and one long line comes out. Use the line-level pass beside each of them when the list has to stay a list.

The shuffle is a Fisher-Yates, which is the only common shuffle that is actually uniform: every ordering equally likely. The obvious alternative, sorting by a random comparator, is subtly biased and the bias is large enough to matter. Microsoft used it in a browser-ballot screen in 2010 and one browser landed in the last position about half the time. Randomness comes from the browser’s cryptographic source rather than Math.random, which costs nothing here. What it does not do is take a seed, so a shuffle cannot be reproduced afterwards; the random order generator is the tool when a draw has to be shown to be fair.

Splitting and joining are the small piece of plumbing that sits between every two systems that disagree about how a list is written: a spreadsheet gives a column, an API wants a comma string; a config file has a comma string, a form wants one per line. Both directions trim each item and drop empty entries, which matters because splitting a, b, c naively leaves a leading space on b, and that space is invisible and causes exactly the class of bug where a value looks right and does not match. The separator is matched literally, character for character, which is why it starts at a bare comma: that splits a,b,c and a, b, c alike, since the space is trimmed off afterwards, while a separator of comma-and-space would find nothing in the first of those. Change it to a semicolon, a pipe or a pasted tab for anything else. What none of this does is quoting: a comma inside a value makes the joined line ambiguous, and Smith, John splits into two items. Data like that is CSV rather than a comma list, and that distinction is the whole reason CSV has a specification.

The prefix, suffix and bullet passes work on whole lines, never on the characters between them, which is why they have no ends to get wrong. Doing it with find-and-replace on the line break almost works and then does not: replacing every break with a break plus your prefix misses the first line, and the same trick for a suffix misses the last, because that line has no break after it. The field keeps exactly what you type, and the space is usually the point: - item is a Markdown bullet and -item is not. Nothing checks whether a prefix is already there, so running it twice gives you two.

Two more, and both are splitters of a sort. Chunking breaks at whitespace — a space or a line break, whichever falls last inside the budget — so words stay whole, with one exception it cannot avoid: a single word longer than the whole chunk, which in practice means a URL, is cut mid-word, because the alternative is an oversized chunk the box you are filling will reject. Chunks are trimmed, so a break landing exactly on a boundary is lost while breaks inside a chunk are kept. The default is 280, because that is one post on X; 160 is one SMS, 2,200 an Instagram caption, 3,000 a LinkedIn post. The SMS number is the one that catches people, because it is about encoding rather than characters: a single message holds 160 while the text stays inside the GSM alphabet, and one emoji or curly apostrophe switches the whole message to 16-bit encoding, where a single message holds 70. And the comparisons treat two lists as sets: what is in the first only, what is in both, both combined without duplicates. That answers the three questions that come up constantly with two exports and that a diff tool answers badly, because a diff cares about order and position and an export has neither.

What people use it for

  • Alphabetising the words in a paragraph rather than the lines
  • Reversing or de-duplicating a pasted column of lines
  • Scrambling a sentence into a word bank for an exercise
  • Turning a semicolon-delimited string back into one item per line
  • Joining a spreadsheet column into one comma-separated line
  • Converting a comma list into a properly escaped JSON array
  • Quoting every line for a SQL IN clause with a prefix and a suffix
  • Finding which addresses are in the old export and not the new one
  • Breaking a long post into 280-character chunks
  • Working as a line splitter: one delimited string in, one item per line out

Questions

Text comparison goes character by character and 1 precedes 9. Use the numeric sort for a column of numbers.

MDN, Intl.CollatorMDN, String
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding