List sorter and line tools
Converted in your browser · nothing is uploaded
Twenty-one operations on a list of lines: sort A–Z, Z–A or numerically, remove duplicates, reverse or shuffle the order, split a comma list into lines and join it back, number or prefix every line, chunk long text, and compare two lists for difference, intersection or union. Sorted as text, 10 comes before 9.
How to use the list sorter and line tools
The numeric sort exists because text sorting puts 10 before 9. Sorted as strings, comparison is character by character, and 1 comes before 9, so a list of version numbers or file sizes comes out in an order that looks arbitrary. This is the single most common complaint about sorting anywhere, and the reason spreadsheets ask whether a column is text or a number. The numeric sort reads each line from the start and stops at the first character that cannot be part of a number, so 1.234,50 is read as 1.234 and €9 carries no number at all and lands wherever it happens to fall. There is no natural sort here either: item10 still comes before item2, so pad with a leading zero if that matters.
Alphabetical sorting is case-insensitive and locale-aware, which means Apple and apple sort together instead of every capital arriving first. A plain byte comparison puts every uppercase letter before every lowercase one: Zebra before apple, because that is the ASCII order, and it is almost never what anyone wants. Locale awareness also puts ä beside a, not away at the end after z, and because the sort is stable, Émile and Emile compare as equal and keep the order they arrived in. Deduplication sits alongside because sorting is when duplicates become visible: it keeps the first occurrence and trims each line before comparing it, so leading and trailing spaces are ignored and a and a count as the same line. That is what a spreadsheet export needs, since a trailing space there is both common and invisible.
But the sort and the comparison do not agree about capitals, and that is worth knowing before it surprises you. The sorts fold case, so Apple and apple sort together. Everything that compares one line against another — removing duplicate lines, and the three two-list operations — compares them exactly: Apple and apple are two different lines and both survive. That is deliberate rather than an oversight, because a line in a list is usually an identifier, a filename or an address, and those can differ by case and mean something else. For email addresses, where case genuinely does not matter, lowercase both lists first with the case converter and then compare. The word-level passes go the other way and fold case, since a word is a word. And when a duplicate survives a list that really was lowercase throughout, the line is carrying something trimming does not remove: a non-breaking space, or a zero-width character pasted out of a web page. The text cleaner strips those.
The two word-level passes have a shape of their own as well. Sorting the words and removing duplicate words both split on whitespace and rejoin with single spaces, so neither has any line breaks left to put back: a pasted column goes in and one long line comes out. Use the line-level pass beside each of them when the list has to stay a list.
The shuffle is a Fisher-Yates, which is the only common shuffle that is actually uniform: every ordering equally likely. The obvious alternative, sorting by a random comparator, is subtly biased and the bias is large enough to matter. Microsoft used it in a browser-ballot screen in 2010 and one browser landed in the last position about half the time. Randomness comes from the browser’s cryptographic source rather than Math.random, which costs nothing here. What it does not do is take a seed, so a shuffle cannot be reproduced afterwards; the random order generator is the tool when a draw has to be shown to be fair.
Splitting and joining are the small piece of plumbing that sits between every two systems that disagree about how a list is written: a spreadsheet gives a column, an API wants a comma string; a config file has a comma string, a form wants one per line. Both directions trim each item and drop empty entries, which matters because splitting a, b, c naively leaves a leading space on b, and that space is invisible and causes exactly the class of bug where a value looks right and does not match. The separator is matched literally, character for character, which is why it starts at a bare comma: that splits a,b,c and a, b, c alike, since the space is trimmed off afterwards, while a separator of comma-and-space would find nothing in the first of those. Change it to a semicolon, a pipe or a pasted tab for anything else. What none of this does is quoting: a comma inside a value makes the joined line ambiguous, and Smith, John splits into two items. Data like that is CSV rather than a comma list, and that distinction is the whole reason CSV has a specification.
The prefix, suffix and bullet passes work on whole lines, never on the characters between them, which is why they have no ends to get wrong. Doing it with find-and-replace on the line break almost works and then does not: replacing every break with a break plus your prefix misses the first line, and the same trick for a suffix misses the last, because that line has no break after it. The field keeps exactly what you type, and the space is usually the point: - item is a Markdown bullet and -item is not. Nothing checks whether a prefix is already there, so running it twice gives you two.
Two more, and both are splitters of a sort. Chunking breaks at whitespace — a space or a line break, whichever falls last inside the budget — so words stay whole, with one exception it cannot avoid: a single word longer than the whole chunk, which in practice means a URL, is cut mid-word, because the alternative is an oversized chunk the box you are filling will reject. Chunks are trimmed, so a break landing exactly on a boundary is lost while breaks inside a chunk are kept. The default is 280, because that is one post on X; 160 is one SMS, 2,200 an Instagram caption, 3,000 a LinkedIn post. The SMS number is the one that catches people, because it is about encoding rather than characters: a single message holds 160 while the text stays inside the GSM alphabet, and one emoji or curly apostrophe switches the whole message to 16-bit encoding, where a single message holds 70. And the comparisons treat two lists as sets: what is in the first only, what is in both, both combined without duplicates. That answers the three questions that come up constantly with two exports and that a diff tool answers badly, because a diff cares about order and position and an export has neither.
What people use it for
- Alphabetising the words in a paragraph rather than the lines
- Reversing or de-duplicating a pasted column of lines
- Scrambling a sentence into a word bank for an exercise
- Turning a semicolon-delimited string back into one item per line
- Joining a spreadsheet column into one comma-separated line
- Converting a comma list into a properly escaped JSON array
- Quoting every line for a SQL IN clause with a prefix and a suffix
- Finding which addresses are in the old export and not the new one
- Breaking a long post into 280-character chunks
- Working as a line splitter: one delimited string in, one item per line out
Questions
Text comparison goes character by character and 1 precedes 9. Use the numeric sort for a column of numbers.
The same reason: the digits are compared as characters. There is no natural sort here, so pad with a leading zero if you need one.
No. Apple and apple sort together; a plain byte comparison would put every capital ahead of every lowercase letter. Note that the comparisons are the other way round: they match lines exactly.
Yes, and so is removing duplicate lines. Apple and apple are two different lines and both survive. That is deliberate: a line in a list is usually an identifier, a filename or an address, and those can differ by case and mean something else. For email addresses, lowercase both lists first with the case converter.
No. The two word-level passes split on whitespace and rejoin with single spaces, so a pasted column comes back as one long line. Use the line-level sort or de-duplication to keep a list a list.
Only when a single word is longer than the whole chunk, which in practice means a URL. Everything else breaks at the last space or line break inside the budget.
Next to their base letter, following locale rules: ä beside a, not after z. Émile and Emile compare as equal and keep their original order.
It reads from the start of the line and stops at the first character that is not part of a number, so it reads 1.234 from the first and nothing from the second. Clean the column first.
Usually a capital. Lines are compared exactly, so Apple and apple are two lines; lowercase the list first if you want them folded. It is not spaces at the ends — each line is trimmed before the comparison. If the list really is one case throughout, the survivor is carrying a character trimming leaves alone: a non-breaking space, or a zero-width character out of a web page. Run it through the text cleaner. Words, unlike lines, are compared without case.
Yes, it is a Fisher-Yates shuffle: every ordering equally likely. Sorting by a random comparator is biased and noticeably so.
Not here; the randomness is the browser’s cryptographic source and takes no seed. The random order generator does take one, so a draw can be re-run.
A blank line between them. Everything before it is the first list, everything after it the second.
A diff cares about order and position. These comparisons treat each list as a set, which is what two exports need.
Yes, any string: a semicolon, a pipe, a tab pasted in. It is matched exactly rather than as a set of characters, so a two-character separator only splits where both characters appear in that order.
No. It splits on the separator wherever it appears, so a quoted comma inside a value will split it. Use a CSV parser for that.
Yes, in both directions, and an empty item is dropped rather than carried through. A leading space on " b" is what makes a lookup fail invisibly.
A bullet needs the space: "- item", not "-item". The field keeps exactly what you type.
Run an apostrophe as a prefix, run it again as a suffix, then join the lines with a comma.
Yes. Nothing checks whether one is already there, so paste the original back in before you run it again.
280 for a post on X, 160 for one SMS, 2,200 for an Instagram caption, 3,000 for a LinkedIn post. An emoji counts as two characters here, so leave headroom.
It depends on the operation. The sorts, the de-duplications, the shuffles, the numbering and the prefix passes drop them, because a numbered or prefixed empty line is never wanted. Reverse line order keeps them where they fall, and the three comparisons read one as the boundary between the two lists.