Developer Comparing

Text compare

Changed +3 −1
7 unchanged
1 1 Dear
2 2 Sam,
3 3 Thanks
4 4 for
5 5 the
6 + quick
6 7 update.
7 − Best,
8 + Kind
9 + regards,
8 10 Alex

Compared in your browser · nothing is uploaded

Local · words, not lines · formatting is not compared

Compares two drafts of the same piece of writing a word at a time, so an edit inside a sentence marks the two or three words that actually changed instead of replacing the paragraph. A word is whatever sits between two spaces, which means where the lines happen to break stops mattering. This page opens in word mode; the diff checker is the line-by-line one.

How to use the text compare

1 Paste the earlier draft on the left and the later one on the right.
2 Read the marked words. By word is already on: a word here is whatever sits between two spaces, so punctuation travels with the word it is attached to, and only the wording is compared — bold, italics and headings are not in the text you pasted.
3 Turn on ignore case where one draft has been through a style pass and capitalisation is not the question.
4 Switch By word off for anything where a line is a unit rather than a sentence — a config file, or two exported lists.

By word is the setting this page is about, which is why it is already on when you arrive, and splitting on whitespace rather than on newlines is what makes it usable on prose. The reason is the wrapping. The same paragraph hard-wrapped at 72 columns in an editor and soft-wrapped in a document is the identical sentence with the newlines in different places; compared line by line every line differs, compared word by word nothing does. That is also why a paste out of a PDF, which carries a line break at the end of every printed line, only makes sense in this mode.

What trips a prose comparison up is not the writing, it is the characters a word processor substitutes while you type. Word and Google Docs turn a straight apostrophe into a curly one, straight quotes into typographic pairs, and two hyphens into an em dash. None of the substitutions are visible at reading size and each of them is a different character, so a sentence typed in a document and the same sentence typed in a code editor differ in ways nobody can see. A curly apostrophe sits inside the word, so it shows as that word changed; a non-breaking space between words does not, because any run of whitespace is just a separator here. If a comparison is full of single words marked as changed and they look identical, that is what happened, and running both sides through the text cleaner first is the fix.

Nothing about the layout survives the paste. Bold, italics, heading levels, comments and tracked changes are formatting rather than text, so a draft where the only change was making a sentence bold reads as identical, correctly. The same goes for a move: swapping two paragraphs shows as a deletion in one place and an addition in another, because a longest-common-subsequence match has no idea of a paragraph as a thing that can travel.

One thing this deliberately does not produce is a patch. A patch is addressed by line number and this comparison is not thinking in lines at all; for code, config or two exported lists, where a line is a unit and a patch is the output you want, the diff checker is the page.

What people use it for

  • Finding the sentence an editor rewrote between two drafts
  • Checking a contract against the version that was agreed
  • Seeing which wording changed after a legal or compliance pass
  • Comparing a draft against what was actually published
  • Spotting an unintended edit in boilerplate copied between two documents

Questions

Whatever sits between two spaces. Punctuation travels with the word it is attached to, so removing a comma marks that word as changed.

Myers: an O(ND) difference algorithm
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding