PDF Size

Compress PDF

Saved —

Your file never leaves this device

Local · images re-encoded, structure rewritten

Most of the size of a large PDF is images. This decodes each embedded JPEG, scales it down if it is larger than the target, re-encodes it at the quality you choose, then rewrites the file with object streams and the document information fields cleared. Scanned documents typically shrink by half or more; a text-only PDF saves a few per cent. Drop in several files at once and the results come back as a ZIP.

How to compress a PDF

1 Drop the PDF in. Compression starts straight away. Several files at once is fine.
2 Move the quality and size sliders if the result is still too big, or if it looks too soft.
3 Download the smaller version, or take the whole batch as one ZIP.
4 If nothing could be saved, the tool says so rather than handing you a worse file.

Being precise about what this does is what lets you predict whether it will help. It walks the document looking for embedded images, and it re-encodes exactly one kind: baseline JPEG data, in plain colour or plain greyscale, with no custom decode array attached. Each of those is decoded, scaled down if its longest side exceeds the target, re-encoded at the chosen quality, and put back only if the result is genuinely smaller. Then the whole file is rewritten using object streams and with its title, author, subject, keywords, producer and creator fields cleared.

Everything else is left exactly as it was. Vector artwork, fonts, form fields and text are untouched. So are images stored in any other way, and that group is larger than people expect: PNG-style lossless images inside a PDF, JPEG 2000 images, images in a CMYK colour space, and the fax-style bilevel encoding that many document scanners produce for black-and-white pages. A PDF made entirely of those will come back with only the structural saving.

Which makes the answer to "why did nothing happen" nearly always the same. A text-heavy report is already efficiently compressed and there is no image data to work on. A greyscale scan is a partial case: re-encoding it goes through a colour canvas, so the replacement is often no smaller than the original and gets discarded on the spot. When the output would be bigger than the input, the whole file is rejected rather than offered, because handing back a larger file labelled "compressed" is worse than saying nothing could be done.

What it deliberately does not do is the list of things desktop tools do that lose information without telling you: it does not re-run text through OCR, subset or drop fonts, flatten form fields, downsample vector artwork to bitmaps, or discard layers. Every one of those saves space and every one of them can throw away something you needed.

One side effect looks like a privacy benefit and is only half of one. What gets cleared is the document information dictionary — the six fields a reader lists under Properties, including Author and Producer. Most PDFs written by a word processor, a layout application or Acrobat also carry a second copy of the same information as an XMP metadata packet attached to the document catalogue, and that packet is left exactly as it was: a test file whose XMP named its author still named that author after compression. The properties panel can therefore go blank while the name is still sitting in the bytes. Treat this as tidying rather than as redaction, and use something that rewrites XMP if a document genuinely has to be scrubbed before it leaves.

What people use it for

  • Getting a scan under an email attachment limit
  • Meeting an upload cap on a government or bank portal
  • Shrinking a phone-scanned contract before sending it on
  • Cutting a photo-heavy report down for a slow connection
  • Reducing a file without handing the document to a web service
  • Clearing the Author and Producer fields a reader shows under Properties

Questions

Because it is mostly text and vectors, which are already compressed efficiently. With no eligible image data to re-encode, only the structural saving of a few per cent is available.

ISO 32000-1:2008, the PDF 1.7 specification
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding