Compress PDF
Your file never leaves this device
Most of the size of a large PDF is images. This decodes each embedded JPEG, scales it down if it is larger than the target, re-encodes it at the quality you choose, then rewrites the file with object streams and the document information fields cleared. Scanned documents typically shrink by half or more; a text-only PDF saves a few per cent. Drop in several files at once and the results come back as a ZIP.
How to compress a PDF
Being precise about what this does is what lets you predict whether it will help. It walks the document looking for embedded images, and it re-encodes exactly one kind: baseline JPEG data, in plain colour or plain greyscale, with no custom decode array attached. Each of those is decoded, scaled down if its longest side exceeds the target, re-encoded at the chosen quality, and put back only if the result is genuinely smaller. Then the whole file is rewritten using object streams and with its title, author, subject, keywords, producer and creator fields cleared.
Everything else is left exactly as it was. Vector artwork, fonts, form fields and text are untouched. So are images stored in any other way, and that group is larger than people expect: PNG-style lossless images inside a PDF, JPEG 2000 images, images in a CMYK colour space, and the fax-style bilevel encoding that many document scanners produce for black-and-white pages. A PDF made entirely of those will come back with only the structural saving.
Which makes the answer to "why did nothing happen" nearly always the same. A text-heavy report is already efficiently compressed and there is no image data to work on. A greyscale scan is a partial case: re-encoding it goes through a colour canvas, so the replacement is often no smaller than the original and gets discarded on the spot. When the output would be bigger than the input, the whole file is rejected rather than offered, because handing back a larger file labelled "compressed" is worse than saying nothing could be done.
What it deliberately does not do is the list of things desktop tools do that lose information without telling you: it does not re-run text through OCR, subset or drop fonts, flatten form fields, downsample vector artwork to bitmaps, or discard layers. Every one of those saves space and every one of them can throw away something you needed.
One side effect looks like a privacy benefit and is only half of one. What gets cleared is the document information dictionary — the six fields a reader lists under Properties, including Author and Producer. Most PDFs written by a word processor, a layout application or Acrobat also carry a second copy of the same information as an XMP metadata packet attached to the document catalogue, and that packet is left exactly as it was: a test file whose XMP named its author still named that author after compression. The properties panel can therefore go blank while the name is still sitting in the bytes. Treat this as tidying rather than as redaction, and use something that rewrites XMP if a document genuinely has to be scrubbed before it leaves.
What people use it for
- Getting a scan under an email attachment limit
- Meeting an upload cap on a government or bank portal
- Shrinking a phone-scanned contract before sending it on
- Cutting a photo-heavy report down for a slow connection
- Reducing a file without handing the document to a web service
- Clearing the Author and Producer fields a reader shows under Properties
Questions
Because it is mostly text and vectors, which are already compressed efficiently. With no eligible image data to re-encode, only the structural saving of a few per cent is available.
Baseline JPEG data in plain colour or plain greyscale. Anything else in the file is left alone.
Many document scanners store black-and-white pages with a fax-style bilevel encoding rather than as JPEG. That is already very compact and is not touched here.
Skipped. Re-encoding them through a browser canvas would convert them to screen colour and change how the document prints, which is not a trade worth making silently.
Re-encoding goes through a colour canvas, so the replacement is often no smaller than the greyscale original. When that happens the original is kept.
At quality 70 and 1600 pixels on the longest side, a scanned page stays readable on screen and in print at normal sizes. Raise both sliders if you need archival quality.
Quality runs from 30 to 95, and the longest side from 600 to 3000 pixels. Changing either re-runs the whole batch.
Quality 85 and at least 2400 on the longest side, which is about 205 dpi across an A4 page. The 1600 default works out near 137 dpi, fine on screen and visibly soft on paper.
Yes. Drop in as many as you like. They are processed one after another and come back as a single ZIP, because browsers block a run of separate downloads.
It is discarded and the file is reported as already about as small as it goes. Your original is the better file in that case.
No. Text, fonts, vector artwork, form fields and page geometry are all untouched. Only image data and the file structure change.
Six fields of the document information dictionary: title, author, subject, keywords, producer and creator. Nothing visible on any page is removed.
No. Only the information dictionary is cleared. A PDF written by a word processor or a layout application usually carries the same details a second time in an XMP metadata packet, and that is left untouched — so the Properties panel can read blank while the author name is still in the file. Use a tool that rewrites XMP if the document has to be genuinely scrubbed.
Rewriting the file invalidates any existing signature, as it would with any tool that changes the bytes. Sign after compressing, not before.
No. An encrypted PDF has to be unlocked first, which is something only you can do.
The PDF parsing library downloads the first time you choose a file. It is about 175 KB, it is loaded once for the whole batch, and after that it is cached.
That rasterises the whole document, so the text stops being text and becomes a picture of text. It is smaller and it is no longer searchable or selectable.
It is bounded by memory rather than by a fixed limit, since the whole document is held in the browser while it is rewritten. A phone gives out well before a laptop does.
No. The whole thing happens in this page. The parser is downloaded to your browser when you pick a file; the file itself goes nowhere.