Open a scanned document and the first thing you notice is how much of it is nothing. A sheet of A4 fed through an office scanner comes back as a page where the text sits in the middle and a wide band of empty paper surrounds it on all four sides. On a monitor that band is wasted screen; on a phone it is most of the screen.
Cropping fixes that. What it does not do is most of the things people expect it to.
The margins are larger than they look
We put a scanned report page through Crop PDF and asked the tool where it thought the content was. On a 596 × 842 pt page it placed the content box 96 pt from the left edge, 110 pt from the top, 96 pt from the right — and found 362 pt of empty paper below the last line.
Put together, the text occupied 30% of the sheet. Seventy per cent of that page was margin. That is not an unusual document; it is what a one-page letter looks like when it is printed on a page sized for a full one.
This is why automatic detection is worth having. Dragging a box by hand on a page whose content you can barely see at thumbnail size is guesswork, and guesswork on the tight side clips descenders off the bottom line in a way nobody notices until it is printed.
Cropping moves a boundary. It does not re-render the page
A PDF page carries two relevant rectangles: the MediaBox, which is the sheet of paper, and the CropBox, which is the region a reader actually displays. Cropping sets the second one.
Nothing inside the box is touched. The text is not re-flowed, the fonts are not re-embedded, images are not re-compressed, and a page that was selectable before is selectable afterwards at exactly the same quality. In our run the page went from 596 × 842 pt to 404 × 370 pt and the words inside it were byte-for-byte the same words.
That is the useful mental model: you are changing the window, not the view.
It barely changes the file size, and here is the number
This is the expectation that most often goes wrong. If a page loses 70% of its area, it feels as though the file should lose something like 70% of its weight.
It does not. The same run measured 38,653 bytes before and 36,419 bytes after — a reduction of 5.8%, while the visible page area fell by roughly 70%.
The reason is that empty margin was never costing anything. A blank area of a page is the absence of drawing instructions; there is nothing stored there to remove. What a PDF actually weighs is its fonts, its images and its structure, and cropping changes none of those.
If the goal is a smaller file, cropping is the wrong instrument. Compression is the right one, and it works by reducing image resolution — which is why it transforms a document full of photographs and does almost nothing to a document that is only type.
What happens to the content you crop away
Here the PDF format allows two quite different behaviours, and it is worth knowing which one you are getting.
A tool can crop by setting the CropBox alone. The page then displays smaller while every object outside the new boundary remains in the file, recoverable by anyone who resets the box or extracts the text. The page looks cropped and is not.
We measured ours. Cropping a page so that its lower half fell outside the box took the extractable text from 140 words to 54 — the 86 words outside the box were removed from the text layer, not hidden behind it.
That is the behaviour you want, and it is still not a reason to use cropping for privacy. Cropping is a layout operation that happens to discard what it excludes; it has no concept of sensitive content, it works in rectangles at page edges, and it cannot touch anything in the middle of a page. For removing information that must not survive, use Redact PDF, which exists for that job and which you can verify by searching the finished file for the string you removed.
The general lesson applies to any tool, including ours: if you crop something away because it was confidential, check by extracting the text from the result. Do not take the appearance of the page as evidence.
Scanned pages and generated pages behave differently
A PDF exported from a word processor has clean, empty margins, and finding the content box is close to arithmetic.
A scan does not. A flatbed leaves speckle, a faint grey cast where the lid did not sit flat, and often a thin dark line down one edge where the paper ended. Every one of those is ink as far as detection is concerned, and a detector that treats them as content will find a box barely smaller than the page and appear to have done nothing.
So if automatic detection returns a disappointing box on a scan, the page is usually telling you there is something at its edge. Cleaning the scan — or simply dragging the box in past the artefact — is the fix. On a run of pages from the same scanner the same margins almost always apply, which is why cropping every page to one box is the sensible default and per-page adjustment is the exception.
What to check before you keep the result
Three things, and they take less time to do than to read about.
Look at the bottom line of text. Clipping shows up first on descenders — the tails of g, j, p, q and y. If those look shaved, the box is a few points too tight.
Check a page you did not look at. Cropping every page to one box is right until one page has a wide table or a footer that sits lower than the rest. Open the longest page in the document, not the first.
Decide whether it is going to print. A cropped page has a smaller paper size, so a printer will either scale it to fit or centre it on a sheet with new, different margins. For reading on screen that is irrelevant; for a document going back onto A4 it is not.
In short
Cropping a PDF changes what a page shows. It does not re-render the content, it does not meaningfully change the file size, and it is not a privacy tool even when it genuinely discards what it excludes. Used for what it is — getting rid of dead paper around a scan so the text fills the screen — it is one of the few PDF operations that is lossless and instant at the same time.
You can run it on your own document with Crop PDF; the automatic detection shown above is a button on that page.