Why is my PDF so large?
Short answer:Almost always the images. Scanned pages and photographs can be several megabytes each, while hundreds of pages of text and vector graphics might total only a few hundred kilobytes. A 20 MB PDF is usually a few heavy images wearing a small document as a disguise. The fix is to re-encode or downsample those images, not to “compress the PDF” as a whole.
What makes a PDF big
Four kinds of content fill a PDF: text, fonts, vector graphics and images. Text is tiny, so much so that a full novel is under a megabyte. Fonts add a little. Vector graphics, like charts and logos, are compact and scale without loss. Images are the outlier: a single colour photograph at high resolution can outweigh everything else in the file.
Scanned documents are the worst case, because a scan is just a photograph of a page. A scan at 300 DPI stores an image of the whole page, so even a text-only scan carries megabytes of picture data where a few kilobytes of real text would do.
How to tell what is heavy
If the file is a scan or contains photos, images are the weight. If it is a generated report with a few charts, the charts are usually small and the file may already be close to its floor. A quick hint: open the PDF and zoom in. If text stays crisp but images soften, the file is image-heavy and will compress well.
What actually shrinks it
Two levers, both applied to images. First, re-encode them more efficiently, since modern encoders store the same picture in fewer bytes. Second, downsample them to the resolution they are actually shown at, which throws away detail the eye cannot see on the page.
Text, fonts and vector graphics should not be touched at all; they are recompressed losslessly so they stay identical. A tool that claims to “compress everything” is usually doing far less than one that targets images honestly.
The limit you cannot cross
File size is information, so there is a floor. A sharp photograph contains a lot of detail, and keeping it sharp means keeping its size. The only ways down are to store it more cleverly or to keep less of it. If a document must be both tiny and photographic, something has to give, and it should be resolution, not your text.
More questions
Why is my PDF so large?
Almost always images, especially scans. Text, fonts and vector graphics are small; a scan is a photograph of a page and can be several megabytes on its own. Re-encoding or downsampling those images is what actually reduces the size.
How do I reduce a PDF’s size?
Compress it with a tool that targets images. Encoding them more efficiently and, at higher settings, reducing their resolution shrinks the file the most, while text and vector graphics are recompressed losslessly and stay identical.
Can a PDF be too compressed?
Images can lose visible detail if pushed too far, which is why levels exist: a light setting barely changes them, a balanced one is hard to tell apart from the original, and a maximum setting softens photos deliberately. Text stays crisp at any setting.
Why is a text-only PDF already small?
Because text and vector graphics are compact by nature. A document that is mostly words has very little image data to remove, so it is already near its minimum and compression will save less.