Find out what's making the file large
Before changing any setting, work out what kind of PDF you have.
- Text is compact: Text itself is remarkably compact: a page of typed words stored as characters often takes just a few kilobytes.
- Images are the weight: Images are where the megabytes live. A scanned document is the extreme case, because every page is a full-page photograph of paper, even when all it shows is black text. Photos taken on a phone add several megabytes each, and a slide deck exported with high-resolution pictures can be larger than the presentation file it came from.
- The ten-second test: A quick test tells you which case you're in: try to select a word. If the cursor draws a box instead of highlighting text, the page is an image, and image compression will make a real difference. If the text selects normally and the file is still large, look for embedded photos, screenshots and diagrams.
- Megabytes per page: Dividing the file size by the page count is another clue: a 40 MB file with 20 pages averages 2 MB per page, which almost always means scanned or photographic pages.
Resolution is the setting that matters most
Image resolution is measured in DPI, the number of pixels used for each inch of the printed page.
- Why pixels multiply: It has an outsized effect on size because pixels grow with area, not width. An A4 page scanned at 300 DPI is about 2,480 by 3,508 pixels, roughly 8.7 million pixels. The same page at 150 DPI is about 1,240 by 1,754 pixels, around 2.2 million, a quarter of the data. Halving the resolution therefore removes about three quarters of the pixels before any compression is applied.
- Choosing a target: For reading on a screen, 150 DPI is comfortable for ordinary text. Small print, footnotes and signatures stay clearer around 200 DPI. And 300 DPI is a common target for printing.
- Nothing is upscaled: Good compression tools only reduce images that are above the target, so a page already at 120 DPI isn't made worse by a 150 DPI setting. Compress PDF works this way, with levels that correspond to screen, everyday and print use.
Why JPEG compression blurs text first
- How JPEG saves space: Most photographic images inside PDFs use JPEG compression, which saves space by discarding detail the eye is unlikely to miss. It divides an image into small blocks and simplifies each one. Photographs tolerate that well, because they're full of soft gradients.
- Why letters suffer first: Text doesn't, because letters are hard black edges on white, and aggressive JPEG settings leave faint halos, smudged serifs and blocky noise around every character. That's why a scanned contract degrades so much faster than a holiday photo at the same quality setting.
- Two habits that help: Two habits help. First, compress from the original file instead of from an already compressed copy, since each round of lossy compression stacks new artifacts on the old ones. Second, judge quality on the smallest text in the document, not on the overall look of the page.
- Line art and screenshots are better served by lossless compression, which keeps edges exact, although it produces larger files for photographs.
Color, grayscale and black and white
- What grayscale saves: A color image stores three channels of information for every pixel, while a grayscale image stores one. As a rule of thumb, converting an image to grayscale cuts its raw data to about a third before compression, and scans of plain paperwork lose nothing meaningful in the process. Grayscale keeps the full range of tones, so shading, photos and stamps still look natural.
- When color carries meaning: It's the wrong choice when color carries information: charts with color-coded series, highlighted passages, colored annotations, or signatures where blue ink is part of what proves an original.
- How the tool converts: Grayscale PDF converts the whole document. On a server with Ghostscript it keeps text as real text, while the fallback renders pages as grayscale images.
- Best for scanned paperwork: For scanned letters, forms and statements, grayscale followed by moderate compression is often the single most effective combination for reaching a small size while keeping everything readable.
Fonts, duplicates and hidden overhead
After images, the remaining weight in a PDF usually comes from embedded fonts and from how the file was saved.
- Embedded fonts: Well-made PDFs embed only the characters actually used, called a subset, but some documents embed complete fonts, and a file that uses many fonts can carry a surprising amount of font data.
- Repeated and leftover objects: Logos and backgrounds repeated on every page may be stored once and reused, or stored again on each page, depending on the software that created the file. Documents that have been edited and saved many times can also accumulate unused objects left over from earlier versions.
- Cleanup on every save: Rewriting the file with cleanup, which VelloDoc does whenever it saves a PDF, removes that dead weight. These gains are real but modest compared with image compression, so they rarely fix a file that's several times over a limit on their own.
Hitting a hard size limit
- Aim below the stated limit: Upload portals care about one number, and even that number is ambiguous: some count a megabyte as 1,048,576 bytes and others as 1,000,000. Aim slightly below the stated limit, like 1.9 MB for a 2 MB form, to avoid rejection.
- How target-size compression works: Compress to Target Size is built for this situation. It tries several times, lowering image resolution and JPEG quality a little each time. It starts around 150 DPI and steps down toward 55 DPI. It keeps the first result that fits, because that one keeps the most detail.
- When nothing fits: If no attempt fits, it returns the smallest version instead of an error, so always check the final size.
- Read the result before sending: Then open the file and read the smallest text on the page. A file that meets the limit but can't be read hasn't solved the problem, it has only moved it to the person reviewing your application.
When compressing is the wrong answer
Some files can't reach a limit without becoming unreadable, and some shouldn't be compressed at all.
- Send less instead: If a 60-page scan has to fit into 2 MB, the honest answer is often to send less. Take out pages the reader doesn't need with Remove Pages. Drop empty scanned pages with Remove Blank Pages. Or split the document into parts you upload one by one with Split PDF.
- Keep the original for archiving: For documents you're archiving, keep the full-quality original and compress only the copy you send.
- Faster to open isn't smaller: For PDFs you put on a website, making a file open faster isn't the same as making it smaller. Optimize PDF for Web reorders a file so it can start showing sooner, but it doesn't shrink it. That's why the two work best together.
A practical checklist
- Start from the original file, not a copy that has already been compressed.
- Remove pages you don't need: Remove pages that don't need to be sent, including blank backs of duplex scans.
- Convert paperwork to grayscale: If the pages are black-and-white paperwork, convert to grayscale.
- Compress, then compare at 200 percent: Compress at the balanced level and compare the result with the original at 200 percent zoom, paying attention to the smallest text and any signatures.
- Match a portal's limit: If a portal enforces a limit, use target-size compression with a value just under that limit, then check readability again.
- Split when legibility suffers: If the target still can't be met without damaging legibility, split the document or ask the recipient for another way to send it.
- Keep the original until accepted: Keep the original until the smaller version has been accepted. Following this order avoids the two common failures: files that are small but unreadable, and hours spent fighting a limit that no amount of compression can reasonably meet. The guide to optimization tools covers the related tools in more depth, and the scan quality guide shows how to capture smaller, cleaner scans in the first place.
Questions
Why is my scanned PDF so much bigger than a typed document?
Every scanned page is a full-page image, while a typed document stores its words as compact characters. A page of text may take a few kilobytes as characters and hundreds of kilobytes as an image.
What resolution should I use for a PDF sent by email?
Around 150 DPI suits most documents read on screen. Choose about 200 DPI if the document has small print or signatures that must stay legible, and keep 300 DPI for printing.
Does compressing a PDF twice make it smaller?
Usually only slightly, and each round of lossy compression adds artifacts. Compress the original once at the level you need instead of compressing a compressed copy.
Will compression remove pages or text?
No. Compression changes how images are stored. Pages, text and document structure are unchanged, although scanned text can look softer at strong settings.
Why did compression barely change my file?
The file is probably mostly text or already compressed. VelloDoc returns your original when the result wouldn't be smaller. Splitting or removing pages may be the better route.