VelloDoc
Optimize PDF

Optimizing PDFs: smaller files, cleaner scans, searchable text

An optimized PDF is one that does its job without friction: it fits under the upload limit, opens quickly, reads clearly, can be searched and doesn't fail in a strict viewer. When a PDF falls short, the cause usually traces back to how it was made, most often a scanner or a phone camera. This guide explains what makes PDFs heavy, why scans behave differently from digital documents, and how the optimization tools fit together. You can fix the underlying cause instead of guessing at settings.

Optimize PDF tools

Tools in this area: 10. Each one links to a detailed page with its settings, limits and common questions.

Where PDF size comes from

Text is tiny inside a PDF: a hundred pages of typed words can take less space than a single photograph. What makes files heavy is images, and in particular images stored at more resolution or color depth than anyone needs. A scanned document is the extreme case, because every page is a full-page picture, even when all it shows is black text on white paper. Embedded fonts and large numbers of vector drawings add some weight, but rarely enough to matter. So the first question for any oversized PDF is whether it's image-heavy. If it is, compression and grayscale conversion will help a lot. If it's mostly text, they will help very little, and splitting may be the better answer. The guide to reducing PDF size goes into resolution, JPEG quality and fonts in detail.

Choosing a compression tool

Compress PDF offers three levels. On a server with Ghostscript they downsample images to about 72, 150 or 300 DPI. Without it, only images above each level's resolution threshold are recompressed. Text and vector graphics are never turned into images, and if the result wouldn't be smaller than your original, the original is returned. Compress to Target Size works toward a hard limit instead, trying progressively stronger settings from about 150 DPI down toward 55 DPI and stopping at the first attempt that fits. If nothing fits, you receive the smallest version, so always check the final size. For black-and-white paperwork scanned in color, Grayscale PDF removes data that adds nothing, and when a limit is simply too small for readable pages, splitting the document is the honest solution.

Scanned documents are pictures of text

A scan looks like a document but behaves like a photograph: you can't search it, copy a sentence or have a screen reader read it aloud. OCR PDF changes that by recognizing the words on each page and laying an invisible text layer over them. The pages keep their own picture, and pages that already contain text are left alone, so a mixed file can go through as it is. Recognition uses a language model, so choosing the document's language matters more than any other setting; two can be combined for mixed documents. The searchable PDF guide explains text layers, why they matter for archives and accessibility, and how to tell whether a PDF already has one.

Cleaning up scans before recognition

OCR accuracy depends on clean input, and two tools prepare it. Remove Blank Pages finds nearly empty pages by measuring brightness and ink coverage, always keeping pages that already contain text, and lets you choose how strict the detection is. Deskew PDF measures the tilt of each page and rotates it level, which noticeably improves recognition, because OCR expects lines of text to run straight. Deskewing turns each page rather than redrawing it, so the scan keeps its detail, and it belongs before OCR so recognition reads level lines. A reliable sequence for a stack of scans is: remove blank pages, deskew, run OCR, and compress last. The scan quality guide covers the capture side, where most problems can be prevented.

Repairing files that won't open

A PDF that shows an error, blank pages or a damage warning often has a broken internal index instead of missing content. Repair PDF opens the file with MuPDF, which rebuilds that index by scanning the file, and writes a fresh, consistent copy. When MuPDF can't read the file at all and Ghostscript is available, a second attempt rewrites the whole document. Repair has a hard limit, though: it works with the data that's present. A download that stopped halfway is missing its remaining pages, and no tool can reconstruct them, so fetching the file again is usually faster than any repair. Repaired files are a good starting point for archiving with PDF/A conversion.

Preparing PDFs for the web

Documents published on a website have extra requirements. Large files should open quickly. Optimize PDF for Web helps by reordering the file so browsers can show the first page before the rest arrives. This works as long as the web server can send files in parts, and almost all can. Size still matters, so compress first. Published PDFs should also be readable by search engines and screen readers. This means scans need OCR, and they should carry a sensible title instead of a leftover template name. Set that with View / Edit Metadata. Together these steps make a document faster to open, easier to find and more accessible, without changing what it says.

Checking the result

Optimization always involves trade-offs, so check the output the way a recipient would. Zoom to 200 percent on the smallest text and on any signature after compression. After OCR, search for a word you can see on a page. After removing blank pages, confirm the page count and skim for anything missing. Keep originals until you're satisfied, because every tool produces a new copy and the processed file on the server is deleted once your download finishes. Some optimization tools rely on engines that not every server has installed, like Ghostscript or Tesseract. When one is missing, the tool says so plainly instead of returning a file that silently skipped the work.

Common workflows

Shrink a scan for an upload portal

Remove the blank backs of a duplex scan, convert color paperwork to grayscale, then compress to the portal's exact size limit.

Make an archive of scans searchable

Straighten tilted pages, add a searchable text layer in the right language, and compress the result so the archive stays manageable.

Rescue and publish a broken report

Rebuild a PDF that won't open, reduce its image weight, and linearize it so it opens quickly from your website.

Get a large document through email

Compress the file first, and if it's still over the attachment limit, split it into parts that each fit.

Questions about optimizing PDFs

Why is my PDF so large?

Almost always because of images, especially scanned pages or high-resolution photos. Text-only PDFs are small, so a large file is usually image-heavy.

Will compression make my text blurry?

Real text stays sharp because only images are recompressed. Text that's part of a scan is an image, so it softens at lower settings.

Should I run OCR before or after compressing?

Run OCR first, then compress, because recognition works best on the sharpest version of the scan.

Can I make a scan searchable without changing how it looks?

Yes. OCR keeps each page exactly as it was and adds the recognized words over it as invisible text.

Why are some optimization tools unavailable?

Target-size compression, OCR and web optimization depend on Ghostscript or Tesseract. A server without the engine reports which one is missing.

Does optimizing change my original file?

No. Each tool returns a new copy, and the temporary copy on the server is deleted after your download finishes.

Related guides