VelloDoc
Scanning & OCR

What makes a PDF searchable?

Two PDFs can look identical on screen and behave completely differently. In one you can search for a name, copy a paragraph and have a screen reader read the page aloud. In the other nothing happens, because the page is only a picture of text. The difference is a text layer. This guide explains what a text layer is and how to check if a PDF has one. It shows how OCR (text recognition) builds one from a scan, and what makes the result accurate enough to trust.

By VelloDocUpdated Reading time: 7 min

Try it now: OCR PDF

Recognize the words in scanned pages and add an invisible text layer so the PDF can be searched and copied.

Open the full OCR PDF page

Choose files

Drop files here or click to browse


Accepts: PDFMax 80 MBDeleted after delivery

Recognizes 39 languages, two at a time. Check recognized text before relying on it.

Text layers versus pictures of text

  • Digital-born PDFs: A PDF created from a word processor, a spreadsheet or a web browser stores text as characters: each letter is recorded with its font, size and position on the page. Because the file knows which characters are there, viewers can search them, select them, copy them and hand them to assistive technology.
  • Scanned pages: A scanner or a phone camera works differently. It records a picture of the page, a grid of colored pixels with no idea that some of those pixels form the word "invoice". The result can look perfectly sharp while containing no text at all.
  • Mixed documents: Many real documents mix both kinds of page: a digitally produced contract with a scanned signature page attached, or a report with photographed appendices. That's why search sometimes finds a word on page 3 but not on page 12, even though you can plainly read it on both.

How to check a PDF in ten seconds

  • Try selecting a word: The quickest test is to try selecting a word. If the words highlight as you drag, the page has a text layer. If the cursor draws a rectangle across the page instead, you're looking at an image.
  • Search for a word you can see: A second test is to search for a word you can see. No match on a visible word means that page has no usable text.
  • Zoom in: Zooming in is a useful third clue: real text stays crisp at any magnification, while text in an image turns soft and pixelated.
  • Check more than the first page: For mixed documents, check a few pages instead of just the first.
  • Extract the text to be sure: If you want to see exactly what text a PDF contains, PDF to TXT extracts it. A file that comes back empty, or empty for certain pages, is telling you where the text layer is missing.

What OCR actually does

  • Recognition, step by step: Optical character recognition turns pictures of text back into characters. The software finds the lines of text on a page and breaks them into words and letters. Then it compares each shape with what it knows about the language and picks the most likely letters and words.
  • Text with positions: The output isn't just a stream of text but text with positions, so each recognized word can be placed exactly over its image.
  • The sandwich PDF: OCR PDF uses the Tesseract engine. Each page is drawn at about 250 DPI and recognized, and the recognized words are laid invisibly over the original page, which keeps its own picture. This arrangement is often called a sandwich PDF.
  • Why the page looks unchanged: When you search, the viewer finds the hidden word and highlights the spot where the visible word appears, so the page looks unchanged while behaving like a digital document.

Why the language setting matters

  • Language models do the work: OCR engines don't recognize letters in isolation. They rely on language models that know which characters exist, how they combine and which words are likely. Picking the right language makes a big difference.
  • The wrong model garbles text: An English model reading a French letter will garble accented letters and misread words it doesn't know. A model for Latin letters can't read Arabic or Urdu at all, because those use different letters written right to left.
  • Languages VelloDoc offers: VelloDoc offers 39 languages, including Arabic, Hebrew, Hindi, Chinese, Japanese and Korean, and two can be combined for mixed documents. Each needs its language data installed on the server.
  • Documents that mix languages: For documents that mix languages, choose the language used for most of the text and expect the rest to need more checking.
  • Missing language data: If a language isn't installed, the tool says so instead of returning a PDF with no usable text layer.

Getting accurate results from scans

Most OCR errors are decided when the page is captured, not when it's recognized.

  • Resolution matters: as a rule of thumb, scanning ordinary printed text at around 300 DPI gives recognition enough detail, while very small print benefits from more.
  • Contrast matters too, so dark text on clean white paper recognizes far better than faint text on gray or patterned backgrounds.
  • Straight lines matter because recognition expects level rows of text; Deskew PDF corrects small tilts before recognition.
  • Blank pages add nothing but processing time, and Remove Blank Pages clears them out.
  • Phone-scan problems: Shadows, curved pages near a book spine and glare from glossy paper are the most common problems with phone scans. The guide to cleaner phone scans explains how to avoid them.

Checking and correcting recognized text

Even good OCR makes mistakes, and the mistakes follow patterns.

  • Look-alike characters: Similar-looking characters are confused: the letters r and n side by side can become m, the digit 0 and the letter O swap places, and 1, l and I are easily mixed up.
  • Tables and columns: Text in tables can come out in an unexpected order, and multi-column pages may be read across columns instead of down them.
  • Names, numbers and dates: Names, reference numbers, dates and amounts deserve the closest attention, because a single wrong digit changes their meaning while looking plausible.
  • A practical spot check: A practical check is to search the OCR output for several words and numbers you can see on the page.
  • When exact wording matters: For documents where exact wording matters, like legal or financial records, treat recognized text as a convenience for searching and always rely on the page image itself as the authoritative version.

Searchable isn't the same as accessible

  • What a text layer gives you: A text layer is a big step for accessibility, because screen readers can finally read the words on a scanned page aloud. A fully accessible PDF needs more than that, though.
  • What accessible PDFs add: Accessible documents carry tags that describe their structure: which text is a heading, which is a list and which image needs a description. They also set a reading order, so screen readers present the content in a sensible sequence.
  • OCR doesn't add tags: OCR doesn't add those tags. The PDF/UA standard describes what full accessibility requires.
  • A web page alternative: If the content needs to be easy for everyone to read, including on phones, another option is to turn the searchable PDF into a web page with PDF to HTML. Web page text fits any screen and works with the reading tools built into browsers.

When OCR isn't the right tool

  • Don't run it on digital PDFs: OCR is for pages that are images. VelloDoc leaves pages that already have text alone, so running it on a digital PDF simply gives it back unchanged.
  • The one exception: There's one important exception: some digital PDFs contain text that looks normal but copies out as nonsense, because the file stores characters without information about which letters they represent. OCR fixes that by reading the visible page instead.
  • Prefer the original document: If you have the original document, exporting a new PDF from it's always better than recognizing a scan.
  • A sensible archive order: For long-term archives, a sensible order is to scan, clean up, run OCR, and then convert the result with PDF to PDF/A so the searchable document remains readable for years. The guide to optimization tools shows how these steps fit together.
  • Arabic-script documents: For Arabic-script documents, see the guides to OCR for Arabic PDFs and OCR for Urdu PDFs.

Questions

Why can I see the text but not search it?

The page is an image of text instead of text itself, which is typical of scans and phone photos. Run OCR PDF to add a searchable text layer.

Does OCR change how my PDF looks?

No. VelloDoc keeps each page exactly as it was and lays the recognized words over it as invisible text.

Can OCR read handwriting?

Generally not reliably. Tesseract is made for printed text. Neat block capitals are sometimes recognized, but joined-up handwriting usually isn't.

Is OCR text accurate enough for legal documents?

Use it to find and navigate documents, but verify names, numbers and dates against the page image. The image remains the authoritative record.

Should I run OCR on a PDF that's already searchable?

No, unless its text copies out as garbled characters. OCR on a digital PDF replaces real text with an image and a hidden text layer.

Tools for this topic