VelloDoc
Optimize PDF

OCR PDF

Add a searchable, selectable text layer to scanned PDFs with OCR in 39 languages. Pages keep their look, and you can export the recognized text too.

Choose files

Drop files here or click to browse


Accepts: PDFMax 80 MBDeleted after delivery

Recognizes 39 languages, two at a time. Check recognized text before relying on it.

Related tools

Next steps people often take after this one.

A scanned PDF looks like a document but behaves like a photograph: you can't search it, copy a sentence from it or have a screen reader read it aloud. Optical character recognition fixes that. OCR PDF reads the words on each page and lays them over the page as invisible text, so it looks exactly as before but can be searched, selected and indexed. It reads 39 languages, two at a time for documents that mix them, and can hand you the text as a plain file instead.

How to use OCR PDF

  1. Add the scanned PDF. Straighten crooked pages first with Deskew PDF for better accuracy.
  2. Choose the document's language, and a second language if the pages mix two.
  3. Choose a searchable PDF or a text file, then select Run OCR.
  4. Download the result and test it by searching for a word you can see on a page.

How the searchable layer is made

  • Read, then laid over the page: Each page is drawn at 250 DPI and the Tesseract engine finds the words in that drawing. The words are then placed on the original page as invisible text, in the same positions as the words you see.
  • The page itself is untouched: The scan keeps its own resolution, and any real text, drawings or links on the page stay as they were. Only the invisible layer is added.
  • Pages that are already searchable: Pages that already carry text, typed or from an earlier OCR run, are left alone, so running OCR twice doesn't double the text.

Choosing the right language

  • Why the model matters: OCR recognizes characters with a model for a specific language, so matching the document's language matters more than any other setting.
  • 39 languages: The list covers Arabic, Bulgarian, Catalan, Chinese (simplified and traditional), Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Latvian, Lithuanian, Norwegian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swedish, Thai, Turkish, Ukrainian, Urdu and Vietnamese.
  • Two languages at once: For documents that mix languages, such as a French contract with an English annex, choose a second language and both models are used together.
  • The wrong model garbles text: Scanning a French letter with English selected produces lost accents and misread words.
  • Arabic script: For scans in Arabic script, the guides to Arabic OCR and Urdu OCR explain scan preparation and what accuracy to expect.

Getting accurate results

Recognition is only as good as the scan.

  • What works best: Straight, evenly lit pages with good contrast work best, and Deskew PDF helps when pages are tilted.
  • Where errors appear: Small print, handwriting, stamps over text and multi-column layouts are where errors appear, so check names, numbers and dates before relying on the text.
  • Getting the text out: Choose Text file to get the recognized words as a .txt file, page after page, or follow up with PDF to Word for an editable document.
  • Where to read more: The searchable PDF guide explains text layers in detail.

When people use OCR PDF

Searchable document archives

Make years of scanned letters and statements findable by keyword instead of opening files one by one.

Copying text from a scan

Select and copy a paragraph, address or reference number from a scanned page instead of retyping it.

Readable by screen readers

A text layer lets screen readers read scanned pages aloud, which is the first step toward an accessible document.

Limitations to know

OCR PDF recognizes printed text in 39 languages, up to two per run, and needs Tesseract with the matching language data on the server. Pages that already contain text are left as they are. Expect recognition errors on small print, handwriting and complex layouts, and check names, numbers and dates.

OCR PDF: common questions

How do I know if my PDF needs OCR?

Try to select a word with your cursor or search for a word you can see. If nothing is found, or you can only draw a box over the page, the text is part of an image and OCR will help.

Does OCR change how the document looks?

No. Each page is kept as it was, picture quality included, and the recognized words are added over it as invisible text.

Can OCR read handwriting?

Tesseract is made for printed text. Neat block capitals are sometimes recognized, but joined-up handwriting generally isn't, so don't rely on OCR for handwritten notes.

Can I get just the text?

Yes. Choose Text file as the output and the recognized words come back as a .txt file, page by page. Pages that already had text contribute that text as it is.

Why does OCR report that a language isn't installed?

Each language needs its Tesseract language data on the server. If it's missing, VelloDoc says so instead of returning a file without a text layer.

Learn more