VelloDoc
Scanning & OCR

How to OCR Arabic PDFs online: scan quality and accuracy tips

A scanned Arabic document is only a picture of text: you can't search it for a name, copy a paragraph or have a screen reader read it aloud. OCR (text recognition) turns those images back into text. Arabic is harder than languages written in Latin letters, though. Its letters join and change shape, and its dots and marks are easy to lose in a poor scan. This guide explains how OCR handles Arabic, how to prepare pages for a better result and how to check what was recognized.

By VelloDocUpdated Reading time: 7 min

Try it now: OCR PDF

Recognize the words in scanned pages and add an invisible text layer so the PDF can be searched and copied.

Open the full OCR PDF page

Choose files

Drop files here or click to browse


Accepts: PDFMax 80 MBDeleted after delivery

Recognizes 39 languages, two at a time. Check recognized text before relying on it.

Why Arabic is harder for OCR

  • Joined, shape-shifting letters: Arabic is written right to left, and most letters join the letters next to them. Each letter can take a different shape at the start, middle or end of a word, or when it stands alone.
  • Dots change the word: Several letters share the same basic shape and differ only in their dots, like ب, ت and ث. A speck of dust or a missing dot can change the word.
  • Short-vowel marks: Many texts also leave out short-vowel marks, while religious, poetic and educational texts may include them in full, adding small marks above and below the line.
  • Mixed directions: Numbers and Latin words run left to right inside right-to-left sentences.
  • What the engine has to do: An OCR engine has to separate joined letters, keep every dot and mark, and put everything back in the right order, which is why scan quality matters even more for Arabic than for English.

How VelloDoc recognizes Arabic text

  • The engine and its language data: OCR PDF uses the open-source Tesseract engine with its Arabic language data. When you choose Arabic as the OCR language, each page is drawn at about 250 DPI and Tesseract finds the words on it.
  • How the text is added: The recognized text is then laid invisibly over the words on the original page, which keeps its own picture. The page looks the same, but you can now search, select and copy its text.
  • Arabic is preselected here: On this guide, the tool above starts with Arabic already selected.
  • Two languages at once: For a document that mixes Arabic and English, choose Arabic and add English as the second language, so both are recognized.
  • Only run it on scans: Pages that already contain selectable text are left alone, so running OCR on a mixed file only reads the scanned pages.

Scanning Arabic documents for better recognition

Start with the sharpest, flattest page you can get.

  • On a scanner, 300 DPI in grayscale is a good default for printed Arabic. Very small type or text with full vowel marks can benefit from a higher setting, as long as the result stays sharp.
  • Why a clean source still matters: Although pages are drawn again at about 250 DPI for recognition, a clean source still matters, because blur, noise and shadows carry through.
  • With a phone, fill the frame with the page, hold the camera parallel to the paper and use even light without shadows across the text.
  • What Scan to PDF does not do: Scan to PDF combines phone photos into one PDF, but it doesn't detect page edges or correct perspective, so crop and straighten your shots first.
  • Where to read more: The guide to cleaner phone scans covers lighting and framing in detail.

Straighten pages before running OCR

  • Why tilt hurts Arabic most: Tilted lines confuse recognition, and joined Arabic letters are especially sensitive because the engine follows the line the letters sit on.
  • What Deskew PDF does: Deskew PDF detects the angle of each page and straightens it automatically.
  • Straighten first, then recognize: The order matters: Tesseract reads level lines far more reliably than tilted ones, so straighten first, then run OCR on the straightened file.
  • Remove blank backs: If a document has blank backs of pages from a duplex scan, remove them before recognition with Remove Blank Pages, which saves time and avoids empty pages in the result.

Checking and reusing the recognized text

  • Search for a word you can see: Open the result and search for a few words you can see on the page, like a name or a place. If the search finds them, the text layer is working.
  • Check the details that matter: Then check the details that matter most, because OCR mistakes are rarely obvious. Look at letters that differ only by dots, letters at the end of words like ة and ه or ي and ى, and all dates and numbers.
  • If copied text looks out of order: Right-to-left text in an invisible layer is handled differently by different PDF viewers, so copied text can appear out of order in some applications. Trying another viewer often helps.
  • Getting the text out: To work with the text itself, PDF to TXT saves it as a plain text file you can open in any editor that supports Arabic. PDF to Word gives you an editable document, without the original layout.

Documents OCR will struggle with

Some material is simply beyond what general-purpose OCR reads reliably.

  • Handwriting and old prints: Handwriting, decorative calligraphy, very old or damaged prints and faint photocopies usually produce many errors or no text at all.
  • Stamps, tables and vowel marks: Stamps and signatures over text hide letters, tables and multi-column layouts can be read in the wrong order, and fully vocalized text may lose or misplace its marks.
  • Use it to find, not to quote: For documents like these, the searchable layer can still help you find the right page, but important passages should be checked against the image or retyped.
  • Where to read more: The guide to searchable PDFs explains text layers and when OCR is the right tool.

Questions

Can VelloDoc read Arabic handwriting?

Not reliably. The Arabic language data is made for printed text, so handwriting usually produces many errors or no text at all.

Why does copied Arabic text look reversed or broken?

PDF viewers handle right-to-left text in invisible text layers differently. Try another viewer, or extract the text with PDF to TXT and open it in an editor that supports Arabic.

Should I choose Arabic or English for a mixed document?

Choose Arabic, and add English as the second language. Both are then used together, so English words in the document are read far more reliably.

Does OCR change how my pages look?

No. The pages keep their own picture, and the recognized text is added over them as an invisible layer. Pages that already contain selectable text are left as they are.

What resolution should I scan at?

300 DPI is a good default for printed Arabic. Very small print or text with full vowel marks can benefit from a higher resolution, as long as the scan stays sharp.

Tools for this topic