Nastaliq and why it matters for OCR
Nastaliq is a flowing, calligraphic style.
- Words slope and overlap: Instead of sitting on one flat line, the letters of a word slope downward from right to left, joined letters stack and overlap, and dots crowd into small spaces between them.
- Gaps inside words: Words are often made of several separate groups of joined letters, so the gaps inside a word can look like the gaps between words.
- Why engines struggle: Every one of these features makes it harder for an OCR engine to separate letters, keep their dots and decide where words begin and end.
- Naskh recognizes better: Documents printed in a Naskh-style Urdu font, which some forms and digital documents use, generally recognize better than traditional Nastaliq print. It's worth knowing which style you have before judging the results.
How VelloDoc's Urdu OCR works
- The engine and how text is added: OCR PDF uses the open-source Tesseract engine with its Urdu language data. When you choose Urdu, each page is drawn at about 250 DPI and the words are recognized. The recognized text is then laid invisibly over the original page.
- The page looks the same, but you can now search and copy its text.
- Urdu is preselected here: On this guide, the tool above starts with Urdu already selected.
- Choose Urdu, not Arabic: Choose Urdu instead of Arabic for Urdu documents: Urdu uses letters that Arabic doesn't, like ٹ, ڈ, ڑ, ں and ے, and the Urdu data is built around Urdu words.
- Urdu with English: For Urdu documents with English words in them, add English as the second language so both scripts are recognized.
Preparing Urdu scans
Nastaliq's small, crowded details are the first thing a poor scan destroys, so give the engine as much clarity as you can.
- On a scanner: Scan printed pages at 300 DPI or higher in grayscale, keep the pages flat so lines don't curve near the spine, and clean the scanner glass.
- With a phone, photograph each page straight on in even light, fill the frame with the paper and avoid shadows.
- Don't compress first: Don't compress scans before recognition: heavy JPEG compression smears the dots and joins that tell letters apart.
- Compress afterwards: Run OCR on the best version you have, and compress the finished file afterwards if it needs to be smaller.
A reliable order of steps
- Combine the photos: For a stack of scanned pages, a good sequence is: combine phone photos into one file with Scan to PDF, or start from your scanner's PDF.
- Drop empty backs of duplex scans with Remove Blank Pages.
- Straighten tilted pages with Deskew PDF.
- Run OCR: Then run OCR.
- Why the order matters: Keep this order, because level lines of Urdu script are recognized far more reliably than tilted ones.
- Check, then compress: Finally, check the result and, if needed, reduce its size with Compress PDF.
Checking and editing the recognized Urdu text
- Search for words you can see: Search the result for a few words you can see on the page. Expect some searches to miss even when the word is there, because a single misread letter breaks the match.
- Check names, numbers and dates carefully. Urdu documents may use Eastern Arabic-Indic digits, and how those are recognized can vary.
- Getting the text out to edit: To edit the text, PDF to Word recovers it as an editable document, and PDF to TXT saves it as plain text. Open either in a program that supports right-to-left Urdu text, and plan time for corrections.
- If copied text looks out of order: Right-to-left text in invisible layers is handled differently by different viewers, so if copied text looks out of order, try another application.
When OCR isn't the right answer
- What it isn't built for: General-purpose OCR isn't built for handwriting, decorative calligraphy, faint carbon copies or badly damaged pages, and Urdu poetry set in ornate styles is especially difficult.
- Use it to find the page: For material like this, the searchable layer may still help you find the right page, but the text itself isn't dependable.
- When accuracy is critical: If accuracy is critical, for example for legal or academic quotation, compare recognized passages against the page image or retype them.
- Where to read more: The guide to searchable PDFs explains text layers, and the Arabic OCR guide covers the related challenges of Arabic print.
Questions
Should I choose Urdu or Arabic for Urdu text?
Choose Urdu. Urdu uses letters that Arabic doesn't, and the Urdu language data is built around Urdu words.
Why are there so many mistakes in my recognized Urdu text?
Nastaliq print is difficult for OCR engines, and small type, low resolution, tilt and noise add errors. Rescan at 300 DPI or higher, straighten pages with Deskew PDF and check the result carefully.
Can I edit the Urdu text after OCR?
Yes. Extract it with PDF to Word or PDF to TXT and edit it in a program that supports right-to-left text. Expect to correct some recognition errors.
Does OCR work on handwritten Urdu?
Generally not. The language data is made for printed text, so handwriting produces unreliable results.
Is my document kept after OCR?
No. The file is processed in a temporary job folder and deleted after the result is delivered. The file security page explains the details.