Prompt
How do I OCR mixed English and non-English documents?
Latest observation
To OCR documents that mix English with other languages, you want an OCR setup that can detect multiple scripts/languages in the same page and preserve layout well enough for review.
Practical approach
1) Identify the languages/scripts involved
Examples:
- English + Spanish/French/German
- English + Chinese/Japanese/Korean
- English + Arabic/Hebrew
- English + Cyrillic, Devanagari, Thai, etc.
This matters because some OCR engines handle multiple Latin languages easily, while mixed scripts may need special configuration.
2) Use an OCR engine that supports multilingual OCR
Good options:
- Google Cloud Vision / Document AI
- Microsoft Azure AI Vision / Document Intelligence
- ABBYY FineReader
- Tesseract
- PaddleOCR
- Amazon Textract (good for documents, but language flexibility varies)
For open-source:
- Tesseract is the classic choice and supports many languages.
- PaddleOCR often performs better on complex layouts and some non-Latin scripts.
3) Enable all relevant language models
If using Tesseract, specify multiple languages, for example:
tesseract input.png output -l eng+fra+deu
For mixed scripts:
tesseract input.png output -l eng+chi_sim
The key is: include every likely language the document contains.
4) Preprocess the images before OCR
Quality improves a lot if you:
- deskew pages
- remove noise
- increase contrast
- binarize carefully
- crop margins if needed
- dewarp scanned pages
- rotate correctly
This is especially important for bilingual or low-quality scans.
5) Use page segmentation or region-based OCR
If the document has:
- columns
- tables
- headers/footers
- alternating languages by section
then OCR works better if you first detect text regions and run OCR on each region separately. Some tools do this automatically; otherwise, use layout detection first.
6) Post-process by language
After OCR:
- run spell-check or language-specific correction
- normalize punctuation
- fix encoding issues
- keep original text if you need fidelity
- optionally detect which language each block belongs to
7) For PDFs, choose between searchable PDF and text extraction
If your source is a PDF:
- If it already has selectable text, extract text first rather than OCR.
- If it’s scanned, use OCR to create a searchable PDF so you keep the original image plus an invisible text layer.
Recommendations by use case
Best accuracy, minimal setup
Use a cloud OCR service:
- Google Document AI
- Azure Document Intelligence
- ABBYY
Best open-source option
Use:
- Tesseract for straightforward multilingual OCR
- PaddleOCR for tougher layouts or more complex documents
For mixed Latin languages
Usually easy:
- set all relevant language packs
- OCR the whole page
For mixed scripts
Better to:
- detect script/language by region
- OCR each region with the right model if possible
Example workflow
- Convert PDF pages to high-resolution images.
- Preprocess images.
- Detect layout regions.
- OCR with all relevant languages enabled.
- Post-correct per language.
- Export as text, DOCX, or searchable PDF.
Tesseract example in Python
import pytesseract
from PIL import Image
img = Image.open("page.png")
text = pytesseract.image_to_string(img, lang="eng+fra+chi_sim")
print(text)
Common pitfalls
- Using only
engwhen the document has non-English text - OCRing low-resolution scans
- Not handling columns/tables/layout
- Mixing scripts without script detection
- Expecting perfect accuracy on handwritten text
If you want, I can also recommend the best OCR tool for your specific language pair or give you a Python pipeline for mixed-language PDFs.