Prompt
Why is OCR missing text on low-resolution scans?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
OCR often misses text on low-resolution scans because the image no longer contains enough reliable detail for the OCR system to distinguish characters accurately.
Main reasons:
- Blurred character shapes: At low resolution, letters lose sharp edges and fine features like serifs, holes in letters (e.g.,
a,e,o), and small gaps. - Merged characters: Adjacent letters may touch or overlap, making one word look like a blob.
- Lost thin strokes: Parts of letters like
i,j, punctuation, accents, and small marks can disappear entirely. - Bad line/word segmentation: OCR first tries to find text lines and individual characters. Low resolution makes that layout analysis less accurate.
- Compression artifacts: If the scan is also heavily compressed, blockiness or noise can confuse the OCR.
- Font size matters: Small text at low DPI may become effectively unreadable even if the page looks okay to a human when zoomed out.
Typical remedies:
- Scan at 300 DPI minimum, often 600 DPI for small text.
- Use grayscale or black-and-white with good contrast.
- Apply de-skewing, denoising, and sharpening before OCR.
- If possible, rescan from the original at a higher resolution.
If you want, I can also explain how OCR engines like Tesseract decide whether a character is “good enough” to recognize.