Prompt
My OCR output is missing text on skewed pages, how do I fix that?
Latest observation
Oct 6, 2026
OpenAI APIWeb search: off
Skewed pages are a common OCR failure mode: if the text lines are tilted or the page is rotated, the OCR engine may miss characters or entire lines. The fix is usually to deskew the image before OCR.
What to do
-
Detect the page angle
- Estimate the tilt of the text or document borders.
- Often done with Hough lines, projection profiles, or OCR/layout tools.
-
Rotate the image to straighten it
- Correct small skew angles (e.g. ±1–15°).
- If pages are rotated 90/180/270°, rotate those first.
-
Run OCR on the corrected image
- OCR accuracy typically improves a lot after deskewing.
Practical fixes
- Use a preprocessing library
- OpenCV
- Pillow
- scikit-image
- Use OCR software with built-in orientation/deskew
- Tesseract (
--psm, OSD/orientation detection) - EasyOCR / PaddleOCR sometimes handle rotation better
- Tesseract (
- Improve image quality
- Increase contrast
- Threshold/binarize
- Remove noise
- Sharpen if needed
- Ensure sufficient DPI (ideally 300 DPI or higher for documents)
Example approach with OpenCV
Typical pipeline:
- Convert to grayscale
- Threshold to isolate text
- Compute skew angle
- Rotate image by the negative angle
- OCR the rotated image
If OCR still misses text
- Check if text is too small or blurred
- Try different page segmentation modes
- Crop margins, headers, and footers separately
- For heavily skewed or warped pages, use dewarping as well as deskewing
If you want, I can give you:
- a Python/OpenCV deskew script
- a Tesseract setup example
- or help you debug your specific OCR pipeline and sample output.