Prompt
How do I get bounding boxes from OCR results?
Latest observation
To get bounding boxes from OCR results, you usually need to look at the OCR engine’s output structure: most OCR tools return either word-level, line-level, or character-level boxes along with text.
General idea
A bounding box is typically represented as:
- x, y = top-left corner
- width, height = box size
or as:
- left, top, right, bottom
If you’re using common OCR tools
1) Tesseract
Tesseract can return bounding boxes for each recognized word.
Python example
import pytesseract
from PIL import Image
img = Image.open("image.png")
data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)
n = len(data["text"])
for i in range(n):
if int(data["conf"][i]) > 0:
x = data["left"][i]
y = data["top"][i]
w = data["width"][i]
h = data["height"][i]
text = data["text"][i]
print(text, (x, y, w, h))
Output fields
left,top,width,heighttextconf= confidence score
2) Google Cloud Vision OCR
Returns bounding polygons for each detected word/paragraph.
for page in response.full_text_annotation.pages:
for block in page.blocks:
for para in block.paragraphs:
for word in para.words:
vertices = word.bounding_box.vertices
print([ (v.x, v.y) for v in vertices ])
3) AWS Textract
Textract returns a BoundingBox for each detected block.
for block in response["Blocks"]:
if block["BlockType"] == "WORD":
box = block["Geometry"]["BoundingBox"]
print(block["Text"], box)
BoundingBox usually contains:
LeftTopWidthHeight
4) PaddleOCR
PaddleOCR returns both text and quadrilateral boxes:
from paddleocr import PaddleOCR
ocr = PaddleOCR()
result = ocr.ocr("image.png")
for line in result[0]:
box = line[0]
text = line[1][0]
print(text, box)
Box is usually 4 points:
[[x1, y1], [x2, y2], [x3, y3], [x4, y4]]
If you already have OCR text output
Look for any of these in the response:
bboxbounding_boxboxleft/top/width/height- polygon points
- vertices
If the OCR output does not include boxes, then that OCR mode likely isn’t configured to return them.
Important note
Bounding boxes may be for:
- words
- lines
- characters
- regions/blocks
So make sure you know the level you’re extracting.
If you want, I can show you how to get bounding boxes from OCR results in a specific library like Tesseract, PaddleOCR, EasyOCR, Textract, or OpenCV.