Prompt

How do I get bounding boxes from OCR results?

Artificial Intelligence / Computer Vision1 observationLast seen Oct 11, 2026

Latest observation

Oct 11, 2026OpenAI APIWeb search: off

To get bounding boxes from OCR results, you usually need to look at the OCR engine’s output structure: most OCR tools return either word-level, line-level, or character-level boxes along with text.

General idea

A bounding box is typically represented as:

  • x, y = top-left corner
  • width, height = box size

or as:

  • left, top, right, bottom

If you’re using common OCR tools

1) Tesseract

Tesseract can return bounding boxes for each recognized word.

Python example

import pytesseract
from PIL import Image

img = Image.open("image.png")
data = pytesseract.image_to_data(img, output_type=pytesseract.Output.DICT)

n = len(data["text"])
for i in range(n):
    if int(data["conf"][i]) > 0:
        x = data["left"][i]
        y = data["top"][i]
        w = data["width"][i]
        h = data["height"][i]
        text = data["text"][i]
        print(text, (x, y, w, h))

Output fields

  • left, top, width, height
  • text
  • conf = confidence score

2) Google Cloud Vision OCR

Returns bounding polygons for each detected word/paragraph.

for page in response.full_text_annotation.pages:
    for block in page.blocks:
        for para in block.paragraphs:
            for word in para.words:
                vertices = word.bounding_box.vertices
                print([ (v.x, v.y) for v in vertices ])

3) AWS Textract

Textract returns a BoundingBox for each detected block.

for block in response["Blocks"]:
    if block["BlockType"] == "WORD":
        box = block["Geometry"]["BoundingBox"]
        print(block["Text"], box)

BoundingBox usually contains:

  • Left
  • Top
  • Width
  • Height

4) PaddleOCR

PaddleOCR returns both text and quadrilateral boxes:

from paddleocr import PaddleOCR

ocr = PaddleOCR()
result = ocr.ocr("image.png")

for line in result[0]:
    box = line[0]
    text = line[1][0]
    print(text, box)

Box is usually 4 points:

[[x1, y1], [x2, y2], [x3, y3], [x4, y4]]

If you already have OCR text output

Look for any of these in the response:

  • bbox
  • bounding_box
  • box
  • left/top/width/height
  • polygon points
  • vertices

If the OCR output does not include boxes, then that OCR mode likely isn’t configured to return them.

Important note

Bounding boxes may be for:

  • words
  • lines
  • characters
  • regions/blocks

So make sure you know the level you’re extracting.

If you want, I can show you how to get bounding boxes from OCR results in a specific library like Tesseract, PaddleOCR, EasyOCR, Textract, or OpenCV.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.