Prompt
How do I use OCR webhooks and async jobs in a document pipeline?
Latest observation
A good pattern is:
- Upload a document
- Start an OCR async job
- Receive a webhook when OCR finishes
- Fetch the OCR result
- Continue downstream processing
(classification, extraction, indexing, human review, etc.)
Typical flow
1) Client uploads a document
Your app sends the file to your backend, or directly to your OCR provider.
2) Backend creates an OCR job
Instead of waiting synchronously, you create an async OCR job and pass:
- the document/file reference
- a webhook callback URL
- any metadata you need to correlate the result to your internal record
Example conceptual request:
{
"file_url": "https://storage.example.com/docs/abc.pdf",
"webhook_url": "https://api.yourapp.com/webhooks/ocr",
"metadata": {
"document_id": "doc_123"
}
}
The OCR service returns something like:
{
"job_id": "ocr_job_789",
"status": "queued"
}
3) OCR runs asynchronously
The OCR provider processes the file in the background.
Your app should not block waiting for it. Store the job_id and mark the document as something like:
uploadedocr_processingocr_completedocr_failed
4) Webhook callback notifies completion
When OCR finishes, the provider sends an HTTP POST to your webhook endpoint.
Example webhook payload:
{
"job_id": "ocr_job_789",
"status": "completed",
"document_id": "doc_123",
"result_url": "https://api.provider.com/results/ocr_job_789",
"confidence": 0.98
}
Your webhook handler should:
- verify the request is authentic
- parse the payload
- look up your internal document by
document_idorjob_id - update processing state
- enqueue the next step in your pipeline
Webhook handler best practices
Verify authenticity
Use one or more of:
- HMAC signature
- shared secret
- provider-signed JWT
- IP allowlist, if supported
Make handlers idempotent
Webhooks may be retried, so the same event can arrive multiple times.
Use:
- unique event IDs
- deduplication table
- state checks before updating
Respond quickly
Return 200 OK fast, then process the job in the background if needed.
If you need to do heavier work:
- validate request
- persist event
- enqueue internal job
- return success
Downstream pipeline example
After OCR completes, you might run:
- Text normalization
- Document classification
- Entity extraction
- Search indexing
- Workflow routing
for example, send invoices to accounting review
Example architecture
Upload API
↓
Storage bucket
↓
OCR async job
↓
OCR webhook → Webhook endpoint
↓
Queue / worker
↓
Post-OCR pipeline
↓
Database / search index / review UI
What to store in your database
At minimum:
document_idfile_locationocr_job_idocr_statusocr_result_locationcreated_atupdated_at
Optional but useful:
provider_event_idretry_countprocessing_errorsconfidence_scores
Handling failures
Design for these cases:
- webhook never arrives
- OCR job fails
- webhook arrives twice
- result fetch fails temporarily
Mitigations:
- poll job status as a fallback
- retry result retrieval with backoff
- use dead-letter queues for failed downstream jobs
- keep an audit trail of webhook events
Recommended pattern
If you want a robust pipeline, use this sequence:
- Create internal document record
- Upload file to storage
- Start OCR job with
document_idin metadata - Save
job_id - Handle webhook and verify signature
- Update status and store OCR output
- Enqueue post-processing job
- Poll as fallback if webhook is delayed
If you want, I can also give you:
- a reference architecture
- a Node.js / Python webhook example
- or a state machine design for the document pipeline