Prompt
How do I support batch inference for large document jobs?
Latest observation
To support batch inference for large document jobs, design the system so documents are processed asynchronously, in chunks, and with retries/checkpointing rather than one huge synchronous request.
Core pattern
1) Split intake from processing
- Accept a job request with many documents or a large archive.
- Immediately return a job ID.
- Store the job metadata in a database or queue.
- Process documents in the background.
2) Chunk the work
Depending on your model and document type:
- Per document: each file becomes one task.
- Per page / section / chunk: for very large docs, split into smaller units.
- Per batch: group small documents to improve throughput.
A common approach is:
- ingest job
- extract text
- segment into chunks
- enqueue inference tasks per chunk
- aggregate results
3) Use a queue + worker architecture
Typical components:
- API service: receives jobs
- Object storage: stores files and intermediate artifacts
- Queue: SQS, Kafka, RabbitMQ, Redis queue, etc.
- Workers: run inference
- Result store: DB, blob storage, search index
This gives you:
- horizontal scaling
- backpressure
- retry handling
- isolation of failures
4) Make inference idempotent
Each unit of work should have:
- a stable task ID
- deterministic input versioning
- deduplication keys
So if a worker retries, you don’t double-write results.
5) Track job state carefully
Maintain:
job_iddocument_idchunk_id- status:
queued,running,succeeded,failed - attempts
- timestamps
- error message
Use checkpoints so a partially completed job can resume.
6) Aggregate results after all chunks finish
For large docs, you may need a final reduce step:
- combine chunk predictions
- merge entities
- reconcile overlaps
- run postprocessing
This is especially important if chunks overlap for context.
Practical design options
Option A: Simple batch processing
Best when documents are moderate in size.
- Upload documents
- Enqueue one task per document
- Worker runs inference
- Store results
- Mark job complete
Option B: Map-reduce pipeline
Best for huge docs or lots of docs.
- Map: split each doc into chunks and infer independently
- Reduce: merge chunk outputs into final doc output
Option C: Streaming pipeline
Best when documents arrive continuously.
- Extract and chunk on the fly
- Process each chunk as soon as available
- Emit partial results early
Important engineering considerations
Throughput and cost
- Batch requests where your model supports it
- Tune batch size based on latency and GPU memory
- Use dynamic batching if available
Model constraints
If the model has a token/context limit:
- chunk text by tokens, not just pages
- keep overlaps between chunks if needed
- preserve document structure metadata
Fault tolerance
- retries with exponential backoff
- dead-letter queue for poison jobs
- partial result persistence
- timeout limits per task
Observability
Track:
- queue depth
- worker utilization
- inference latency
- failure rates
- per-job progress
Data consistency
- write results atomically per chunk
- only mark job complete after all required chunks are persisted
- handle duplicate messages safely
Example workflow
- User uploads 5,000 PDFs.
- API stores files and creates
batch_job_123. - A splitter extracts text and creates 200,000 chunk tasks.
- Workers pull chunk tasks and run inference.
- Results are stored per chunk.
- A reducer aggregates results per document.
- Job status becomes
complete.
If you’re building with an ML model API
If your model endpoint supports batch inference:
- send a list of inputs in one request
- respect max batch size and token limits
- use async batch submission for long-running jobs
- poll for completion or use callbacks/webhooks
If it does not:
- implement batching at the application layer
- run many concurrent single-item calls with worker pools
Recommended baseline architecture
A solid default stack is:
- API gateway/service
- Postgres for job metadata
- S3/GCS/Azure Blob for documents and outputs
- Redis/SQS/Kafka for task queueing
- worker autoscaling
- reducer service for final aggregation
If you want, I can also provide:
- a reference architecture diagram,
- a database schema for batch jobs, or
- sample pseudocode for the worker pipeline.