Prompt
I'm building a workflow that needs both batch and real-time inference
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
That’s a very common pattern. The key is to design one model-serving contract and then expose it through two delivery modes:
- Batch inference for high-throughput, offline scoring
- Real-time inference for low-latency requests
A good architecture usually looks like this:
1) Share the same feature logic
Keep training, batch scoring, and online scoring using the same:
- preprocessing
- feature definitions
- model version
This avoids training/serving skew.
2) Separate compute paths
- Batch path: runs on a schedule or event trigger, processes large datasets, writes predictions to a table, file, or cache.
- Real-time path: exposes an API or model endpoint for per-request predictions.
3) Use the same model artifact
Deploy the same serialized model to both systems, or better:
- register model in a model registry
- deploy latest approved version to online endpoint
- use same version for batch jobs
4) Choose storage/output by use case
- Batch predictions: data warehouse, object storage, feature store, or downstream table
- Real-time predictions: API response, cache, or event stream
5) Watch for latency and freshness tradeoffs
- Batch is cheaper and scales well, but predictions can be stale
- Real-time is fresh, but more expensive and harder to scale
6) Common pattern: hybrid
Use batch for:
- daily/weekly re-scoring
- backfills
- periodic risk/propensity scores
Use real-time for:
- interactive user experiences
- fraud checks
- dynamic ranking
- request-time personalization
If you want, I can help you design this for a specific stack, like:
- AWS SageMaker
- GCP Vertex AI
- Azure ML
- Databricks
- FastAPI + batch jobs
- Kafka/streaming + online inference