Prompt
What should I use for pay-per-request inference?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want pay-per-request inference, the best choice depends on your priorities:
- Lowest operational hassle: use a serverless inference API from a managed provider.
- Lowest latency / production control: use a containerized endpoint with autoscaling that can scale to zero or near-zero.
- Occasional or bursty usage: pay-per-request is ideal because you avoid paying for always-on GPUs.
Good options
- AWS Bedrock / SageMaker Serverless Inference
- Google Vertex AI online prediction
- Azure Machine Learning managed online endpoints
- Hugging Face Inference Endpoints (if supported for your model)
- Replicate / Together / Fireworks / OpenRouter for model-specific APIs
What to choose
- For custom models: managed serverless inference on your cloud provider.
- For LLMs: a hosted model API is usually easiest and truly request-based.
- For strict cost control: compare:
- request price
- token/output pricing
- cold start time
- max concurrency
- GPU availability
Rule of thumb
- Small, infrequent workloads → serverless
- Heavy but spiky workloads → autoscaling endpoints
- LLM apps → model API / hosted inference provider
If you tell me:
- the model type (LLM, vision, embedding, etc.),
- traffic pattern,
- cloud provider,
- latency target,
I can recommend the best specific setup.