Prompt
ChatGPT: I want to serve an LLM and embeddings for a SaaS app. Recommend an architecture that keeps latency low and costs predictable.
Latest observation
Here’s a pragmatic architecture for serving both an LLM and embeddings for a SaaS app with low latency and predictable cost.
Recommended high-level architecture
Use a two-tier inference setup:
-
Real-time path for user-facing requests
- Dedicated model serving on GPU instances
- Autoscaled but with minimum warm capacity
- Separate services for:
- LLM generation
- Embedding generation
-
Async / batch path for non-interactive work
- Queue-backed jobs for:
- document ingestion
- embedding backfills
- cache warming
- summarization / indexing tasks
- Queue-backed jobs for:
This split is the biggest lever for keeping latency low without letting costs spike unpredictably.
Core design
1) Separate LLM and embedding services
Do not run embeddings through the same runtime as your chat/completions model unless you are extremely small.
Why:
- Embeddings are usually high-throughput, short requests.
- LLM inference is token-heavy, latency-sensitive, and more variable.
- Independent scaling prevents one workload from starving the other.
Suggested setup:
- LLM service: GPU-backed inference server
- Embedding service: smaller GPU or even CPU depending on model size and throughput needs
- Both exposed behind an internal load balancer/API gateway
2) Use an inference server optimized for batching and KV cache
For the LLM, use a serving stack such as:
- vLLM
- TGI (Text Generation Inference)
- TensorRT-LLM if you want maximum performance and can accept more operational complexity
Key features to require:
- continuous batching
- paged attention / KV cache efficiency
- streaming tokens
- request cancellation
- quantization support if model quality allows
This lowers GPU waste and improves tail latency.
3) Keep warm replicas for predictable latency
For SaaS, cold starts are painful and expensive to users.
Recommendation:
- Maintain a minimum number of always-on replicas
- Use autoscaling only above that baseline
- Scale on:
- queued requests
- GPU utilization
- token throughput
- p95 latency
This gives predictable baseline cost and controlled burst behavior.
4) Put embeddings behind caching and deduplication
Embedding calls are often repeated for the same content.
Use:
- content hash cache: identical text → same vector
- document chunk cache
- job deduplication for ingestion pipelines
- bulk embedding endpoints to amortize overhead
Store embeddings in:
- a vector DB such as pgvector, Pinecone, Weaviate, Milvus, or OpenSearch vector
- plus an object store or cache for raw vectors if needed
Suggested deployment topology
Frontend/API layer
- API gateway / load balancer
- Auth, rate limiting, tenant quotas
- Request routing to appropriate service
App services
- Chat orchestration service
- Retrieval service
- Ingestion worker service
- Background job queue
Inference layer
- LLM inference cluster
- GPU nodes
- model server with batching
- optional streaming gateway
- Embedding inference cluster
- smaller GPU pool or CPU pool
- batch-friendly endpoint
Data layer
- Relational DB for tenants, conversations, metadata
- Vector DB for embeddings
- Redis for:
- response cache
- embedding cache
- rate limit counters
- session state
Async layer
- Queue: SQS / PubSub / RabbitMQ / Kafka
- Workers for indexing, document chunking, re-embedding, eval jobs
Latency optimization tactics
For LLM serving
- Use a smaller model when possible
- Prefer RAG over stuffing long prompts
- Keep prompt context trimmed
- Stream tokens immediately
- Use speculative decoding if supported
- Pin model replicas close to your app region
- Avoid cross-region inference for interactive traffic
For embeddings
- Batch requests from ingestion pipelines
- Precompute embeddings for known content
- Cache by hash
- Use a smaller embedding model unless retrieval quality suffers materially
For retrieval
- Keep the vector DB and app in the same region
- Use approximate nearest neighbor indexes
- Limit top-k aggressively and rerank only when needed
Cost predictability tactics
1) Baseline + burst model
Set a fixed baseline of GPU capacity for normal traffic, then allow autoscaling for peaks.
This avoids surprise bills while still handling traffic spikes.
2) Model tiering
Use different models for different tasks:
- small/cheap model for classification, routing, extraction
- larger model only for final user generation
- small embedding model unless recall requires more
3) Token budgets and quotas
Control cost through:
- max input tokens
- max output tokens
- per-tenant rate limits
- monthly usage caps
- request prioritization
4) Queue non-urgent work
Never let ingestion or reindexing compete with live user chats.
5) Right-size GPU types
- Use the cheapest GPU that meets your latency target
- Quantize models where acceptable
- Favor throughput-optimized hosting if your traffic is steady
A good default implementation
If you want a solid starting point:
- LLM serving: vLLM on GPU nodes
- Embedding service: separate vLLM/TGI endpoint or lightweight embedding server
- Queue: SQS/RabbitMQ
- Cache: Redis
- Vector store: pgvector if modest scale, Milvus/Pinecone if larger
- Autoscaling: Kubernetes HPA/KEDA or cloud-native autoscaling
- Observability: per-tenant metrics, p95 latency, token usage, GPU utilization, queue depth
Practical architecture pattern
Request flow for chat
- User sends message
- API gateway authenticates and rate-limits
- App service retrieves relevant context from vector DB
- App builds prompt
- LLM inference service generates response
- Stream response back to user
- Log tokens, latency, tenant usage
Request flow for document ingestion
- Document uploaded
- App chunks document
- Worker computes hashes
- Worker requests embeddings in batches
- Store vectors in vector DB
- Update index metadata
- Notify completion
What I would avoid
- Serving embeddings and generation from one shared GPU pool without isolation
- Fully serverless LLM inference for latency-sensitive chat unless traffic is very spiky
- Over-scaling to zero for interactive models
- Huge prompts instead of retrieval
- No caching for repeated embeddings or system prompts
- Cross-region inference for every request
If you want the simplest reliable version
For an early-stage SaaS:
- 1 app service
- 1 Redis cache
- 1 vector DB
- 1 GPU inference service for LLM
- 1 smaller embedding service
- 1 queue + worker pool for ingestion
- minimum always-on replicas, then autoscale
That gives you:
- low latency for live traffic
- predictable baseline spend
- room to scale without re-architecting
If you want, I can also provide:
- a reference architecture diagram,
- a Kubernetes deployment layout, or
- a cost-optimized option for AWS/GCP/Azure.