Prompt
RAG infrastructure on Kubernetes
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A solid RAG (Retrieval-Augmented Generation) infrastructure on Kubernetes usually has these main parts:
1) Core components
- API / App layer
- Your chat or query service
- Handles user requests, auth, session state
- Embedding service
- Converts text into vectors
- Can be self-hosted or external
- Vector database
- Stores embeddings + metadata
- Examples: Milvus, Weaviate, Qdrant, pgvector
- Document ingestion pipeline
- Pulls from S3, Git, databases, file uploads, etc.
- Splits/chunks text, cleans it, generates embeddings
- LLM inference layer
- Generates final answers using retrieved context
- Could be hosted via vLLM, TGI, Ollama, OpenAI API, etc.
- Cache
- For prompts, retrieval results, embeddings, and sessions
- Usually Redis
- Storage
- Object store for raw docs/chunks/index artifacts
- Often S3 / MinIO
- Observability
- Logs, metrics, tracing, evaluation
- Prometheus, Grafana, Loki, OpenTelemetry
2) Typical Kubernetes deployment pattern
Namespace layout
rag-apprag-ingestionrag-inferencerag-dataobservability
Workloads
- Deployments
- API gateway / backend
- embedding workers
- retrieval service
- web UI
- Redis
- StatefulSets
- vector DB if self-hosted
- object store if self-hosted
- Jobs / CronJobs
- batch ingestion
- reindexing
- document cleanup
- HPA / VPA
- scale API and workers based on CPU/memory/queue depth
- GPU node pools
- for LLM inference and sometimes embeddings
3) Request flow
- User sends a question to the API
- API embeds the query
- Vector DB returns top-k relevant chunks
- Reranker optionally improves results
- Context + prompt sent to LLM
- LLM generates answer
- Response returned, with citations if needed
4) Recommended architecture choices
Simple production stack
- API: FastAPI / Node.js service
- Vector DB: Qdrant or pgvector
- Cache: Redis
- Inference: vLLM on GPU nodes
- Storage: S3/MinIO
- Ingestion: Celery workers or Kubernetes Jobs
More scalable stack
- Kafka / NATS for ingestion events
- Argo Workflows or KEDA for event-driven scaling
- Service mesh only if you need advanced traffic policy
- Dedicated reranker service
- Prompt/version registry
5) Important Kubernetes concerns
- Resource isolation
- Separate CPU and GPU node pools
- Autoscaling
- HPA for API and workers
- KEDA for queue-based ingestion
- Secrets management
- External Secrets / Vault / Sealed Secrets
- Network policies
- Restrict access to vector DB and inference endpoints
- Persistent volumes
- Needed for stateful services
- Rolling updates
- Ensure embedding model/index compatibility
- Backup and restore
- Especially for vector DB and object store metadata
6) Operational best practices
- Version embeddings by model version
- Store document IDs, chunk IDs, source URLs, timestamps
- Add citation metadata to each chunk
- Monitor:
- retrieval latency
- top-k hit quality
- token usage
- answer correctness
- index freshness
- Use canary deployments for new models/prompts
- Add rate limiting and request size limits
7) Example minimal architecture
User -> Ingress -> RAG API -> Redis
|-> Embedding service -> Qdrant
|-> LLM inference (vLLM)
Ingestion: S3/Git -> Parser -> Chunker -> Embeddings -> Qdrant + metadata
If you want, I can also provide:
- a reference Kubernetes YAML/Helm architecture,
- a production-grade diagram, or
- a step-by-step implementation plan for RAG on K8s.