Prompt
AI infrastructure for RAG apps
Latest observation
Building AI infrastructure for RAG apps usually means creating the stack that lets your application ingest data, retrieve relevant context, generate answers, and operate reliably at scale.
Core layers of RAG infrastructure
1. Data ingestion
You need pipelines to bring in content from:
- PDFs, docs, web pages
- databases
- APIs
- tickets, chats, knowledge bases
- internal systems like Drive, Slack, Notion, Confluence
Key components:
- ETL/ELT jobs
- document parsing/OCR
- chunking and metadata extraction
- deduplication and normalization
2. Embedding generation
Convert content into vectors using an embedding model.
Infrastructure needs:
- embedding service or batch jobs
- GPU/CPU compute depending on scale
- retry logic and queueing
- versioning for models and embeddings
3. Vector storage / retrieval
Store embeddings in a vector database or search engine.
Common options:
- vector DBs: Pinecone, Weaviate, Milvus, Qdrant, pgvector
- hybrid search: Elasticsearch/OpenSearch + vectors
- metadata filters and access control
Important features:
- fast ANN search
- filtering by tenant, source, date, permissions
- sharding/replication
- index refresh and reindexing
4. Retrieval orchestration
The app must decide what to fetch and how.
This often includes:
- query rewriting
- hybrid retrieval
- reranking
- multi-hop retrieval
- context compression
- source citation handling
5. LLM inference layer
This is the generation engine.
Infrastructure considerations:
- hosted API vs self-hosted model
- latency and throughput
- caching
- fallback models
- prompt management
- token budgeting
6. Application layer
Where the user-facing product lives:
- chat UI or API
- session memory
- auth and tenancy
- tool use / function calling
- audit logs
- conversation state
7. Observability and evaluation
Critical for RAG quality.
Track:
- retrieval hit rate
- answer faithfulness
- groundedness
- latency
- cost per request
- hallucination rate
- user feedback
Tools/practices:
- offline eval sets
- automated regression tests
- tracing across retrieval and generation
- prompt/version tracking
8. Security and governance
For enterprise RAG, this is essential:
- document-level permissions
- row-level security
- encryption at rest/in transit
- secrets management
- PII redaction
- auditability
- data retention controls
Typical RAG infrastructure architecture
A common flow looks like:
- Content sources →
- Ingestion pipeline →
- Chunking + embedding →
- Vector store / search index →
- Retriever + reranker →
- Prompt assembly →
- LLM inference →
- Response + citations →
- Logging/monitoring/evaluation
Infrastructure choices by scale
MVP / startup
- Managed embeddings API
- pgvector or hosted vector DB
- one LLM provider
- simple ingestion pipeline
- basic logging and evaluation
Mid-scale product
- async ingestion workers
- hybrid search
- reranking
- caching
- tracing/monitoring
- permission-aware retrieval
Enterprise
- multi-tenant architecture
- private networking
- governance and compliance
- self-hosted or VPC-isolated models
- strict observability and offline eval pipelines
- disaster recovery and SLOs
What matters most in practice
If you’re building RAG infrastructure, the highest-leverage areas are usually:
- Document ingestion quality
- Chunking strategy
- Retrieval quality
- Permissioning/security
- Evaluation/monitoring
Bad retrieval infrastructure will make even a strong LLM perform poorly.
If you’re deciding what to build
A practical stack might be:
- Ingestion: Airflow / Dagster / cron + workers
- Parsing: Unstructured, OCR, custom parsers
- Embeddings: OpenAI / Voyage / Cohere / local model
- Vector store: pgvector / Qdrant / Pinecone
- Search: hybrid keyword + vector
- Reranking: cross-encoder or hosted reranker
- LLM: API or self-hosted
- Observability: tracing + dashboards + eval harness
If you want, I can also give you:
- a reference architecture diagram,
- a cloud architecture for AWS/GCP/Azure, or
- a startup-grade RAG infrastructure stack.