Prompt
I'm building a RAG app and want the simplest reliable infrastructure stack
Latest observation
If you want the simplest reliable RAG infrastructure stack, I’d optimize for:
- few moving parts
- managed services where possible
- easy observability
- something you can scale later without a rewrite
Recommended “simple but solid” stack
1) App/API
- Python + FastAPI
- Deploy on:
- Render / Fly.io / Railway for simplicity, or
- AWS ECS/Fargate if you want more “real production” control
Why: FastAPI is easy for RAG workflows, streaming responses, and tool integration.
2) Document ingestion / background jobs
- Simple queue + worker
- Celery + Redis if you want a common setup
- or RQ + Redis if you want even simpler
If you want the absolute simplest ops:
- use one background worker process alongside the app initially
3) Vector store
Pick one of these:
Best “simple and reliable” choice
- Postgres + pgvector
Why:
- one database for app data + vectors
- easy backups
- mature and familiar
- good enough for many RAG apps
Use managed Postgres:
- Supabase
- Neon
- RDS
- Crunchy Bridge
If you expect very large scale or want dedicated vector infra later, move to a vector DB later. But for a simplest reliable stack, pgvector is the default recommendation.
4) Document storage
- Object storage
- S3 / GCS / R2
Store:
- raw files
- parsed text
- chunk metadata
- ingestion artifacts
Why: don’t put documents directly in the app server filesystem.
5) LLM provider
- OpenAI or Anthropic
- Use whichever matches your latency/cost/quality needs
For a simple stack:
- OpenAI for embeddings + generation is straightforward
- Anthropic is excellent for answer quality
6) Embeddings
- Use a managed embedding model from your LLM provider
- Store embeddings in pgvector
Keep one embedding model initially to avoid complexity.
7) Retrieval / ranking
Start with:
- basic semantic search
- optional hybrid search if needed later
- optional reranker once relevance matters
For simplicity:
- chunk docs
- embed chunks
- vector search top-k
- optionally rerank
- pass context to LLM
8) Caching
- Redis
Use it for:
- query/result caching
- rate limiting
- job queues
- session state if needed
If you’re keeping it minimal, Redis can also be omitted at first unless you need background jobs.
9) Observability
Don’t skip this.
Minimal set:
- structured logs
- request IDs
- error tracking: Sentry
- basic metrics: provider dashboards or OpenTelemetry later
For RAG-specific visibility:
- log retrieved chunk IDs, scores, prompt version, model, latency, and token usage
The simplest “production-ready” version
If I had to pick one stack for most small-to-mid RAG apps:
- Frontend: Next.js
- Backend: FastAPI
- DB: Managed Postgres + pgvector
- Storage: S3/R2
- Queue: Redis + RQ or Celery
- LLM/Embeddings: OpenAI or Anthropic
- Monitoring: Sentry + logs
- Deploy: Render/Fly/Railway or ECS/Fargate
Even simpler MVP stack
If you want the absolute minimum to ship fast:
- Next.js
- Supabase Postgres + pgvector
- Supabase Storage or S3
- No queue initially; ingest synchronously for small volumes
- OpenAI embeddings + chat model
- Sentry
- Vercel or Render
This is often enough for:
- internal tools
- prototypes
- early customer demos
- low/medium traffic apps
What I would avoid at first
To keep it simple and reliable, avoid initially:
- running your own vector database cluster
- multiple queues and workers
- overcomplicated microservices
- custom retrieval infrastructure before you need it
- complicated hybrid search pipelines on day 1
- self-hosted models unless you have a strong cost/privacy reason
My default recommendation
If you want one concrete answer:
FastAPI + Postgres(pgvector) + S3 + Redis + OpenAI + Sentry
That’s the simplest stack I’d still call reliable.
If you want, I can also give you:
- a bare-minimum architecture diagram,
- a cheap startup version, or
- a more scalable AWS version.