Prompt
Which vector search setup should I use for a small production app?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For a small production app, the best vector search setup is usually:
Recommended default
PostgreSQL + pgvector
Why this is a good choice
- Simple ops: one database for both app data and vectors
- Production-ready: reliable, mature, easy to back up and monitor
- Good enough performance for small-to-medium scale
- Flexible: supports metadata filtering, joins, transactions
- Cost-effective: no extra vector database to run
Best if your app has:
- up to a few hundred thousand vectors, sometimes more depending on latency needs
- moderate query volume
- need for metadata filtering
- a small team or solo developer
- an existing PostgreSQL stack
When to use a dedicated vector database instead
Consider a vector DB like Qdrant, Pinecone, Weaviate, or Milvus if:
- you expect millions+ vectors
- you need low-latency high-QPS search
- you want hybrid search and advanced retrieval features
- you don’t want to manage tuning Postgres indexes for vector workloads
For a small production app, though, these are often more complexity than you need.
Practical setup I’d choose
If starting today:
- Database: PostgreSQL 15+ with pgvector
- Index:
HNSWif your pgvector version supports it and you want fast approximate search - Fallback:
IVFFLATif HNSW isn’t available or if your workload fits it better - Metadata: store in same table and filter in SQL
- Embeddings: use one consistent model and normalize if using cosine similarity
Rule of thumb
- < 1M vectors and want simplicity → Postgres + pgvector
- > 1M–5M vectors or heavy traffic → consider Qdrant/Pinecone
- Need only semantic search in a simple app → pgvector almost always first choice
If you want, I can recommend a setup more precisely based on:
- expected vector count,
- query latency target,
- whether you need metadata filters, and
- your current stack.