Prompt
What is the best way to store and query embeddings for a chatbot knowledge base?
Latest observation
The best approach depends on your scale and latency needs, but for most chatbot knowledge bases the winning pattern is:
Store embeddings in a vector database, keep metadata in a regular database, and use hybrid retrieval.
Recommended architecture
1) Chunk your knowledge base
Split documents into semantically meaningful chunks:
- 200–800 tokens per chunk is common
- Keep overlap between chunks to preserve context
- Store chunk-level metadata: source doc, section, URL, timestamps, permissions, etc.
2) Generate embeddings
Create embeddings for each chunk using a model suited to your language/domain.
Store:
chunk_idembedding_vectortext- metadata fields
3) Use a vector database for similarity search
Good options:
- Pinecone: managed, easy scaling
- Weaviate: flexible, supports hybrid search
- Qdrant: strong open-source choice
- Milvus: great at scale
- Postgres + pgvector: excellent if your corpus is moderate and you want simplicity
4) Keep metadata in SQL/NoSQL
Use a standard database for:
- source documents
- access control
- tags
- freshness/versioning
- audit logs
Then join metadata with vector search results.
5) Query with hybrid retrieval
For best chatbot performance:
- Vector search finds semantically similar chunks
- Keyword/BM25 search catches exact terms, names, IDs, acronyms
- Combine both, then rerank results
This usually outperforms pure vector search.
Query flow for a chatbot
- User asks a question
- Embed the query
- Run vector search with metadata filters
- Optionally run keyword search in parallel
- Merge and rerank top results
- Pass top chunks into the LLM as context
What to optimize for
If your KB is small to medium
Use:
- Postgres + pgvector
- simple metadata tables
- hybrid search if possible
This is often the best balance of cost and simplicity.
If your KB is large or high-QPS
Use:
- dedicated vector DB
- caching
- reranking
- sharding/replication as needed
Storage tips
- Store the raw text alongside embeddings so retrieval is easy
- Use versioning for documents and embeddings
- Re-embed when you change models
- Normalize metadata fields for filtering
- Add ACL/security filters at query time if knowledge is permissioned
Practical best practice
For most chatbot knowledge bases:
Postgres for documents + pgvector for embeddings + full-text search + reranker is the best “default” choice.
It’s simple, cheap, and effective until you hit very large scale.
If you want, I can also give you:
- a sample schema,
- a query example in SQL/pgvector,
- or a comparison of Pinecone vs Qdrant vs pgvector.