Prompt
I'm building semantic search for a knowledge base, what should I use?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
For semantic search over a knowledge base, the usual best choice is:
Recommended setup
Vector embeddings + a vector database + reranking
1) Create embeddings for your KB content
Convert each document/chunk into an embedding using an embedding model.
Good options:
- OpenAI:
text-embedding-3-smallfor lower cost,text-embedding-3-largefor better quality - Cohere Embed
- Sentence Transformers if you want self-hosted/open-source
2) Store them in a vector database
Use a database that supports fast similarity search:
- Pinecone – easy managed service
- Weaviate – strong features, hybrid search
- Qdrant – great open-source/managed option
- Milvus – scalable, more infra-heavy
- pgvector – if you already use PostgreSQL and your scale is modest
3) Retrieve with semantic similarity
At query time:
- Embed the user query
- Find nearest chunks/documents
- Return top-k results
4) Add reranking for better relevance
A reranker improves quality a lot, especially for knowledge bases:
- Cohere Rerank
- bge-reranker
- Cross-encoder rerankers
If you want the simplest practical answer
If you’re starting today, I’d suggest:
- Embeddings: OpenAI
text-embedding-3-smallor-large - Vector DB: Qdrant or Pinecone
- Chunking: split docs into ~300–800 tokens with overlap
- Reranker: add later if search quality isn’t enough
When to use hybrid search
If your KB contains:
- product names
- IDs
- exact phrases
- code/error messages
Then hybrid search is better:
- lexical search (BM25 / keyword)
-
- vector search
-
- reranking
This helps because pure semantic search can miss exact matches.
Rule of thumb
- Small/simple KB:
pgvector+ embeddings - Production knowledge base: Qdrant/Pinecone + embeddings + reranker
- Need exact + semantic: hybrid search
If you tell me:
- your KB size,
- whether you need self-hosting,
- your stack (Python/JS/Postgres/etc.), I can recommend a concrete architecture and model choice.