Prompt
I'm unhappy with my search stack because embeddings and keyword search are split apart
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
That’s a very common pain point. When keyword search and vector/embedding search live in separate systems, you end up juggling:
- two indexes and two relevance models
- hard-to-tune result merging
- inconsistent filtering/faceting
- higher latency and operational overhead
- awkward debugging when results “feel wrong”
What you usually want is a hybrid search stack where lexical and semantic retrieval happen in one query path, with one ranking pipeline.
Common ways to fix it
1) Use a search engine that supports both
Examples:
- Elasticsearch / OpenSearch: BM25 + vector fields + hybrid scoring
- Vespa: very strong for hybrid ranking and large-scale retrieval
- Solr: can do both, though vector capabilities are less commonly used than ES/OpenSearch
- Postgres + pgvector: workable for smaller systems, but not a full search engine replacement
- Weaviate / Qdrant / Milvus: strong vector DBs, but keyword search and ranking sophistication vary
2) Move to a two-stage hybrid ranking design
A practical pattern is:
-
Candidate retrieval
- fetch top N by BM25
- fetch top N by vector similarity
- union or blend candidates
-
Re-ranking
- combine signals: text match, vector similarity, freshness, popularity, filters, business rules
- optionally use a cross-encoder or LLM-based reranker for top 50–200 results
This avoids “keyword versus embedding” as separate user-facing systems.
3) Make keyword and semantic signals part of one relevance score
You can combine scores like:
final_score = a * bm25 + b * cosine_similarity + c * recency + d * clicks- or use rank fusion methods like:
- Reciprocal Rank Fusion (RRF)
- weighted normalization
- learning-to-rank
RRF is often a great first step because it’s robust and simple.
If you want the least painful migration
A pragmatic path is:
- Keep your current keyword index.
- Add vector fields to the same documents.
- Implement hybrid retrieval in one search API.
- Add reranking later.
- Measure on a small labeled query set before tuning weights.
What to look for in a replacement
If your main complaint is “split brain,” prioritize:
- single query API
- hybrid retrieval and ranking
- filtering/facets integrated with both retrieval modes
- good relevance tuning tools
- fast indexing and incremental updates
- observability for why a result ranked where it did
My recommendation
- For enterprise search / general app search: Elasticsearch or OpenSearch
- For advanced hybrid relevance at scale: Vespa
- For mostly vector-first with some keyword needs: Weaviate or Qdrant
- For small/simple stacks: Postgres + pgvector, if your search needs aren’t too demanding
If you want, I can help you choose a concrete architecture based on:
- your current stack
- data size
- query volume
- whether you need filters/facets
- and how important relevance quality is versus simplicity.