Prompt
What should I use for hybrid search over structured and unstructured data?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For hybrid search over structured + unstructured data, the best choice usually depends on how much of each you have and how tightly you need to combine them. In practice, these are the common options:
Good architecture options
1. Search engine with both keyword + vector + filters
Best when you want one system to handle:
- full-text search over documents
- semantic/vector search
- structured filters/facets on metadata
Examples:
- Elasticsearch / OpenSearch
- Azure AI Search
- Meilisearch or Typesense for simpler setups
Use this if:
you want a single retrieval layer and your structured data can be represented as filters/fields alongside text.
2. Vector database + SQL/OLAP database
Best when structured data is truly relational or analytical.
Pattern:
- store documents/embeddings in a vector DB like:
- Pinecone
- Weaviate
- Milvus
- pgvector
- store structured entities in Postgres / MySQL / Snowflake / BigQuery
- query both, then merge/rerank
Use this if:
your structured data has joins, transactions, or complex filtering logic.
3. Postgres with pgvector
Best for smaller to medium systems or when simplicity matters.
You can do:
- SQL filters on structured columns
- vector similarity with
pgvector - keyword search with Postgres full-text search
Use this if:
you want a simpler stack and don’t need massive scale.
What I’d recommend
If you want the most practical “hybrid search” stack:
- OpenSearch / Elasticsearch if search relevance is the top priority
- Postgres + pgvector if you want simplicity and already use Postgres
- Weaviate / Azure AI Search if you want built-in hybrid retrieval with less engineering
Rule of thumb
- Mostly search / documents / metadata filters → use Elasticsearch/OpenSearch
- Mostly app data / relational structured records → use Postgres + pgvector
- Need both strong semantic and structured retrieval with minimal plumbing → use Weaviate or Azure AI Search
If you tell me:
- your data types,
- expected scale, and
- whether you need joins/filters/facets,
I can suggest a specific stack.