Prompt
I’m building a document search app and need recommendations on embeddings models and vector storage
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Here’s a practical guide for choosing embeddings models and vector storage for a document search app.
1) Embeddings models: what to pick
Good default choices
If you want strong quality with minimal tuning:
-
OpenAI
text-embedding-3-small- Good cost/performance balance
- Solid for semantic search over documents
- Lower latency and cheaper than larger models
-
OpenAI
text-embedding-3-large- Better retrieval quality, especially for nuanced queries
- More expensive, but worth it if search quality is critical
Strong open-source options
If you want to self-host or avoid vendor lock-in:
- bge-large-en / bge-base-en
- Very good general-purpose retrieval embeddings
- Popular for document search
- e5-large / e5-base
- Strong for semantic retrieval
- Works well with query/document prefixing patterns
- jina-embeddings-v2
- Good multilingual support and practical performance
How to choose
- English-only, general doc search:
text-embedding-3-smallorbge-base-en - Highest retrieval quality:
text-embedding-3-largeorbge-large-en - Multilingual search:
text-embedding-3-largeor multilingual models likejina-embeddings-v2 - Self-hosted / open-source: BGE or E5 family
2) Vector storage: what to use
Best hosted/vector DB options
- Pinecone
- Very easy to operationalize
- Good scaling and managed experience
- Weaviate
- Strong feature set, hybrid search support
- Good if you want metadata filtering + vector search
- Qdrant
- Great performance, simple API, strong filtering
- Excellent open-source and managed options
If you already use a database
- Postgres + pgvector
- Best for simpler apps or when you want everything in one DB
- Great if your scale is moderate
- Easy metadata + relational queries
- MongoDB Atlas Vector Search
- Useful if your app already lives in MongoDB
- Elasticsearch / OpenSearch
- Good if you need strong keyword + vector hybrid search at scale
If you want local/prototyping
- FAISS
- Fast and simple for experiments
- Not a full database; you manage persistence and metadata separately
3) Recommended pairings by scenario
Small-to-medium app, fast to ship
- Embeddings:
text-embedding-3-small - Storage:
pgvectoror Qdrant - Why: simple, cheap, good enough for many document search apps
Higher search quality, production use
- Embeddings:
text-embedding-3-largeorbge-large-en - Storage: Qdrant, Pinecone, or Weaviate
- Why: better retrieval and robust vector indexing
Enterprise / hybrid keyword + semantic search
- Embeddings:
text-embedding-3-largeor BGE - Storage: OpenSearch or Elasticsearch
- Why: combines lexical and vector ranking well
4) Practical architecture tips
Chunking matters as much as embeddings
For documents:
- Split into chunks of about 200–800 tokens
- Use overlap of ~10–20%
- Store:
- chunk text
- document ID
- section/page info
- metadata like title, author, date, tags
Use hybrid retrieval
Combine:
- vector similarity
- keyword/BM25 search
- optional reranker
This usually improves relevance a lot for document search.
Add a reranker if quality matters
A reranker can greatly improve top-k results:
- Cross-encoder rerankers
- Hosted rerankers
- Open-source rerankers like BGE reranker models
5) My simple recommendation
If you want a safe default:
- Embeddings:
text-embedding-3-small - Vector store:
Qdrantorpgvector - Search strategy: hybrid keyword + vector
- Optional: reranker for top 20 results
If search quality is top priority:
- Embeddings:
text-embedding-3-large - Vector store:
Qdrant,Pinecone, orWeaviate - Add reranking
6) Quick decision rule
- Need cheapest and simplest:
pgvector - Need managed scale: Pinecone
- Need best open-source option: Qdrant
- Need hybrid enterprise search: Elasticsearch/OpenSearch
- Need best overall retrieval quality: larger embeddings + reranker
If you want, I can also give you:
- a decision matrix comparing these options,
- a reference architecture for ingestion/search,
- or a code example for one stack like Postgres + embeddings + retrieval.