Prompt
How do I build semantic search over documents with vector embeddings?
Latest observation
To build semantic search over documents with vector embeddings, you generally do four things:
- Split documents into chunks
- Convert each chunk into an embedding vector
- Store those vectors in a vector index/database
- Embed the user query and retrieve the nearest chunks by similarity
1) Chunk your documents
Instead of embedding an entire document at once, split it into smaller pieces like paragraphs or 200–500 token chunks.
Why:
- Better retrieval granularity
- Less noise
- Easier to fit model token limits
Example chunking strategies:
- fixed size with overlap
- paragraph-based
- sentence-window chunks
2) Generate embeddings
Use an embedding model to turn each chunk into a vector, e.g.:
- OpenAI embeddings
- sentence-transformers
- Cohere embeddings
- local models like
bge,e5,all-MiniLM
Each chunk becomes something like:
"Refunds are processed within 5 business days." -> [0.12, -0.03, ...]
The vector captures semantic meaning, so similar ideas land near each other in vector space.
3) Store vectors in a vector database or index
Save:
- chunk text
- embedding vector
- metadata such as doc id, title, section, timestamp, permissions
Common options:
- FAISS
- Pinecone
- Weaviate
- Milvus
- Chroma
- Elasticsearch / OpenSearch vector search
- PostgreSQL with pgvector
4) Search with a query embedding
When a user searches:
- embed the query
- find nearest chunk vectors using cosine similarity or dot product
- return the top-k chunks
- optionally rerank them with a cross-encoder or LLM
Example:
Query: "How long does a refund take?"
Returns chunks about "Refunds are processed within 5 business days."
Basic architecture
Documents -> Chunking -> Embedding -> Vector Store
User Query -> Embedding -> Similarity Search -> Top Results
Minimal Python example with sentence-transformers + FAISS
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
# Example chunks
chunks = [
"Refunds are processed within 5 business days.",
"You can reset your password from the account settings page.",
"Shipping takes 3 to 7 business days depending on location."
]
# Load embedding model
model = SentenceTransformer("all-MiniLM-L6-v2")
# Embed chunks
embeddings = model.encode(chunks, convert_to_numpy=True)
# Normalize for cosine similarity
faiss.normalize_L2(embeddings)
# Build FAISS index
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim) # inner product on normalized vectors = cosine similarity
index.add(embeddings)
# Search
query = "When will I get my refund?"
query_vec = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(query_vec)
k = 2
scores, indices = index.search(query_vec, k)
for score, idx in zip(scores[0], indices[0]):
print(score, chunks[idx])
Important design tips
1. Use the right embedding model
Choose a model trained for semantic retrieval, not just classification.
Good choices:
bge-*e5-*all-MiniLM-*- provider embeddings from OpenAI/Cohere/etc.
2. Chunk size matters
Too large:
- less precise matches
Too small:
- loses context
A common starting point:
- 200–400 tokens
- 10–20% overlap
3. Use metadata filters
If documents are scoped by:
- user
- department
- language
- date
Filter before or during retrieval to improve relevance and access control.
4. Consider hybrid search
Combine:
- keyword search (BM25)
- vector search
This often works better than embeddings alone, especially for:
- exact terms
- IDs
- product codes
- rare names
5. Rerank results
Vector search is usually a first-pass retrieval. For better quality:
- retrieve top 20–100
- rerank with a cross-encoder or LLM
- return top 3–10
Production pattern
A common production pipeline:
Ingestion
- parse docs
- clean text
- chunk
- embed
- store vectors + metadata
Query
- embed query
- retrieve top-k
- filter by metadata/permissions
- rerank
- return results with citations
If you want the shortest path
If you’re building a prototype:
- use
sentence-transformers - use FAISS or Chroma
- store chunks + metadata
- retrieve top-k by cosine similarity
If you’re building production:
- use a managed vector DB
- add metadata filters
- add hybrid search
- add reranking
- evaluate recall and relevance
If you want, I can also give you:
- a FastAPI + vector DB example,
- a RAG pipeline example,
- or a production-ready schema for document embeddings.