Prompt
How do I generate embeddings from documents and search them by meaning?
Latest observation
To search documents by meaning rather than exact keywords, you typically use embeddings + vector search.
Basic idea
-
Split documents into chunks
Large documents are broken into smaller pieces (paragraphs, sections, sentences). -
Create embeddings for each chunk
An embedding model converts each chunk into a numeric vector that captures semantic meaning. -
Store vectors in a vector database / index
Along with metadata like document ID, title, source, etc. -
Embed the user query
Convert the search query into a vector using the same embedding model. -
Find nearest vectors
Use similarity search (cosine similarity, dot product, or Euclidean distance) to return the most semantically similar chunks. -
Optionally rerank results
A reranker can improve relevance by considering the query and candidate chunks together.
Typical workflow
1) Chunk your documents
Example:
- 500–1,000 tokens per chunk
- 10–20% overlap between chunks
This helps retrieval because embeddings work best on focused text segments.
2) Generate embeddings
Use an embedding model such as:
- OpenAI embeddings
- Sentence Transformers
- Cohere embeddings
- Azure OpenAI embeddings
Each chunk becomes a vector like:
[0.12, -0.44, 0.88, ...]
3) Store in a vector database
Popular options:
- FAISS
- Chroma
- Pinecone
- Weaviate
- Milvus
- pgvector
Store:
- vector
- chunk text
- metadata
4) Search by query
Embed the query:
“How do I reset my password?”
Then retrieve chunks close in vector space, even if they don’t contain the exact words “reset password” but say things like:
- “change account credentials”
- “recover login access”
- “set a new password”
Minimal Python example
Using sentence-transformers + faiss:
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
# Example chunks
chunks = [
"Our refund policy allows returns within 30 days.",
"To reset your password, click 'Forgot Password' on the login page.",
"Shipping usually takes 3 to 5 business days."
]
# Load embedding model
model = SentenceTransformer("all-MiniLM-L6-v2")
# Create embeddings
embeddings = model.encode(chunks, normalize_embeddings=True)
# Build vector index
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim) # inner product works well with normalized vectors
index.add(np.array(embeddings, dtype=np.float32))
# Query
query = "How do I change my password?"
query_embedding = model.encode([query], normalize_embeddings=True)
# Search
scores, indices = index.search(np.array(query_embedding, dtype=np.float32), k=2)
for score, idx in zip(scores[0], indices[0]):
print(score, chunks[idx])
Tips for better semantic search
Chunking
- Don’t embed huge documents as one block
- Use chunking with overlap
- Keep related text together
Use the same embedding model
- Document embeddings and query embeddings must come from the same model
Normalize if using cosine similarity
- Common practice for semantic search
Add metadata filters
You can filter by:
- document type
- date
- author
- category
Rerank top results
If you need high precision, use a reranker after vector search.
When embeddings are a good fit
Use them when you want:
- semantic search
- FAQ retrieval
- document Q&A
- recommendation by meaning
- duplicate or similar text detection
Common architecture for RAG
Embeddings are often used in retrieval-augmented generation:
- User asks a question
- Retrieve relevant chunks with embeddings
- Send chunks to an LLM to generate an answer
If you want, I can also give you:
- a complete working example in Python,
- an example using OpenAI embeddings, or
- a design for search over PDFs / documents at scale.