Prompt
How do I generate embeddings from PDFs and make them searchable?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To generate embeddings from PDFs and make them searchable, the usual pipeline is:
- Extract text from PDFs
- Split text into chunks
- Create embeddings for each chunk
- Store embeddings in a vector database
- Embed the user query and search for nearest chunks
- Return matching passages, optionally with metadata
1) Extract text from PDFs
Use a PDF parser such as:
- PyMuPDF (
fitz) - pdfplumber
- pypdf
Example with PyMuPDF:
import fitz # PyMuPDF
def extract_text_from_pdf(pdf_path):
doc = fitz.open(pdf_path)
text = ""
for page in doc:
text += page.get_text()
return text
2) Chunk the text
Embeddings work best on smaller passages, not whole documents. Split into chunks of around 200–1000 tokens, with some overlap.
Simple character-based example:
def chunk_text(text, chunk_size=1000, overlap=200):
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end])
start = end - overlap
return chunks
For better results, chunk by paragraphs or tokens.
3) Generate embeddings
You can use an embedding model from:
- OpenAI
- SentenceTransformers
- Hugging Face
- Cohere, etc.
Example with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
chunks = chunk_text(extract_text_from_pdf("myfile.pdf"))
embeddings = model.encode(chunks, normalize_embeddings=True)
4) Store in a vector index
For small projects, use FAISS. For larger systems, use:
- Pinecone
- Weaviate
- Qdrant
- Milvus
- Chroma
Example with FAISS:
import faiss
import numpy as np
# embeddings should be a numpy array of shape (num_chunks, dim)
embeddings = np.array(embeddings).astype("float32")
dimension = embeddings.shape[1]
index = faiss.IndexFlatIP(dimension) # cosine-like similarity if normalized
index.add(embeddings)
Also keep the original chunk text and metadata in a side list or database:
chunk_store = [
{"text": chunk, "page": 1, "source": "myfile.pdf"}
for chunk in chunks
]
5) Search with a query
Embed the query the same way, then search nearest vectors:
def search(query, model, index, chunk_store, top_k=5):
q_emb = model.encode([query], normalize_embeddings=True)
q_emb = np.array(q_emb).astype("float32")
scores, ids = index.search(q_emb, top_k)
results = []
for score, idx in zip(scores[0], ids[0]):
results.append({
"score": float(score),
"text": chunk_store[idx]["text"],
"source": chunk_store[idx]["source"],
"page": chunk_store[idx]["page"]
})
return results
6) Make it searchable in an app
You can wrap this in:
- a CLI
- a FastAPI/Flask backend
- a chatbot/RAG pipeline
A typical flow:
- Upload PDF
- Extract and chunk text
- Compute embeddings
- Store vectors + metadata
- On query, embed query and retrieve top matches
Practical tips
- OCR: If PDFs are scanned images, first run OCR with Tesseract, Azure OCR, Google Vision, etc.
- Chunk carefully: Keep chunks semantically coherent; avoid splitting mid-sentence if possible.
- Metadata: Store page number, section, filename, and chunk index.
- Normalize embeddings if using cosine similarity.
- Hybrid search: Combine keyword search (BM25) + vector search for better recall.
Minimal end-to-end example
import fitz
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer
def extract_text_from_pdf(pdf_path):
doc = fitz.open(pdf_path)
pages = []
for i, page in enumerate(doc):
pages.append((i + 1, page.get_text()))
return pages
def chunk_text(text, chunk_size=1000, overlap=200):
chunks = []
start = 0
while start < len(text):
end = min(start + chunk_size, len(text))
chunks.append(text[start:end])
if end == len(text):
break
start = end - overlap
return chunks
# Load model
model = SentenceTransformer("all-MiniLM-L6-v2")
# Build index
pages = extract_text_from_pdf("myfile.pdf")
chunk_store = []
all_embeddings = []
for page_num, text in pages:
for chunk in chunk_text(text):
chunk_store.append({"page": page_num, "text": chunk})
emb = model.encode(chunk, normalize_embeddings=True)
all_embeddings.append(emb)
embeddings = np.array(all_embeddings).astype("float32")
index = faiss.IndexFlatIP(embeddings.shape[1])
index.add(embeddings)
# Search
query = "What does the document say about cancellation policy?"
q_emb = model.encode([query], normalize_embeddings=True).astype("float32")
scores, ids = index.search(q_emb, 5)
for score, idx in zip(scores[0], ids[0]):
print(score, chunk_store[idx]["page"], chunk_store[idx]["text"][:300])
If you want, I can also give you:
- a production-ready Python script
- a FastAPI example
- or a version using OpenAI embeddings + Chroma/Qdrant.