Prompt
how to index PDFs for semantic search
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To index PDFs for semantic search, the usual pipeline is:
- Extract text from PDFs
- Clean and chunk the text
- Generate embeddings for each chunk
- Store embeddings in a vector database / index
- Query by embedding the search text and retrieve nearest chunks
1) Extract text from PDFs
Use a PDF parser depending on the PDF type:
- Text-based PDFs:
pypdf,pdfplumber,pymupdf - Scanned PDFs: OCR needed (
tesseract, AWS Textract, Google Document AI, Azure OCR)
Example with pymupdf:
import fitz # pymupdf
def extract_text(pdf_path):
doc = fitz.open(pdf_path)
pages = []
for i, page in enumerate(doc):
text = page.get_text("text")
pages.append({"page": i + 1, "text": text})
return pages
2) Clean and chunk the text
Semantic search works better on smaller chunks rather than whole documents.
Typical chunk sizes:
- 200–500 tokens
- 10–20% overlap
You can split by:
- paragraphs
- sentences
- fixed token windows
Example:
def chunk_text(text, chunk_size=1000, overlap=200):
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end])
start += chunk_size - overlap
return chunks
For better quality, split on sentence boundaries or use token-based chunking.
3) Create embeddings
Use an embedding model to convert each chunk into a vector.
Options:
- OpenAI embeddings
- SentenceTransformers
- Cohere
- Azure OpenAI
- local models like
all-MiniLM-L6-v2
Example with SentenceTransformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(chunks, normalize_embeddings=True)
4) Store vectors in a vector index
Common choices:
- FAISS for local/simple indexing
- Chroma
- Pinecone
- Weaviate
- Qdrant
- Milvus
- Elasticsearch/OpenSearch with vector support
Example with FAISS:
import faiss
import numpy as np
dim = len(embeddings[0])
index = faiss.IndexFlatIP(dim) # inner product for normalized embeddings
index.add(np.array(embeddings, dtype=np.float32))
Also store metadata alongside each chunk:
- document name
- page number
- chunk text
- chunk id
5) Search
Embed the query and retrieve nearest vectors.
query = "What is the refund policy?"
q_vec = model.encode([query], normalize_embeddings=True)
scores, ids = index.search(np.array(q_vec, dtype=np.float32), k=5)
Then map ids back to your stored chunk metadata and return the matched text.
Recommended practical pipeline
For production, do this:
- Extract per page
- Keep page metadata
- Chunk by paragraphs or tokens
- Embed chunks
- Store in vector DB
- Optionally add keyword/BM25 search too
- Use hybrid retrieval for better results
- Rerank top results with a cross-encoder or LLM if needed
If PDFs are scanned
If the PDF contains images instead of real text:
- Render pages to images
- OCR each page
- Clean OCR output
- Chunk and embed as above
Important tips
- Don’t index the whole PDF as one vector
- Preserve metadata like
filename,page,section - Use overlap between chunks to avoid losing context
- Normalize text (remove headers/footers if repeated)
- For better search, combine semantic search + keyword search
Minimal example architecture
- Parser: PyMuPDF
- Embedding model: SentenceTransformers
- Vector store: FAISS or Qdrant
- Retriever: top-k nearest chunks
- Optional reranker: cross-encoder
If you want, I can give you:
- a complete Python example for indexing PDFs with FAISS, or
- an example using LangChain/LlamaIndex, or
- a production-ready setup with Qdrant/Pinecone.