Prompt

How do I generate embeddings from PDFs and make them searchable?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To generate embeddings from PDFs and make them searchable, the usual pipeline is:

  1. Extract text from PDFs
  2. Split text into chunks
  3. Create embeddings for each chunk
  4. Store embeddings in a vector database
  5. Embed the user query and search for nearest chunks
  6. Return matching passages, optionally with metadata

1) Extract text from PDFs

Use a PDF parser such as:

  • PyMuPDF (fitz)
  • pdfplumber
  • pypdf

Example with PyMuPDF:

import fitz  # PyMuPDF

def extract_text_from_pdf(pdf_path):
    doc = fitz.open(pdf_path)
    text = ""
    for page in doc:
        text += page.get_text()
    return text

2) Chunk the text

Embeddings work best on smaller passages, not whole documents. Split into chunks of around 200–1000 tokens, with some overlap.

Simple character-based example:

def chunk_text(text, chunk_size=1000, overlap=200):
    chunks = []
    start = 0
    while start < len(text):
        end = start + chunk_size
        chunks.append(text[start:end])
        start = end - overlap
    return chunks

For better results, chunk by paragraphs or tokens.

3) Generate embeddings

You can use an embedding model from:

  • OpenAI
  • SentenceTransformers
  • Hugging Face
  • Cohere, etc.

Example with sentence-transformers:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")

chunks = chunk_text(extract_text_from_pdf("myfile.pdf"))
embeddings = model.encode(chunks, normalize_embeddings=True)

4) Store in a vector index

For small projects, use FAISS. For larger systems, use:

  • Pinecone
  • Weaviate
  • Qdrant
  • Milvus
  • Chroma

Example with FAISS:

import faiss
import numpy as np

# embeddings should be a numpy array of shape (num_chunks, dim)
embeddings = np.array(embeddings).astype("float32")

dimension = embeddings.shape[1]
index = faiss.IndexFlatIP(dimension)  # cosine-like similarity if normalized
index.add(embeddings)

Also keep the original chunk text and metadata in a side list or database:

chunk_store = [
    {"text": chunk, "page": 1, "source": "myfile.pdf"}
    for chunk in chunks
]

5) Search with a query

Embed the query the same way, then search nearest vectors:

def search(query, model, index, chunk_store, top_k=5):
    q_emb = model.encode([query], normalize_embeddings=True)
    q_emb = np.array(q_emb).astype("float32")
    scores, ids = index.search(q_emb, top_k)

    results = []
    for score, idx in zip(scores[0], ids[0]):
        results.append({
            "score": float(score),
            "text": chunk_store[idx]["text"],
            "source": chunk_store[idx]["source"],
            "page": chunk_store[idx]["page"]
        })
    return results

6) Make it searchable in an app

You can wrap this in:

  • a CLI
  • a FastAPI/Flask backend
  • a chatbot/RAG pipeline

A typical flow:

  • Upload PDF
  • Extract and chunk text
  • Compute embeddings
  • Store vectors + metadata
  • On query, embed query and retrieve top matches

Practical tips

  • OCR: If PDFs are scanned images, first run OCR with Tesseract, Azure OCR, Google Vision, etc.
  • Chunk carefully: Keep chunks semantically coherent; avoid splitting mid-sentence if possible.
  • Metadata: Store page number, section, filename, and chunk index.
  • Normalize embeddings if using cosine similarity.
  • Hybrid search: Combine keyword search (BM25) + vector search for better recall.

Minimal end-to-end example

import fitz
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer

def extract_text_from_pdf(pdf_path):
    doc = fitz.open(pdf_path)
    pages = []
    for i, page in enumerate(doc):
        pages.append((i + 1, page.get_text()))
    return pages

def chunk_text(text, chunk_size=1000, overlap=200):
    chunks = []
    start = 0
    while start < len(text):
        end = min(start + chunk_size, len(text))
        chunks.append(text[start:end])
        if end == len(text):
            break
        start = end - overlap
    return chunks

# Load model
model = SentenceTransformer("all-MiniLM-L6-v2")

# Build index
pages = extract_text_from_pdf("myfile.pdf")
chunk_store = []
all_embeddings = []

for page_num, text in pages:
    for chunk in chunk_text(text):
        chunk_store.append({"page": page_num, "text": chunk})
        emb = model.encode(chunk, normalize_embeddings=True)
        all_embeddings.append(emb)

embeddings = np.array(all_embeddings).astype("float32")
index = faiss.IndexFlatIP(embeddings.shape[1])
index.add(embeddings)

# Search
query = "What does the document say about cancellation policy?"
q_emb = model.encode([query], normalize_embeddings=True).astype("float32")
scores, ids = index.search(q_emb, 5)

for score, idx in zip(scores[0], ids[0]):
    print(score, chunk_store[idx]["page"], chunk_store[idx]["text"][:300])

If you want, I can also give you:

  • a production-ready Python script
  • a FastAPI example
  • or a version using OpenAI embeddings + Chroma/Qdrant.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.