Prompt

I need an embeddings pipeline for RAG over PDFs, including chunking and re-embedding

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Here’s a solid embeddings pipeline for RAG over PDFs, with chunking, incremental updates, and re-embedding support.

1) High-level architecture

PDFs → text extraction → cleaning → structure-aware chunking → embeddings → vector DB → retrieval

You’ll want the pipeline to support:

  • Initial ingestion
  • Reprocessing changed PDFs
  • Chunk updates
  • Embedding model upgrades
  • Metadata tracking for traceability

2) Recommended pipeline stages

A. PDF ingestion

Store the source PDF in object storage or a file system and assign:

  • document_id
  • source_uri
  • version
  • checksum or etag
  • ingested_at

This lets you detect changes and re-embed only when needed.


B. Text extraction

Use a robust PDF parser depending on the document type:

  • Text-based PDFs: pdfplumber, pymupdf, unstructured
  • Scanned PDFs: OCR via Tesseract, AWS Textract, Azure Form Recognizer, Google Document AI

Recommended approach:

  1. Try native text extraction first
  2. If low text density, fall back to OCR
  3. Preserve layout info if possible:
    • page number
    • section headings
    • paragraph boundaries
    • tables separately if needed

C. Cleaning and normalization

Before chunking:

  • remove repeated headers/footers
  • normalize whitespace
  • de-hyphenate line breaks
  • fix broken paragraphs
  • optionally preserve code blocks, tables, and bullet lists as atomic units

Keep the raw text and cleaned text separately.


3) Chunking strategy

For PDFs, chunking should be structure-aware, not just fixed-length splitting.

Best practice

Chunk by:

  1. Section / heading
  2. Paragraph
  3. Sentence-aware fallback
  4. Token limit cap

Suggested chunk targets

A practical default:

  • Chunk size: 300–800 tokens
  • Overlap: 50–150 tokens

Use smaller chunks for dense technical content, larger chunks for narrative content.

Chunking rules

  • Don’t split mid-table if avoidable
  • Keep headings with the following content
  • Avoid splitting numbered lists across chunks if possible
  • If a chunk is too large, recursively split by paragraph, then sentence

Chunk metadata

Store for each chunk:

  • document_id
  • chunk_id
  • chunk_index
  • page_start
  • page_end
  • section_title
  • char_start, char_end or token offsets
  • text
  • checksum
  • embedding_model
  • embedding_version

4) Embedding generation

Use a consistent embedding model and version it.

Store embedding metadata

For each embedded chunk:

  • embedding_model_name
  • embedding_model_version
  • dimension
  • created_at

Re-embedding triggers

Re-embed when:

  • the source PDF changes
  • chunking logic changes
  • embedding model changes
  • preprocessing changes materially
  • OCR/text extraction improves

5) Re-embedding strategy

This is the key part for maintainability.

Option 1: Full re-embed

Use when:

  • model changes significantly
  • chunking policy changes
  • document structure changes a lot

Process:

  1. detect affected documents
  2. regenerate chunks
  3. embed all chunks
  4. replace old vectors

Option 2: Incremental re-embed

Use when only some documents changed.

Process:

  1. compare document checksum/version
  2. extract and chunk only changed PDFs
  3. compute chunk checksums
  4. only embed new/modified chunks
  5. delete stale chunks from vector DB

Option 3: Chunk-level diffing

Best for large documents.

Store a stable chunk fingerprint based on:

  • normalized chunk text
  • page range
  • section path

Then:

  • unchanged chunk fingerprint → reuse existing embedding
  • changed chunk fingerprint → re-embed
  • removed fingerprint → delete from index

6) Vector database schema

A simple schema:

Document table

  • document_id
  • source_uri
  • version
  • checksum
  • title
  • ingested_at
  • updated_at

Chunk table

  • chunk_id
  • document_id
  • chunk_index
  • page_start
  • page_end
  • section_title
  • text
  • text_checksum
  • embedding_model
  • embedding_version
  • status (active, stale, deleted)

Vector store payload

  • chunk_id
  • document_id
  • page_start
  • page_end
  • section_title
  • source_uri

7) Retrieval flow for RAG

At query time:

  1. embed the user query
  2. retrieve top-k chunks from vector DB
  3. optionally rerank with a cross-encoder or LLM
  4. assemble context with citations
  5. generate answer

Good retrieval practices

  • hybrid search: dense + keyword/BM25
  • metadata filtering:
    • document type
    • date
    • source
    • section
  • reranking for better precision
  • context window packing with deduplication

8) Suggested implementation pattern

Pipeline components

  • Extractor
  • Cleaner
  • Chunker
  • Embedder
  • Indexer
  • Change detector
  • Reindexer

Event-driven flow

When a PDF is added/updated:

  1. upload PDF
  2. compute checksum
  3. if checksum unchanged, skip
  4. extract text
  5. chunk text
  6. hash each chunk
  7. compare against prior chunk hashes
  8. embed only new/changed chunks
  9. delete obsolete chunk vectors
  10. update document status

9) Practical chunking algorithm

A robust approach:

  1. parse PDF into page text blocks
  2. detect headings using font/size/layout if available
  3. group blocks into sections
  4. form chunks by accumulating paragraphs until token limit
  5. overlap neighboring chunks
  6. if a chunk exceeds limit, split at paragraph/sentence boundaries

This gives much better retrieval than naive fixed windows.


10) Example pseudo-flow

PDF uploaded
  ↓
checksum compare
  ↓
if changed:
  ↓
extract text/OCR
  ↓
clean text
  ↓
structure-aware chunking
  ↓
chunk fingerprinting
  ↓
compare with stored chunk fingerprints
  ↓
embed new/changed chunks
  ↓
upsert vectors
  ↓
delete stale vectors

11) Best practices

  • Keep raw text, cleaned text, and chunked text separately
  • Version everything:
    • source document
    • extraction method
    • chunking policy
    • embedding model
  • Store page references for citations
  • Use overlap, but not too much, or retrieval gets redundant
  • Don’t embed massive chunks; retrieval becomes less precise
  • Don’t over-chunk; context gets fragmented
  • Re-embed on model upgrades, not just on content changes

12) Recommended defaults

If you need a starting point:

  • Chunk size: 500 tokens
  • Overlap: 100 tokens
  • Embedding model: choose a strong general-purpose text embedding model
  • Store metadata: page range + section title + document version
  • Update policy: checksum-based incremental re-embedding

13) If you want a production-ready stack

A common setup:

  • Extraction: PyMuPDF / unstructured / Textract
  • Chunking: custom structure-aware Python logic
  • Embeddings: OpenAI embeddings or a local sentence-transformer
  • Vector DB: pgvector / Qdrant / Pinecone / Weaviate
  • Orchestration: Airflow / Prefect / Celery
  • Metadata store: Postgres

If you want, I can turn this into:

  1. a Python reference implementation,
  2. a LangChain/LlamaIndex pipeline, or
  3. a production architecture diagram + schema.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.