Prompt
I need an embeddings pipeline for RAG over PDFs, including chunking and re-embedding
Latest observation
Here’s a solid embeddings pipeline for RAG over PDFs, with chunking, incremental updates, and re-embedding support.
1) High-level architecture
PDFs → text extraction → cleaning → structure-aware chunking → embeddings → vector DB → retrieval
You’ll want the pipeline to support:
- Initial ingestion
- Reprocessing changed PDFs
- Chunk updates
- Embedding model upgrades
- Metadata tracking for traceability
2) Recommended pipeline stages
A. PDF ingestion
Store the source PDF in object storage or a file system and assign:
document_idsource_uriversionchecksumoretagingested_at
This lets you detect changes and re-embed only when needed.
B. Text extraction
Use a robust PDF parser depending on the document type:
- Text-based PDFs:
pdfplumber,pymupdf,unstructured - Scanned PDFs: OCR via
Tesseract,AWS Textract,Azure Form Recognizer,Google Document AI
Recommended approach:
- Try native text extraction first
- If low text density, fall back to OCR
- Preserve layout info if possible:
- page number
- section headings
- paragraph boundaries
- tables separately if needed
C. Cleaning and normalization
Before chunking:
- remove repeated headers/footers
- normalize whitespace
- de-hyphenate line breaks
- fix broken paragraphs
- optionally preserve code blocks, tables, and bullet lists as atomic units
Keep the raw text and cleaned text separately.
3) Chunking strategy
For PDFs, chunking should be structure-aware, not just fixed-length splitting.
Best practice
Chunk by:
- Section / heading
- Paragraph
- Sentence-aware fallback
- Token limit cap
Suggested chunk targets
A practical default:
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
Use smaller chunks for dense technical content, larger chunks for narrative content.
Chunking rules
- Don’t split mid-table if avoidable
- Keep headings with the following content
- Avoid splitting numbered lists across chunks if possible
- If a chunk is too large, recursively split by paragraph, then sentence
Chunk metadata
Store for each chunk:
document_idchunk_idchunk_indexpage_startpage_endsection_titlechar_start,char_endor token offsetstextchecksumembedding_modelembedding_version
4) Embedding generation
Use a consistent embedding model and version it.
Store embedding metadata
For each embedded chunk:
embedding_model_nameembedding_model_versiondimensioncreated_at
Re-embedding triggers
Re-embed when:
- the source PDF changes
- chunking logic changes
- embedding model changes
- preprocessing changes materially
- OCR/text extraction improves
5) Re-embedding strategy
This is the key part for maintainability.
Option 1: Full re-embed
Use when:
- model changes significantly
- chunking policy changes
- document structure changes a lot
Process:
- detect affected documents
- regenerate chunks
- embed all chunks
- replace old vectors
Option 2: Incremental re-embed
Use when only some documents changed.
Process:
- compare document checksum/version
- extract and chunk only changed PDFs
- compute chunk checksums
- only embed new/modified chunks
- delete stale chunks from vector DB
Option 3: Chunk-level diffing
Best for large documents.
Store a stable chunk fingerprint based on:
- normalized chunk text
- page range
- section path
Then:
- unchanged chunk fingerprint → reuse existing embedding
- changed chunk fingerprint → re-embed
- removed fingerprint → delete from index
6) Vector database schema
A simple schema:
Document table
document_idsource_uriversionchecksumtitleingested_atupdated_at
Chunk table
chunk_iddocument_idchunk_indexpage_startpage_endsection_titletexttext_checksumembedding_modelembedding_versionstatus(active,stale,deleted)
Vector store payload
chunk_iddocument_idpage_startpage_endsection_titlesource_uri
7) Retrieval flow for RAG
At query time:
- embed the user query
- retrieve top-k chunks from vector DB
- optionally rerank with a cross-encoder or LLM
- assemble context with citations
- generate answer
Good retrieval practices
- hybrid search: dense + keyword/BM25
- metadata filtering:
- document type
- date
- source
- section
- reranking for better precision
- context window packing with deduplication
8) Suggested implementation pattern
Pipeline components
- Extractor
- Cleaner
- Chunker
- Embedder
- Indexer
- Change detector
- Reindexer
Event-driven flow
When a PDF is added/updated:
- upload PDF
- compute checksum
- if checksum unchanged, skip
- extract text
- chunk text
- hash each chunk
- compare against prior chunk hashes
- embed only new/changed chunks
- delete obsolete chunk vectors
- update document status
9) Practical chunking algorithm
A robust approach:
- parse PDF into page text blocks
- detect headings using font/size/layout if available
- group blocks into sections
- form chunks by accumulating paragraphs until token limit
- overlap neighboring chunks
- if a chunk exceeds limit, split at paragraph/sentence boundaries
This gives much better retrieval than naive fixed windows.
10) Example pseudo-flow
PDF uploaded
↓
checksum compare
↓
if changed:
↓
extract text/OCR
↓
clean text
↓
structure-aware chunking
↓
chunk fingerprinting
↓
compare with stored chunk fingerprints
↓
embed new/changed chunks
↓
upsert vectors
↓
delete stale vectors
11) Best practices
- Keep raw text, cleaned text, and chunked text separately
- Version everything:
- source document
- extraction method
- chunking policy
- embedding model
- Store page references for citations
- Use overlap, but not too much, or retrieval gets redundant
- Don’t embed massive chunks; retrieval becomes less precise
- Don’t over-chunk; context gets fragmented
- Re-embed on model upgrades, not just on content changes
12) Recommended defaults
If you need a starting point:
- Chunk size: 500 tokens
- Overlap: 100 tokens
- Embedding model: choose a strong general-purpose text embedding model
- Store metadata: page range + section title + document version
- Update policy: checksum-based incremental re-embedding
13) If you want a production-ready stack
A common setup:
- Extraction: PyMuPDF / unstructured / Textract
- Chunking: custom structure-aware Python logic
- Embeddings: OpenAI embeddings or a local sentence-transformer
- Vector DB: pgvector / Qdrant / Pinecone / Weaviate
- Orchestration: Airflow / Prefect / Celery
- Metadata store: Postgres
If you want, I can turn this into:
- a Python reference implementation,
- a LangChain/LlamaIndex pipeline, or
- a production architecture diagram + schema.