Prompt
How should I design an embeddings pipeline for PDFs, images, and metadata filtering if I’m optimizing for retrieval quality and low cost?
Latest observation
A good low-cost, high-quality embeddings pipeline for PDFs, images, and metadata filtering should be designed around one principle:
Use the cheapest representation that preserves retrieval signal, and only add richer processing where it measurably improves recall/precision.
Here’s a practical architecture.
1) Separate the problem into 3 retrieval layers
A. Metadata first
Use metadata for hard filtering:
- document type
- source
- date range
- author
- tenant/customer
- language
- page number
- product/category/tags
- security labels
This is very cheap and should happen before vector search whenever possible.
B. Embedding retrieval
Use vector search for semantic matching:
- query text
- PDF text chunks
- OCR text from images
- captions / extracted labels from images
C. Reranking
Use a stronger but more expensive model only on the top-k candidates:
- cross-encoder reranker
- LLM-based reranker if needed
- hybrid score combining metadata, vector similarity, and rerank score
This usually gives the best quality-per-dollar.
2) Ingest PDFs in two representations
PDFs are often mixed-content, so don’t treat them as one blob.
Extract:
-
Text layer
- Parse native text from the PDF first.
- Preserve structure if possible:
- headings
- paragraphs
- tables
- page number
- section titles
-
Layout-aware chunks
- Chunk by semantic boundaries, not fixed tokens only.
- Include:
- title
- heading path
- page index
- nearby context
-
Fallback OCR only for text-poor pages
- If a page has little or no extractable text, run OCR.
- This saves cost versus OCR-ing everything.
Recommended storage per chunk:
chunk_textdoc_idpage_start,page_endsection_pathchunk_type(native_text,ocr,table,caption)metadata
Quality tips:
- Keep chunks around 200–500 tokens with overlap only when needed.
- For tables, store both:
- flattened text representation
- optionally a compact structured form
- For scanned PDFs, OCR output should be chunked by page and paragraph.
3) Treat images as multimodal assets, but index them cheaply
Images can be retrieved in several ways. Use a tiered strategy.
Option 1: OCR + caption + tags
For most image documents, the cheapest strong baseline is:
- OCR text
- auto-caption
- detected objects / labels
- surrounding document text if image appears in a PDF
Then embed the resulting text.
Option 2: Image embeddings
If visual similarity matters:
- use an image-text embedding model
- store image embeddings separately from text embeddings
- optionally add generated captions so text queries can still match images
Best practice:
For each image, index both:
- a text surrogate: OCR/caption/alt-text/tags
- a visual embedding: for image-to-image or image-to-text retrieval
This improves recall without forcing all queries into expensive multimodal search.
4) Use one canonical chunk record format
Whether the source is PDF text, OCR text, or image caption, normalize everything into a common record.
Example schema
{
"id": "chunk_123",
"doc_id": "doc_45",
"source_type": "pdf|image|ocr|table",
"text": "normalized searchable content",
"embedding_text": "text used for embedding",
"image_embedding_id": "optional",
"metadata": {
"tenant": "acme",
"language": "en",
"date": "2025-01-12",
"page": 7,
"section": "Pricing > Enterprise",
"tags": ["invoice", "contract"]
}
}
The important idea:
- keep raw text
- keep normalized text
- keep metadata
- keep source pointers back to the original file/page/image
5) Embed with purpose-built fields, not everything concatenated blindly
To optimize quality and cost:
Good practice
Create separate text fields for different retrieval intents:
titlesection_headingchunk_textcaptionocr_texttags
Then either:
- embed a controlled concatenation, or
- create multiple embeddings per record for different fields
Cheap, strong default
For most records:
embedding_input = title + section_heading + chunk_text + caption + key metadata
Avoid stuffing huge metadata blobs into the embedding text. Metadata should usually be used for filtering, not semantic embedding.
6) Use hybrid retrieval
For retrieval quality, don’t rely on vectors alone.
Recommended approach
Combine:
- BM25 / keyword search
- vector search
- metadata filters
Why:
- keyword search is excellent for exact terms, IDs, error codes, names
- vector search is excellent for semantic matching
- filters narrow the candidate set cheaply
A simple hybrid score can work well:
- 50% vector similarity
- 30% lexical match
- 20% reranker score
Exact weights depend on your data.
7) Metadata filtering strategy
Use metadata in two ways:
Hard filters
Before retrieval:
- tenant
- permissions
- document type
- language
- date range
This reduces index load and improves precision.
Soft ranking features
After retrieval:
- recency
- popularity
- source trust score
- document freshness
- user-specific relevance
This is especially useful if multiple documents are semantically similar.
Important:
Keep filters indexed and normalized:
- dates as dates
- enums as enums
- tags as arrays
- IDs as exact-match fields
Don’t embed these into the vector unless they are part of the semantic content.
8) Choose chunking carefully
Chunking has a huge effect on retrieval quality.
PDFs
- Prefer structure-aware chunks.
- Use headings and paragraphs as boundaries.
- Split very long sections.
- Add overlap only at semantic boundaries.
Images
- One image = one record, unless it’s a multi-panel image.
- For multi-panel figures, split by panel if you can detect it.
- Attach nearby caption/legend text.
Tables
- Keep row groups together if possible.
- If the table is large, index:
- table summary
- row-level chunks
- optionally column context
9) Reduce OCR cost aggressively
OCR is often the biggest cost center.
Cost-saving rules
- OCR only pages with low native text coverage.
- Use layout/text detection before OCR to decide.
- Cache OCR outputs permanently.
- Deduplicate repeated pages/images.
- Skip OCR on pages where native extraction is good enough.
If possible:
- use a cheaper OCR model first
- fall back to a better one only when confidence is low
10) Improve retrieval quality with query-time routing
Not all queries should use the same pipeline.
Examples
- Exact lookup queries: use lexical + metadata
- Conceptual questions: use vector + reranking
- Image-specific queries: use image embeddings + OCR/captions
- Date/tenant constrained queries: apply filters first
A simple query router can lower cost and improve performance.
11) Deduplicate aggressively
Duplication hurts both cost and ranking.
Deduplicate:
- identical pages
- repeated boilerplate
- near-duplicate documents
- repeated headers/footers
- repeated figure captions
This reduces indexing cost and avoids noisy retrieval.
12) A practical low-cost architecture
Ingestion
- Detect file type
- For PDFs:
- extract native text
- detect scanned pages
- OCR only needed pages
- chunk structurally
- For images:
- OCR if text exists
- generate caption
- optionally generate image embedding
- Normalize metadata
- Store raw assets + chunk records + embeddings
Indexes
- Metadata index for filtering
- Lexical index for keyword search
- Vector index for semantic search
- Optional image vector index for visual similarity
Retrieval
- Apply metadata filters
- Run hybrid lexical + vector retrieval
- Merge results
- Rerank top 20–100
- Return top N with source citations
13) Recommended “good default” setup
If you want the best balance of quality and cost:
- PDF text extraction first
- OCR only scanned/low-text pages
- Chunk by headings/paragraphs
- Embed chunk text plus title/heading
- Use separate image captions and image embeddings
- Use metadata as filters, not embedded content
- Hybrid lexical + vector retrieval
- Rerank top candidates
- Deduplicate and cache everything
14) Common mistakes to avoid
- OCR-ing every PDF page regardless of need
- Embedding giant unstructured blobs
- Ignoring metadata filters
- Using only vector search
- Not preserving page/source references
- Over-chunking into tiny fragments
- Not deduplicating boilerplate
- Mixing text and image retrieval without normalization
15) If you want the shortest version
For best retrieval quality at low cost:
- Extract native PDF text first.
- OCR only when needed.
- Chunk by semantic structure.
- Store metadata separately and filter with it.
- Use hybrid lexical + vector search.
- Add reranking on top-k.
- For images, index both captions/OCR text and image embeddings.
- Deduplicate and cache aggressively.
If you want, I can also give you:
- a reference architecture diagram
- a database/schema design
- or a concrete implementation plan using FAISS / pgvector / Pinecone / Elasticsearch / OpenSearch.