Prompt

How should I design an embeddings pipeline for PDFs, images, and metadata filtering if I’m optimizing for retrieval quality and low cost?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good low-cost, high-quality embeddings pipeline for PDFs, images, and metadata filtering should be designed around one principle:

Use the cheapest representation that preserves retrieval signal, and only add richer processing where it measurably improves recall/precision.

Here’s a practical architecture.


1) Separate the problem into 3 retrieval layers

A. Metadata first

Use metadata for hard filtering:

  • document type
  • source
  • date range
  • author
  • tenant/customer
  • language
  • page number
  • product/category/tags
  • security labels

This is very cheap and should happen before vector search whenever possible.

B. Embedding retrieval

Use vector search for semantic matching:

  • query text
  • PDF text chunks
  • OCR text from images
  • captions / extracted labels from images

C. Reranking

Use a stronger but more expensive model only on the top-k candidates:

  • cross-encoder reranker
  • LLM-based reranker if needed
  • hybrid score combining metadata, vector similarity, and rerank score

This usually gives the best quality-per-dollar.


2) Ingest PDFs in two representations

PDFs are often mixed-content, so don’t treat them as one blob.

Extract:

  1. Text layer

    • Parse native text from the PDF first.
    • Preserve structure if possible:
      • headings
      • paragraphs
      • tables
      • page number
      • section titles
  2. Layout-aware chunks

    • Chunk by semantic boundaries, not fixed tokens only.
    • Include:
      • title
      • heading path
      • page index
      • nearby context
  3. Fallback OCR only for text-poor pages

    • If a page has little or no extractable text, run OCR.
    • This saves cost versus OCR-ing everything.

Recommended storage per chunk:

  • chunk_text
  • doc_id
  • page_start, page_end
  • section_path
  • chunk_type (native_text, ocr, table, caption)
  • metadata

Quality tips:

  • Keep chunks around 200–500 tokens with overlap only when needed.
  • For tables, store both:
    • flattened text representation
    • optionally a compact structured form
  • For scanned PDFs, OCR output should be chunked by page and paragraph.

3) Treat images as multimodal assets, but index them cheaply

Images can be retrieved in several ways. Use a tiered strategy.

Option 1: OCR + caption + tags

For most image documents, the cheapest strong baseline is:

  • OCR text
  • auto-caption
  • detected objects / labels
  • surrounding document text if image appears in a PDF

Then embed the resulting text.

Option 2: Image embeddings

If visual similarity matters:

  • use an image-text embedding model
  • store image embeddings separately from text embeddings
  • optionally add generated captions so text queries can still match images

Best practice:

For each image, index both:

  • a text surrogate: OCR/caption/alt-text/tags
  • a visual embedding: for image-to-image or image-to-text retrieval

This improves recall without forcing all queries into expensive multimodal search.


4) Use one canonical chunk record format

Whether the source is PDF text, OCR text, or image caption, normalize everything into a common record.

Example schema

{
  "id": "chunk_123",
  "doc_id": "doc_45",
  "source_type": "pdf|image|ocr|table",
  "text": "normalized searchable content",
  "embedding_text": "text used for embedding",
  "image_embedding_id": "optional",
  "metadata": {
    "tenant": "acme",
    "language": "en",
    "date": "2025-01-12",
    "page": 7,
    "section": "Pricing > Enterprise",
    "tags": ["invoice", "contract"]
  }
}

The important idea:

  • keep raw text
  • keep normalized text
  • keep metadata
  • keep source pointers back to the original file/page/image

5) Embed with purpose-built fields, not everything concatenated blindly

To optimize quality and cost:

Good practice

Create separate text fields for different retrieval intents:

  • title
  • section_heading
  • chunk_text
  • caption
  • ocr_text
  • tags

Then either:

  • embed a controlled concatenation, or
  • create multiple embeddings per record for different fields

Cheap, strong default

For most records:

  • embedding_input = title + section_heading + chunk_text + caption + key metadata

Avoid stuffing huge metadata blobs into the embedding text. Metadata should usually be used for filtering, not semantic embedding.


6) Use hybrid retrieval

For retrieval quality, don’t rely on vectors alone.

Recommended approach

Combine:

  • BM25 / keyword search
  • vector search
  • metadata filters

Why:

  • keyword search is excellent for exact terms, IDs, error codes, names
  • vector search is excellent for semantic matching
  • filters narrow the candidate set cheaply

A simple hybrid score can work well:

  • 50% vector similarity
  • 30% lexical match
  • 20% reranker score

Exact weights depend on your data.


7) Metadata filtering strategy

Use metadata in two ways:

Hard filters

Before retrieval:

  • tenant
  • permissions
  • document type
  • language
  • date range

This reduces index load and improves precision.

Soft ranking features

After retrieval:

  • recency
  • popularity
  • source trust score
  • document freshness
  • user-specific relevance

This is especially useful if multiple documents are semantically similar.

Important:

Keep filters indexed and normalized:

  • dates as dates
  • enums as enums
  • tags as arrays
  • IDs as exact-match fields

Don’t embed these into the vector unless they are part of the semantic content.


8) Choose chunking carefully

Chunking has a huge effect on retrieval quality.

PDFs

  • Prefer structure-aware chunks.
  • Use headings and paragraphs as boundaries.
  • Split very long sections.
  • Add overlap only at semantic boundaries.

Images

  • One image = one record, unless it’s a multi-panel image.
  • For multi-panel figures, split by panel if you can detect it.
  • Attach nearby caption/legend text.

Tables

  • Keep row groups together if possible.
  • If the table is large, index:
    • table summary
    • row-level chunks
    • optionally column context

9) Reduce OCR cost aggressively

OCR is often the biggest cost center.

Cost-saving rules

  • OCR only pages with low native text coverage.
  • Use layout/text detection before OCR to decide.
  • Cache OCR outputs permanently.
  • Deduplicate repeated pages/images.
  • Skip OCR on pages where native extraction is good enough.

If possible:

  • use a cheaper OCR model first
  • fall back to a better one only when confidence is low

10) Improve retrieval quality with query-time routing

Not all queries should use the same pipeline.

Examples

  • Exact lookup queries: use lexical + metadata
  • Conceptual questions: use vector + reranking
  • Image-specific queries: use image embeddings + OCR/captions
  • Date/tenant constrained queries: apply filters first

A simple query router can lower cost and improve performance.


11) Deduplicate aggressively

Duplication hurts both cost and ranking.

Deduplicate:

  • identical pages
  • repeated boilerplate
  • near-duplicate documents
  • repeated headers/footers
  • repeated figure captions

This reduces indexing cost and avoids noisy retrieval.


12) A practical low-cost architecture

Ingestion

  1. Detect file type
  2. For PDFs:
    • extract native text
    • detect scanned pages
    • OCR only needed pages
    • chunk structurally
  3. For images:
    • OCR if text exists
    • generate caption
    • optionally generate image embedding
  4. Normalize metadata
  5. Store raw assets + chunk records + embeddings

Indexes

  • Metadata index for filtering
  • Lexical index for keyword search
  • Vector index for semantic search
  • Optional image vector index for visual similarity

Retrieval

  1. Apply metadata filters
  2. Run hybrid lexical + vector retrieval
  3. Merge results
  4. Rerank top 20–100
  5. Return top N with source citations

13) Recommended “good default” setup

If you want the best balance of quality and cost:

  • PDF text extraction first
  • OCR only scanned/low-text pages
  • Chunk by headings/paragraphs
  • Embed chunk text plus title/heading
  • Use separate image captions and image embeddings
  • Use metadata as filters, not embedded content
  • Hybrid lexical + vector retrieval
  • Rerank top candidates
  • Deduplicate and cache everything

14) Common mistakes to avoid

  • OCR-ing every PDF page regardless of need
  • Embedding giant unstructured blobs
  • Ignoring metadata filters
  • Using only vector search
  • Not preserving page/source references
  • Over-chunking into tiny fragments
  • Not deduplicating boilerplate
  • Mixing text and image retrieval without normalization

15) If you want the shortest version

For best retrieval quality at low cost:

  1. Extract native PDF text first.
  2. OCR only when needed.
  3. Chunk by semantic structure.
  4. Store metadata separately and filter with it.
  5. Use hybrid lexical + vector search.
  6. Add reranking on top-k.
  7. For images, index both captions/OCR text and image embeddings.
  8. Deduplicate and cache aggressively.

If you want, I can also give you:

  • a reference architecture diagram
  • a database/schema design
  • or a concrete implementation plan using FAISS / pgvector / Pinecone / Elasticsearch / OpenSearch.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.