Prompt
embedding pipeline for PDFs and wiki pages
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Here’s a practical embedding pipeline for PDFs and wiki pages you can use for RAG / semantic search.
1) Ingest sources
- PDFs
- Extract text with a PDF parser/OCR if needed
- Preserve page numbers, headings, tables, and figures if possible
- Wiki pages
- Pull via API / dump / HTML scrape
- Keep page title, section headings, URL, last updated time
2) Clean and normalize
- Remove boilerplate:
- headers/footers
- navigation menus
- duplicate citations
- page numbers
- Normalize:
- whitespace
- unicode
- broken hyphenation from PDF line wraps
- Optional:
- language detection
- OCR cleanup for scanned PDFs
3) Structure the text
Split into logical units:
- Wiki: by page → section → subsection
- PDF: by page → heading blocks → paragraphs
Keep metadata for each chunk:
source_type: pdf/wikisource_id: file path / wiki page IDtitlesection_pathpage_numberurlchunk_index
4) Chunking strategy
Use chunk sizes that fit your embedding model and retrieval needs:
- Typical starting point: 300–800 tokens per chunk
- Overlap: 50–150 tokens
- Prefer semantic boundaries:
- headings
- paragraphs
- list items
- table rows when possible
Good practice:
- avoid splitting in the middle of a sentence
- keep related content together
- if a section is long, recursively split it
5) Create embeddings
For each chunk:
- pass text to embedding model
- store vector + metadata
You may want:
- one embedding per chunk
- optionally a separate embedding for:
- title
- section heading
- summary
6) Store in a vector database
Store:
- embedding vector
- chunk text
- metadata
Common choices:
- pgvector
- Pinecone
- Weaviate
- Milvus
- FAISS
- Elasticsearch/OpenSearch vector fields
7) Retrieval flow
At query time:
- embed the user query
- retrieve top-k similar chunks
- rerank with cross-encoder or LLM if needed
- assemble context with citations
Helpful retrieval features:
- metadata filters:
- source type
- date
- wiki namespace
- document type
- hybrid search:
- keyword + vector
- deduplication:
- avoid repeated adjacent chunks
8) PDF-specific tips
- Use OCR for scanned docs
- Extract tables separately if they matter
- Preserve page numbers for citations
- Watch for:
- multi-column layouts
- footnotes
- text extraction order issues
9) Wiki-specific tips
- Strip nav/sidebar content
- Expand internal links or keep resolved titles
- Capture page hierarchy
- Reindex on page updates using timestamps or change events
10) Recommended architecture
A simple pipeline:
Source connectors → Text extraction → Cleaning → Chunking → Embedding → Vector store → Retrieval API
Example metadata schema
{
"id": "doc123_chunk045",
"source_type": "pdf",
"source_id": "manual_v2.pdf",
"title": "API Reference",
"section_path": ["Authentication", "Token Refresh"],
"page_number": 12,
"chunk_index": 45,
"url": null,
"text": "..."
}
Example chunking rule of thumb
- PDFs: start with 1–2 paragraphs per chunk
- Wiki pages: chunk by subsection
- If a chunk exceeds your target size, split again with overlap
If you want, I can also give you:
- a Python pipeline example
- a LangChain/LlamaIndex version
- a production architecture diagram
- a recommended chunking strategy for your specific document types