Prompt
I’m trying to build semantic search over product docs and need a practical plan for chunking, indexing, and re-embedding over time
Latest observation
Here’s a practical, production-oriented plan for building semantic search over product docs, with a focus on chunking, indexing, and re-embedding over time.
1) Overall architecture
A good setup is:
- Ingest docs from your sources
- Normalize and structure them into sections
- Chunk the text into retrievable pieces
- Embed each chunk
- Index chunks in a vector store + keyword index
- Search with hybrid retrieval
- Re-rank results if needed
- Re-embed only what changed as docs evolve
A strong default is hybrid search:
- Vector search for semantic similarity
- BM25 / keyword search for exact matches, product names, error codes, API paths, etc.
2) Chunking strategy
Chunking is the most important design choice for doc search.
Goals for chunks
Each chunk should be:
- Self-contained enough to understand
- Small enough to retrieve precisely
- Large enough to preserve context
Recommended chunk size
A good starting point:
- 300–800 tokens per chunk
- 10–20% overlap between adjacent chunks
If your docs are highly structured, use section-based chunks rather than arbitrary fixed windows.
Prefer structure-aware chunking
Split by:
- Headings
- Subheadings
- Lists
- Tables
- Code blocks
- FAQ entries
- Steps/procedures
This usually beats naive sliding windows.
Practical rules
- Keep one topic per chunk
- Don’t split in the middle of a code sample or numbered procedure if possible
- If a section is too long, subdivide it by paragraph or subheading
- If a section is too short, merge with adjacent content until it’s useful
Add metadata to every chunk
Store:
doc_iddoc_titlesection_headingsubheadingchunk_indexsource_urlproductversionlast_updatedcontent_hash- optionally
permissions/ACL
This helps filtering, deduping, and reindexing later.
3) Chunk types by doc format
Product docs / knowledge base articles
Use section-based chunks:
- Title + intro
- Each major heading as a chunk
- Bulleted steps as their own chunk if they are a coherent procedure
API docs
Chunk by:
- Endpoint
- Parameter table
- Request example
- Response example
- Errors / edge cases
For APIs, exact keyword search matters a lot, so hybrid retrieval is especially useful.
Troubleshooting docs
Chunk by:
- Symptom
- Cause
- Resolution
- Related errors
This makes retrieval cleaner than a large “troubleshooting” blob.
Long guides
Use chapter/section boundaries first, then subchunk within sections.
4) Indexing design
Use two indexes:
A. Vector index
Stores embeddings for semantic retrieval.
Common options:
- Pinecone
- Weaviate
- Milvus
- Qdrant
- pgvector
B. Lexical index
Use BM25 / full-text search for exact match retrieval. Options:
- Elasticsearch / OpenSearch
- Postgres full-text
- Lucene-based systems
Why hybrid?
Semantic search is great for meaning, but docs often contain:
- Product names
- Error codes
- API routes
- Version strings
- Config keys
These are often better handled by lexical search.
Retrieval flow
A practical retrieval pipeline:
- Query arrives
- Run vector search
- Run keyword search
- Merge results
- Deduplicate
- Re-rank top results
- Return final chunks
A reranker improves quality a lot if you can afford the latency.
5) Embedding choices
Choose an embedding model based on:
- Language coverage
- Cost
- Latency
- Context length
- Domain fit
General recommendations:
- Use a strong modern embedding model
- Keep embedding dimension stable if possible
- Test on your own docs with sample queries
Important
Use the same embedding model for all chunks in a given index version.
If you change models later, treat it as a new index version or re-embed everything.
6) Re-embedding strategy over time
This is where many systems get messy. The key is to make indexing incremental.
Store a content hash
For each chunk, store a hash of the canonical text:
- If hash unchanged, skip re-embedding
- If hash changed, re-embed that chunk
Version your documents
Track:
doc_versionchunk_versionembedding_model_versionindex_version
This lets you know what needs rebuilding.
Re-embed only what changed
On update:
- Detect changed source docs
- Re-parse and re-chunk
- Compare new chunks to old via hashes
- Re-embed only new/changed chunks
- Delete obsolete chunks from indexes
- Insert new chunks
When to do a full re-embed
Do a full rebuild if:
- You switch embedding models
- You change chunking strategy significantly
- You alter normalization substantially
- Your corpus undergoes large-scale structural changes
7) Handling doc updates safely
Docs change in real systems. Use an ingestion pipeline that supports:
Upserts
Each chunk gets a stable ID, such as:
{doc_id}:{section_id}:{chunk_index}:{content_hash_prefix}
Or:
- stable logical ID + version field
Deletes
If a doc is removed or a section disappears:
- Mark old chunks as deleted
- Remove them from search indexes
- Keep tombstones if you need auditability
Rechunking
If a section changes enough that chunk boundaries shift:
- Treat resulting chunks as new
- Delete obsolete chunk IDs
- Don’t try to force old embeddings to stay valid
8) Practical chunking algorithm
A good default approach:
- Parse document into a tree:
- title
- headings
- paragraphs
- lists
- tables
- Create chunks by walking the tree top-down
- Keep chunks near target token size
- Merge small siblings
- Split large sections on paragraph boundaries
- Add overlap only when necessary
Example heuristic
- Target: 500 tokens
- Min: 200 tokens
- Max: 800 tokens
- Overlap: 50–100 tokens
For lists or steps:
- Keep the whole list if it fits
- Otherwise split by logical subgroups, not arbitrary token count
9) Metadata filtering
Metadata is crucial for relevance and precision.
Useful filters:
- product
- doc type
- version
- region
- language
- audience
- access scope
- release date
Examples:
- “Only search docs for Product A”
- “Prefer latest version”
- “Exclude deprecated docs”
- “Restrict to user’s permission scope”
This prevents irrelevant retrieval from unrelated product lines.
10) Ranking and answer quality
Even with good chunking, raw vector search can return noisy matches.
Improve with:
- Hybrid search
- Metadata filters
- Re-ranking
- Query rewriting for known aliases
- Synonym dictionaries
Good retrieval behavior
Return:
- The best chunk
- Its parent section title
- A nearby context chunk if needed
- Source links
This helps users trust results.
11) Observability and evaluation
You should measure retrieval quality continuously.
Track:
- Click-through rate
- Query success rate
- Zero-result rate
- Top-k relevance
- Latency
- Freshness lag
- Re-embedding backlog
Build a test set
Create a small gold set of:
- Queries
- Expected relevant docs/chunks
Examples:
- “How do I reset API keys?”
- “What does error 429 mean?”
- “How to configure SSO for org accounts?”
- “Which endpoint returns billing usage?”
Use this set to compare chunking and retrieval strategies.
12) Recommended operational workflow
Initial build
- Crawl docs
- Normalize content
- Chunk by structure
- Generate hashes
- Embed
- Index in vector + lexical store
- Validate with test queries
Daily incremental updates
- Fetch changed docs
- Re-parse/re-chunk
- Compare hashes
- Re-embed changed chunks only
- Upsert indexes
- Remove deleted chunks
Periodic maintenance
- Rebuild if model changes
- Re-evaluate chunk sizes
- Remove stale/deprecated docs
- Tune hybrid weights and reranker settings
13) A sensible default configuration
If you want a practical starting point:
- Chunk by heading/section first
- Target chunk size: 500 tokens
- Overlap: 75 tokens
- Store full metadata and content hash
- Use hybrid retrieval
- Re-rank top 20–50 results
- Re-embed only changed chunks
- Full reindex only when model/chunking strategy changes
14) Common mistakes to avoid
- Chunking by fixed size only, ignoring structure
- Making chunks too large
- Losing headings/context in chunks
- Using vector search without keyword fallback
- Re-embedding everything on every update
- Not tracking content hashes
- Not versioning your embedding model
- Ignoring permissions/access control
- Failing to evaluate with real queries
15) Simple implementation blueprint
A clean data model:
Chunk record
chunk_iddoc_idtitlesection_pathtextembeddingcontent_hashsource_urlversionupdated_atmetadata
Pipeline
parse_doc()chunk_doc()hash_chunk()embed_chunk()upsert_vector_index()upsert_keyword_index()delete_removed_chunks()
If you want, I can also give you:
- a reference architecture diagram,
- a Python pseudocode implementation for chunking + incremental reindexing, or
- a recommended stack for Postgres/pgvector vs Pinecone/OpenSearch.