Prompt

I’m trying to build semantic search over product docs and need a practical plan for chunking, indexing, and re-embedding over time

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Here’s a practical, production-oriented plan for building semantic search over product docs, with a focus on chunking, indexing, and re-embedding over time.


1) Overall architecture

A good setup is:

  1. Ingest docs from your sources
  2. Normalize and structure them into sections
  3. Chunk the text into retrievable pieces
  4. Embed each chunk
  5. Index chunks in a vector store + keyword index
  6. Search with hybrid retrieval
  7. Re-rank results if needed
  8. Re-embed only what changed as docs evolve

A strong default is hybrid search:

  • Vector search for semantic similarity
  • BM25 / keyword search for exact matches, product names, error codes, API paths, etc.

2) Chunking strategy

Chunking is the most important design choice for doc search.

Goals for chunks

Each chunk should be:

  • Self-contained enough to understand
  • Small enough to retrieve precisely
  • Large enough to preserve context

Recommended chunk size

A good starting point:

  • 300–800 tokens per chunk
  • 10–20% overlap between adjacent chunks

If your docs are highly structured, use section-based chunks rather than arbitrary fixed windows.

Prefer structure-aware chunking

Split by:

  • Headings
  • Subheadings
  • Lists
  • Tables
  • Code blocks
  • FAQ entries
  • Steps/procedures

This usually beats naive sliding windows.

Practical rules

  • Keep one topic per chunk
  • Don’t split in the middle of a code sample or numbered procedure if possible
  • If a section is too long, subdivide it by paragraph or subheading
  • If a section is too short, merge with adjacent content until it’s useful

Add metadata to every chunk

Store:

  • doc_id
  • doc_title
  • section_heading
  • subheading
  • chunk_index
  • source_url
  • product
  • version
  • last_updated
  • content_hash
  • optionally permissions/ACL

This helps filtering, deduping, and reindexing later.


3) Chunk types by doc format

Product docs / knowledge base articles

Use section-based chunks:

  • Title + intro
  • Each major heading as a chunk
  • Bulleted steps as their own chunk if they are a coherent procedure

API docs

Chunk by:

  • Endpoint
  • Parameter table
  • Request example
  • Response example
  • Errors / edge cases

For APIs, exact keyword search matters a lot, so hybrid retrieval is especially useful.

Troubleshooting docs

Chunk by:

  • Symptom
  • Cause
  • Resolution
  • Related errors

This makes retrieval cleaner than a large “troubleshooting” blob.

Long guides

Use chapter/section boundaries first, then subchunk within sections.


4) Indexing design

Use two indexes:

A. Vector index

Stores embeddings for semantic retrieval.

Common options:

  • Pinecone
  • Weaviate
  • Milvus
  • Qdrant
  • pgvector

B. Lexical index

Use BM25 / full-text search for exact match retrieval. Options:

  • Elasticsearch / OpenSearch
  • Postgres full-text
  • Lucene-based systems

Why hybrid?

Semantic search is great for meaning, but docs often contain:

  • Product names
  • Error codes
  • API routes
  • Version strings
  • Config keys

These are often better handled by lexical search.

Retrieval flow

A practical retrieval pipeline:

  1. Query arrives
  2. Run vector search
  3. Run keyword search
  4. Merge results
  5. Deduplicate
  6. Re-rank top results
  7. Return final chunks

A reranker improves quality a lot if you can afford the latency.


5) Embedding choices

Choose an embedding model based on:

  • Language coverage
  • Cost
  • Latency
  • Context length
  • Domain fit

General recommendations:

  • Use a strong modern embedding model
  • Keep embedding dimension stable if possible
  • Test on your own docs with sample queries

Important

Use the same embedding model for all chunks in a given index version.
If you change models later, treat it as a new index version or re-embed everything.


6) Re-embedding strategy over time

This is where many systems get messy. The key is to make indexing incremental.

Store a content hash

For each chunk, store a hash of the canonical text:

  • If hash unchanged, skip re-embedding
  • If hash changed, re-embed that chunk

Version your documents

Track:

  • doc_version
  • chunk_version
  • embedding_model_version
  • index_version

This lets you know what needs rebuilding.

Re-embed only what changed

On update:

  1. Detect changed source docs
  2. Re-parse and re-chunk
  3. Compare new chunks to old via hashes
  4. Re-embed only new/changed chunks
  5. Delete obsolete chunks from indexes
  6. Insert new chunks

When to do a full re-embed

Do a full rebuild if:

  • You switch embedding models
  • You change chunking strategy significantly
  • You alter normalization substantially
  • Your corpus undergoes large-scale structural changes

7) Handling doc updates safely

Docs change in real systems. Use an ingestion pipeline that supports:

Upserts

Each chunk gets a stable ID, such as: {doc_id}:{section_id}:{chunk_index}:{content_hash_prefix}

Or:

  • stable logical ID + version field

Deletes

If a doc is removed or a section disappears:

  • Mark old chunks as deleted
  • Remove them from search indexes
  • Keep tombstones if you need auditability

Rechunking

If a section changes enough that chunk boundaries shift:

  • Treat resulting chunks as new
  • Delete obsolete chunk IDs
  • Don’t try to force old embeddings to stay valid

8) Practical chunking algorithm

A good default approach:

  1. Parse document into a tree:
    • title
    • headings
    • paragraphs
    • lists
    • tables
  2. Create chunks by walking the tree top-down
  3. Keep chunks near target token size
  4. Merge small siblings
  5. Split large sections on paragraph boundaries
  6. Add overlap only when necessary

Example heuristic

  • Target: 500 tokens
  • Min: 200 tokens
  • Max: 800 tokens
  • Overlap: 50–100 tokens

For lists or steps:

  • Keep the whole list if it fits
  • Otherwise split by logical subgroups, not arbitrary token count

9) Metadata filtering

Metadata is crucial for relevance and precision.

Useful filters:

  • product
  • doc type
  • version
  • region
  • language
  • audience
  • access scope
  • release date

Examples:

  • “Only search docs for Product A”
  • “Prefer latest version”
  • “Exclude deprecated docs”
  • “Restrict to user’s permission scope”

This prevents irrelevant retrieval from unrelated product lines.


10) Ranking and answer quality

Even with good chunking, raw vector search can return noisy matches.

Improve with:

  • Hybrid search
  • Metadata filters
  • Re-ranking
  • Query rewriting for known aliases
  • Synonym dictionaries

Good retrieval behavior

Return:

  • The best chunk
  • Its parent section title
  • A nearby context chunk if needed
  • Source links

This helps users trust results.


11) Observability and evaluation

You should measure retrieval quality continuously.

Track:

  • Click-through rate
  • Query success rate
  • Zero-result rate
  • Top-k relevance
  • Latency
  • Freshness lag
  • Re-embedding backlog

Build a test set

Create a small gold set of:

  • Queries
  • Expected relevant docs/chunks

Examples:

  • “How do I reset API keys?”
  • “What does error 429 mean?”
  • “How to configure SSO for org accounts?”
  • “Which endpoint returns billing usage?”

Use this set to compare chunking and retrieval strategies.


12) Recommended operational workflow

Initial build

  1. Crawl docs
  2. Normalize content
  3. Chunk by structure
  4. Generate hashes
  5. Embed
  6. Index in vector + lexical store
  7. Validate with test queries

Daily incremental updates

  1. Fetch changed docs
  2. Re-parse/re-chunk
  3. Compare hashes
  4. Re-embed changed chunks only
  5. Upsert indexes
  6. Remove deleted chunks

Periodic maintenance

  • Rebuild if model changes
  • Re-evaluate chunk sizes
  • Remove stale/deprecated docs
  • Tune hybrid weights and reranker settings

13) A sensible default configuration

If you want a practical starting point:

  • Chunk by heading/section first
  • Target chunk size: 500 tokens
  • Overlap: 75 tokens
  • Store full metadata and content hash
  • Use hybrid retrieval
  • Re-rank top 20–50 results
  • Re-embed only changed chunks
  • Full reindex only when model/chunking strategy changes

14) Common mistakes to avoid

  • Chunking by fixed size only, ignoring structure
  • Making chunks too large
  • Losing headings/context in chunks
  • Using vector search without keyword fallback
  • Re-embedding everything on every update
  • Not tracking content hashes
  • Not versioning your embedding model
  • Ignoring permissions/access control
  • Failing to evaluate with real queries

15) Simple implementation blueprint

A clean data model:

Chunk record

  • chunk_id
  • doc_id
  • title
  • section_path
  • text
  • embedding
  • content_hash
  • source_url
  • version
  • updated_at
  • metadata

Pipeline

  • parse_doc()
  • chunk_doc()
  • hash_chunk()
  • embed_chunk()
  • upsert_vector_index()
  • upsert_keyword_index()
  • delete_removed_chunks()

If you want, I can also give you:

  1. a reference architecture diagram,
  2. a Python pseudocode implementation for chunking + incremental reindexing, or
  3. a recommended stack for Postgres/pgvector vs Pinecone/OpenSearch.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.