Prompt

I'm building semantic search over product docs and need advice on storage and indexing

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For semantic search over product docs, the main decision is usually where to store the source docs, embeddings, and metadata, and what to index for retrieval. A practical setup looks like this:

Recommended architecture

1) Store raw documents separately

Keep the original docs in durable object storage or a document store:

  • S3 / GCS / Azure Blob
  • PostgreSQL / MongoDB if docs are smaller and structured
  • A search index should not be your only source of truth

Store:

  • document ID
  • title
  • body / chunks
  • version
  • source URL
  • access control info
  • last updated timestamp

2) Chunk before embedding

Don’t embed whole docs unless they’re small. Chunk by:

  • headings / sections
  • paragraphs
  • ~200–500 tokens per chunk as a starting point

Each chunk should carry metadata:

  • doc_id
  • chunk_id
  • section title
  • product/version
  • language
  • permissions/tags

3) Store embeddings in a vector index

Use a vector-capable database or search engine:

  • pgvector if you want simplicity and already use Postgres
  • Pinecone / Weaviate / Milvus / Qdrant for dedicated vector search
  • Elasticsearch / OpenSearch if you want hybrid lexical + vector in one system

Index fields typically:

  • chunk embedding
  • chunk text
  • metadata filters
  • maybe sparse lexical fields for BM25

Indexing strategy

Best practice: hybrid retrieval

Semantic search works better when combined with keyword search:

  • Vector search for meaning
  • BM25/keyword search for exact terms, part numbers, error codes, product names

Typical flow:

  1. Query comes in
  2. Run vector similarity search
  3. Run lexical search
  4. Merge/rerank results

This is especially important for product docs because users often search for:

  • exact feature names
  • version numbers
  • error codes
  • acronyms

Use metadata filters

Index metadata so you can filter by:

  • product line
  • doc type
  • language
  • version
  • access scope

This helps both relevance and security.

Consider reranking

If you need higher precision:

  • retrieve top 20–100 candidates
  • rerank with a cross-encoder or LLM reranker
  • return top 5–10

Storage choices by scale

Small to medium scale

If you’re starting out:

  • PostgreSQL + pgvector
  • separate table for docs/chunks
  • GIN/B-tree indexes on metadata
  • full-text search in Postgres for lexical fallback

Good when:

  • < a few million chunks
  • you want simpler ops
  • you need transactional consistency

Medium to large scale

If search performance is more important:

  • OpenSearch/Elasticsearch for hybrid search
  • or Qdrant/Weaviate/Milvus for vectors + separate keyword layer

Good when:

  • many millions of chunks
  • heavy query traffic
  • multiple filters and facets

Practical schema example

Documents table

  • doc_id
  • source
  • title
  • url
  • product
  • version
  • permissions
  • updated_at

Chunks table

  • chunk_id
  • doc_id
  • chunk_text
  • chunk_order
  • section_title
  • embedding
  • language
  • metadata JSONB

Indexes:

  • vector index on embedding
  • B-tree on doc_id, product, version
  • GIN on metadata JSONB if using Postgres
  • full-text index on chunk_text for keyword search

Operational advice

Re-embed on content changes

Track document versions. When docs change:

  • re-chunk if structure changed
  • re-embed affected chunks
  • delete old chunk vectors

Deduplicate

Product docs often repeat boilerplate. Remove or down-rank:

  • navigation text
  • legal disclaimers
  • repeated headers/footers

Monitor retrieval quality

Track:

  • click-through
  • query success rate
  • zero-result queries
  • top query terms with poor recall

If you want the simplest solid stack

A good default is:

  • PostgreSQL for doc/chunk storage
  • pgvector for embeddings
  • Postgres full-text search for lexical matching
  • metadata filters in JSONB
  • optional reranker later

This is usually enough to get a strong semantic search MVP without too much infrastructure.

If you tell me:

  1. your expected corpus size,
  2. update frequency,
  3. whether you need strict ACLs,
  4. and your preferred stack,

I can recommend a specific schema and indexing plan.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.