Prompt

How do I build semantic search over documents with vector embeddings?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To build semantic search over documents with vector embeddings, you generally do four things:

  1. Split documents into chunks
  2. Convert each chunk into an embedding vector
  3. Store those vectors in a vector index/database
  4. Embed the user query and retrieve the nearest chunks by similarity

1) Chunk your documents

Instead of embedding an entire document at once, split it into smaller pieces like paragraphs or 200–500 token chunks.

Why:

  • Better retrieval granularity
  • Less noise
  • Easier to fit model token limits

Example chunking strategies:

  • fixed size with overlap
  • paragraph-based
  • sentence-window chunks

2) Generate embeddings

Use an embedding model to turn each chunk into a vector, e.g.:

  • OpenAI embeddings
  • sentence-transformers
  • Cohere embeddings
  • local models like bge, e5, all-MiniLM

Each chunk becomes something like:

"Refunds are processed within 5 business days." -> [0.12, -0.03, ...]

The vector captures semantic meaning, so similar ideas land near each other in vector space.

3) Store vectors in a vector database or index

Save:

  • chunk text
  • embedding vector
  • metadata such as doc id, title, section, timestamp, permissions

Common options:

  • FAISS
  • Pinecone
  • Weaviate
  • Milvus
  • Chroma
  • Elasticsearch / OpenSearch vector search
  • PostgreSQL with pgvector

4) Search with a query embedding

When a user searches:

  1. embed the query
  2. find nearest chunk vectors using cosine similarity or dot product
  3. return the top-k chunks
  4. optionally rerank them with a cross-encoder or LLM

Example: Query: "How long does a refund take?" Returns chunks about "Refunds are processed within 5 business days."


Basic architecture

Documents -> Chunking -> Embedding -> Vector Store
User Query -> Embedding -> Similarity Search -> Top Results

Minimal Python example with sentence-transformers + FAISS

from sentence_transformers import SentenceTransformer
import faiss
import numpy as np

# Example chunks
chunks = [
    "Refunds are processed within 5 business days.",
    "You can reset your password from the account settings page.",
    "Shipping takes 3 to 7 business days depending on location."
]

# Load embedding model
model = SentenceTransformer("all-MiniLM-L6-v2")

# Embed chunks
embeddings = model.encode(chunks, convert_to_numpy=True)

# Normalize for cosine similarity
faiss.normalize_L2(embeddings)

# Build FAISS index
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim)  # inner product on normalized vectors = cosine similarity
index.add(embeddings)

# Search
query = "When will I get my refund?"
query_vec = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(query_vec)

k = 2
scores, indices = index.search(query_vec, k)

for score, idx in zip(scores[0], indices[0]):
    print(score, chunks[idx])

Important design tips

1. Use the right embedding model

Choose a model trained for semantic retrieval, not just classification.

Good choices:

  • bge-*
  • e5-*
  • all-MiniLM-*
  • provider embeddings from OpenAI/Cohere/etc.

2. Chunk size matters

Too large:

  • less precise matches

Too small:

  • loses context

A common starting point:

  • 200–400 tokens
  • 10–20% overlap

3. Use metadata filters

If documents are scoped by:

  • user
  • department
  • language
  • date

Filter before or during retrieval to improve relevance and access control.

4. Consider hybrid search

Combine:

  • keyword search (BM25)
  • vector search

This often works better than embeddings alone, especially for:

  • exact terms
  • IDs
  • product codes
  • rare names

5. Rerank results

Vector search is usually a first-pass retrieval. For better quality:

  • retrieve top 20–100
  • rerank with a cross-encoder or LLM
  • return top 3–10

Production pattern

A common production pipeline:

Ingestion

  • parse docs
  • clean text
  • chunk
  • embed
  • store vectors + metadata

Query

  • embed query
  • retrieve top-k
  • filter by metadata/permissions
  • rerank
  • return results with citations

If you want the shortest path

If you’re building a prototype:

  • use sentence-transformers
  • use FAISS or Chroma
  • store chunks + metadata
  • retrieve top-k by cosine similarity

If you’re building production:

  • use a managed vector DB
  • add metadata filters
  • add hybrid search
  • add reranking
  • evaluate recall and relevance

If you want, I can also give you:

  1. a FastAPI + vector DB example,
  2. a RAG pipeline example,
  3. or a production-ready schema for document embeddings.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.