Prompt

How do I generate embeddings from documents and search them by meaning?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To search documents by meaning rather than exact keywords, you typically use embeddings + vector search.

Basic idea

  1. Split documents into chunks
    Large documents are broken into smaller pieces (paragraphs, sections, sentences).

  2. Create embeddings for each chunk
    An embedding model converts each chunk into a numeric vector that captures semantic meaning.

  3. Store vectors in a vector database / index
    Along with metadata like document ID, title, source, etc.

  4. Embed the user query
    Convert the search query into a vector using the same embedding model.

  5. Find nearest vectors
    Use similarity search (cosine similarity, dot product, or Euclidean distance) to return the most semantically similar chunks.

  6. Optionally rerank results
    A reranker can improve relevance by considering the query and candidate chunks together.


Typical workflow

1) Chunk your documents

Example:

  • 500–1,000 tokens per chunk
  • 10–20% overlap between chunks

This helps retrieval because embeddings work best on focused text segments.

2) Generate embeddings

Use an embedding model such as:

  • OpenAI embeddings
  • Sentence Transformers
  • Cohere embeddings
  • Azure OpenAI embeddings

Each chunk becomes a vector like:

[0.12, -0.44, 0.88, ...]

3) Store in a vector database

Popular options:

  • FAISS
  • Chroma
  • Pinecone
  • Weaviate
  • Milvus
  • pgvector

Store:

  • vector
  • chunk text
  • metadata

4) Search by query

Embed the query:

“How do I reset my password?”

Then retrieve chunks close in vector space, even if they don’t contain the exact words “reset password” but say things like:

  • “change account credentials”
  • “recover login access”
  • “set a new password”

Minimal Python example

Using sentence-transformers + faiss:

from sentence_transformers import SentenceTransformer
import faiss
import numpy as np

# Example chunks
chunks = [
    "Our refund policy allows returns within 30 days.",
    "To reset your password, click 'Forgot Password' on the login page.",
    "Shipping usually takes 3 to 5 business days."
]

# Load embedding model
model = SentenceTransformer("all-MiniLM-L6-v2")

# Create embeddings
embeddings = model.encode(chunks, normalize_embeddings=True)

# Build vector index
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim)  # inner product works well with normalized vectors
index.add(np.array(embeddings, dtype=np.float32))

# Query
query = "How do I change my password?"
query_embedding = model.encode([query], normalize_embeddings=True)

# Search
scores, indices = index.search(np.array(query_embedding, dtype=np.float32), k=2)

for score, idx in zip(scores[0], indices[0]):
    print(score, chunks[idx])

Tips for better semantic search

Chunking

  • Don’t embed huge documents as one block
  • Use chunking with overlap
  • Keep related text together

Use the same embedding model

  • Document embeddings and query embeddings must come from the same model

Normalize if using cosine similarity

  • Common practice for semantic search

Add metadata filters

You can filter by:

  • document type
  • date
  • author
  • category

Rerank top results

If you need high precision, use a reranker after vector search.


When embeddings are a good fit

Use them when you want:

  • semantic search
  • FAQ retrieval
  • document Q&A
  • recommendation by meaning
  • duplicate or similar text detection

Common architecture for RAG

Embeddings are often used in retrieval-augmented generation:

  1. User asks a question
  2. Retrieve relevant chunks with embeddings
  3. Send chunks to an LLM to generate an answer

If you want, I can also give you:

  1. a complete working example in Python,
  2. an example using OpenAI embeddings, or
  3. a design for search over PDFs / documents at scale.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.