Prompt

How do I generate embeddings from text documents for semantic search?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To generate embeddings from text documents for semantic search, the basic pipeline is:

  1. Split documents into chunks
  2. Convert each chunk into a vector embedding
  3. Store embeddings in a vector database or index
  4. Embed the user query the same way
  5. Find nearest vectors to the query embedding
  6. Return the most similar chunks

1) Chunk your documents

Large documents should usually be split into smaller pieces, such as:

  • paragraphs
  • 200–500 word chunks
  • chunks with overlap

This improves retrieval because semantic search works better on focused text segments.

Example chunking strategy:

  • chunk size: 300–500 tokens
  • overlap: 50–100 tokens

2) Generate embeddings

Use an embedding model to turn each chunk into a numeric vector.

Example with OpenAI embeddings in Python:

from openai import OpenAI

client = OpenAI()

text = "Your document chunk goes here."

response = client.embeddings.create(
    model="text-embedding-3-small",
    input=text
)

embedding = response.data[0].embedding
print(len(embedding))  # vector dimension

If you have many chunks, embed them in batches for efficiency.

3) Store embeddings

Store:

  • the vector embedding
  • the original text chunk
  • metadata like document ID, page number, title, etc.

Common storage options:

  • FAISS for local/in-memory search
  • Pinecone
  • Weaviate
  • Milvus
  • pgvector in PostgreSQL
  • Elasticsearch/OpenSearch with vector support

4) Embed the query

When a user searches, embed the query using the same model:

query = "How do I reset my password?"

query_embedding = client.embeddings.create(
    model="text-embedding-3-small",
    input=query
).data[0].embedding

5) Compute similarity

Compare the query embedding to stored embeddings using:

  • cosine similarity
  • dot product
  • Euclidean distance

Cosine similarity is common for semantic search.

6) Retrieve top matches

Return the top-k most similar chunks, optionally with reranking.


Minimal example using FAISS

import faiss
import numpy as np
from openai import OpenAI

client = OpenAI()

docs = [
    "Cats are small domesticated mammals.",
    "Dogs are loyal and often kept as pets.",
    "Embeddings convert text into vectors for search."
]

# Generate embeddings
embeddings = []
for doc in docs:
    emb = client.embeddings.create(
        model="text-embedding-3-small",
        input=doc
    ).data[0].embedding
    embeddings.append(emb)

vectors = np.array(embeddings, dtype="float32")

# Build FAISS index
dimension = vectors.shape[1]
index = faiss.IndexFlatIP(dimension)  # inner product
faiss.normalize_L2(vectors)            # normalize for cosine similarity
index.add(vectors)

# Search
query = "How do text vectors help search?"
q_emb = client.embeddings.create(
    model="text-embedding-3-small",
    input=query
).data[0].embedding

q_vec = np.array([q_emb], dtype="float32")
faiss.normalize_L2(q_vec)

scores, ids = index.search(q_vec, k=2)

print("Top matches:")
for score, idx in zip(scores[0], ids[0]):
    print(score, docs[idx])

Best practices

  • Use a modern embedding model
  • Chunk documents before embedding
  • Normalize text formatting
  • Store metadata for filtering
  • Use batching for scale
  • Consider reranking for better final results
  • Evaluate search quality with real queries

Simple architecture

Documents → chunking → embeddings → vector DB
User query → embedding → vector similarity search → results

If you want, I can also give you:

  • a full production-style Python example
  • an example using Pinecone / pgvector / FAISS
  • a guide on choosing chunk sizes
  • help designing a semantic search pipeline for your specific data

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.