Prompt

How do I store embeddings and search them by similarity?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To store embeddings and search by similarity, you usually do this:

1) Generate embeddings

Use an embedding model to convert each item into a vector, for example:

  • text → vector
  • image → vector
  • audio → vector

Example:

  • "apple" → [0.12, -0.44, 0.98, ...]

2) Store them with metadata

Save each vector along with:

  • an ID
  • the original text/item
  • optional metadata like timestamp, category, author, etc.

Common storage options:

  • Vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma
  • Postgres with pgvector
  • FAISS for local/in-memory search
  • Elasticsearch/OpenSearch with vector support

3) Search using nearest-neighbor similarity

When you have a query:

  1. embed the query into a vector
  2. compare it against stored vectors
  3. return the most similar ones

Typical similarity metrics:

  • cosine similarity (very common)
  • dot product
  • Euclidean distance

Cosine similarity is usually:

[ \text{cosine_sim}(a,b) = \frac{a \cdot b}{||a||,||b||} ]

Higher = more similar.


Simple workflow

Insert

  • Text: "I like apples"
  • Embedding: [0.21, 0.77, ...]
  • Store in vector DB with ID and metadata

Query

  • Query text: "fruit I enjoy"
  • Embed query
  • Retrieve nearest vectors
  • Return top-k results

Example with Python + FAISS

import faiss
import numpy as np

# Example embeddings (must all have same dimension)
vectors = np.array([
    [0.1, 0.2, 0.3],
    [0.9, 0.8, 0.7],
    [0.15, 0.25, 0.35]
]).astype("float32")

# If using cosine similarity, normalize vectors first
faiss.normalize_L2(vectors)

index = faiss.IndexFlatIP(3)  # inner product on normalized vectors = cosine similarity
index.add(vectors)

query = np.array([[0.12, 0.22, 0.32]], dtype="float32")
faiss.normalize_L2(query)

scores, ids = index.search(query, k=2)
print(scores, ids)

Example with PostgreSQL + pgvector

CREATE TABLE items (
    id SERIAL PRIMARY KEY,
    content TEXT,
    embedding VECTOR(1536)
);

CREATE INDEX ON items USING ivfflat (embedding vector_cosine_ops);

Query similar items:

SELECT id, content
FROM items
ORDER BY embedding <=> '[0.12, 0.22, ...]'  -- cosine distance
LIMIT 5;

Best practices

  • Use the same embedding model for all items and queries
  • Normalize vectors if using cosine similarity
  • Store metadata to filter results
  • Use approximate nearest-neighbor search for large datasets
  • Periodically re-index if embeddings or model change

Which tool should you choose?

  • Small/local prototype: FAISS or Chroma
  • SQL-based app: PostgreSQL + pgvector
  • Large-scale production: Pinecone, Weaviate, Qdrant, Milvus

If you want, I can show you:

  1. a minimal end-to-end example in Python, or
  2. the best storage/search setup for your specific use case.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.