Prompt
How do I store embeddings and search them by similarity?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
To store embeddings and search by similarity, you usually do this:
1) Generate embeddings
Use an embedding model to convert each item into a vector, for example:
- text → vector
- image → vector
- audio → vector
Example:
"apple"→[0.12, -0.44, 0.98, ...]
2) Store them with metadata
Save each vector along with:
- an ID
- the original text/item
- optional metadata like timestamp, category, author, etc.
Common storage options:
- Vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma
- Postgres with pgvector
- FAISS for local/in-memory search
- Elasticsearch/OpenSearch with vector support
3) Search using nearest-neighbor similarity
When you have a query:
- embed the query into a vector
- compare it against stored vectors
- return the most similar ones
Typical similarity metrics:
- cosine similarity (very common)
- dot product
- Euclidean distance
Cosine similarity is usually:
[ \text{cosine_sim}(a,b) = \frac{a \cdot b}{||a||,||b||} ]
Higher = more similar.
Simple workflow
Insert
- Text:
"I like apples" - Embedding:
[0.21, 0.77, ...] - Store in vector DB with ID and metadata
Query
- Query text:
"fruit I enjoy" - Embed query
- Retrieve nearest vectors
- Return top-k results
Example with Python + FAISS
import faiss
import numpy as np
# Example embeddings (must all have same dimension)
vectors = np.array([
[0.1, 0.2, 0.3],
[0.9, 0.8, 0.7],
[0.15, 0.25, 0.35]
]).astype("float32")
# If using cosine similarity, normalize vectors first
faiss.normalize_L2(vectors)
index = faiss.IndexFlatIP(3) # inner product on normalized vectors = cosine similarity
index.add(vectors)
query = np.array([[0.12, 0.22, 0.32]], dtype="float32")
faiss.normalize_L2(query)
scores, ids = index.search(query, k=2)
print(scores, ids)
Example with PostgreSQL + pgvector
CREATE TABLE items (
id SERIAL PRIMARY KEY,
content TEXT,
embedding VECTOR(1536)
);
CREATE INDEX ON items USING ivfflat (embedding vector_cosine_ops);
Query similar items:
SELECT id, content
FROM items
ORDER BY embedding <=> '[0.12, 0.22, ...]' -- cosine distance
LIMIT 5;
Best practices
- Use the same embedding model for all items and queries
- Normalize vectors if using cosine similarity
- Store metadata to filter results
- Use approximate nearest-neighbor search for large datasets
- Periodically re-index if embeddings or model change
Which tool should you choose?
- Small/local prototype: FAISS or Chroma
- SQL-based app: PostgreSQL + pgvector
- Large-scale production: Pinecone, Weaviate, Qdrant, Milvus
If you want, I can show you:
- a minimal end-to-end example in Python, or
- the best storage/search setup for your specific use case.