Prompt

Can I do vector search with only embeddings and no vector DB?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Yes. You can absolutely do vector search with just embeddings and no vector database.

How it works

  1. Generate embeddings for your documents/items.
  2. Store them somewhere simple, like:
    • a JSON/CSV file
    • a database table
    • a Python list / NumPy array
  3. At query time:
    • embed the user query
    • compute similarity against all stored embeddings
    • return the top matches

Common similarity methods

  • Cosine similarity — most common for text embeddings
  • Dot product — often equivalent if vectors are normalized
  • Euclidean distance — sometimes used, less common for text search

Example in Python

import numpy as np
from sklearn.metrics.pairwise import cosine_similarity

docs = ["cat on a mat", "deep learning basics", "best pizza recipes"]
doc_embeddings = np.array([
    [0.1, 0.2, 0.3],
    [0.9, 0.8, 0.7],
    [0.2, 0.1, 0.4]
])

query_embedding = np.array([[0.15, 0.18, 0.28]])

scores = cosine_similarity(query_embedding, doc_embeddings)[0]
top_idx = np.argsort(scores)[::-1]

for i in top_idx:
    print(docs[i], scores[i])

When this is a good idea

  • Small to medium datasets
  • Prototyping
  • Simple apps
  • When you want to avoid extra infrastructure

When a vector DB helps

A vector DB becomes useful when you need:

  • very large scale
  • fast approximate nearest-neighbor search
  • filtering/metadata queries
  • persistence and replication
  • managed indexing
  • low-latency production workloads

Rule of thumb

  • Up to tens of thousands of vectors: plain arrays/files/SQL are often fine
  • Hundreds of thousands to millions: a vector DB or ANN index is usually better
  • Very large or high-traffic systems: definitely consider a vector search engine

If you want, I can show you:

  • a pure Python example,
  • a SQL-based approach,
  • or how to do it efficiently with NumPy / FAISS without a vector DB.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.