Prompt

Why is my vector search slow after adding more embeddings?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Vector search usually gets slower as you add more embeddings because the system has more work to do per query. Common reasons:

  1. More vectors to scan or compare

    • If your index is doing brute-force or partially brute-force search, query time grows with the number of embeddings.
  2. Index type doesn’t scale well

    • Some ANN methods get slower or less accurate as the dataset grows if they’re not tuned.
    • Examples:
      • HNSW: may need higher efSearch / better parameters.
      • IVF / PQ: may need more clusters, retraining, or tuning nprobe.
      • Flat index: exact search slows linearly with data size.
  3. Memory pressure / cache misses

    • As embeddings increase, the index may no longer fit in RAM or CPU caches, causing disk access or poor locality.
  4. Too many dimensions or unnormalized vectors

    • High-dimensional vectors increase compute cost.
    • If your similarity metric or preprocessing is off, the search may do extra work or return poor candidates, making tuning ineffective.
  5. Metadata filtering overhead

    • If you’re applying filters after or during vector search, larger datasets can make candidate filtering more expensive.
  6. Index fragmentation or stale indexes

    • Incremental inserts can degrade performance if the index wasn’t rebuilt or compacted.
  7. Concurrency / load

    • More embeddings may coincide with more users, making the slowdown look like it’s from the index when it’s actually contention.

What to check

  • What index type are you using?
  • Is the search exact or approximate?
  • Are the vectors stored in RAM?
  • Are your ANN parameters tuned for the new size?
  • Did query latency rise because of more data, or because of more traffic?
  • Are you filtering heavily on metadata?

Typical fixes

  • Use an ANN index if you’re using exact search.
  • Tune parameters like efSearch, nprobe, top_k, and shard count.
  • Rebuild/retrain the index after major data growth.
  • Keep the index memory-resident if possible.
  • Normalize vectors if your similarity metric expects it.
  • Pre-filter candidates where possible.

If you tell me what database/index you’re using (FAISS, Pinecone, Milvus, pgvector, Elasticsearch, etc.) and roughly how many embeddings you added, I can suggest specific tuning steps.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.