Prompt

How do I migrate embeddings from Elasticsearch to a vector database?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

Migrating embeddings from Elasticsearch to a dedicated vector database is usually straightforward if you separate it into data extraction, schema mapping, bulk loading, and validation.

1) Inventory what you have in Elasticsearch

First, identify:

  • Index name(s)
  • Document ID
  • Embedding field name
  • Vector dimension
  • Similarity metric used for search, if applicable
  • Any metadata fields you want to keep

Example ES document:

{
  "_id": "doc-123",
  "_source": {
    "title": "Example",
    "body": "Some text",
    "embedding": [0.12, -0.03, 0.88, ...],
    "category": "docs"
  }
}

2) Map Elasticsearch fields to the vector DB schema

Typical mapping:

  • ES _id → vector DB point/vector ID
  • ES embedding array → vector field
  • ES metadata fields → payload/properties/metadata
  • ES index → collection/index/table in the vector DB

Example:

  • Pinecone: namespace + id + vector + metadata
  • Qdrant: collection + point id + vector + payload
  • Weaviate: class/collection + uuid + vector + properties
  • Milvus: collection + primary key + vector + scalar fields
  • pgvector: table row with embedding column + metadata columns

3) Export documents from Elasticsearch

Use one of:

  • Scroll / search_after
  • Elasticsearch reindex to an intermediate store
  • Bulk export via Python client

For large indexes, prefer search_after or scrolling in batches.

Example using Python:

from elasticsearch import Elasticsearch

es = Elasticsearch("http://localhost:9200")

resp = es.search(
    index="my-index",
    body={
        "size": 1000,
        "query": {"match_all": {}},
        "_source": ["title", "body", "embedding", "category"]
    },
    scroll="2m"
)

scroll_id = resp["_scroll_id"]
hits = resp["hits"]["hits"]

4) Transform documents into vector DB format

For each doc:

  • Extract the vector
  • Ensure it’s the right type (float32 commonly preferred)
  • Ensure the vector has the expected dimension
  • Convert metadata into the target DB’s payload format
  • Preserve the ID if possible

Important checks:

  • No missing vectors
  • No mismatched dimensions
  • No NaNs/inf values
  • Similarity metric compatibility:
    • Elasticsearch may store vectors for cosine, dot product, or L2
    • Your destination DB should use the same or equivalent metric

5) Bulk insert into the vector database

Do this in batches for speed and reliability.

Example pseudocode:

batch = []
for hit in hits:
    src = hit["_source"]
    vector = src["embedding"]
    payload = {
        "title": src.get("title"),
        "body": src.get("body"),
        "category": src.get("category")
    }
    batch.append({
        "id": hit["_id"],
        "vector": vector,
        "payload": payload
    })

# send batch to vector DB

Most vector DBs support upserts:

  • Use upsert if you may rerun the migration
  • Use insert only if the target is empty and IDs won’t conflict

6) Validate the migration

Compare source and target:

  • Count of documents
  • Random sample of IDs
  • Vector dimensions
  • Metadata integrity
  • Search results on a few known queries

Validation examples:

  • Check that a specific ID exists in the new DB
  • Run a nearest-neighbor query in both systems and compare top-k overlap
  • Confirm cosine scores behave similarly

7) Rebuild your application layer

Update your app to:

  • Query the vector DB instead of Elasticsearch
  • Store non-vector text/metadata where appropriate
  • Keep Elasticsearch only if you still need full-text search or hybrid retrieval

Often the best architecture is:

  • Elasticsearch/OpenSearch for keyword/full-text search
  • Vector DB for semantic search
  • Optional hybrid search orchestration in the app

Example migration script outline

from elasticsearch import Elasticsearch
# import your vector DB client

es = Elasticsearch("http://localhost:9200")

def transform(hit):
    src = hit["_source"]
    return {
        "id": hit["_id"],
        "vector": src["embedding"],
        "payload": {
            "title": src.get("title"),
            "body": src.get("body"),
            "category": src.get("category")
        }
    }

scroll = es.search(
    index="my-index",
    body={"query": {"match_all": {}}, "_source": True, "size": 1000},
    scroll="2m"
)

scroll_id = scroll["_scroll_id"]
while True:
    hits = scroll["hits"]["hits"]
    if not hits:
        break

    batch = [transform(h) for h in hits]
    # vector_db.upsert(batch)

    scroll = es.scroll(scroll_id=scroll_id, scroll="2m")
    scroll_id = scroll["_scroll_id"]

Common pitfalls

  • Wrong vector dimension in the destination schema
  • Non-normalized vectors when the metric assumes normalization
  • Embedding field stored as strings instead of numbers
  • Batch size too large causing timeouts/memory pressure
  • Missing metadata that your app relies on
  • Different distance metric behavior causing ranking changes

If you want zero-downtime migration

  1. Start dual-writing new embeddings to both systems
  2. Backfill old data into the vector DB
  3. Validate search parity
  4. Switch reads to the vector DB
  5. Keep Elasticsearch as fallback until stable

If you tell me which vector database you’re migrating to (Pinecone, Qdrant, Milvus, Weaviate, pgvector, etc.), I can give you a concrete migration script and schema mapping.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.