Prompt
How do I compare recall across different vector indexes?
Latest observation
To compare recall across different vector indexes, you want to measure how often each index returns the true nearest neighbors (or a trusted approximation of them) for the same set of queries.
Basic approach
-
Choose a ground truth
- Ideally, use an exact brute-force search over the full dataset.
- For each query, record the top-k true nearest neighbors.
-
Run the same queries on each vector index
- Use the same embeddings, distance metric, and top-k value.
-
Compute recall@k
- For each query: [ \text{recall@k} = \frac{|\text{returned} \cap \text{ground truth}|}{k} ]
- Then average across queries.
Example
If the true top-10 neighbors are:
- Ground truth:
{A, B, C, D, E, F, G, H, I, J} - Index returns:
{A, B, X, Y, Z, F, G, H, K, L}
Intersection = {A, B, F, G, H} = 5 items
So:
[ \text{recall@10} = 5/10 = 0.5 ]
Important considerations
1. Use the same metric
Make sure all indexes are evaluated with the same similarity measure:
- cosine
- dot product
- Euclidean / L2
2. Compare at the same k
If one index returns top-10 and another top-20, normalize using the same recall@k definition.
3. Use enough queries
A small query set can be misleading. Use a representative sample of queries from your workload.
4. Consider latency and memory too
Recall alone doesn’t tell the whole story. In practice, compare:
- recall@k
- latency
- index size / memory
- build time
- throughput
5. Tune each index fairly
If you compare ANN indexes, evaluate them across a similar latency budget or parameter sweep, since recall depends on search parameters:
- HNSW:
efSearch - IVF:
nprobe - PQ / SQ settings
- tree depth / candidate counts
Recommended benchmarking workflow
For each index:
- build the index on the same dataset
- run the same query set
- measure recall@k against exact ground truth
- plot recall vs latency
This gives you a clear accuracy/performance tradeoff curve.
If you want, I can also show you:
- a Python example for computing recall@k, or
- a benchmark template for comparing FAISS / HNSW / Milvus / Pinecone-style indexes.