Prompt
What's the most cost-effective way to run semantic search using an LLM API provider at scale?
Latest observation
The most cost-effective pattern is usually:
1) Don’t use an LLM for retrieval if you don’t have to
For semantic search at scale, the cheapest architecture is typically:
- Generate embeddings once for your corpus
- Store them in a vector database / ANN index
- At query time, embed the user query
- Do a vector similarity search
- Optionally use an LLM only for reranking, summarization, or answer generation
This is much cheaper than calling a chat/completion model over large chunks of content for every search.
2) Use a small, cheap embedding model
Most providers offer embedding models that are far less expensive than generation models. For scale, you want:
- Low-cost embedding model
- Batching for corpus ingestion
- Caching for repeated queries
- Short, normalized text chunks to reduce token usage
If your provider offers multiple embedding tiers, choose the smallest one that meets quality requirements. In many real-world cases, a “small” embedding model is enough.
3) Keep the search pipeline two-stage
A cost-effective setup is:
Stage A: Cheap candidate retrieval
- Vector search over embeddings
- Return top 20–100 candidates
Stage B: Optional LLM rerank or answer
- Use a stronger model only on the small candidate set
- Or skip LLM entirely if you only need results
This keeps the expensive model usage tightly bounded.
4) Minimize token costs aggressively
If you use an LLM in the loop, costs are mostly driven by tokens. Reduce them by:
- Chunking documents into small, semantically coherent pieces
- Removing boilerplate, HTML, nav text, duplicates
- Storing metadata separately from content
- Sending only the top-ranked snippets
- Truncating long context before LLM calls
5) Cache everything you can
At scale, caching matters a lot:
- Query embedding cache for repeated or near-repeated queries
- Result cache for common searches
- Rerank cache if the same candidate set recurs
- Document embedding cache during reindexing
Even modest cache hit rates can cut costs significantly.
6) Use hybrid retrieval if relevance matters
A purely semantic approach can be more expensive if you need many candidates. Often best practice is:
- BM25 / keyword search + vector search
- Merge results
- Rerank top hits
This improves recall and may let you retrieve fewer items with the LLM.
7) Optimize index and serving costs too
Provider API cost is only part of total cost. You should also consider:
- Vector DB pricing
- Storage footprint
- Query latency
- Reindexing costs
- Throughput limits / rate limits
Sometimes a slightly more expensive embedding model is worth it if it reduces downstream reranking or improves precision enough to save money overall.
8) A practical “best cost/performance” setup
If you want the usual sweet spot:
- Chunk documents into 200–800 tokens
- Generate embeddings once with a cheap embedding model
- Store in a vector index
- At query time:
- embed the query
- retrieve top 20–50 results
- optionally rerank top 10 with a small LLM
- Only call a larger LLM when you need a natural-language answer
9) If you’re doing huge volume
For very large scale, the best cost-saving levers are:
- Use the cheapest adequate embedding model
- Batch embedding requests
- Avoid LLM reranking unless necessary
- Cache query embeddings and results
- Use approximate nearest-neighbor search
- Precompute and deduplicate aggressively
- Use tiered retrieval: cheap first pass, expensive second pass only when needed
Bottom line
Most cost-effective approach:
Embedding-based vector search + cheap embeddings + caching + optional small reranker
—not direct LLM calls for each search.
If you want, I can also give you:
- a cost-optimized architecture diagram,
- a provider-by-provider comparison, or
- a rough cost model for your expected query/document volume.
Brands