Prompt
Google Vertex AI embeddings cost too much for batch jobs
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If Vertex AI embeddings are too expensive for batch jobs, the fastest way to cut cost is usually to reduce the number of calls and the amount of text you send, or to switch to a cheaper embedding path/model.
Cost-cutting options
1) Chunk smarter
- Don’t embed every tiny record separately.
- Deduplicate repeated text before embedding.
- Merge very short adjacent segments into larger chunks.
- Remove boilerplate, headers, signatures, HTML, etc.
2) Use a cheaper embedding model
If you’re using a large/high-quality model, try a smaller one first.
- Compare retrieval quality on a sample set.
- For many batch workloads, a smaller model is “good enough.”
3) Batch requests efficiently
- Send the maximum practical batch size supported by the API.
- Group texts by similar length to reduce padding/overhead.
- Avoid retrying whole batches when only one item fails.
4) Cache embeddings
- If the same content appears again, reuse the stored vector.
- Hash the normalized input text and use that as a cache key.
- This is often the biggest savings for recurring batch jobs.
5) Pre-filter content
Only embed text that’s likely to matter:
- Drop empty/low-signal records
- Skip near-duplicates
- Use rules or lightweight models to filter before embedding
6) Consider alternative providers/models
Depending on your quality needs and scale, other embedding APIs or open-source models hosted on your own infrastructure may be cheaper for large batch workloads.
7) Run open-source embeddings yourself
For very large volumes, self-hosted models can be much cheaper:
sentence-transformers- bge / e5 family models
- Hosted on CPU or GPU depending on throughput needs
This trades off:
- lower per-token cost
- more ops/maintenance
- more setup complexity
Practical approach
A good pattern for batch pipelines is:
- Normalize text
- Deduplicate by hash
- Filter low-value rows
- Chunk to a target size
- Batch embed
- Cache results
- Reuse embeddings on reruns
If you want, I can help with one of these:
- estimate Vertex AI embedding cost for your workload
- redesign your batching strategy
- suggest cheaper embedding models
- show a Python example for caching + batching
Brands