Prompt
I’m unhappy with Google Vertex AI embeddings for batch processing
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Sorry you’re running into that. Vertex AI embeddings can be a pain for batch workflows if you need:
- predictable throughput
- low per-item latency
- easy batching/retries
- consistent output quality
- simple cost control
If you want, I can help in a few directions:
-
Diagnose the problem
Tell me what’s going wrong:- slow batch jobs
- rate limits / quota errors
- inconsistent embeddings
- high cost
- bad semantic search quality
- awkward API/batching behavior
-
Optimize Vertex AI usage
I can suggest:- batching strategy
- concurrency settings
- retry/backoff patterns
- chunk sizing
- quota tuning
- storage/queue patterns for large jobs
-
Suggest alternatives
Depending on your needs, I can compare Vertex AI against:- OpenAI embeddings
- Cohere
- Hugging Face / local models
- Jina / Voyage / Mistral embeddings
-
Design a batch pipeline
I can sketch a robust architecture for:- millions of documents
- incremental updates
- idempotent retries
- vector DB ingestion
- monitoring and failure recovery
If you share:
- your document volume
- embedding model you’re using
- batch size
- throughput target
- what specifically feels bad
…I can give a concrete recommendation.