Prompt
Why is search latency spiking under load?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Search latency usually spikes under load because one or more parts of the search path stop scaling linearly once traffic, data size, or concurrency rises. Common causes:
1) Query thread pool saturation
- Too many concurrent searches
- Requests queue up waiting for worker threads
- Latency jumps even if CPU is not fully pegged
2) CPU bottlenecks
- Expensive queries, scoring, aggregations, regex/wildcards, sorting
- High GC overhead in JVM-based systems
- Uneven shard distribution causing one node to do most of the work
3) Disk I/O pressure
- Cold cache or frequent cache evictions
- Segment reads causing random I/O
- Merge activity competing with search for disk bandwidth
- Slow storage under mixed read/write load
4) Memory pressure / garbage collection
- Heap too small or too full
- Large result sets, heavy aggregations, fielddata/doc values pressure
- GC pauses stall query processing
5) Cache misses
- Query cache, shard request cache, filesystem cache not warm
- Under load, working set exceeds cache size
- Latency increases as searches hit disk more often
6) Hot shards or skewed routing
- A few shards get most requests
- Those shards become overloaded while others sit idle
- Response time tracks the slowest shard
7) Merge and indexing interference
- Heavy indexing triggers segment merges
- Merges consume CPU, disk, and memory
- Search latency rises during write bursts
8) Network or coordination overhead
- Cross-node shard fan-out adds hop latency
- Large responses or slow nodes delay overall search completion
- Coordinator node becomes a bottleneck
9) Inefficient queries
- Leading wildcards, large offsets, deep pagination
- Broad filters, unbounded aggregations
- Sorting on unanalyzed fields without proper indexes/doc values
How to narrow it down
Check:
- CPU, load average, and run queue
- Search thread pool queue/rejections
- GC pause times
- Disk latency and IOPS
- Cache hit ratios
- Merge activity and indexing rate
- Shard latency distribution, especially p95/p99
- Whether latency is on all nodes or just a subset
Quick interpretation
- High queue/rejections → thread pool saturation
- High CPU / GC → compute or memory pressure
- High disk latency → storage bottleneck
- Only some shards slow → hot shard / skew
- Spikes during writes → merge/index interference
If you want, I can help you debug this for a specific system like Elasticsearch/OpenSearch/Solr/Postgres full-text search, or help interpret metrics/logs you already have.