Prompt
I'm frustrated with Scrapy setup for high-volume crawling
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
Totally fair — Scrapy can feel great for “normal” crawling, then get annoying once you push into high-volume, high-concurrency, or lots-of-site diversity.
A few common pain points with high-volume Scrapy setups:
- Memory growth / leaks when queues get huge or responses are large
- Throttle tuning across many domains
- Retry storms causing duplicate load instead of recovery
- Pipeline bottlenecks (DB writes, item processing, exporting)
- Too many open connections / file descriptors
- Proxy / ban handling becoming messy
- Scheduler pressure when the frontier gets massive
- Spider overlap and coordination across multiple jobs
If you want, I can help in one of these ways:
-
Diagnose your current setup
Share:- Scrapy version
- number of spiders / domains
- concurrency settings
- download delay / AutoThrottle
- pipeline/export setup
- what “high-volume” means for you
- the specific failure mode (OOM, slow crawl, bans, timeouts, crashes)
-
Review your settings.py
I can suggest concrete settings for:- high throughput
- stability
- polite crawling
- proxy rotation
- distributed crawling
-
Design a scalable architecture For example:
- Scrapy + Redis frontier
- Scrapy Cluster
- Kafka/RabbitMQ task queue
- per-domain concurrency controls
- async item export
- separation of crawl, parse, and storage
-
Help simplify the stack Sometimes the best fix is:
- reduce middleware/pipeline complexity
- move heavy parsing outside Scrapy
- batch writes
- use separate workers for fetching vs storing
If you paste your settings.py and describe the bottleneck, I can give you a concrete tuning plan instead of generic advice.