Prompt
I'm unhappy with Scrapy for scheduling recurring jobs
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you’re unhappy with Scrapy for scheduling recurring jobs, that’s pretty understandable—Scrapy is a great crawling framework, but it’s not really a scheduler or job orchestrator.
A few better options depend on what you mean by “recurring jobs”:
If you just need “run this scraper every X minutes/hours”
- cron / systemd timers: simplest and very reliable
- GitHub Actions / CI schedules: good for lightweight periodic runs
- Airflow / Prefect / Dagster: better if you want visibility, retries, dependencies, and monitoring
If you need distributed job scheduling
- Celery beat with Celery workers
- RQ Scheduler if you’re already using Redis/RQ
- APScheduler for in-app scheduling, though it’s more limited for serious production orchestration
If the issue is with scraping itself
You might keep Scrapy for the crawl logic, but use:
- a scheduler to trigger runs
- a queue/worker system to manage execution
- storage for deduping and incremental crawling
Common architecture
- Scheduler: cron/Airflow/Prefect
- Crawler: Scrapy
- Queue/runner: Celery/RQ or just a subprocess
- State store: DB/Redis/S3
If you want, I can suggest:
- the best replacement for Scrapy scheduling based on your setup, or
- a simple architecture for recurring crawls using Scrapy plus a proper scheduler.