Prompt
What should I use for a crawler with retries and scheduling?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you want a crawler that supports retries and scheduling, the best choice depends on how much control you need and how big the crawl is.
Good options
1. Scrapy + a scheduler/queue
Best if you want a flexible Python crawler.
- Scrapy handles crawling well
- Add Celery, RQ, or APScheduler for retries/scheduling
- Good for custom logic, rate limiting, and distributed jobs
Use this if:
- You want Python
- You need custom retry rules
- You may later scale to multiple workers
2. Apify
Best if you want something ready-made and managed.
- Built-in scheduling
- Automatic retries
- Easy deployment
- Good for production crawling without a lot of infrastructure
Use this if:
- You want speed to production
- You don’t want to maintain your own queue system
3. Playwright + job queue
Best for sites that need a browser.
- Use Playwright for dynamic pages
- Use Celery / RQ / Temporal for retries and scheduling
Use this if:
- The site is heavily JavaScript-based
- You need browser automation
4. Temporal
Best for robust retry/scheduling/workflow control.
- Strong retry policies
- Excellent scheduling/workflow orchestration
- More complex, but very reliable
Use this if:
- You need guaranteed workflows
- You have many crawl tasks and complex retry logic
My recommendation
- Simple Python crawler: Scrapy + Celery
- Dynamic pages: Playwright + Celery
- Managed platform: Apify
- Enterprise-grade workflows: Temporal
If you tell me:
- your language preference,
- whether the site is JS-heavy,
- and roughly how many pages/tasks per day,
I can recommend a specific stack.