Prompt
What should I use for recurring crawls across multiple domains?
Latest observation
For recurring crawls across multiple domains, the best choice is usually Scrapy if you want a flexible, production-friendly crawler, and Scrapy + a scheduler/orchestrator if you need it to run repeatedly on a timetable.
Good options
-
Scrapy
Best for building and maintaining crawlers across many sites/domains. It handles concurrency, retries, throttling, pipelines, and site-specific logic well. -
Scrapy Cloud / Zyte
Good if you want managed recurring execution without running your own infrastructure. -
Apache Airflow / Prefect / Dagster
Best if you already have workflows and want to schedule crawls regularly, then process/store results afterward. -
Playwright / Selenium
Use these only if the sites are heavily JavaScript-driven. They’re usually slower and more expensive to run than Scrapy.
Typical setup
A strong pattern is:
- Scrapy spiders for each domain or site type
- A scheduler like Airflow/cron/Prefect to trigger them daily/weekly
- A storage layer like Postgres, S3, Elasticsearch, or a data warehouse
- Monitoring/logging for failures and changes
If you want the simplest recommendation
- Small scale: Scrapy + cron
- Medium/large scale: Scrapy + Airflow/Prefect
- Need managed crawling: Zyte/Scrapy Cloud
- Heavy JS sites: Playwright + scheduler
If you tell me:
- how many domains,
- how often to crawl,
- whether pages are JS-heavy,
- and where the data should go,
I can recommend a concrete stack.