Prompt
Can you help me choose a crawling setup for monitoring competitor sites, where I need scheduled revisits, retries, deduplication, and JSON output?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
Yes — for that use case, you want a crawler framework that gives you:
- scheduled revisits of pages
- retry logic for failures/timeouts
- deduplication so you don’t reprocess the same URLs/content unnecessarily
- structured JSON output for downstream analysis
Best fit: Scrapy + a scheduler layer
If you’re comfortable with Python, Scrapy is usually the strongest choice.
Why it fits
- Retries: built-in retry middleware
- Deduplication: built-in duplicate request filtering
- JSON output: easy with Scrapy’s feed exports
- Scheduled revisits: add a queue/scheduler on top, or run spiders on a schedule
Common setup
- Scrapy for crawling
- Redis / RabbitMQ / a DB to track URLs and revisit intervals
- Cron / Airflow / Celery beat to trigger recurring crawls
- PostgreSQL / S3 / JSON lines files for storing results
Good if:
- you need control and flexibility
- sites are mostly static or moderately dynamic
- you want a Python-based pipeline
If you want distributed scheduled crawling: Apache Nutch or Heritrix
These are more “crawler infrastructure” than lightweight app frameworks.
Apache Nutch
- good for large-scale crawling
- supports scheduling/revisit logic better than many lightweight tools
- can integrate with Hadoop ecosystem
- less convenient than Scrapy for custom extraction
Heritrix
- strong for archival-style crawling
- good on politeness and breadth crawling
- less ideal if you want custom JSON extraction pipelines quickly
Good if:
- you expect scale to grow significantly
- you need a more crawler-platform approach
- you have ops resources to maintain it
If pages are heavily JavaScript-rendered: Playwright + a crawl orchestrator
For competitor sites that depend on client-side rendering:
- Playwright or Selenium for rendering
- pair with a scheduler/orchestrator like Airflow, Celery, or Scrapy-Playwright
Tradeoff
- more resource intensive
- slower than plain HTTP crawling
- but necessary for dynamic sites
Recommended practical setups
Option A: Best general-purpose choice
Scrapy + Redis + cron/Airflow + JSON Lines
- easiest to build and maintain
- retries and dedupe are straightforward
- revisit scheduling is handled externally
- good balance of simplicity and power
Option B: More production-oriented
Scrapy + Celery beat + PostgreSQL + S3/JSON
- better if you need frequent scheduled runs
- easier to track crawl history and revisits
- useful when crawling many competitor domains
Option C: Enterprise-scale
Apache Nutch + scheduler + storage pipeline
- if you’re crawling many sites at high volume
- more operational overhead
My recommendation
For most competitor monitoring projects, I’d choose:
Scrapy + scheduled jobs + a database for crawl state + JSON output
That gives you:
- revisit control
- retries
- dedupe
- clean structured output
- low complexity compared with full crawler platforms
A few questions to narrow it down
- Are the competitor sites mostly static HTML or JavaScript-heavy?
- Roughly how many pages/domains do you need to monitor?
- Do you need near-real-time checks, or is daily/weekly enough?
- Do you want to run this yourself, or use a managed service?
If you answer those, I can recommend a specific stack and architecture.