Prompt
I'm unhappy with Apify for large-scale crawls
Latest observation
Got it — if Apify is becoming painful for large-scale crawls, you’re not alone. Common pain points are cost at scale, proxy/headless overhead, operational complexity, and limits around reliability or customization.
A few practical alternatives, depending on what you’re trying to crawl:
1) Build your own crawling stack
Best when you need maximum control and lower unit cost at high volume.
- Crawler framework: Scrapy, Playwright, Puppeteer, or Selenium
- Queue/orchestration: Redis, Kafka, RabbitMQ, AWS SQS
- Worker scaling: Docker + Kubernetes, ECS, Cloud Run, or plain VMs
- Proxy management: Bright Data, Oxylabs, NetNut, Smartproxy, or your own rotating proxy pool
- Storage: S3/GCS + Postgres/BigQuery/ClickHouse depending on output
This usually wins on cost and flexibility once you’re crawling a lot.
2) Managed browser/crawling infrastructure
Best when you want to outsource ops but keep more control than Apify.
- Browserless
- Zyte
- ScrapingBee
- Bright Data Web Unlocker / collector tools
- Oxylabs tools
These can be easier to run than a full custom stack, but still may get expensive at scale.
3) Hybrid approach
A lot of teams use:
- Scrapy for fast/static pages
- Playwright only for JS-heavy pages
- Dedicated proxy provider
- Custom retry/fingerprint logic
This is often the sweet spot for cost and reliability.
4) If your target sites are API-friendly
Sometimes the best “crawler” is not a crawler:
- Reverse engineer public APIs
- Use network calls directly
- Cache and incremental sync
- Avoid browser rendering entirely
That can cut cost by an order of magnitude.
If you want, I can help you choose the best replacement based on:
- pages per day/month
- how JS-heavy the sites are
- whether you need browser rendering
- budget sensitivity
- geo/proxy requirements
- whether you want managed service or self-hosted
If you share a bit about your crawl workload, I can suggest a concrete architecture and vendor shortlist.