Prompt
I'm unhappy with Browserless for long-running crawl jobs
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
Understood — if Browserless is struggling with long-running crawl jobs, that usually means you’re hitting one or more of these issues:
- session time limits
- unstable Chrome instances over long durations
- memory leaks from large crawls
- rate limits / throttling
- difficulty scaling across many pages or domains
A few better directions, depending on what you need:
-
Run your own browser workers
- Use Playwright or Puppeteer on dedicated VMs/containers.
- Best if you want full control over timeouts, retries, concurrency, and browser lifecycle.
-
Use a crawl platform built for long jobs
- Tools like Crawlee + Playwright/Puppeteer for orchestration.
- Add your own queue, persistence, and checkpointing so jobs can resume.
-
Split crawl jobs into resumable chunks
- Instead of one long browser session, break work into batches.
- Store progress in Redis/Postgres/S3 and restart from the last checkpoint.
-
Separate fetching from rendering
- Use plain HTTP fetch where possible.
- Only use a browser for pages that truly need JS rendering.
-
Use managed infrastructure, not managed browser sessions
- For example: Kubernetes jobs, ECS tasks, or serverless workers running headless Chrome.
- More reliable for multi-hour or multi-day crawls.
If you want, I can help you with one of these next steps:
- compare Browserless vs self-hosted Playwright
- design a reliable long-running crawl architecture
- suggest a stack for resumable crawling
- help diagnose why your current crawls fail