Prompt
How do I crawl JavaScript-heavy websites reliably?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To crawl JavaScript-heavy sites reliably, you usually need a browser-based crawler rather than a plain HTTP scraper. The key is to choose the lightest approach that still executes the page’s JavaScript and waits for the right content.
Recommended approach
1) Use a real browser engine
Tools:
- Playwright — best default choice today
- Puppeteer — good if you only need Chromium
- Selenium — older, broader ecosystem, usually less convenient
Why:
- Pages render dynamically
- Data may load via API calls after initial HTML
- Content may depend on client-side routing, hydration, or infinite scroll
2) Prefer intercepting API calls over scraping rendered DOM
If the site loads data via XHR/fetch requests:
- Open DevTools and inspect network traffic
- Identify the JSON endpoints
- Crawl the endpoints directly when possible
This is usually:
- Faster
- More stable
- Easier to scale
- Less brittle than DOM scraping
3) Wait for the right condition
Avoid arbitrary sleeps if possible. Use:
waitForSelectorwaitForLoadState("networkidle")in Playwright- Custom waits for specific text, element counts, or API responses
Example idea:
- Wait for the product cards to appear
- Then extract data
4) Handle infinite scroll and lazy loading
Common patterns:
- Scroll incrementally
- Wait for new items to load
- Stop when item count stops increasing
Be careful not to scroll too fast.
5) Make your crawler resilient
Add:
- Retries with backoff
- Timeout handling
- Session/page restarts after failures
- Logging of failed URLs and screenshots/HTML snapshots
6) Scale carefully
Browser automation is expensive. For larger crawls:
- Reuse browser contexts where possible
- Limit concurrency
- Use request blocking for images/fonts/media if not needed
- Consider a queue-based architecture
7) Reduce fingerprinting issues
Some sites detect automation. To reduce problems:
- Use realistic user agents
- Set viewport and locale
- Enable cookies and sessions
- Avoid obviously robotic interaction patterns
- Respect robots.txt and site terms where appropriate
Practical strategy
A good workflow is:
- Inspect network calls
- If data is in APIs, crawl APIs directly
- If not, use Playwright
- Wait for specific content
- Extract structured data
- Add retries, throttling, and monitoring
Example stack
- Playwright for rendering
- BeautifulSoup / lxml for parsing after render
- Redis / RabbitMQ for job queues
- PostgreSQL / MongoDB for storage
Example Playwright pattern
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com", wait_until="networkidle")
page.wait_for_selector(".product-card")
items = page.locator(".product-card").all_text_contents()
print(items)
browser.close()
Important caveats
- Some sites require login, tokens, or anti-bot handling
- Some content is rendered only after user interaction
- Some sites block headless browsers
- Always check legal/ethical constraints
If you want, I can also give you:
- a Playwright-based crawler template,
- an API-first scraping workflow, or
- a crawler architecture for scale.