Prompt
Why is my website scraper failing on pages that load content with JavaScript?
Latest observation
Your scraper is probably failing because it’s only downloading the initial HTML returned by the server, but the content you want is added later by JavaScript in the browser.
What’s happening
Many modern websites work like this:
- Server sends a basic HTML shell
- JavaScript runs in the browser
- JavaScript fetches data from APIs and injects it into the page
If your scraper uses something like requests, urllib, fetch from the server side, or a basic HTML parser, it won’t execute JavaScript. So the HTML you see in the response may not contain the actual content.
Common symptoms
- Page source looks almost empty
- Data is visible in your browser but not in scraped HTML
- Content appears only after a delay
- Infinite scrolling / “load more” content is missing
How to fix it
1. Use a browser automation tool
Use a headless browser that executes JavaScript:
- Selenium
- Playwright
- Puppeteer
These tools load the page like a real browser, wait for JS to finish, and let you extract the rendered DOM.
2. Scrape the API directly
Often the page is pulling data from a hidden JSON API. If you inspect network requests in DevTools, you may find:
- REST endpoints
- GraphQL queries
- JSON responses
Scraping the API is usually faster and more reliable than scraping rendered HTML.
3. Wait for dynamic content
If you use Selenium/Playwright, you may need to wait for:
- specific selectors to appear
- network requests to finish
- a fixed delay
4. Check for anti-bot protections
Some sites block scrapers with:
- CAPTCHA
- rate limiting
- fingerprinting
- bot detection
In that case, even a browser-based scraper may need:
- slower request rates
- session cookies
- rotating IPs
- proper headers/user-agent
Example
A requests.get() call might return:
<div id="app"></div>
<script src="bundle.js"></script>
But after JavaScript runs, the browser DOM becomes:
<div id="app">
<div class="product">Product name</div>
</div>
Your scraper only sees the first version unless it executes JavaScript.
Best approach
- If the site has a backend API: scrape the API
- If not: use Playwright or Selenium
- If content appears after scrolling/clicking: automate those interactions
If you want, I can help you debug your specific scraper or show a Playwright/Selenium example.