Prompt
Oxylabs scraping API for many domains
Latest observation
If you want to use Oxylabs Web Scraper API / Scraper API to scrape many different domains, the main thing to plan for is how you route each request, because different sites may need different:
- target endpoint / source
- parameters
- rendering mode
- geo-location
- anti-bot handling
- parsing logic
Recommended approach
1) Build a domain-to-config map
Create a configuration object for each domain, for example:
{
"amazon.com": {
"source": "universal",
"render": "html",
"parse_type": "product"
},
"linkedin.com": {
"source": "universal",
"render": "html",
"geo_location": "United States"
},
"example-news.com": {
"source": "universal",
"render": "false",
"parse_type": "article"
}
}
This lets your app decide:
- which Oxylabs source to use
- whether to use browser rendering
- whether to send location/session settings
2) Use a single wrapper function
Instead of writing separate code per site, make one function that accepts:
- the domain
- the URL
- an optional config
Example in Python:
import requests
OXYLABS_USER = "your_username"
OXYLABS_PASS = "your_password"
def scrape_url(url, domain_config=None):
domain_config = domain_config or {}
payload = {
"url": url,
"source": domain_config.get("source", "universal"),
"render": domain_config.get("render", False),
}
if "geo_location" in domain_config:
payload["geo_location"] = domain_config["geo_location"]
response = requests.post(
"https://realtime.oxylabs.io/v1/queries",
auth=(OXYLABS_USER, OXYLABS_PASS),
json=payload,
timeout=60
)
response.raise_for_status()
return response.json()
3) Add site-specific parsing after scraping
Oxylabs gives you the fetched content, but you still need custom parsing depending on the domain.
For example:
- product pages → title, price, rating
- news pages → headline, author, publish date
- directories → list of records
Use BeautifulSoup / lxml / regex / Playwright depending on the HTML structure.
4) Handle high volume with queues
If you need many domains and many URLs:
- store URLs in a queue
- rate limit per domain
- retry failed requests
- log response codes and errors
- cache successful results
Suggested architecture:
- Job queue: Redis / RabbitMQ / SQS
- Workers: send Oxylabs requests
- Parser layer: per-domain extraction
- Storage: DB / parquet / S3
5) Use browser rendering only when needed
Rendering is slower and more expensive, so:
- use
render: trueonly for JS-heavy sites - keep
render: falsefor static pages
6) Normalize output across domains
If you scrape many domains, it helps to convert all results into a common schema, e.g.:
{
"domain": "example.com",
"url": "https://example.com/item/123",
"title": "...",
"price": "...",
"scraped_at": "2026-10-04T12:00:00Z",
"raw_html": "..."
}
That makes downstream processing easier.
Practical tips
- Maintain a per-domain allowlist and config
- Detect when a site changes layout
- Use sessions/cookies only when required
- Keep separate parsers from scraper logic
- Monitor success rate by domain
If you want, I can give you:
- a Python framework for scraping many domains with Oxylabs,
- a Node.js version, or
- a config-driven architecture example for production use.