Prompt
How do I build a geo-aware scraping system for international sites?
Latest observation
Building a geo-aware scraping system means designing your crawler so it can reliably access and interpret sites as they appear to users in different countries, languages, or regulatory regions.
Here’s a practical blueprint.
1) Define what “geo-aware” means for your use case
Common geo factors:
- IP geolocation: content changes based on the requester’s IP country/region.
- Language/localization: different translations, currencies, measurements, or date formats.
- Region-specific routing: different subdomains, paths, or domains like
example.com,example.co.uk,example.de. - Legal/access restrictions: some content is blocked or altered by jurisdiction.
- Personalization: site behavior changes with cookies, account region, browser locale, or Accept-Language.
Decide which of these you need to emulate.
2) Build a region abstraction layer
Create a configuration model for each target region:
{
"region": "de-DE",
"country": "DE",
"language": "de",
"currency": "EUR",
"timezone": "Europe/Berlin",
"base_url": "https://example.de",
"accept_language": "de-DE,de;q=0.9,en;q=0.8",
"proxy_pool": "de-residential",
"search_params": {
"hl": "de",
"gl": "DE"
}
}
This lets your scraping jobs run as “region profiles” instead of hardcoding country logic everywhere.
3) Use the right network identity
If the site geofences by IP, you’ll need traffic that originates from the target country.
Options:
- Datacenter proxies: cheaper, often blocked more often.
- Residential proxies: more realistic, usually better for geo-sensitive sites.
- Mobile proxies: strong trust signal, expensive.
Best practices:
- Maintain a proxy pool per country/region.
- Rotate proxies carefully; avoid changing IP too often within a single session.
- Keep session affinity for sites that bind cart/session/login state to IP.
4) Match browser locale and headers to region
Geo-aware scraping is not just about IP. Send consistent browser signals:
Accept-LanguageUser-AgentSec-CH-UA*headers if needed- Timezone
- Locale
- Cookie region preferences
- Currency / country parameters in URLs
Example headers:
Accept-Language: fr-FR,fr;q=0.9,en;q=0.8
User-Agent: Mozilla/5.0 ...
If using a browser automation tool, also set:
- browser language
- timezone override
- geolocation permissions if relevant
- viewport and platform consistency
5) Model site variants explicitly
International sites often vary by:
- domain:
site.com,site.fr - path:
/en/,/fr/ - subdomain:
fr.site.com - query params:
?lang=fr&country=FR - server-side content negotiation
Store these rules in a “site registry”:
example:
regions:
US:
url: https://example.com
headers:
Accept-Language: en-US,en;q=0.9
FR:
url: https://example.fr
headers:
Accept-Language: fr-FR,fr;q=0.9
Your crawler should select the right combination automatically.
6) Normalize and compare content across locales
Content will differ by:
- text language
- decimal separators:
1,234.56vs1 234,56 - dates:
24/09/2026vs09/24/2026 - currency:
€,£,¥ - units: kilometers vs miles
Build a normalization pipeline:
- detect language
- parse locale-aware numbers/dates
- convert currencies to a canonical form if needed
- store original text alongside normalized fields
This is crucial if you’re comparing prices or availability across regions.
7) Detect geo-blocking and fallback behavior
Build detectors for:
- HTTP 403/451
- redirects to a country selector
- CAPTCHA or bot challenges
- “not available in your region” banners
- different markup patterns by country
When detected, automatically:
- switch to a region-appropriate proxy
- update language headers
- retry with the correct domain/version
- log the block reason for observability
8) Session and cookie strategy
Many sites set locale via cookies, for example:
country=DElocale=fr_FRcurrency=EUR
Approach:
- start each region with a fresh cookie jar
- persist cookies per region/session if site requires it
- don’t mix cookies across countries
- if login is needed, keep account region aligned with proxy region
9) Use browser automation for dynamic sites
For JS-heavy international sites, use:
- Playwright
- Selenium
- Puppeteer
Playwright is often a good fit because it handles:
- browser contexts
- locale/timezone settings
- network interception
- storage state per region
Example conceptually:
- one browser context per region
- set locale/timezone
- route traffic through the matching proxy
- capture rendered DOM after hydration
10) Build a scheduler and queue around regions
Treat each scrape as a job with region metadata:
- target site
- region
- proxy class
- retry policy
- rate limit bucket
This helps you:
- avoid overusing one country proxy pool
- distribute workload across regions
- separate failures by locale/site version
Example queue key:
site:example | region:FR | priority:normal
11) Add observability and validation
Track:
- success rate by region
- block rate by proxy type
- response time by geography
- content drift by locale
- selector breakage per site version
Validation checks:
- page title matches expected locale
- currency symbol matches region
- language detection matches requested locale
- country selector absent when it shouldn’t appear
Store HTML snapshots or screenshots for failed jobs.
12) Respect legal and ethical constraints
This is important for international scraping:
- obey applicable laws and site terms
- respect robots.txt where appropriate
- don’t scrape personal data without a lawful basis
- avoid bypassing protections that you’re not authorized to circumvent
- rate-limit to minimize load
If you’re scraping public data, still be careful with jurisdiction-specific rules like GDPR or local consumer-protection laws.
A simple architecture
Components:
-
Region Registry
- region configs, headers, URLs, proxy pools
-
Job Scheduler
- queues tasks by site and region
-
Fetcher
- HTTP client or browser automation
- proxy + locale + cookie handling
-
Parser
- region-specific selectors or DOM rules
-
Normalizer
- dates, currencies, language, units
-
Validator
- checks for geo-blocks and content correctness
-
Storage
- raw HTML, structured records, screenshots, logs
Example workflow
- Job requests product data for France.
- Scheduler selects
FRprofile. - Fetcher uses a French proxy and
Accept-Language: fr-FR. - Browser context sets timezone
Europe/Paris. - Page loads French variant.
- Parser extracts price, title, availability.
- Normalizer converts
12,99 €to12.99 EUR. - Validator confirms language and currency.
- Data is stored with region metadata.
Recommended implementation stack
- Orchestration: Airflow, Temporal, Celery, or simple queue workers
- HTTP scraping: Requests, httpx, Scrapy
- Browser automation: Playwright
- Proxy management: provider SDK + internal pool manager
- Parsing: BeautifulSoup, lxml, selectolax
- Normalization: Babel, dateparser, pytz/zoneinfo
- Monitoring: Prometheus, Grafana, Sentry, ELK
Common pitfalls
- Using the right proxy but wrong language headers
- Reusing cookies across regions
- Ignoring timezone differences in rendered content
- Assuming one selector works across all locales
- Over-rotating IPs and triggering anti-bot systems
- Failing to normalize currency/date formats
- Not logging which region version produced a record
Minimal design principle
If you remember only one thing:
Make region a first-class input to every layer: routing, headers, browser settings, parsing, normalization, and storage.
That’s what makes a scraper truly geo-aware.
If you want, I can also give you:
- a reference architecture diagram,
- a Playwright example with per-country profiles, or
- a Scrapy-based implementation pattern.