Prompt
How do I set up a web scraping tool for tracking pricing changes across merchants?
Latest observation
To set up a web scraping tool for tracking pricing changes across merchants, build it as a small pipeline: discover pages → fetch product data → normalize → store snapshots → compare over time → alert/report.
1) Define what you want to track
Start with a clear schema:
- Merchant name
- Product identifier/SKU
- Product name
- URL
- Price
- Sale price / list price
- Availability
- Timestamp
- Currency
- Shipping price, if relevant
If you’ll compare across merchants, also create a canonical product ID for matching the same item from different sites.
2) Choose a scraping approach
Simple sites
Use:
- Requests + BeautifulSoup for static HTML
JavaScript-heavy sites
Use:
- Playwright or Selenium
- Prefer Playwright for reliability and speed
Large-scale, long-running monitoring
Use:
- A queue + workers
- A database
- Scheduler like cron, Airflow, or Celery beat
3) Respect legal and technical constraints
Before scraping:
- Check the site’s Terms of Service
- Review robots.txt
- Avoid aggressive request rates
- Identify yourself with a user agent if appropriate
- Don’t try to bypass paywalls, anti-bot protections, or login restrictions unless you have permission
If merchants offer feeds, APIs, or affiliate product catalogs, those are usually better than scraping.
4) Build the scraper
Basic example with Requests + BeautifulSoup
import requests
from bs4 import BeautifulSoup
url = "https://example.com/product/123"
headers = {
"User-Agent": "PriceMonitorBot/1.0"
}
resp = requests.get(url, headers=headers, timeout=20)
resp.raise_for_status()
soup = BeautifulSoup(resp.text, "html.parser")
name = soup.select_one(".product-title").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)
print(name, price)
Example with Playwright for dynamic pages
from playwright.sync_api import sync_playwright
url = "https://example.com/product/123"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="networkidle")
name = page.locator(".product-title").inner_text()
price = page.locator(".price").inner_text()
print(name, price)
browser.close()
5) Normalize the data
Prices come in many formats:
$1,299.991.299,99 €From $49Sale: $29.99
Convert them into:
- numeric value
- currency code
- whether it’s sale/list price
Example normalization steps:
- remove currency symbols
- parse locale-specific decimal separators
- store as decimal, not float if possible
- keep original raw text too
6) Store snapshots
Use a database table like:
products
- id
- merchant
- merchant_sku
- canonical_name
- product_url
price_snapshots
- id
- product_id
- scraped_at
- price
- currency
- availability
- raw_price_text
This lets you track changes over time instead of only the latest value.
7) Detect price changes
Compare the newest snapshot with the previous one:
- If price changed, log the delta
- Trigger alerts if price drops below threshold
- Track percentage change
Example logic:
- previous = 100.00
- current = 89.99
- change = -10.01
- percent = -10.01%
8) Schedule runs
Depending on your needs:
- Every hour for fast-moving merchants
- Daily for slower catalogs
- More often during promotions
Options:
cron- GitHub Actions
- Celery beat
- Airflow
- Kubernetes CronJobs
9) Add monitoring and retries
Scraping breaks often, so add:
- timeout handling
- retries with backoff
- logging
- error alerts
- HTML change detection
- fallback selectors
Useful practices:
- Save raw HTML for failed runs
- Track selector success rate by merchant
- Alert when scrape coverage drops suddenly
10) Present or alert on changes
Possible outputs:
- Email alerts
- Slack notifications
- Dashboard
- CSV export
- BI tool integration
Example alert rule:
- notify when price drops by more than 10%
- notify when an item becomes out of stock
- notify when competitor undercuts your price
11) Suggested architecture
A practical setup:
- Config file with merchant URLs and selectors
- Scraper workers for each merchant
- Database for products and price history
- Comparison job to detect changes
- Notification service to send alerts
- Dashboard for viewing trends
12) Recommended stack
For a Python-based solution:
- Scraping:
requests,BeautifulSoup,playwright - Parsing:
pandas,price-parser - Storage: PostgreSQL or SQLite for small setups
- Scheduling: cron or Celery
- Alerts: Slack webhook, email via SMTP
- Deployment: Docker
13) Minimal project structure
price-monitor/
scrapers/
merchant_a.py
merchant_b.py
config/
merchants.yaml
db/
models.py
migrations/
jobs/
scrape_all.py
compare_prices.py
alerts/
slack.py
requirements.txt
14) Keep it maintainable
- Write merchant-specific parsers
- Use shared helper functions
- Keep selectors in config where possible
- Version your scraper rules
- Test against saved HTML samples
If you want, I can also provide:
- a full Python starter project,
- a database schema, or
- a Playwright-based scraper template for multiple merchants.