Prompt

How do I scrape product prices from multiple websites into a spreadsheet?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

You can do this in a few different ways, depending on how many sites you need, how often you want updates, and whether the sites allow scraping.

1) Decide the data flow

Typical pipeline:

  1. List the product pages you want to track
  2. Extract the price from each page
  3. Normalize the data into columns like:
    • Website
    • Product name
    • Product URL
    • Price
    • Currency
    • Date/time scraped
  4. Export to a spreadsheet like CSV, Excel, or Google Sheets
  5. Repeat on a schedule if needed

2) Choose an approach

Option A: No-code / low-code

Good if you want something quick.

Tools:

  • Octoparse
  • ParseHub
  • Apify
  • Browse AI
  • Web Scraper browser extension

Workflow:

  • Set up a scraper for each site
  • Select the price element
  • Export to CSV or Google Sheets

Best for:

  • Few sites
  • Simple pages
  • Non-developers

Option B: Python script

Best if you want flexibility and automation.

Common libraries:

  • requests for fetching pages
  • BeautifulSoup for parsing HTML
  • pandas for spreadsheet export
  • openpyxl or xlsxwriter for Excel
  • Selenium or Playwright for sites that render prices with JavaScript

Best for:

  • Multiple sites
  • Custom logic
  • Scheduled scraping
  • Cleaning and merging data

3) Basic Python example

Here’s a simple pattern for websites where the price is in the HTML:

import requests
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime

products = [
    {"site": "SiteA", "url": "https://example.com/product1"},
    {"site": "SiteB", "url": "https://example.com/product2"},
]

rows = []

for item in products:
    try:
        r = requests.get(item["url"], headers={"User-Agent": "Mozilla/5.0"}, timeout=10)
        r.raise_for_status()

        soup = BeautifulSoup(r.text, "html.parser")

        # Example selectors — change these per site
        price_el = soup.select_one(".price")
        name_el = soup.select_one("h1")

        price = price_el.get_text(strip=True) if price_el else None
        name = name_el.get_text(strip=True) if name_el else None

        rows.append({
            "site": item["site"],
            "product_name": name,
            "url": item["url"],
            "price": price,
            "scraped_at": datetime.utcnow().isoformat()
        })

    except Exception as e:
        rows.append({
            "site": item["site"],
            "product_name": None,
            "url": item["url"],
            "price": None,
            "scraped_at": datetime.utcnow().isoformat(),
            "error": str(e)
        })

df = pd.DataFrame(rows)
df.to_csv("prices.csv", index=False)
df.to_excel("prices.xlsx", index=False)

4) If prices load dynamically

If the price isn’t in the page HTML and appears only after JavaScript runs, use:

  • Playwright (preferred)
  • Selenium

Example idea:

  • Open page in a browser automation tool
  • Wait for the price element to appear
  • Read the text
  • Save it to your sheet

5) Put data into a spreadsheet

You can export to:

CSV

  • Easy to open in Excel or Google Sheets

Excel

  • Better formatting and multiple tabs

Google Sheets

Use the Google Sheets API or:

  • gspread in Python
  • Zapier/Make to push rows automatically

6) Handle multiple websites cleanly

Since each site has a different structure, it’s best to create a scraper per site:

def scrape_site_a(url):
    # parse Site A selectors
    return {"name": ..., "price": ...}

def scrape_site_b(url):
    # parse Site B selectors
    return {"name": ..., "price": ...}

Then combine all results into one dataframe.


7) Important considerations

  • Check the site’s Terms of Service
  • Respect robots.txt
  • Don’t overload sites with requests
  • Use rate limiting and retries
  • Some sites may block bots or require login
  • Prices may vary by location, currency, cookies, or account status

8) Recommended setup

If you want a practical start:

  • Python + BeautifulSoup + pandas for simple sites
  • Playwright for JavaScript-heavy sites
  • Export to CSV or Google Sheets
  • Schedule with:
    • Windows Task Scheduler
    • cron on Linux/macOS
    • GitHub Actions
    • a cloud server

9) If you want a fast path

Tell me:

  1. The websites
  2. Whether the pages need login
  3. Whether prices load after page load
  4. Whether you want CSV, Excel, or Google Sheets

And I can help you build a scraper template for your specific sites.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.