Prompt
How to export scraped website data to BigQuery
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To export scraped website data to BigQuery, the usual flow is:
- Scrape the website
- Clean/normalize the data
- Write the data to a file or stream
- Load it into BigQuery
Common ways to do it
Option 1: Save as JSON/CSV and load into BigQuery
This is the simplest approach.
Example with Python
import pandas as pd
from google.cloud import bigquery
# Example scraped data
data = [
{"title": "Product A", "price": 19.99, "url": "https://example.com/a"},
{"title": "Product B", "price": 29.99, "url": "https://example.com/b"},
]
# Convert to DataFrame
df = pd.DataFrame(data)
# Save locally
df.to_csv("scraped_data.csv", index=False)
# Upload to BigQuery
client = bigquery.Client()
table_id = "your-project.your_dataset.your_table"
job = client.load_table_from_dataframe(df, table_id)
job.result()
print("Loaded to BigQuery successfully")
Option 2: Append rows directly via BigQuery API
Good if you want to load data continuously without creating files.
from google.cloud import bigquery
client = bigquery.Client()
table_id = "your-project.your_dataset.your_table"
rows_to_insert = [
{"title": "Product A", "price": 19.99, "url": "https://example.com/a"},
{"title": "Product B", "price": 29.99, "url": "https://example.com/b"},
]
errors = client.insert_rows_json(table_id, rows_to_insert)
if errors:
print("Errors:", errors)
else:
print("Rows inserted successfully")
Option 3: Export to Google Cloud Storage, then load to BigQuery
Best for large datasets.
- Scrape data and save as CSV/JSON
- Upload file to GCS
- Use BigQuery load job from GCS
Example:
from google.cloud import bigquery
client = bigquery.Client()
table_id = "your-project.your_dataset.your_table"
uri = "gs://your-bucket/scraped_data.csv"
job_config = bigquery.LoadJobConfig(
source_format=bigquery.SourceFormat.CSV,
skip_leading_rows=1,
autodetect=True,
)
load_job = client.load_table_from_uri(uri, table_id, job_config=job_config)
load_job.result()
print("Loaded from GCS to BigQuery")
Recommended schema design
Before exporting, make sure your scraped data is consistent:
- Use stable column names
- Convert dates to ISO format (
YYYY-MM-DDor RFC3339) - Flatten nested objects if possible
- Handle missing values
- Use correct data types:
- strings
- integers
- floats
- booleans
- timestamps
Typical workflow
- Scrape with
requests + BeautifulSoup,Selenium, orPlaywright - Store data in a list of dictionaries
- Convert to pandas DataFrame
- Load into BigQuery using
google-cloud-bigquery
Setup requirements
Install the BigQuery client:
pip install google-cloud-bigquery pandas
Set authentication:
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
Important notes
- BigQuery tables must exist or be auto-created depending on your method
- For production scraping, add:
- retries
- deduplication
- schema validation
- batching
- If scraping frequently, consider using:
- Cloud Run / Cloud Functions
- Cloud Scheduler
- Pub/Sub or Dataflow
If you want, I can also give you:
- a full end-to-end Python scraper → BigQuery example, or
- a workflow for large-scale scraping with Cloud Storage + BigQuery.