Prompt
How do I move Salesforce and S3 data into BigQuery on a schedule?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To move Salesforce and Amazon S3 data into BigQuery on a schedule, you generally have 3 good patterns depending on how much control, reliability, and freshness you need.
Best options
1) Use a managed ELT/ETL tool
This is the easiest approach.
Common tools:
- Fivetran
- Stitch
- Airbyte
- Matillion
- Talend
- Informatica
How it works:
- Connect Salesforce and S3 as sources
- Connect BigQuery as the destination
- Configure sync frequency, usually every 15 min, hourly, or daily
- The tool handles schema changes, incremental loads, retries, and API limits
Best for:
- Minimal engineering effort
- Reliable scheduled syncs
- Ongoing maintenance handled by the vendor
2) Build a custom pipeline with GCP services
This is best if you want more control or lower cost at scale.
Typical architecture:
Salesforce
- Use Salesforce REST/Bulk API or a connector
- Extract data on a schedule
S3
- Read files from S3 using:
- a Cloud Run / Cloud Functions job
- Dataflow
- a Python script in Composer/Airflow
- Transfer Service or external tooling
Landing and loading
- Store raw data in GCS first, or stream/load directly into BigQuery
- Use:
- BigQuery load jobs for batch files
- BigQuery Data Transfer Service where supported
- Airflow/Cloud Composer for orchestration
Scheduling
- Cloud Scheduler
- Cloud Composer (Airflow)
- Workflows
- cron in a container/job environment
Best for:
- Full customization
- Complex transformations
- Precise control over incremental logic
3) Use BigQuery-native or Google Cloud transfer tools where possible
Google has some built-in connectors, but support varies.
- BigQuery Data Transfer Service supports some SaaS sources, but Salesforce support may depend on region/product availability and not all use cases are covered.
- For S3, the usual approach is to copy files from S3 to GCS first, then load into BigQuery.
Best for:
- Simpler setups if your source is supported
- Less code than a fully custom pipeline
Recommended approach
If you want the fastest reliable setup:
Use Fivetran or Airbyte to sync:
- Salesforce → BigQuery
- S3 → BigQuery
If you want a Google Cloud-native scheduled pipeline:
Use:
- Cloud Scheduler to trigger
- Cloud Run / Cloud Functions / Airflow to extract from Salesforce and S3
- GCS as a staging area
- BigQuery load jobs to ingest on schedule
Example architecture
Salesforce → BigQuery
- Schedule a job every hour
- Call Salesforce Bulk API
- Pull changed records since last sync
- Write results as CSV/JSON/Parquet to GCS
- Load into BigQuery staging table
- Merge into final table with
MERGE
S3 → BigQuery
- Schedule a job every day/hour
- List new files in S3
- Copy files to GCS or process directly
- Load into BigQuery using load jobs
- Merge or append as needed
Key design decisions
Incremental vs full refresh
- Incremental is usually preferred
- Track:
updated_atlast_modified- file timestamps
- object version IDs or ETags for S3
File format
Best formats for BigQuery:
- Parquet
- Avro
- CSV or JSON if necessary
Parquet is usually best for performance and schema fidelity.
Data modeling
Typical pattern:
rawtables: exact source datastagingtables: cleaned/normalizedmarttables: analytics-ready
Scheduling
- Hourly, daily, or near-real-time depending on needs
- For Salesforce, API limits often make hourly/daily more practical
- For S3, schedule depends on how often files arrive
Example with Airbyte
If you want a simple scheduled sync:
- Deploy Airbyte
- Add Salesforce source
- Add S3 source
- Add BigQuery destination
- Set sync frequency
- Choose incremental sync mode
- Monitor sync jobs and schema evolution
Example with Cloud Composer
If you need orchestration:
- DAG task 1: extract Salesforce delta
- DAG task 2: copy new S3 objects
- DAG task 3: load to BigQuery staging tables
- DAG task 4: run
MERGESQL - DAG task 5: validate row counts/logs
Things to watch out for
- Salesforce API limits and pagination
- OAuth token refresh
- Schema drift in Salesforce objects and S3 files
- Duplicate records from reprocessing
- Timezone handling
- BigQuery partitioning and clustering
- Permissions/IAM for S3 access and BigQuery writes
If you want, I can give you one of these next:
- a recommended architecture diagram
- a step-by-step Airbyte setup
- a custom Python + Cloud Scheduler + BigQuery example
- a Cloud Composer DAG example