Prompt
I need cloud-native pipelines for Snowflake and BigQuery
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
Absolutely — for Snowflake and BigQuery, “cloud-native pipelines” usually means ETL/ELT pipelines that run in managed services, scale elastically, and minimize ops overhead.
Typical cloud-native pipeline patterns
1) ELT into the warehouse
Best for most analytics workloads.
Flow
- Ingest from apps / databases / SaaS / files
- Land raw data in object storage or directly in the warehouse
- Transform using warehouse-native SQL/compute
- Publish modeled tables for BI / ML
Snowflake stack
- Ingest: Fivetran, Airbyte, Debezium, Kafka, Snowpipe
- Storage: S3 / GCS / Azure Blob
- Transform: dbt, Snowpark, stored procedures
- Orchestration: Dagster, Airflow, Prefect, native tasks
- BI: Tableau, Power BI, Looker
BigQuery stack
- Ingest: Dataflow, Datastream, Pub/Sub, Fivetran, Airbyte
- Storage: GCS
- Transform: dbt, Dataform, BigQuery SQL/UDFs
- Orchestration: Cloud Composer, Workflows, Cloud Run jobs, Dagster
- BI: Looker, Tableau, Power BI
2) Streaming pipeline
Best for near-real-time use cases.
Flow
- Events in Kafka / Pub/Sub / Kinesis
- Stream processing / validation
- Write to warehouse tables
- Downstream aggregates and dashboards
Snowflake
- Kafka Connector / Snowpipe Streaming
- Tasks + Streams for incremental processing
- Optional Flink/Spark for enrichment
BigQuery
- Pub/Sub + Dataflow streaming
- BigQuery streaming inserts or Storage Write API
- Materialized views / scheduled queries for serving
3) Lakehouse-style pipeline
Best when you want cheap raw storage plus flexible compute.
Flow
- Raw data in cloud storage
- Catalog + governance
- ETL/ELT into curated warehouse models
This works well when you want:
- replayability
- decoupled storage/compute
- multiple downstream consumers
Recommended architecture
For Snowflake
- Ingestion: Fivetran/Airbyte + Snowpipe
- Raw layer: External stage or internal landing tables
- Transform: dbt + Snowflake Tasks/Streams
- Orchestration: Dagster or Airflow
- Observability: Monte Carlo, Great Expectations, dbt tests
- CI/CD: GitHub Actions + Terraform
For BigQuery
- Ingestion: Pub/Sub/Dataflow + Datastream + Fivetran
- Raw layer: GCS + BigQuery raw datasets
- Transform: dbt or Dataform
- Orchestration: Cloud Composer or Dagster on GCP
- Observability: Dataform assertions, dbt tests, Great Expectations
- CI/CD: GitHub Actions + Terraform
Cloud-native best practices
- Use ELT where possible
- Keep raw, staging, and mart layers separate
- Make pipelines idempotent
- Prefer incremental models for cost efficiency
- Use schema evolution handling
- Add data quality checks early
- Automate secrets, IAM, and infra with Terraform
- Use environment promotion: dev → staging → prod
- Monitor freshness, volume, and anomaly detection
If you want a simple modern setup
Snowflake
- Fivetran → Snowflake raw
- dbt → curated models
- Dagster → orchestration
- GitHub Actions → CI/CD
- Terraform → infra
BigQuery
- Fivetran / Datastream → BigQuery raw
- dbt or Dataform → curated models
- Cloud Composer or Dagster → orchestration
- GitHub Actions → CI/CD
- Terraform → infra
If you want, I can also provide:
- a reference architecture diagram for both platforms,
- a sample pipeline folder structure,
- or a production-ready implementation plan for Snowflake or BigQuery.
Brands