Prompt
I'm building a CDC pipeline for operational data into BigQuery, what should I use?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
For a CDC pipeline into BigQuery, the best choice depends on your source systems, latency needs, and how much you want to operate yourself. Here are the common options, with a practical recommendation.
Best default choices
1) Managed CDC tool (recommended for most teams)
Use a CDC platform like:
- Fivetran
- Airbyte
- Matillion
- Qlik Replicate
- Striim
Why:
- Handles database log reading, schema evolution, retries, and backfills
- Lowest operational overhead
- Good if you want reliable ingestion without building custom infrastructure
Best for:
- Teams that want to move fast
- Standard relational sources like Postgres, MySQL, SQL Server, Oracle
- Moderate to high volume operational data
2) Debezium + Kafka / Pub/Sub + BigQuery
If you want more control and already run streaming infrastructure:
- Debezium captures changes from source DB logs
- Send events through Kafka or Google Pub/Sub
- Load into BigQuery using Dataflow, Beam, or custom consumers
Why:
- Flexible and extensible
- Good for complex event-driven architectures
- Can support multiple downstream consumers
Tradeoff:
- More engineering and ops work
- You own schema handling, ordering, deduplication, retries, and operational recovery
Best for:
- Platform teams
- Large-scale or multi-consumer architectures
- Need for near-real-time processing and custom logic
3) Native BigQuery ingestion patterns
For some Google Cloud-centered setups:
- Datastream for CDC from Oracle/MySQL/Postgres into BigQuery or Cloud Storage
- Dataflow for transformation and loading
- Pub/Sub if your app emits change events directly
Why:
- Deep integration with GCP
- Less glue code than building everything yourself
- Good path if your stack is already mostly on Google Cloud
Best for:
- GCP-native environments
- Teams wanting managed replication with lower maintenance
My practical recommendation
If you want the simplest and most reliable path:
Use a managed CDC tool or Google Datastream.
- Datastream if you’re on GCP and your source DB is supported
- Fivetran/Airbyte if you want broad source support and minimal engineering
- Debezium if you need full control and can operate the pipeline yourself
What to consider when choosing
Source database support
Make sure the tool supports:
- Your DB type
- Log-based CDC
- Schema changes
- Large initial snapshots
Delivery semantics
BigQuery loads are usually effectively at-least-once, so you’ll want:
- Primary keys
- Deduplication strategy
- Merge/upsert pattern into final tables
Latency
- Minutes: most managed CDC tools
- Seconds to sub-minute: Debezium/Kafka/Dataflow or Datastream depending on config
Data modeling in BigQuery
A strong pattern is:
- Raw landing tables for CDC events
- Curated tables built via
MERGE - Use partitioning/clustering for performance
Common architecture pattern
- Source DB emits changes via binlog/WAL
- CDC tool captures inserts/updates/deletes
- Land raw change events into BigQuery or Cloud Storage
- Transform into final tables with
MERGE - Handle deduplication, deletes, and schema changes
Quick recommendation by scenario
- Small team, want fast results: Fivetran or Airbyte
- GCP-first: Datastream
- Need custom streaming and multiple consumers: Debezium + Kafka/Pub/Sub + BigQuery
- Very large enterprise replication needs: Qlik Replicate / Striim
If you tell me:
- your source DBs,
- whether you’re on GCP,
- desired latency,
- expected volume,
I can recommend a specific architecture and tooling stack.