Prompt
I'm building near-real-time replication from MySQL to a warehouse, what should I use?
Latest observation
For near-real-time MySQL → warehouse replication, the most common pattern is:
Best option: CDC-based streaming
Use Change Data Capture (CDC) from MySQL binlogs, then stream changes into your warehouse.
Typical architecture
MySQL → CDC connector → message/stream layer → warehouse loader
Examples:
- Debezium + Kafka / Kafka Connect
- AWS DMS
- Fivetran / Airbyte / Hevo / Stitch for managed ELT
- Google Datastream if you’re in GCP
- Azure Data Factory / Fabric options in Microsoft ecosystems
What I’d recommend by scenario
1) You want open source / control
Debezium + Kafka
Best if you need:
- low latency
- schema-change handling
- replayability
- operational control
Good warehouses to land in:
- Snowflake
- BigQuery
- Redshift
- Databricks
- ClickHouse
2) You want simplest managed service
Fivetran or Airbyte Cloud
Best if you want:
- minimal ops
- fast setup
- automatic schema handling
- decent reliability
Tradeoff: higher cost, less control.
3) You’re on AWS and want native-ish
AWS DMS
Good for:
- MySQL binlog replication
- near-real-time sync
- simple setup into S3, Redshift, Aurora, etc.
Tradeoff: less elegant for complex transforms and some edge cases.
4) You’re on GCP
Datastream → BigQuery / GCS
Very solid for MySQL CDC into GCP.
Important decision criteria
You should choose based on:
- Latency target: seconds vs minutes
- Data volume / change rate
- Schema evolution frequency
- Need for exactly-once or at-least-once
- Warehouse type
- Operational burden tolerance
- Cost
My practical recommendation
If you’re starting fresh and want a robust, scalable setup:
- Open-source/engineering-heavy: Debezium → Kafka → warehouse sink
- Low-ops/production-fast: Fivetran or Airbyte Cloud
- AWS-native: AWS DMS
- GCP-native: Datastream
If you want “near-real-time” with minimal fuss
I’d usually say:
MySQL binlog CDC + managed connector into your warehouse
because it avoids polling, scales better, and is the standard approach for low-latency replication.
If you tell me:
- your warehouse (Snowflake/BigQuery/Redshift/etc.)
- cloud (AWS/GCP/Azure/on-prem)
- acceptable latency
- budget/ops preference
…I can recommend a specific stack.