Prompt
I'm building a data movement layer for analysts and engineers, what should I use?
Latest observation
It depends on what kind of “data movement” you mean, because the best tool changes a lot by use case.
Quick recommendation by scenario
1) You want a general-purpose, reliable ELT/ETL layer for analysts + engineers
Use: Airbyte + dbt + your warehouse
- Airbyte for moving data from SaaS/apps/databases into your warehouse
- dbt for transformations, modeling, tests, and documentation
- Warehouse like Snowflake, BigQuery, Redshift, or Databricks
This is a strong default if you want:
- lots of connectors
- SQL-friendly transformations
- separation of ingestion and transformation
- something your analysts can understand and your engineers can operate
2) You need orchestration of complex pipelines, dependencies, and scheduling
Use: Dagster or Airflow
- Dagster if you want modern developer experience, strong asset-based modeling, and better ergonomics
- Airflow if you need a very mature ecosystem and already have Python-heavy pipeline orchestration
Usually this sits on top of ingestion tools, not instead of them.
3) You want low-code / no-code data integration for business users
Use: Fivetran, Workato, Make, or Zapier
- Fivetran for managed ELT
- Workato for enterprise automation and integrations
- Make/Zapier for lightweight workflow automation
Good when analyst/product ops teams need speed, less engineering ownership, and you can accept vendor cost.
4) You need event streaming or real-time movement
Use: Kafka / Redpanda / Kinesis / Pub/Sub
- Kafka/Redpanda for event-driven architecture
- Kinesis if you’re on AWS
- Pub/Sub if you’re on GCP
This is for operational data movement, not just analytics sync.
5) You need custom application-to-application or database-to-database sync
Use: Python + SQL + orchestration If the use case is specific and not connector-heavy, a custom service can be best:
- Python scripts or services
- API clients
- retries/idempotency
- queues
- observability
This is often the right answer for edge cases, but not as a universal platform.
My default stack recommendation
If you’re building a data movement layer for both analysts and engineers, I’d usually recommend:
- Airbyte for ingestion
- dbt for transformation
- Dagster for orchestration
- Snowflake/BigQuery/Databricks as the destination
- Great Expectations or dbt tests for quality checks
- OpenLineage/Marquez or built-in metadata for lineage if needed
That gives you a good balance of:
- usability
- extensibility
- analyst friendliness
- engineering control
- lower lock-in
How to choose
Ask these questions:
Choose managed tools like Fivetran if:
- you want minimal ops
- your team is small
- budget is less of a concern
- connectors matter more than customization
Choose Airbyte if:
- you want broad connector coverage
- you want open-source or hybrid control
- you need a more cost-efficient option at scale
- you expect custom sources/destinations
Choose Dagster/Airflow if:
- your movement layer needs dependencies and workflow logic
- you need retries, SLAs, backfills, and scheduling
- you want a platform engineers can extend
Choose Kafka/Redpanda if:
- you need real-time/event-based movement
- multiple downstream consumers need the same data
- durability and decoupling matter
If you want one concrete answer
For most teams today: Airbyte + dbt + Dagster + a cloud warehouse is a very solid stack.
If you want the simplest managed experience: Fivetran + dbt + a cloud warehouse.
If you tell me:
- your data sources
- batch vs real-time
- team size
- cloud provider
- whether analysts will self-serve
…I can recommend a much more specific stack.