Prompt
What should I use instead of custom Python scripts for ETL?
Latest observation
If you’re moving away from custom Python scripts for ETL, the best replacement depends on how much complexity, scale, and governance you need. Here are the main options, from simplest to most robust:
1. Managed ETL/ELT platforms
Best if you want to reduce maintenance and move quickly.
- Airbyte / Fivetran / Stitch — good for pulling data from SaaS apps, databases, and APIs into a warehouse.
- AWS Glue / Azure Data Factory / Google Cloud Data Fusion — good if you’re already in a cloud ecosystem.
- Databricks — strong if your pipelines involve heavy transformation and large-scale processing.
Use these if: your ETL is mostly standard connectors, scheduled loads, and straightforward transformations.
2. Workflow orchestration tools
Best if the issue isn’t the Python itself, but the lack of structure around scripts.
- Apache Airflow
- Prefect
- Dagster
These don’t replace transformation logic entirely, but they replace “cron + scripts + retries + manual monitoring” with proper orchestration.
Use these if: you still need custom code, but want reliability, scheduling, retries, lineage, and observability.
3. SQL-based ELT
Best if your data already lands in a warehouse.
- dbt for transformations in SQL
- Pair with Airbyte/Fivetran for ingestion
This is one of the most common modern patterns: extract/load first, transform in-warehouse.
Use these if: most of your work is transforming relational data and you want version-controlled, testable transformations.
4. Low-code / no-code ETL tools
Best for small teams or less technical users.
- Talend
- Informatica
- Microsoft Power Automate / SSIS
- Zapier / Make for lightweight app-to-app workflows
Use these if: you want less engineering overhead and your workflows are not very complex.
5. Data integration platforms with governance
Best for enterprises with compliance and lineage needs.
- Informatica
- Talend
- Matillion
- Boomi
Use these if: you need enterprise controls, metadata management, and standardized operations.
A practical recommendation
For many teams, the modern replacement for custom Python ETL is:
Airbyte or Fivetran + dbt + Airflow/Prefect if needed
That gives you:
- easier ingestion
- SQL-based transformations
- better scheduling/monitoring
- less code to maintain
Rule of thumb
- Simple app/database syncs: Fivetran/Airbyte
- Warehouse transformations: dbt
- Complex scheduling and dependencies: Airflow/Prefect/Dagster
- Heavy compute / big data: Databricks / Spark
- Enterprise governance: Informatica / Talend / Matillion
If you want, I can suggest the best replacement based on your setup:
- data sources,
- target warehouse,
- volume,
- and whether you prefer low-code or code-first tools.