A data orchestrator is the system that runs a pipeline's tasks in the right order, on a schedule, retries the ones that fail, and records what ran and when. Apache Airflow is the archetype: you describe tasks and their dependencies as a DAG, and the orchestrator's scheduler, workers, and metadata database do the rest. Dagster and Prefect are the modern alternatives. The question a small team should ask first is not which orchestrator but whether they need one, because a pipeline framework that resolves its own dependency graph, such as Bruin, removes the layer entirely for pipelines made of loads, transformations, and checks.
What an orchestrator actually does
Four jobs, and every orchestrator does all four:
| Job | What it means | Airflow's implementation |
|---|---|---|
| Dependency resolution | Run B after A, C after both | The DAG, written in Python |
| Scheduling | Start the DAG at 02:00 daily, or when a file lands | The scheduler process |
| Execution | Run the tasks, in parallel where possible, with retries | Workers, plus an executor to distribute them |
| Bookkeeping | Which run did what, when, with what result | The metadata database and the web UI |
The value is real. Without those four, a pipeline is a cron entry and a prayer. The cost is also real: each of the four is a component someone operates.
Why it became a layer of its own
Orchestrators exist because the modern data stack was assembled from single-purpose tools. Fivetran loads, dbt transforms, Great Expectations checks, a BI tool reads. None of them knows about the others, so a fifth tool has to know about all of them and call them in order. That is what most Airflow DAGs in data teams are: glue that runs an ingestion sync, then a dbt run, then a check, then a dashboard refresh.
That glue has a price. A self-hosted Airflow is a scheduler, workers, a metadata database, a web server, a deployment image, secrets handling, and someone who understands all of it. Managed Airflow removes the servers but keeps the concepts and the bill: Amazon MWAA's own pricing example for a small environment with typical retention comes to about $449 a month before the warehouse and the tools it orchestrates. Astronomer is the other managed route. Either way the team is paying to run a tool whose job is to call other tools.
When you do not need one
If the DAG mostly runs data work, loads, SQL and Python transformations, quality checks, and a schedule, then the orchestrator is doing dependency resolution for tasks that could declare their dependencies themselves. A pipeline framework that owns all of those steps does not need a separate system to order them.
That is the design of Bruin. Each asset is a file that declares what it depends on, and the SQL is parsed for the rest:
/* @bruin
name: mart.orders
type: sf.sql
depends: [raw.orders, mart.customers]
materialization:
type: table
columns:
- name: order_id
checks:
- name: not_null
- name: unique
@bruin */
bruin run ./pipeline.yml resolves the graph across ingestion assets, SQL and Python models, and checks, and runs them in order with parallelism where the graph allows. The schedule is one line in pipeline.yml. Locally or in CI that is the whole orchestrator; Bruin Cloud runs the same project on the schedule with retries, backfills, alerts to Slack or Teams, lineage, and a catalog. There is no scheduler, worker pool, or metadata database to operate, because the dependency graph is a property of the pipeline rather than a separate program.
When you do
Keep, or adopt, a general orchestrator when the DAG coordinates things that are not data pipelines: infrastructure jobs, ML training runs, calls to external systems, multi-day business workflows. Dagster fits data platforms that want software-defined assets across many tools, Prefect fits Python-heavy teams that want the lightest operational footprint, and Airflow fits teams with hundreds of working DAGs and someone who owns it. Temporal and Argo are for application workflows and Kubernetes batch, respectively, and are not data orchestrators at all.
The rule of thumb: count what the DAGs do. If most tasks are load, transform, check, and schedule, a framework that owns those four removes the layer. If most tasks are something else, an orchestrator is the right tool and the data pipeline becomes one of the things it calls.
For the tool comparison see the best data pipeline tools in 2026, and for the migration path, alternatives to Airflow.