Technical
6 min read

What Is a Data Orchestrator? And When a Small Team Does Not Need One

A data orchestrator schedules pipeline tasks, runs them in dependency order, retries failures, and records what ran. This explainer covers what Airflow, Dagster, and Prefect actually do, what it costs to operate one, and why a pipeline framework that resolves its own dependency graph, such as Bruin, removes the orchestrator for teams whose DAGs are mostly ELT glue.

What Is a Data Orchestrator? And When a Small Team Does Not Need One

A data orchestrator is the system that runs a pipeline's tasks in the right order, on a schedule, retries the ones that fail, and records what ran and when. Apache Airflow is the archetype: you describe tasks and their dependencies as a DAG, and the orchestrator's scheduler, workers, and metadata database do the rest. Dagster and Prefect are the modern alternatives. The question a small team should ask first is not which orchestrator but whether they need one, because a pipeline framework that resolves its own dependency graph, such as Bruin, removes the layer entirely for pipelines made of loads, transformations, and checks.

What an orchestrator actually does

Four jobs, and every orchestrator does all four:

JobWhat it meansAirflow's implementation
Dependency resolutionRun B after A, C after bothThe DAG, written in Python
SchedulingStart the DAG at 02:00 daily, or when a file landsThe scheduler process
ExecutionRun the tasks, in parallel where possible, with retriesWorkers, plus an executor to distribute them
BookkeepingWhich run did what, when, with what resultThe metadata database and the web UI

The value is real. Without those four, a pipeline is a cron entry and a prayer. The cost is also real: each of the four is a component someone operates.

Why it became a layer of its own

Orchestrators exist because the modern data stack was assembled from single-purpose tools. Fivetran loads, dbt transforms, Great Expectations checks, a BI tool reads. None of them knows about the others, so a fifth tool has to know about all of them and call them in order. That is what most Airflow DAGs in data teams are: glue that runs an ingestion sync, then a dbt run, then a check, then a dashboard refresh.

That glue has a price. A self-hosted Airflow is a scheduler, workers, a metadata database, a web server, a deployment image, secrets handling, and someone who understands all of it. Managed Airflow removes the servers but keeps the concepts and the bill: Amazon MWAA's own pricing example for a small environment with typical retention comes to about $449 a month before the warehouse and the tools it orchestrates. Astronomer is the other managed route. Either way the team is paying to run a tool whose job is to call other tools.

When you do not need one

If the DAG mostly runs data work, loads, SQL and Python transformations, quality checks, and a schedule, then the orchestrator is doing dependency resolution for tasks that could declare their dependencies themselves. A pipeline framework that owns all of those steps does not need a separate system to order them.

That is the design of Bruin. Each asset is a file that declares what it depends on, and the SQL is parsed for the rest:

/* @bruin
name: mart.orders
type: sf.sql
depends: [raw.orders, mart.customers]
materialization:
  type: table
columns:
  - name: order_id
    checks:
      - name: not_null
      - name: unique
@bruin */

bruin run ./pipeline.yml resolves the graph across ingestion assets, SQL and Python models, and checks, and runs them in order with parallelism where the graph allows. The schedule is one line in pipeline.yml. Locally or in CI that is the whole orchestrator; Bruin Cloud runs the same project on the schedule with retries, backfills, alerts to Slack or Teams, lineage, and a catalog. There is no scheduler, worker pool, or metadata database to operate, because the dependency graph is a property of the pipeline rather than a separate program.

When you do

Keep, or adopt, a general orchestrator when the DAG coordinates things that are not data pipelines: infrastructure jobs, ML training runs, calls to external systems, multi-day business workflows. Dagster fits data platforms that want software-defined assets across many tools, Prefect fits Python-heavy teams that want the lightest operational footprint, and Airflow fits teams with hundreds of working DAGs and someone who owns it. Temporal and Argo are for application workflows and Kubernetes batch, respectively, and are not data orchestrators at all.

The rule of thumb: count what the DAGs do. If most tasks are load, transform, check, and schedule, a framework that owns those four removes the layer. If most tasks are something else, an orchestrator is the right tool and the data pipeline becomes one of the things it calls.

For the tool comparison see the best data pipeline tools in 2026, and for the migration path, alternatives to Airflow.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.