Technical
6 min read

What Is Agentic Data Engineering? AI Agents That Build and Repair Pipelines

Agentic data engineering is the practice of letting an AI coding agent build, run, test, and repair data pipelines, with a human reviewing the diff. This explainer covers what an agent can actually do for a pipeline, what the pipeline has to look like for that to be safe, how MCP connects the agent, and where Bruin, dbt's MCP server, Databricks Genie, and Dagster Compass fit.

What Is Agentic Data Engineering? AI Agents That Build and Repair Pipelines

Agentic data engineering is the practice of letting an AI coding agent do the work of a data engineer, building pipelines, running them, testing them, and repairing them when they break, with a human reviewing the change rather than typing it. It became practical in 2026 for one reason: the agents got good enough, and the pipeline frameworks that keep everything as plain files gave them something to read. The agent is Claude Code, Cursor, or Codex. The framework is what makes the agent safe. Bruin is built to be operated by agents, with an MCP server that exposes the whole pipeline, and is the worked example below.

What an agent can actually do for a pipeline

Four jobs, in the order teams adopt them.

  1. Build. "Add a daily load of Stripe charges into the warehouse, model net revenue by plan, and add a not-null check on the charge id." The agent writes the ingestion asset, the SQL model, and the checks, runs them against a development target, and opens a pull request with the diff.
  2. Document. Descriptions for every asset and column, generated from the SQL and the upstream context, written into the asset file for review.
  3. Debug. A check failed at 03:00. The agent reads the failure, walks the lineage graph upstream to the asset whose input changed, and explains the cause before anyone is awake.
  4. Repair. The agent proposes the fix, re-runs the failed checks to prove it works, and either applies it or waits for a human, depending on how much autonomy the team has granted that pipeline.

None of this is a chatbot suggesting SQL. It is an agent with tools: read the project, run an asset, run the checks, read lineage, edit a file.

What the pipeline has to look like

An agent can only operate what it can read and verify. Three properties decide whether agentic data engineering works on a given stack or produces confident nonsense.

Everything is a file. Loads, transformations, checks, and the schedule are text in a repository with declared dependencies. An agent edits files well and clicks through UIs badly. A pipeline configured in a console is invisible to it.

Everything is testable. Quality checks that run with the asset and fail it are how the agent knows whether its change worked. Without checks, "the pipeline ran" is the only feedback, and that is not enough to trust a repair.

Everything is connected. Lineage derived from the code is how the agent finds the upstream cause of a downstream failure. Without it, debugging is guessing.

Bruin has all three because it was designed around them: assets are SQL, Python, and YAML files with a depends list; checks are declared on the columns and block by default; lineage is parsed from the SQL. The consequence is that bruin run, bruin validate, and bruin lineage are the same commands for a person and for an agent.

How the agent connects: MCP

The Model Context Protocol is the standard way a coding agent gets tools. A pipeline framework that ships an MCP server gives the agent structured access to the project instead of a file system it has to guess about. Bruin's MCP server lets an agent list assets, read their columns, checks, and dependencies, run assets and checks, read lineage, and query the semantic layer. Setup is one command in the project, after which Claude Code, Cursor, or Codex can be pointed at the pipeline.

The equivalents elsewhere: dbt's MCP server exposes models, tests, and runs for the transformation layer; Databricks Genie and the Lakeflow assistants operate inside the lakehouse; Dagster Compass answers questions over Dagster assets but does not build. The scope of the MCP server is the scope of what the agent can do.

How much autonomy

The mistake is treating this as all or nothing. Autonomy is a setting per pipeline. A reasonable ladder:

  • Propose only. The agent opens a pull request; a human merges. Start here for every pipeline.
  • Auto-apply with review. The agent applies fixes for a known class of failure, such as a schema drift that only needs a column mapping, and notifies the channel.
  • Auto-apply. For pipelines where a wrong fix costs less than a delayed one, the agent repairs and re-runs, with the diff logged.

The guardrails are the same at every rung: the checks have to pass before a change counts as fixed, and lineage has to show what else the change touched. The detail of the detect, diagnose, fix, verify loop is in what is a self-healing data pipeline, and the step-by-step setup is in how to build and maintain data pipelines with AI agents.

What it does not replace

The agent does not decide what the business needs, does not own the semantic definitions, and does not make a pipeline trustworthy that has no checks. Agentic data engineering moves the data engineer from typing the pipeline to reviewing it and from being paged to being informed. The team still owns the design, and the framework still has to make the design legible. That is the whole reason to pick one that keeps ingestion, transformation, checks, and lineage in one repository: it is the version an agent can operate end to end.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.