Self-healing pipelines
The pipeline broke at 2am. Fixed before standup.
It finds the cause, tests the fix in dev, and opens a PR.
stripe_paymentspipeline recovered
Root cause: source API was down
Re-ranBackfilled
revenue_dailymetadata fixed
Missing owner and team tags
Tags addedReview PR #128
Trusted by forward-thinking teams
How it works
Something breaks. Now Bruin's on it.
$ bruin run --tag quality-check-investigate self-heal-demo/demo-pipeline
finance.order_margin
└── net_amount.positive - column 'net_amount' has 1 non-positive values
causeorder_margin.sql groups adjustments by customer_id instead of order_id
fixgroup and join on order_id, on a branch
verifySELECT order_id, net_amount FROM finance.order_margin WHERE net_amount <= 0;
no rows
It catches the failure
A failed check or a failed run.
It finds the cause
Reading run logs, following lineage, querying tables.
It tests the fix in dev
The smallest safe fix, tested outside production.
You approve, it verifies
A pull request to review, then a check on the data, not the exit code.
It leaves a lesson
Every incident should leave behind a check, a description, or a rule.
Customer results
Numbers from teams on Bruin.
Guardrails
You set the limits. It stays inside them.
Outside production
A database it can write to that is not production, plus read-only access to your production sources.
By blast radius
Low-risk, reversible changes can auto-open a pull request. Finance metrics, definitions and write-back fields ask for approval.
Credentials decide
Prompts and AGENTS.md shape what the agent tries to do. Credentials and tokens decide what happens when it tries.
Never silent
Nothing changes production silently.
The platform
Part of the Bruin platform.
It works because checks, column-level lineage and plain-file definitions are built in.
01 · Move
02 · Model & govern
03 · Use
Frequently asked
Questions about self-healing.
What is a self-healing data pipeline?
A self-healing data pipeline is one whose failures are diagnosed and repaired by an AI agent with access to the pipeline, within limits you set. The agent detects a problem, diagnoses the cause, proposes the smallest safe fix, tests it outside production, and opens a pull request or asks for approval. It does not mean an agent silently changing production. The pipeline does not repair itself; an agent with the right context and guardrails does.
Is self-healing just automatic retries?
No. A retry repeats the same action blindly and works for transient failures like a rate limit or dropped connection. A self-healing pipeline first works out what actually broke. Most of that work is evidence gathering - reading run logs, following lineage, querying tables - which is something agents are genuinely good at. It can tell a rate limit from a renamed column and correct the cause rather than rerun the failing step.
How is self-healing different from alerting or observability?
Classic alerting tells you something broke. Self-healing goes further: the agent arrives with the incident already investigated - the data analyzed, the likely cause identified, and a proposed fix in a pull request ready for review. Most observability tools detect and some diagnose, but they sit outside your repo, so they cannot fix, verify, or write the lesson back.
How is self-healing different from Scheduled Agents?
Self-healing works on failures: when a pipeline run or a quality check fails, the agent finds the cause, tests the smallest fix outside production and opens a pull request. Scheduled Agents work on a schedule: they run the recurring work your business asks for, such as briefs, reports, decks and metric alerts.
How much does the agent do on its own?
Autonomy is a setting, not a law. It ranges from diagnosis-only to running the whole loop unattended, and where you sit is governed by permissions, instructions, and risk appetite. A good rule is to escalate by blast radius: low-risk, reversible changes can auto-open a pull request, while anything touching finance metrics, definitions, or write-back fields should ask for approval.
How does the agent fix failures without breaking production?
You give it a database it can write to that is not production, plus read-only access to your production sources. It tests fixes in that isolated space and only then applies them, so it never experiments on live data. You also set the scope it is allowed to work within, so it stays inside the assets and pipelines you assigned it.
What does a pipeline need before Bruin can self-heal it?
Three things in the platform: declared quality checks at the asset level so the agent knows what good data looks like, column-level lineage so it can scope the blast radius of a change, and a definition format the agent can safely read and edit. Where lineage and quality are external bolt-ons, a pipeline cannot self-heal without integrating multiple systems.
Do I need Bruin Cloud, or can I run it locally?
Either. The local path runs the agent in your IDE, terminal, or a coding app such as Cursor over the Bruin MCP. The Cloud path uses a shared agent in Bruin Cloud when the whole team needs one. Both paths use the same project notes and the same testing process.
Can the Bruin Cloud agent open pull requests?
Yes, once you connect Git. Project access is not Git access: branches and pull requests need a separate Git integration with write permission.
How long does it take to set up?
Plan for about 70 minutes. Most of that time goes into building context across the pipelines the agent supports - importing existing tables as assets, enriching the definitions with AI, and adding the business rules only your team knows. The guide uses finance.order_margin as one small follow-along example, but the same process applies to every asset the agent should support.
Wake up to fixes, not failures.
$100 in credits and 50 AI tasks. No credit card.
A demo walks through your own data.
