How do I build self-healing data pipelines?
Quality checks gate each asset, failures stop downstream runs, and retries with alerts recover transient errors, so a pipeline catches and contains its own breakages instead of shipping bad data. Bruin does this in one platform: ingestion, SQL and Python pipelines, quality checks, lineage, and an AI data analyst that answers in Slack, Microsoft Teams, Google Chat, WhatsApp, Discord, Telegram, email and the browser.
Command
bruin runDefined in
SQL + YAML
Works with
Bruin CLI + Bruin Cloud
What you get
How it works in code
columns:
- name: amount
checks:
- name: not_null
- name: non_negative
custom_checks:
- name: row_count_within_bounds
query: SELECT COUNT(*) BETWEEN 1000 AND 100000 FROM {{ this }}Run bruin run and Bruin blocks downstream assets when a check fails.
Give agents governed data
Bruin CLI and ingestr are on GitHub.
A demo walks through your own data.