Data quality

Bad data stops at the check.

Checks run on every load and every model. When one fails, what's downstream waits, and you hear about it before anyone else does.

Built in

The usual checks, one line each.

Declare them on a column, next to the model. They run after every load.

not_null

No empty values in the column.

unique

No duplicates, for keys and IDs.

accepted_values

Only the values you list, like a status.

pattern

Matches a pattern, like an email or SKU.

positive

Greater than zero, for amounts and counts.

non_negative

Zero or more, for balances and quantities.

min

Never below a floor you set.

max

Never above a ceiling you set.

Custom checks

Anything else, in SQL.

Freshness, volume or a business rule is one query away. Template it on the run's dates and it tests only the window that ran, not the whole table.

/* @bruin
name: mart.orders

custom_checks:
  - name: orders arrived in the last 2 hours
    value: 1
    query: |
      select max(created_at) > current_timestamp - interval '2 hours'
      from mart.orders

  - name: daily volume looks normal
    value: 1
    query: |
      select count(*) between 5000 and 50000
      from mart.orders
      where order_date between '{{ start_date }}' and '{{ end_date }}'

  - name: refunds never exceed the order
    value: 0
    query: |
      select count(*) from mart.orders
      where refund_amount > amount
@bruin */

select * from stg.orders

Gates

Blocking by default. Advisory by choice.

A failing check stops everything downstream of it. Mark slow or nice-to-have checks non-blocking and the run carries on, with the warning on record.

/* @bruin
name: mart.orders

columns:
  - name: order_id
    checks:
      - name: unique          # blocking: a failure stops downstream
  - name: email
    checks:
      - name: pattern
        value: '^[^@]+@[^@]+$'
        blocking: false       # advisory: warn and carry on
@bruin */

$ bruin run assets/marts/orders.sql

  1. mart.orders2.4s
  2. unique · order_id
  3. pattern · email12 rows, non-blocking
  4. mart.revenue_dailyran

1 warning, 0 failures. Downstream ran.

Alerts

You hear first. With the reason.

A failed check alerts the channels you set up: which check, how many rows, what it held back and who owns it.

Customer results

Numbers from teams on Bruin.

The platform

Part of the Bruin platform.

Checks guard every layer, from the raw load to the dashboard someone presents on Monday.

Your stack, your call

Replace the modern data stack.

One layer or every layer. Keep what works, swap what doesn't.

Today

On Bruin

Ingestion

On Bruin: Data Ingestionon open-source ingestr

Transformation & orchestration

On Bruin: SQL & PythonBruin Cloud

Quality, lineage & catalog

On Bruin: Data QualityData Governance

BI & dashboards

  • Power BI

On Bruin: AI DashboardsData Apps

AI on your data

On Bruin: AI Data AnalystScheduled Agents

Frequently asked

Questions about data quality.

Which checks are built in?

Column checks such as not_null, unique, accepted_values, pattern, positive, non_negative, min and max, declared on the column in the asset file. The docs list every check and its options.

Can I write my own checks?

Yes. A custom check is a SQL query with an expected result, so freshness, volume, reconciliation between tables or any business rule can be a check. Checks can use the run's start and end dates to test only the new data.

What happens when a check fails?

Checks are blocking by default: the asset is marked failed and everything downstream of it waits, so a bad load never reaches the models and dashboards built on it. Lineage shows exactly what was held.

Can a check warn without stopping the pipeline?

Yes. Mark it non-blocking and the run carries on, with the failure recorded and alerted. It suits slow checks or rules you are still tuning.

Do checks scan the whole table every time?

Only if you write them that way. Checks can be templated with the run's date window, so an incremental load is tested on the rows it just loaded.

Where do alerts go?

To the channels you configure, such as a Slack channel. The alert names the check, the asset, how many rows failed and what was held downstream.

Do we still need a separate data quality tool?

For rule-based testing, usually not: checks live in the same file and run in the same pipeline as the model, so there is no second tool to keep in sync. Bruin also reads the tables your existing tools produce, so you can move over one pipeline at a time.

Is it open source?

Yes. Quality checks are part of the open-source Bruin CLI and run locally or in CI. Bruin Cloud adds scheduling, alerting, run history and lineage.

Catch it before the board deck.

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.