Data quality
Bad data stops at the check.
Checks run on every load and every model. When one fails, what's downstream waits, and you hear about it before anyone else does.
/* @bruin
name: mart.orders
materialization:
type: table
columns:
- name: order_id
checks: [not_null, unique]
- name: amount
checks: [positive]
@bruin */
select order_id, customer_id, amount, status
from raw.orders
where status != 'test'$ bruin run assets/orders.sql
- mart.orders1.8s
- not_null · order_id
- unique · order_id
- positive · amount3 rows below 0
- mart.revenue_dailyskipped, upstream check failed
Built in
The usual checks, one line each.
Declare them on a column, next to the model. They run after every load.
not_null
No empty values in the column.
unique
No duplicates, for keys and IDs.
accepted_values
Only the values you list, like a status.
pattern
Matches a pattern, like an email or SKU.
positive
Greater than zero, for amounts and counts.
non_negative
Zero or more, for balances and quantities.
min
Never below a floor you set.
max
Never above a ceiling you set.
Custom checks
Anything else, in SQL.
Freshness, volume or a business rule is one query away. Template it on the run's dates and it tests only the window that ran, not the whole table.
/* @bruin
name: mart.orders
custom_checks:
- name: orders arrived in the last 2 hours
value: 1
query: |
select max(created_at) > current_timestamp - interval '2 hours'
from mart.orders
- name: daily volume looks normal
value: 1
query: |
select count(*) between 5000 and 50000
from mart.orders
where order_date between '{{ start_date }}' and '{{ end_date }}'
- name: refunds never exceed the order
value: 0
query: |
select count(*) from mart.orders
where refund_amount > amount
@bruin */
select * from stg.ordersGates
Blocking by default. Advisory by choice.
A failing check stops everything downstream of it. Mark slow or nice-to-have checks non-blocking and the run carries on, with the warning on record.
/* @bruin
name: mart.orders
columns:
- name: order_id
checks:
- name: unique # blocking: a failure stops downstream
- name: email
checks:
- name: pattern
value: '^[^@]+@[^@]+$'
blocking: false # advisory: warn and carry on
@bruin */$ bruin run assets/marts/orders.sql
- mart.orders2.4s
- unique · order_id
- pattern · email12 rows, non-blocking
- mart.revenue_dailyran
1 warning, 0 failures. Downstream ran.
Alerts
You hear first. With the reason.
A failed check alerts the channels you set up: which check, how many rows, what it held back and who owns it.
Bruin APP 6:02 AM
Quality check failed: positive on mart.orders.amount
- Rows
- 3 rows below 0
- Held
- mart.revenue_daily, Revenue dashboard
- Owner
- [email protected]
Customer results
Numbers from teams on Bruin.
The platform
Part of the Bruin platform.
Checks guard every layer, from the raw load to the dashboard someone presents on Monday.
Sources
DatabasesWarehousesApps & APIsFiles & storageStreams & webhooksWeb scrapingMove
Data IngestionYour stack, your call
Replace the modern data stack.
One layer or every layer. Keep what works, swap what doesn't.
Today
On Bruin
Layer
What teams run today
On Bruin
Ingestion
On Bruin: Data Ingestionon open-source ingestr
Transformation & orchestration
On Bruin: SQL & Python + Bruin Cloud
Quality, lineage & catalog
On Bruin: Data Quality + Data Governance
BI & dashboards
Power BI
On Bruin: AI Dashboards + Data Apps
AI on your data
On Bruin: AI Data Analyst + Scheduled Agents
Frequently asked
Questions about data quality.
Which checks are built in?
Column checks such as not_null, unique, accepted_values, pattern, positive, non_negative, min and max, declared on the column in the asset file. The docs list every check and its options.
Can I write my own checks?
Yes. A custom check is a SQL query with an expected result, so freshness, volume, reconciliation between tables or any business rule can be a check. Checks can use the run's start and end dates to test only the new data.
What happens when a check fails?
Checks are blocking by default: the asset is marked failed and everything downstream of it waits, so a bad load never reaches the models and dashboards built on it. Lineage shows exactly what was held.
Can a check warn without stopping the pipeline?
Yes. Mark it non-blocking and the run carries on, with the failure recorded and alerted. It suits slow checks or rules you are still tuning.
Do checks scan the whole table every time?
Only if you write them that way. Checks can be templated with the run's date window, so an incremental load is tested on the rows it just loaded.
Where do alerts go?
To the channels you configure, such as a Slack channel. The alert names the check, the asset, how many rows failed and what was held downstream.
Do we still need a separate data quality tool?
For rule-based testing, usually not: checks live in the same file and run in the same pipeline as the model, so there is no second tool to keep in sync. Bruin also reads the tables your existing tools produce, so you can move over one pipeline at a time.
Is it open source?
Yes. Quality checks are part of the open-source Bruin CLI and run locally or in CI. Bruin Cloud adds scheduling, alerting, run history and lineage.
Catch it before the board deck.
$100 in credits and 50 AI tasks. No credit card.
A demo walks through your own data.