A data contract is a formal agreement between the team that produces a dataset and the teams that consume it: these columns, these types, these guarantees, this owner, and this is how changes will be announced. It exists because the schema alone is not a promise. A table can have a customer_id column today and a renamed one tomorrow, and nothing in the warehouse stops that. A contract makes the promise explicit and, if it is written as code, makes breaking it fail a build. In Bruin, the asset definition is the contract and a policy file enforces it; dbt has model contracts; Soda has a contracts specification.
What a contract contains
The useful contracts are short. Five parts:
| Part | What it states | Example |
|---|---|---|
| Schema | Column names and types | order_id integer, status string, order_total float |
| Guarantees | Rules the data will always satisfy | order_id not null and unique; status in four values; order_total non-negative |
| Freshness | How often the table updates and how late is too late | Hourly, stale after two hours |
| Ownership | Who is accountable and where to ask | [email protected] |
| Change policy | What is breaking, how it is announced, how long consumers have | Removing or retyping a column is breaking; announced one release ahead |
Everything except the change policy can be expressed as metadata and checks on the asset that produces the table. That is the whole argument for contracts as code: the document version rots, the code version is enforced.
Why contracts as documents fail
Most first attempts at data contracts are a Confluence page per table. They fail for three reasons. Nobody updates the page when the pipeline changes, so within a quarter the page and the table disagree. Nothing enforces the page, so a producer can break it without knowing. And consumers cannot find the page from the table, so they read the schema instead and the contract is never consulted.
A contract in the pipeline code inverts all three. It changes in the same pull request as the table. It is enforced by the build. And it is where the consumer already looks.
A contract as code
In Bruin, the asset header carries the schema, the guarantees, and the owner:
/* @bruin
name: mart.orders
type: sf.sql
owner: [email protected]
description: One row per order, deduplicated, enriched with customer region. Refunds keep their row with status = 'refunded'.
depends: [raw.orders, mart.customers]
columns:
- name: order_id
type: integer
primary_key: true
description: Order identifier from the source system, stable across updates.
checks:
- name: not_null
- name: unique
- name: customer_id
type: integer
foreign_key:
table: mart.customers
column: customer_id
checks:
- name: not_null
- name: relationships
- name: status
type: string
checks:
- name: accepted_values
value: [placed, paid, shipped, refunded]
custom_checks:
- name: fresh within 2 hours
query: SELECT max(updated_at) > current_timestamp - interval '2 hours' FROM mart.orders
value: 1
@bruin */
SELECT ...
That header is the contract. The relationships check is the contract between two assets: it fails when a foreign key points at a row the referenced table does not have.
Enforcing it in two places
At build time, with policies. A contract that a producer can simply omit is not a contract. Bruin reads a policy.yml at the project root and applies its rules on bruin validate and automatically before every bruin run:
rulesets:
- name: contracts
rules:
- asset-has-owner
- asset-has-description
- asset-has-columns
- asset-has-primary-key
- asset-has-checks
An asset without columns, a primary key, or checks fails validation before it can run, locally and in CI. Custom rules use boolean expressions over the asset metadata, so a team can require, for example, that every asset in the mart schema declares a freshness check. dbt's model contracts (contract: enforced: true with declared column types) and Soda's contract files play the same role in their ecosystems.
At run time, with blocking checks. The guarantees only mean something if breaking them has a consequence. Blocking checks provide it: the producer's own pipeline stops when a promised guarantee fails, before any consumer sees the rows. The bad data is contained in the producer's run rather than discovered in a consumer's dashboard.
On change, with lineage. The change-policy half of the contract needs to know who is affected. If the framework derives lineage from the pipeline code, a pull request that removes or retypes a column can fail validation with the list of downstream assets that read it. Bruin builds that graph from the asset definitions and the SQL it parses, and bruin validate fails the pull request that breaks a dependent asset. That turns "announce breaking changes" from a courtesy into a build step.
Where to start
Pick the five tables the business reads most. Declare columns, types, owner, and the four checks that matter (not null and unique on the key, accepted values on the status column, a freshness check). Turn on a policy that refuses to run an asset without them. That is the lightweight contract, it takes an afternoon, and it removes the most common incident in a small data team: the renamed column that breaks a report a week later. The heavier machinery, schema registries and approval workflows, is for organisations with many producing teams and can wait until you have them.
For the surrounding practice see what is a data quality check, what is data freshness, and the full guide to data quality and testing strategies for modern pipelines.