Impact analysis for a schema change is the step where, before you rename, drop, retype, or repurpose a column or table, you find every model, quality check, dashboard, and downstream consumer that depends on it, so the change ships with its consequences known rather than discovered. It is the data equivalent of "find all references" in a code editor, and it is only possible with lineage at the column level, because the question is never "who reads this table" but "who reads this column". Bruin answers it from the command line on the SQL it parses; dbt, DataHub, Atlan, and Airflow each give a partial version from a different source.
Why table-level lineage is not enough
Say raw.orders has forty columns and twelve downstream models. You are renaming amount to amount_usd. Table-level lineage says twelve models are affected. Column-level lineage says three of them read amount, one check asserts it is non-negative, and two dashboards show it. The other nine models and their dashboards are untouched.
| Table-level | Column-level | |
|---|---|---|
| Says | mart.orders depends on raw.orders | mart.orders.order_total is computed from raw.orders.amount and raw.orders.currency |
| For a column rename | Flags every consumer of the table | Flags only consumers of that column |
| Source | Declared dependencies, or query logs at table grain | Parsed SQL, or query logs at column grain |
| Good enough for | "Can I drop this table?" | "Can I change this column?" |
Most impact analysis in practice is done on the second question, which is why column-level lineage is the prerequisite.
Running it from the command line
In a Bruin project the lineage is derived from the SQL, so the analysis is a command rather than a catalog visit:
bruin lineage raw.orders --full --output json
The output lists every downstream asset, and for each one the columns it reads from raw.orders and the columns it produces from them. Filter for amount and you have the three models, the check, and the assets that feed the two dashboards. Then make the change and let validation find what you missed:
bruin validate ./pipeline.yml
A model that still selects raw.orders.amount fails validation before anything runs. In CI that is a failed pull request check, which is the point: the analysis happens before the merge, not after the dashboard goes blank. The workflow that runs it on every pull request is in running data pipelines in CI/CD with GitHub Actions.
The preventive version: contracts and checks
Impact analysis finds consequences before a change you control. Two related practices handle changes you do not control, or reach consumers outside the repository.
Column checks catch silent changes at run time. If a source system retypes amount from decimal to string, no SQL reference breaks, but non_negative on order_total will, and the run stops before the bad value reaches a dashboard. Declaring checks on the columns downstream teams depend on turns those columns into a tested interface.
A data contract makes the interface explicit. For tables read by other teams or tools, the producing asset declares its columns, types, and checks, and a policy rule requires it:
# policy.yml
rulesets:
- name: tier-one-contracts
selector:
- tag: tier-1
rules:
- asset-has-columns
- asset-has-checks
- asset-has-owner
bruin validate fails if a tier-one asset is missing its columns, so the contract cannot be quietly dropped. The fuller treatment is in what is a data contract.
What the other tools give you
- dbt shows model-level lineage from
ref()in every project and column-level lineage in dbt Cloud; impact analysis on a column is a Cloud feature, or a manual grep in Core. - DataHub, Atlan, Collibra derive column-level lineage from warehouse query logs and show impact in a catalog UI. Accurate for what has run recently; blind to a model that has not run since the logs were collected, and a separate system from the code you are changing.
- Airflow knows task dependencies, not data dependencies. Impact analysis means reading the SQL inside each task.
- Bruin derives the graph from the code, so it is current whenever the code is, runs from the CLI, and fails validation on a broken reference. The trade-off is that it only sees the assets in the project, so consumers outside it, a BI tool reading the warehouse directly, need a contract rather than lineage.
For the tool comparison in full, see the best data lineage tools in 2026.