Column-level data lineage traces every column in a table back to the exact source columns and transformations that produced it, instead of stopping at the table. Table-level lineage says mart.orders depends on raw.orders. Column-level lineage says mart.orders.customer_region is raw.customers.region, reached through a join on customer_id, with nulls replaced by 'unknown'. The second is the one that answers the two questions data teams actually ask: what breaks if I change this column, and where did this wrong number come from. In Bruin, column lineage is derived automatically by parsing the SQL of every asset, so it is available before the code runs; catalogs such as OpenMetadata and DataHub reconstruct it from warehouse query logs afterwards.
Why table-level lineage is not enough
A table-level graph tells you that a dashboard sits three hops downstream of a raw table. It cannot tell you whether the column you are about to rename is one of the twelve that reach the dashboard or one of the forty that do not. So every schema change becomes either a meeting or a gamble.
Column-level lineage removes the gamble. Before the change, it lists exactly which downstream columns and assets read the one you are touching. After a break, it traces a wrong figure in a report back through each transformation to the source column that changed. Both uses are worth the same thing: not finding out from the business.
How column lineage is derived
There are two ways to build the graph, and the difference decides what the lineage is good for.
Parsing the pipeline code. The framework reads each asset's SQL, resolves the SELECT list, joins, and expressions, and records which input columns feed each output column. Python assets contribute through their declared inputs and outputs. Bruin works this way: the dependency graph comes from each asset's depends list and the SQL it parses, with no manual mapping, and bruin lineage prints it.
bruin lineage assets/mart/orders.sql --full
The command lists every upstream dependency the asset relies on and every downstream asset that relies on it, including indirect connections with --full; --output json returns the graph for other tools. Bruin Cloud renders the same graph across every pipeline in the organisation, down to the column, and shows cross-pipeline dependencies. Because it is derived from the code, the lineage for a change exists as soon as the change is written, which is what makes impact analysis a build step rather than a meeting.
Reading warehouse query logs. A catalog connects to Snowflake, BigQuery, or Databricks, reads the history of executed queries, parses them, and reconstructs which columns fed which. OpenMetadata and DataHub do this in open source; Atlan, Collibra, and Alation commercially. The advantage is coverage: any tool that ran a query in the warehouse shows up, including a BI tool or a script on a cron. The limit is timing: the graph describes the code that already ran, so it can show you what a change broke, not what it will break.
Most teams that own their transformations want the first kind, and add the second when they need one graph across tools they do not control.
What it is for
Impact analysis before a change. Remove or retype a column that a downstream asset selects, open a pull request, and run validation:
bruin validate ./pipeline
In Bruin the build fails with the assets that break, in the pull request rather than in the morning dashboard. That is column lineage doing its most valuable job invisibly.
Root cause after a break. A number is wrong in a report. Column lineage walks from the report's column back through each transformation to the source, and the step where the logic or the input changed is usually obvious once the path is in front of you. This is also the path an AI agent follows when diagnosing a failed pipeline, which is why lineage is part of the context layer in how to build a context layer for AI-ready pipelines.
Trust and audit. "Where does this figure come from" is a question finance and auditors ask. A column-level graph answers it with a path from the report to the source table, which is more persuasive than a description.
What column lineage does not do
Lineage derived from one framework only sees assets in that framework. A Looker explore, a notebook, or a legacy Informatica job is invisible to it. It is also not a governance product: there is no stewardship workflow or business glossary with approvals in a pipeline framework's lineage. For an organisation-wide catalog with business owners signing off on definitions, use a catalog; for keeping a pipeline from breaking, lineage derived from the code is the sharper tool. The comparison of both kinds is in the best data lineage and catalog tools in 2026.