A data catalog is an inventory of an organisation's data assets, the tables, columns, dashboards, and pipelines, with what each one means, who owns it, where it came from, and whether it can be trusted right now, searchable in one place. Its job is to answer the questions people otherwise ask in Slack: is there a table for this, which one is the real one, what does this column mean, who do I ask. The catalog products, OpenMetadata, DataHub, Atlan, Collibra, Alation, are built for organisations with many teams and many tools. A small team needs the answers without the product, and a pipeline framework such as Bruin generates them from the pipeline itself.
What a catalog holds
| Field | Answers | Source |
|---|---|---|
| Assets and columns | What exists | The warehouse schema, or the pipeline definitions |
| Descriptions | What it means | Written by people, or generated and reviewed |
| Owners | Who to ask | Declared per asset or pipeline |
| Lineage | Where it came from and what depends on it | Parsed from the pipeline code, or reconstructed from query logs |
| Quality status | Can I trust it today | The checks that ran on the last load |
| Glossary | The business vocabulary and which columns carry it | Maintained by the team |
| Usage | Who queries it and how often | Warehouse query logs |
Search across those fields is the product. Everything else in a catalog is a way of keeping those fields current, and that is where catalogs succeed or die.
How a catalog gets populated
There are two ways, and the choice decides whether the catalog stays accurate.
Crawl the warehouse. The catalog connects to Snowflake, BigQuery, or Databricks, reads the schemas and the query history, and reconstructs assets, usage, and lineage. This is how OpenMetadata, DataHub, Atlan, and Collibra work, and its strength is coverage: every table that exists shows up, whatever produced it. Its weakness is meaning. A crawled catalog knows a column named rev_adj_2 exists and has no idea what it is. Descriptions, owners, and the glossary have to be typed in by people, and in a small team they are not, so the catalog becomes an accurate inventory of undocumented tables.
Derive it from the pipeline. If every table is produced by an asset definition that already carries its columns, description, owner, checks, and dependencies, the catalog is a rendering of the repository. Bruin works this way: the asset header holds the metadata, bruin ai enhance writes the first draft of descriptions and checks for review, lineage is parsed from the SQL, and Bruin Cloud's catalog renders assets, columns, owners, quality status, cross-pipeline lineage, and a glossary stored as YAML in the same repository. Nothing is typed into a second system, because there is no second system.
/* @bruin
name: mart.orders
type: sf.sql
owner: [email protected]
description: One row per order, deduplicated and enriched with customer region.
tags: [orders, finance]
depends: [raw.orders, mart.customers]
columns:
- name: order_id
type: integer
primary_key: true
description: Order identifier from the source system, stable across updates.
checks:
- name: not_null
- name: unique
@bruin */
The trade-off is the mirror of the crawler's. A pipeline-derived catalog only knows the assets in the pipeline. A BI tool's saved reports or a notebook on a cron are invisible to it. That is fine for a team whose data work runs through one framework and a real gap for an organisation with a dozen tools.
Governance without a governance product
Governance for a small team is four questions: does every table have an owner, does every table have a description, do the important tables have checks, and can I see what a change breaks before I make it. A catalog product answers them by asking people to fill in fields. A pipeline framework answers them by refusing to run assets that do not comply. In Bruin that is a policy.yml with rules such as asset-has-owner, asset-has-description, and asset-has-checks, applied by bruin validate before every run and in CI. Governance becomes a property of the build rather than a project.
When to buy a catalog
Buy one when the stack outgrows a single framework: many producing teams, tools nobody controls, a regulated environment that needs stewardship workflows and approval chains, or a business glossary that has to be signed off by owners rather than committed by engineers. OpenMetadata is where to start in open source, DataHub when the metadata model and scale matter more than setup time, Atlan or Collibra when the requirement is a governance programme rather than a tool. Until then, the catalog that costs nothing to maintain is the one derived from the pipeline.
For the tool comparison see the best data lineage and catalog tools in 2026, and for the lineage half in depth, what is column-level lineage.