Technical
6 min read

Does Bruin Lock You In? What You Own, and How You Would Leave

A plain statement of what a Bruin customer owns: pipelines as SQL, Python, and YAML in your own Git repository, data that never leaves your own warehouse, and an Apache 2.0 CLI that runs anywhere, including inside Airflow. It covers what Bruin Cloud adds, what happens if you stop paying, and the one place a dependency exists. Written to be checked, not believed.

Does Bruin Lock You In? What You Own, and How You Would Leave

Bruin does not lock you in, and the way to check that claim is to ask what you would still have the day after you left. With Bruin the answer is everything that matters: your pipelines, which are plain SQL, Python, and YAML files in your own Git repository; your data, which never left your own warehouse; and the ability to keep running the pipelines, because the CLI that runs them is open source under Apache 2.0 and runs anywhere, including as a task inside an Airflow DAG. What you would lose is the managed layer, Bruin Cloud, which is the part you pay for and the part that is optional. This page states that in enough detail to be verified, because "no lock-in" is a claim every vendor makes and few define.

What you own

AssetWhere it livesFormatWho holds it
Pipeline definitionsYour Git repositorySQL, Python, YAML files with a Bruin headerYou
Ingestion configurationSame repositoryYAML ingestr assetsYou
Quality checks and policiesSame repository, on the assetsYAML in the asset header, policy.ymlYou
Semantic layer and dashboards as codeSame repositoryYAML and TSXYou
Documentation and glossarySame repositoryYAML and asset descriptionsYou
The dataYour warehouse or databaseTablesYou
The runtimeAny machineOpen-source CLI, single binaryYou

Nothing in that table is stored in a Bruin system in a form you cannot read. The asset header is a comment block at the top of a SQL file; remove it and you have the SQL. That is the property Bruin's design principles state directly: "Data is core, and lock-in in data should be avoided. This is why Bruin CLI runs on any environment, and supports all the core features out of the box via the open-source product."

What Bruin Cloud adds, and what leaving it costs

Bruin Cloud is the managed service: it runs the schedule with retries and backfills, routes alerts, renders the catalog and cross-pipeline lineage, holds team access controls, SSO, and audit logs, and provides the AI data analyst in Slack, Microsoft Teams, Google Chat, WhatsApp, Discord, Telegram, email, and the browser, plus dashboards generated from prompts.

If you stop paying for it, that is what you lose. The pipelines are unaffected. You run the same project with bruin run from GitHub Actions, a cron, or an Airflow DAG, using the same CLI, against the same warehouse, and the results are the same tables. Teams that self-host from day one never use Cloud at all. The Airflow deployment guide shows the BashOperator and KubernetesPodOperator patterns, which are also the exit route.

The licences, precisely

  • Bruin CLI: open source under the Apache 2.0 licence. Use, modify, and distribute commercially, with an explicit patent grant.
  • ingestr: source-available under the Functional Source License (FSL-1.1-ALv2). Free to use for internal production, development, testing, education, research, and professional services; the one restriction is offering it as a competing commercial service. Each release converts to Apache 2.0 two years after it ships.
  • VS Code extension: open source.
  • Bruin Cloud: commercial, subscription.

An earlier version of our site described the CLI as MIT-licensed. It is Apache 2.0, which is the more explicit of the two on patents, and we have corrected the pages.

Where a dependency does exist

Honesty requires naming it. If you use Bruin's asset header syntax, your SQL files carry a Bruin-specific comment block, and if you use ingestr's incremental strategies, your loads depend on ingestr's implementation of them. Moving off Bruin means either running the open-source CLI yourself, which is the intended path and costs nothing, or translating those headers into another framework's format. That translation is the same work as moving from dbt's YAML to SQLMesh's, and it is work on files you own, not on data held by a vendor. There is no proprietary storage format, no data held hostage, and no export step, because there was never an import.

How this compares

The comparison that matters is with the assembled stack. A team on Fivetran, dbt Cloud, and a managed orchestrator owns its dbt SQL and its warehouse tables; it does not own the connector configurations, which live in Fivetran's UI, or the schedule, which lives in the orchestrator's database. Leaving means rebuilding both. A team on Bruin owns the connector configurations as YAML in Git and the schedule as a line in pipeline.yml. The consolidated platform, counter-intuitively, is the one with less to leave behind.

For the migration path in, see how hard is it to migrate to Bruin. For the open-source components themselves, the Bruin CLI and ingestr repositories are the source of truth for licences and code.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.