Data platform for data teams

Broken at 2am. Fixed in a PR by 8.

Open-source pipelines that run locally, in CI or on Bruin Cloud.

Trusted by forward-thinking teams

In your repo

Ingestion, models and checks, in one repo.

name: raw.orders
type: ingestr
parameters:
  source_connection: postgres
  source_table: 'public.orders'
  destination: snowflake
  incremental_strategy: merge
  incremental_key: updated_at

columns:
  - name: order_id
    primary_key: true
    checks:
      - name: not_null
      - name: unique

$ bruin run pipelines/daily_revenue

  1. raw.orders84,102 rows merged
  2. not_null · order_id
  3. stg.orders1.9s
  4. mart.revenue_daily2.4s
  5. non_negative · revenue

7 assets, 7 checks. Downstream is current.

Column-level lineage

Know the blast radius before you change a column.

From a data lead

“I had an engineer try to export raw Customer.io data to Redshift. It took him about a month and a half to get anything off the ground as he navigated competing priorities. The exact same thing on Snowflake through Bruin, with modern LLM tools, took a day.”

Adam MusialAdam MusialVP of Analytics, ProphetX
runs all data and analytics
1 person
idea to production data source
Month → 1 day
more pipelines, same headcount
3x
Read the ProphetX story

Built on the Bruin platform

Stakeholders ask AI anyway. Make it use your models.

01 · MoveEvery source, in.
Data IngestionAny source in, on open-source ingestr

Raw tables, in your warehouse

02 · Model & trustClean, tested, traced.
SQL & PythonPipelines in SQL and Python
Data QualityChecks on every load
Data LineageEvery column, traced

Tested, governed models

03 · UseAnswers, builds and acts.
AI Data AnalystAnswers in chat, with the query
AI DashboardsDashboards from one prompt
Data AppsLive apps on governed data
Scheduled AgentsBriefs and alerts on your clock
Bruin CloudOrchestration, governance, observability

Frequently asked

Questions from data teams.

Can Bruin replace our Fivetran, dbt and Airflow stack?

Yes. ingestr handles ingestion, SQL models move over as SQL, and dependencies come from the assets, so there are no DAG files. dbt models run beside Bruin assets while you migrate one pipeline at a time.

Which data warehouses does Bruin run on?

Snowflake, BigQuery, Databricks, Redshift, Postgres, ClickHouse, DuckDB, MySQL and SQL Server, among others. SQL runs inside your warehouse, so the data stays where it is.

Can we run and test Bruin pipelines locally and in CI?

Yes. The CLI runs the same pipeline on a laptop, in CI or on Bruin Cloud. Give CI its own environment, and every pull request builds and checks on a CI schema before merge.

Can we run SQL and Python transformations in one Bruin project?

Yes. Python assets sit in the same graph as SQL, with the same checks and lineage, and can return a dataframe for Bruin to materialize as a table.

How does Bruin validate data inside a pipeline?

With column checks such as not_null, unique, accepted_values, positive and non_negative, plus custom checks in SQL. Checks block by default, so a failure holds everything downstream instead of reaching a dashboard.

Can Bruin show the downstream impact of a column change?

Yes. Column-level lineage is built from the pipelines themselves, tracing columns through SQL from the source system to dashboards and AI answers. Pick a column to see everything a change reaches before you merge.

Does Bruin have an MCP server for Claude, ChatGPT or Cursor?

Yes, Bruin MCP. Agents query your models with their descriptions and checks instead of guessing at raw tables, and you choose which assets they can see.

Can business teams get answers from Bruin without waiting on us?

Yes. The AI analyst answers in Slack, Microsoft Teams, Google Chat, WhatsApp, Discord, Telegram, email or the browser, from your models and tested definitions. Every answer shows the query it ran.

Is Bruin open source, and what does Bruin Cloud add?

The Bruin CLI is open source (Apache 2.0) and ingestr is source-available (FSL): ingestion, SQL and Python assets, checks and the dependency graph run anywhere. Bruin Cloud adds managed scheduling, lineage views, the AI analyst, SSO, roles and audit logs.

How much does Bruin cost, and is it per seat?

The Bruin CLI (open source, Apache 2.0) and ingestr (source-available) are free to run yourself. On Bruin Cloud, a daily pipeline of 20 SQL models, a minute each, is about $25 a month, and ingesting 100M rows is about $10. Billed per second, no seats. Committed-use discounts lower the rate as usage grows.

How does Bruin’s pricing compare with Fivetran or Airbyte?

Row-based ELT tools bill by the row, so the bill climbs with volume; Bruin bills compute. At list price, ingesting 100M rows a month is about $10 on Bruin, ~$160 on Fivetran and from ~$200 on Airbyte.

Is our warehouse data safe with Bruin?

Bruin is SOC 2 Type 2 attested and ISO/IEC 27001:2022 certified. Your data is never used to train AI models.

Sleep through the 2am failure.

$100 in credits and 50 AI tasks. No credit card.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.