SQL & Python pipelines

SQL where it fits. Python where it doesn't.

Both in one pipeline and one repo, with dependencies, checks and lineage built in.

/* @bruin
name: mart.revenue_daily
type: bq.sql
materialization:
  type: table
  strategy: delete+insert
  incremental_key: order_date
depends:
  - stg.orders
columns:
  - name: order_date
    checks:
      - name: not_null
@bruin */

select
  order_date,
  sum(amount) as revenue,
  count(distinct customer_id) as customers
from stg.orders
where order_date between '{{ start_date }}' and '{{ end_date }}'
group by order_date

Dependencies

You write the models. Bruin works out the order.

Bruin builds the graph from your assets and runs them in order, across ingestion, SQL and Python.

Quality

Every model, tested on every run.

Checks live in the same file as the model. When one fails, what's downstream waits.

Built in

The boring parts, handled.

Materializations

Tables, views and incremental loads, declared in the asset header.

strategy: delete+insert

Incremental runs

Every run gets a date window, so backfills are a re-run, not a rewrite.

'{{ start_date }}'

Environments

The same code runs against dev and prod connections.

bruin run --environment prod

Runs anywhere

On your laptop, in CI, or scheduled on Bruin Cloud.

bruin run pipelines/

Customer results

Numbers from teams on Bruin.

The platform

Part of the Bruin platform.

Models read what ingestion loads and feed every dashboard, app and answer, tested on the way.

Your stack, your call

Replace the modern data stack.

One layer or every layer. Keep what works, swap what doesn't.

Today

On Bruin

Ingestion

On Bruin: Data Ingestionon open-source ingestr

Transformation & orchestration

On Bruin: SQL & PythonBruin Cloud

Quality, lineage & catalog

On Bruin: Data QualityData Governance

BI & dashboards

  • Power BI

On Bruin: AI DashboardsData Apps

AI on your data

On Bruin: AI Data AnalystScheduled Agents

Frequently asked

Questions about pipelines.

Which languages can a pipeline use?

SQL and Python, side by side in the same pipeline, plus ingestr assets for ingestion. Dependencies can cross languages, so a Python model can read a SQL model and a dashboard can read both.

Which warehouses does it run on?

SQL runs inside your warehouse: Snowflake, BigQuery, Databricks, Redshift, Postgres, ClickHouse, DuckDB, MySQL and SQL Server, among others. Your data stays where it is.

How are dependencies defined?

Each asset lists what it depends on in its @bruin header, and Bruin builds the pipeline graph from those, then runs assets in order. The same graph powers lineage in Bruin Cloud.

Can we keep our dbt models?

Yes. You can run dbt models side by side with Bruin assets and move pipelines over one at a time, or keep dbt for transformation and use Bruin for ingestion, quality and the AI layer.

How do incremental loads and backfills work?

An asset declares a materialization strategy, such as delete+insert with an incremental key, and every run receives a start and end date. Backfilling a period means running the same pipeline for that date range.

Where do Python assets run?

Locally with the Bruin CLI, in CI, or on managed infrastructure in Bruin Cloud, where each asset can pick its Python image. Python assets can return a dataframe for Bruin to materialize as a table.

Is it open source?

Yes. Pipelines, transformations and quality checks run on the open-source Bruin CLI. Bruin Cloud adds scheduling, observability, lineage, governance and the AI layer.

One repo, every model tested.

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.