Bruin for Startups

Your whole data stack, before your first data hire.

Bruin pulls in data from the tools you already use, cleans and models it, and puts it to work - in your team’s decisions and in your AI product.

Bruin Cloud credits

$10,000

Usually lasts an early-stage team 1-2 years

For backed startups

Y CombinatorTechstarsAntler

Under $5m raised and under 20 people

Teams already running on Bruin

Join the startup program.

Who qualifies

  • Backed by YC, Techstars, or Antler
  • Under $5m raised
  • Under 20 people
  • New to Bruin

What you get

  • $10,000 in Bruin Cloud credits*
  • A Slack channel with our engineers
  • Up to 20% off once the credits run out
  • No equity, no card

Not eligible? Apply anyway to join the waitlist, or start free with the open-source tools and guides.

*Credits are estimated to provide 1-2 years of managed pipeline runs (~1,000 GB-hours of compute) or 300+ AI questions and tasks. Credits are active for 2 years, after which any unused balance expires.

Secure from day one.

  • ISO/IEC 27001:2022 certified, by Prescient Security
  • SOC 2 Type 2 tested and attested by Prescient Assurance

Bruin Data Limited is ISO/IEC 27001:2022 certified by Prescient Security and SOC 2 Type 2 attested.

Free and open source

Set up Bruin with your coding agent.

Works with Cursor, Claude Code, Codex, and opencode. No account, no credit card.

Join the Slack community Ask the Bruin team, get answers fast.

Your agent’s starting point

Ready to paste

Set up Bruin in this repo so we can build a data pipeline.

  • 01Install the CLI and set up the docs MCP.
  • 02Choose a template and connect your data.
  • 03Add checks, validate, and run your pipeline.
Read the full prompt
Set up Bruin in this repo so we can build a data pipeline.

1. Install the Bruin CLI:
   curl -LsSf https://getbruin.com/install/cli | sh

2. Set up Bruin's MCP server. Read the setup docs:
   https://getbruin.com/docs/bruin/getting-started/bruin-mcp.html

   Work out which coding agent I am using, then give me step-by-step
   instructions for that client: which config file to edit, exactly what to
   put in it, and whether I need to restart. You cannot enable it yourself,
   so write the steps for me to follow. It runs locally as `bruin mcp` and
   serves Bruin's documentation, so once it is connected you can look up
   sources, asset types, and flags while you work.

3. Run `bruin init`, list the templates, and pick the one closest to my
   stack. Tell me which you chose and why.

4. Read the generated pipeline.yml and .bruin.yml. Ask me for the connection
   details you need - never commit credentials, and keep .bruin.yml in
   .gitignore.

5. Add quality checks to every asset you create, then run:
   bruin validate && bruin run

Explain what each asset does before you run anything that writes data.

Your data stays in your own database

Bruin never keeps a copy. Start with a local DuckDB file, and move to a hosted warehouse when you need one.

  • DuckDB
  • Postgres
  • BigQuery
  • ClickHouse
  • Snowflake
  • and 30+ more

Run it wherever you want

GitHub Actions, Cloud Run, Airflow, or one VM with cron - a guide for each. Or let Bruin Cloud run it for you.

Start on Bruin Cloud free

Frequently asked questions.

  • Who qualifies for the Bruin startup program?
    Teams that have raised under $5m from institutional investors, have fewer than 20 people including founders, are backed by Y Combinator, Techstars, or Antler, and are new to Bruin. The credits are for first-time users, so existing Bruin customers are not eligible. Y Combinator companies can claim it directly from the YC deals page at https://deals.ycombinator.com/deals/14668. Everyone else applies through the form and hears back within a couple of business days.
  • What if my startup does not qualify for the program?
    You can still use everything except the credits. The Bruin CLI is free and open source under Apache 2.0, every template and deployment guide is public, and Bruin Cloud has a free tier. The Bruin Slack community is open to everyone and the team answers questions there, whether or not you are in the program: https://join.slack.com/t/bruindatacommunity/shared_invite/zt-3cymzktqu-bvFxPGyQHpvi~dok_W0L3w
  • Where can I deploy Bruin pipelines?
    Anywhere you already run software: GitHub Actions, AWS Lambda or ECS, Google Cloud Run Jobs, GitLab CI/CD, Apache Airflow, or one VM with cron. There is an official container image at ghcr.io/bruin-data/bruin, and each target has a step-by-step deployment guide in the docs.
  • Do I need a data engineer to use Bruin?
    No, and that is the point for a team of under 20. Pipelines are SQL and Python files in a git repo, so anyone who can write a query can add one, and a coding agent with the Bruin MCP server can scaffold the first ones for you.
  • What is Bruin?
    Bruin is an open-source data ingestion and transformation tool. One CLI covers pulling data from over 150 sources, transforming it with SQL, Python, or R, running quality checks, and tracking column-level lineage - work that teams usually split across separate ingestion, transformation, and orchestration tools.
  • How many data sources and destinations does Bruin support?
    157 sources and 37 destinations as of September 2026. Sources include Stripe, Postgres, MySQL, MongoDB, Shopify, HubSpot, Salesforce, Zendesk, Intercom, PostHog, Amplitude, Mixpanel, Notion, Linear, GitHub, S3, Kafka, Anthropic, and OpenAI. Destinations include BigQuery, Snowflake, ClickHouse, Postgres, DuckDB, MotherDuck, Databricks, Redshift, Athena, and Iceberg.
  • Where does my data live when I use Bruin?
    In your own database or warehouse. Bruin moves and models data into a destination you control and never holds a copy, so if you stop using Bruin your warehouse and everything in it stays exactly where it is. That destination can be a local DuckDB file, your own Postgres, or a hosted warehouse like BigQuery or Snowflake.
  • How is Bruin different from dbt, Fivetran, and Airflow?
    Those are three tools and Bruin is one. dbt transforms data but does not ingest it, Fivetran ingests but does not transform, and Airflow schedules but does neither. Bruin covers ingestion, SQL and Python transformation, quality checks, lineage, and scheduling in a single CLI. The clearest technical difference from dbt is Python: Bruin runs Python assets in isolated uv-managed environments, each with its own dependencies and Python version, and materializes whatever dataframe they return.
Everything else we get asked

The startup program

  • What do you get from the Bruin startup program?
    Three things: $10,000 in Bruin Cloud credits, a dedicated Slack channel shared with Bruin engineers and staff, and up to 20% off once the credits run out. There is no fee to apply and Bruin takes no equity.
  • How far does $10,000 in Bruin Cloud credits go?
    The credits are a fixed $10,000, which we estimate lasts a seed-stage team 1-2 years depending on usage. That is over 1,000 GB-hours of compute, or over 300 AI questions and actions. Bruin Cloud bills on compute, so what burns credits fastest is long-running jobs with large memory footprints rather than the number of pipelines you run. Ordinary daily ingestion and transformation schedules for an early-stage team do not come close to the limit.
  • Do the startup credits expire?
    Yes. The $10,000 in Bruin Cloud credits is active for 2 years from when it is granted. After that, any unused balance expires.
  • What happens when the startup credits run out?
    You get up to 20% off, or you move to the free tier. Nothing is charged automatically and no credit card is required to join the program.
  • Do I need to join the startup program to use Bruin?
    No. The Bruin CLI is free and open source under Apache 2.0, and every template and deployment guide is public, so you can run everything on your own infrastructure indefinitely at no cost. The program only adds Bruin Cloud credits and a shared Slack channel with the Bruin team.
  • Can existing Bruin customers join the startup program?
    No. The credits are for first-time Bruin users, so teams already on a paid Bruin Cloud plan are not eligible. If you are already using Bruin and think your team should qualify, contact Bruin directly rather than applying through the form.
  • Does Bruin take equity from startups in the program?
    No. The credits are non-dilutive, there is no fee to apply, and no credit card is required at any point.

Getting started

  • How do I install Bruin?
    Run `curl -LsSf https://getbruin.com/install/cli | sh`. It installs a single binary and works on macOS and Linux; on Windows it needs Git Bash or WSL. Homebrew installation is deprecated, so uninstall any brew version first.
  • How do I connect Bruin to Claude Code, Cursor, or Codex?
    For Claude Code, run `claude mcp add bruin -- bruin mcp`. Cursor and VS Code take `{"command": "bruin", "args": ["mcp"]}` in their MCP config, and Codex takes an `[mcp_servers.bruin]` block in ~/.codex/config.toml. All three point at the same local stdio server and need no account.
  • What does the Bruin MCP server actually do?
    The local MCP server serves Bruin’s documentation to your coding agent through three tools: an overview, a documentation tree, and individual page content. It does not query your warehouse or run pipelines. Your agent reads the docs, then drives the Bruin CLI from the terminal. Bruin Cloud has a separate hosted MCP server at https://cloud.getbruin.com/mcp that does expose pipeline runs, backfills, and cost, and that one needs an API token.
  • What is the fastest way to get a first Bruin pipeline running?
    Install the CLI, run `bruin init` to scaffold from a template, fill in your connection details in .bruin.yml, then run `bruin validate && bruin run`. Templates cover common starting points like Stripe to BigQuery, Shopify to DuckDB, Notion, and a plain DuckDB project that needs no hosted warehouse at all.

What startups build with it

  • What do early-stage startups use Bruin for?
    Three jobs come up often. Internally, teams pull product events, billing, and CRM together to track growth, acquisition, and spend. For the product, an AI retention tool can join CRM, support, usage, and billing data into a tenant-specific customer health view, while a logistics tool can join orders, inventory, shipping, and external signals into fulfillment context. They also collect bad answers and human reviews into evaluation datasets. Bruin prepares the checked tables and files; your application handles retrieval, prompt assembly, model calls, and actions.
  • Can I run one Bruin pipeline for each of my customers?
    Yes. Pipeline variables and a `variants:` block let one definition run separately for different customers, regions, or environments, with tenant-specific names, filters, schedules, and connection choices. Vault, Doppler, AWS Secrets Manager, and other secret providers can keep credentials outside the repo. Bruin does not provide a tenant registry or application-level authorization, so your product still owns provisioning, access control, and the code that starts each tenant run.
  • Can Bruin prepare grounding data or prompt context for my AI product?
    Yes. Bruin can ingest each tenant’s source data, clean and join it, check freshness and relationships, and publish a product-ready table such as `customer_health` or `fulfillment_context`. Your agent can query those tables when it builds a prompt or decides which action to take. This gives the agent current business context without asking the model to infer relationships from raw CRM exports or support transcripts.
  • What would an AI startup pipeline look like in practice?
    For a retention product, ingest HubSpot or Salesforce, Intercom or Zendesk, product events, and Stripe, then model account health, usage changes, open escalations, and renewal risk. For a logistics product, ingest Shopify, inventory and warehouse data, shipping events, and external weather or traffic data through a Python asset or API export, then model late-delivery risk and stockout risk. In both cases, put the product-facing context tables downstream of quality checks so the agent does not read an unvalidated intermediate table.
  • Can Bruin load data into my application’s database?
    Yes. Postgres and ClickHouse are first-class destinations, alongside BigQuery, Snowflake, DuckDB, Databricks, Redshift, and about thirty more. Writing into the database your product already reads is mechanically supported, though the docs frame destinations as warehouses rather than application backends.
  • Can Bruin build training or evaluation datasets for an AI product?
    Yes, for the data plumbing around them. Python assets run in isolated uv-managed environments with their own dependencies and Python version, so PII scrubbing, dedup, and labelling code lives next to the SQL. Output can land as Parquet or Iceberg on S3 and be snapshotted so a run is reproducible. Bruin does not train models, generate embeddings, or write to vector databases.
  • Does Bruin support vector databases, embeddings, or RAG?
    No. Bruin has no support for Pinecone, Qdrant, Weaviate, Chroma, or pgvector, and it does not generate embeddings. It is an ingestion and transformation tool, so it prepares the structured tables an AI product reads, and anything vector-related belongs to a separate part of your stack.
  • Can Bruin track what our team spends on LLMs?
    Yes. Anthropic, OpenAI, and Cursor are ingestion sources, and the `ai-coding-usage` template pulls the Anthropic and Cursor admin APIs into DuckDB and builds per-user, per-model, per-day cost and token tables. It needs an Anthropic admin API key and a Cursor team admin key.
  • How fresh can Bruin keep the data my agents read?
    As fresh as you schedule it, down to minutes, using incremental materialization strategies like `merge` and `time_interval` so each run only processes new rows. Bruin is a batch tool rather than a streaming one, so the honest framing is frequent incremental runs, not sub-second updates.

Bruin basics

  • Is Bruin open source, and under what licence?
    The Bruin CLI is licensed under Apache 2.0. ingestr, the ingestion engine, is under the Functional Source License (FSL-1.1-ALv2), which is source-available and converts to Apache 2.0 two years after each release. Bruin Cloud is a commercial managed platform and is optional.
  • Do I have to use Bruin Cloud?
    No. Everything except the startup credits runs on the open-source CLI on infrastructure you own. Bruin Cloud adds managed scheduling, run monitoring, backfills, column-level lineage, a cost explorer, and alerting if you would rather not run those yourself.
  • What happens when a Bruin data quality check fails?
    The asset is marked failed, every downstream asset is skipped, and the run exits with a non-zero code. Checks are blocking by default. They run after an asset materializes, so a failing check stops bad data from propagating rather than preventing the write itself - which is why the table your product reads should sit downstream of the checks. Bruin ships ten built-in column checks including not_null, unique, accepted_values, pattern, and relationships, plus custom SQL checks.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.