StarRocks integration

StarRocks data, in your warehouse.

A built-in ingestr connector: add credentials, pick tables, schedule it.

name: raw.table
type: ingestr
parameters:
  source_connection: starrocks
  source_table: '<table>'
  destination: snowflake
  incremental_strategy: merge

$ bruin run assets/raw/starrocks.asset.yml

  1. extract · StarRocks dataincremental
  2. load · snowflake raw.datamerged
  3. checks · not_null, unique

Loaded and checked. Downstream models can run.

How it connects

Connected in three steps.

StarRocks is a high-performance analytical (OLAP) database. Alongside its own internal storage, StarRocks can query open lakehouse table formats — such as Apache Iceberg, Apache Hudi, Apache Hive, and Delta Lake — through external catalogs.

  1. 01

    Add a StarRocks connection with its credentials.

  2. 02

    Pick the tables to load and how: replace, append or merge.

  3. 03

    Bruin runs it on your schedule and checks every load.

Connection parameters

username
the StarRocks user (required)
password
the password for the user (optional, depending on authentication)
host
the StarRocks FE (frontend) hostname or IP address
port
the FE query port that speaks the MySQL protocol (default: 9030)
path
the default catalog and database for unqualified table names. A single segment (/<database>) sets the default database and leaves the catalog at defaultcatalog; two segments (/<catalog>/<database>) set both. Anything specified in the source table name takes priority over these defaults.

Every way StarRocks works with Bruin

  • SourceBuilt-in ingestr sourceDocs →
  • DestinationBuilt-in ingestr destinationDocs →
  • Query engineRuns SQL and Python assetsDocs →

The platform

Part of the Bruin platform.

Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.

01 · Move

Data Ingestion

02 · Model & govern

SQL & Python
Data Quality
Data Governance

03 · Use

AI Data Analyst
AI Dashboards
Data Apps
Self-Healing Pipelines
Bruin Cloudorchestration · governance · observability

Frequently asked

Questions about StarRocks.

Does Bruin have a built-in StarRocks integration?

Yes. Built-in ingestr source. Built-in ingestr destination. Runs SQL and Python assets.

Where can StarRocks data go?

Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Postgres, DuckDB, MotherDuck, Microsoft Fabric and more, plus files on S3 and GCS.

How fresh is the data?

As fresh as your schedule. Incremental loads append, merge or replace a time window, every few minutes if you like.

Do we need Bruin Cloud?

No. The Bruin CLI and ingestr run locally, in CI or in your own orchestrator. Bruin Cloud adds scheduling, lineage, alerts and the AI data analyst on top.

Ready to connect StarRocks?

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.