Azure Data Lake Storage integration

Azure Data Lake Storage data, in your warehouse.

A built-in ingestr connector: add credentials, pick tables, schedule it.

name: raw.table
type: ingestr
parameters:
  source_connection: azure_data_lake_storage
  source_table: '<table>'
  destination: snowflake
  incremental_strategy: merge

$ bruin run assets/raw/azure_data_lake_storage.asset.yml

  1. extract · Azure Data Lake Storage dataincremental
  2. load · snowflake raw.datamerged
  3. checks · not_null, unique

Loaded and checked. Downstream models can run.

How it connects

Connected in three steps.

Azure Data Lake Storage Gen2 is Azure Blob Storage with hierarchical namespace capabilities enabled for data lake workloads. ingestr supports Azure Data Lake Storage Gen2 as both a source and destination.

  1. 01

    Add a Azure Data Lake Storage connection with its credentials.

  2. 02

    Pick the tables to load and how: replace, append or merge.

  3. 03

    Bruin runs it on your schedule and checks every load.

Connection parameters

accountname
Azure storage account name.
tenantid
Microsoft Entra tenant ID for service principal authentication.
clientid
Service principal client ID.
clientsecret
Service principal client secret. URL-encode the value if it contains special characters.
accountkey
Azure storage account key.
sastoken
Shared Access Signature token. URL-encode the token if it contains &.
layout
Destination-only layout template.

The platform

Part of the Bruin platform.

Data in, ready for everything downstream: the models, the checks, the lineage and the AI layer.

01 · Move

Data Ingestion

02 · Model & govern

SQL & Python
Data Quality
Data Governance

03 · Use

AI Data Analyst
AI Dashboards
Data Apps
Self-Healing Pipelines
Bruin Cloudorchestration · governance · observability

Frequently asked

Questions about Azure Data Lake Storage.

Does Bruin have a built-in Azure Data Lake Storage integration?

Yes. Built-in ingestr source.

Where can Azure Data Lake Storage data go?

Snowflake, BigQuery, Databricks, Redshift, ClickHouse, Postgres, DuckDB, MotherDuck, Microsoft Fabric and more, plus files on S3 and GCS.

How fresh is the data?

As fresh as your schedule. Incremental loads append, merge or replace a time window, every few minutes if you like.

Do we need Bruin Cloud?

No. The Bruin CLI and ingestr run locally, in CI or in your own orchestrator. Bruin Cloud adds scheduling, lineage, alerts and the AI data analyst on top.

Ready to connect Azure Data Lake Storage?

$100 in credits and 50 AI tasks. No credit card.

A demo walks through your own data.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.