Comparison
15 min read

Snowflake ETL Tools: How 52,771 Companies Load Data (2026)

Fivetran leads at 12.7% of 52,771 Snowflake companies, then ADF, Informatica, AWS Glue, Airbyte, and Matillion. How the top Snowflake ETL tools compare.

Snowflake ETL Tools: How 52,771 Companies Load Data (2026)

TL;DR: Fivetran is the most common ETL tool among Snowflake companies: 12.7% of 52,771 run it, followed by Azure Data Factory (7.9%), Informatica (6.5%), AWS Glue (2.7%), Airbyte (2.7%), and Matillion (1.5%), per Crustdata data charted by Greybeam. Together the six cover at most a third of Snowflake companies. The rest load data with COPY INTO, Snowpipe, Openflow, smaller vendors, code-first CLIs like Bruin's ingestr, or custom code.

A chart has been doing the rounds in data circles: Greybeam, the company that routes Snowflake queries to DuckDB to cut compute costs, took Crustdata's company data and asked a simple question. When a company runs Snowflake, what does it use to get data into it?

Bar chart of the share of 52,771 Snowflake companies also running each ETL tool: Fivetran 12.7%, Azure Data Factory 7.9%, Informatica 6.5%, AWS Glue 2.7%, Airbyte 2.7%, Matillion 1.5%

This post walks through why each of those six tools ended up in Snowflake stacks, who else is in the market, and how they compare on the criteria that matter when you are picking one in 2026. We build two tools in this space, the standalone ingestion CLI ingestr and the end-to-end Bruin CLI, so we include both in the comparison and try to be clear about where each fits and where it does not.

Which ETL tools do Snowflake companies use?

Fivetran is the most common ETL tool among Snowflake companies, used by roughly 6,700 of the 52,771 in Crustdata's dataset.

RankToolShare of Snowflake companiesApprox. companiesType
1Fivetran12.7%~6,700Managed ELT (SaaS)
2Azure Data Factory7.9%~4,170Cloud-native ETL (Microsoft Azure)
3Informatica6.5%~3,430Enterprise data integration suite
4AWS Glue2.7%~1,420Cloud-native ETL (AWS)
5Airbyte2.7%~1,420Open-source ELT + managed cloud
6Matillion1.5%~790Cloud-warehouse ELT platform

Company counts are derived from the published percentages and rounded. Source: Crustdata, chart by Greybeam.

How to read this data

Three caveats keep the numbers honest:

  1. "Running" means detected, not proven. Crustdata builds its technographic dataset from public web signals and job postings. A job post asking for Informatica experience counts as a signal. That is a good proxy for real usage, but it is not a warehouse audit, and it undercounts tools that nobody hires for by name, like Snowpipe or a Python script.
  2. 52,771 is larger than Snowflake's customer count. Snowflake reported 14,554 customers at the end of July 2026. Crustdata's number counts every company with a Snowflake signal, including subsidiaries, consultancies, and companies that use Snowflake through a parent or partner. Think of it as "companies where Snowflake shows up," not "Snowflake accounts."
  3. The tools overlap. A company can run Fivetran for SaaS sources and Informatica for its mainframe. The percentages add up to 34%, so the real share of Snowflake companies using any of the six is at most a third, and probably less.

Even with those caveats, the ranking lines up with what we see in the field. Here is how each tool got there.

1. Fivetran (12.7%)

What it is: a fully managed ELT service. You connect a source in a web UI, point it at Snowflake, and Fivetran keeps the tables in sync, handling schema changes and API quirks for you.

How it won Snowflake: timing and fit. Fivetran was founded in 2012 and went through Y Combinator in early 2013. It built its business on exactly the pattern Snowflake made popular: load raw data first, transform it later inside the warehouse. As Snowflake grew through the late 2010s, "Fivetran into Snowflake, dbt on top" became the default modern data stack, and both companies promoted it. Snowflake's Partner Connect let new accounts spin up a Fivetran trial in a few clicks. Fivetran already offered log-based replication for the major databases by 2018, and added HVR's high-volume enterprise replication with its acquisition in 2021, and in June 2026 it completed its merger with dbt Labs, so the two halves of that default stack are now one company.

Why teams keep it: the connector catalog is the benchmark for the long tail of SaaS sources, and nobody on the team has to maintain a sync.

Watch-outs: pricing is usage-based on monthly active rows, which is cheap at low volumes and climbs quickly as tables grow or churn. Your Snowflake bill also rises because every sync writes into a warehouse. Fivetran moves and lands data; transformations, tests, and scheduling across the rest of the stack live in other tools (now largely dbt). If cost is the issue, see our Fivetran alternative page and the migration guide from Fivetran to Bruin.

2. Azure Data Factory (7.9%)

What it is: Microsoft's managed data integration and orchestration service on Azure. Pipelines are built in a visual designer (or as JSON), with copy activities, Spark-based mapping data flows, and triggers.

How it won Snowflake: Snowflake has run on Azure since 2018, and a large share of enterprises standardized on Azure for identity, networking, and procurement. For those teams, ADF is already approved, already inside the Microsoft enterprise agreement, and already reachable through the self-hosted integration runtime that connects to on-prem SQL Server behind the firewall. ADF added a native Snowflake connector in 2020, and the current Snowflake V2 connector supports both copy and data flows. Teams moving off SQL Server Integration Services (SSIS) can also lift existing packages into ADF, which kept a lot of Microsoft shops in the ecosystem when they adopted Snowflake.

Why teams keep it: it is the path of least resistance in an Azure organization, and it doubles as the orchestrator.

Watch-outs: it ties you to Azure, the visual pipelines are hard to review and test like code, and the older V1 Snowflake connector was removed in September 2025, so any pipelines still on it must move to V2. Microsoft is steering new work toward Data Factory in Microsoft Fabric, though it has not announced an end date for ADF.

3. Informatica (6.5%)

What it is: the long-standing enterprise data integration suite. PowerCenter is the on-prem ETL product that ran data warehouses for two decades; Intelligent Data Management Cloud (IDMC) is its cloud successor, with integration, data quality, master data management, and a catalog.

How it won Snowflake: Informatica did not win Snowflake users so much as follow its existing customers there. Large banks, insurers, and manufacturers ran PowerCenter against Oracle and Teradata for years. When they moved their warehouse to Snowflake, many kept Informatica for the parts that were hard to replace: mainframe and ERP sources, data quality rules, and governance processes that auditors already understood. IDMC gave them a cloud route without retraining the team.

Why teams keep it: breadth for legacy enterprise sources, and governance and data quality in the same vendor.

Watch-outs: 2025 and 2026 changed the picture. Salesforce completed its acquisition of Informatica in November 2025, and standard support for PowerCenter 10.5 ended on March 31, 2026, with paid extended support for roughly one more year, into 2027. Many of those 3,400-odd companies are deciding right now whether to move to IDMC or to something else. Licensing is consumption-based and complex, and the tooling is GUI-first.

4. AWS Glue (2.7%)

What it is: AWS's serverless ETL service, running Apache Spark jobs with a built-in Data Catalog, crawlers, and a data quality feature.

How it won Snowflake: most Snowflake deployments run on AWS, and many of those companies already land files in S3. Glue is the AWS-native way to clean, reshape, and move that data, billed on the same AWS invoice and governed by the same IAM roles. AWS added a native Snowflake connector to Glue in 2023, so a Glue job can read from and write to Snowflake without a custom JDBC setup.

Why teams keep it: it fits AWS-first teams who already write PySpark and want serverless compute with no cluster to manage.

Watch-outs: it is a Spark platform, not a connector catalog, so pulling from SaaS APIs usually means writing code. Log-based database CDC typically needs a second service (AWS DMS). Spark startup times and DPU-hour costs are overkill for small, frequent loads, where Snowflake could do the transform itself after a plain COPY.

5. Airbyte (2.7%)

What it is: an open-source data movement platform with a very large connector catalog (600+ including community connectors), a low-code connector builder, and Airbyte Cloud as the managed option.

How it won Snowflake: Airbyte launched in 2020 as the open-source alternative to Fivetran, right when Fivetran bills were becoming a line item CFOs asked about. Snowflake was one of its first and most-used destinations. Teams could self-host Airbyte, keep data inside their own network, and build a missing connector themselves instead of waiting on a vendor. That combination made it the default open-source answer for Snowflake loading, and it tied with AWS Glue in this dataset.

Why teams keep it: connector breadth without per-row pricing, and the option to self-host.

Watch-outs: self-hosting Airbyte is real infrastructure to run and upgrade, and community connector quality varies. The platform is licensed under the Elastic License 2.0, which is source-available rather than OSI open source. Airbyte moves data; transformations and tests belong to dbt or another tool. See our Airbyte vs Bruin comparison for a side-by-side.

6. Matillion (1.5%)

What it is: a visual ELT platform built for cloud data warehouses. It generates SQL and pushes the work down to Snowflake, so transformations run on Snowflake compute rather than on a separate engine.

How it won Snowflake: Matillion was founded in Manchester in 2011 as a consultancy, and in 2015 it launched its first software product, Matillion ETL for Amazon Redshift. It then shipped Matillion ETL for Snowflake and became one of the few tools designed from day one for pushdown ELT on a cloud warehouse, which suited teams who wanted a GUI but did not want an external ETL server doing the heavy lifting. Matillion has kept a close relationship with Snowflake, launching its Maia agentic data engineers at Snowflake Summit 2025.

Why teams keep it: visual pipelines that still use Snowflake's compute, with ingestion and transformation in one place.

Watch-outs: pricing is credit-based on top of your Snowflake compute, pipelines are designed in a GUI rather than written as code, and the smaller share here reflects a more focused user base. See our Matillion alternative page for how Bruin approaches the same job as code.

What the data says about the market

A few patterns stand out once you look past the ranking:

  • The cloud providers together rival Fivetran. ADF (7.9%) plus Glue (2.7%) is 10.6%, close to Fivetran's 12.7%. A lot of Snowflake loading is decided by which cloud your company already pays for, not by a data tool evaluation.
  • Legacy is still big, and it is about to move. Informatica at 6.5% is nearly as large as Glue, Airbyte, and Matillion combined (6.9%). With PowerCenter support ending and Informatica now inside Salesforce, expect a migration wave through 2027.
  • Open source has a foothold, not a majority. Airbyte ties Glue, but the long tail of open-source and source-available tools (dlt, Sling, Meltano, ingestr) is too small or too code-first to show up in job-post-based data. That undercounting cuts both ways: engineers rarely list "Snowpipe" or "a Python script" on a job post either.
  • Two-thirds of Snowflake companies are not accounted for. The six tools add up to 34% before removing overlap. Snowflake's native loading, smaller vendors, and in-house code carry the rest.

What other tools load data into Snowflake?

Snowflake's native options

Snowflake keeps adding first-party ways to load data, which is a big reason so many companies do not show up under any third-party tool:

  • COPY INTO bulk-loads files from a stage (S3, GCS, Azure Blob, or an internal stage). Most tools, including ingestr, use it under the hood.
  • Snowpipe loads files automatically as they land in cloud storage, and Snowpipe Streaming writes rows with low latency without staging files.
  • Snowflake Openflow, built on Apache NiFi after Snowflake acquired Datavolo in late 2024, is Snowflake's own managed integration service with CDC connectors for databases like Oracle, PostgreSQL, and MySQL, plus SaaS connectors. It runs in your cloud account or inside Snowflake.

Native options avoid another vendor, but you still need something to schedule them, transform the results, and test the output.

Other commercial tools

  • Qlik (Talend, Stitch, Qlik Replicate): Qlik Replicate is a common log-based CDC choice for enterprise databases, and Stitch is a simple managed ELT service for low-volume syncs.
  • Hevo Data: no-code managed ELT with near real-time syncs.
  • Estuary: managed real-time CDC and streaming into Snowflake.
  • Striim and Oracle GoldenGate: enterprise real-time replication, often for Oracle sources.
  • Boomi (Rivery), SnapLogic, and Coalesce: Boomi and SnapLogic are broad integration platforms; Coalesce is a visual transformation tool that started on Snowflake (now also Databricks and Fabric) that pairs with an ingestion tool.

Open-source and code-first tools

  • dlt: a Python library for writing ingestion as code.
  • Sling: a fast CLI for database-to-database and file loads.
  • Meltano: a DataOps framework on top of Singer taps.
  • Debezium + Kafka: the standard open-source stack for streaming log-based CDC, with the most operational overhead.
  • ingestr and the Bruin CLI: covered next.

For a broader rundown of these tools outside the Snowflake lens, see our best data ingestion tools guide and best CDC tools guide.

Where ingestr and Bruin fit

We build two separate tools that show up in Snowflake stacks in different ways. ingestr is a standalone ingestion CLI, and plenty of teams use it on its own without anything else from Bruin. The Bruin CLI is an end-to-end pipeline tool that uses ingestr for its ingestion steps and adds transformations, quality checks, and lineage around them.

ingestr: standalone ingestion CLI

ingestr moves data from 130+ built-in sources into 20+ destinations, Snowflake included, with one command. There is no server, no UI to click through, and no connector code:

ingestr ingest \
  --source-uri 'postgresql://user:pass@host:5432/app' \
  --source-table 'public.orders' \
  --dest-uri 'snowflake://user:pass@account/ANALYTICS?warehouse=LOADING' \
  --dest-table 'raw.orders' \
  --incremental-strategy merge \
  --incremental-key updated_at \
  --primary-key id \
  --interval-start '2026-09-28'

It bulk-loads through Snowflake stages rather than row-by-row inserts and supports incremental strategies from append and merge to SCD Type 2. Since v1.1 it supports log-based CDC for PostgreSQL, MySQL, MongoDB, and SQL Server; PostgreSQL, MongoDB, and SQL Server CDC can load into Snowflake today, while MySQL CDC does not yet support Snowflake as a destination. CDC runs either as a scheduled batch that picks up every change since the last run, or in continuous streaming mode (--stream), where ingestr subscribes to the database changelog, buffers events, and flushes them to the destination when the buffer fills or a timeout hits. Deletes land as soft deletes, and delivery is at-least-once. In our benchmarks, ingestr v1 was the fastest tool on every 1M-row load into Snowflake we tested, ahead of dlt and Sling.

Because it is a single binary, teams drop it into whatever already runs their jobs: cron, GitHub Actions, Airflow, Dagster, Kubernetes. If all you need is fast, reliable loading into Snowflake from sources it supports, ingestr on its own is the whole answer. ingestr is source-available under the Functional Source License, free for internal use, with each release converting to Apache 2.0 after two years. See the ingestion product page and the database-to-Snowflake replication guide.

Bruin CLI: end-to-end pipelines, including custom ingestion

The Bruin CLI is open source (Apache 2.0) and is for teams that want the load and everything after it in one project. A Bruin pipeline is a folder of files in Git: ingestr assets for built-in sources, SQL assets that run on Snowflake, Python assets that write their results to it, column-level and custom quality checks on any asset, and column-level lineage across the whole flow. It runs locally or in CI.

For sources ingestr does not cover, such as an internal API or a niche SaaS tool, Python materialization is the best option. You write a Python asset whose materialize() function returns a DataFrame, and Bruin loads it into Snowflake with the merge, append, or incremental strategy declared in the asset config. The custom source then gets the same scheduling, checks, and lineage as everything else, without hand-written to_sql calls or credential handling.

Bruin Cloud runs those pipelines as a managed service, adding scheduling, observability, a catalog, access controls, and an AI data analyst that answers questions in Slack or Teams. It runs in Bruin's cloud or in your VPC, which matters for teams that cannot give a SaaS vendor direct access to production databases.

Where they are not the right pick

If you need a maintained connector for an obscure SaaS app tomorrow and have no engineers to write a Python asset, Fivetran's catalog is broader. Continuous CDC from PostgreSQL, MongoDB, or SQL Server into Snowflake is covered by ingestr's streaming mode without Kafka, but if you need exactly-once delivery, streaming from sources ingestr's CDC does not cover, or non-database event streams, look at Snowpipe Streaming, Estuary, or Debezium. If your organization mandates Azure-native or AWS-native services, ADF or Glue will be easier to get approved.

Snowflake ETL tools compared

At a glance

ToolTypeOpen sourceWhere it runsPricing modelBest for
FivetranManaged ELTNoVendor SaaS; hybrid agent optionUsage-based (monthly active rows)Hands-off SaaS syncs with budget to scale
Azure Data FactoryCloud-native ETL + orchestrationNoAzure; self-hosted runtime for private networksPay per activity, data movement, and data flow computeAzure-first organizations and SSIS migrations
Informatica (IDMC)Enterprise integration suiteNoVendor cloud with on-prem Secure AgentConsumption-based (processing units)Regulated enterprises with legacy sources and formal governance
AWS GlueServerless Spark ETLNoAWSPer DPU-hourAWS-first teams writing PySpark
AirbyteOpen-source ELT + cloudSource-available (ELv2)Self-hosted or Airbyte CloudFree to self-host; Cloud is usage/capacity-basedConnector breadth without per-row pricing
MatillionPushdown ELT platformNoVendor SaaS with agent in your cloudCredit-basedVisual pipelines that run on Snowflake compute
Snowflake native (COPY, Snowpipe, Openflow)First-party loadingOpenflow is based on Apache NiFiInside Snowflake or your cloud accountSnowflake creditsTeams that want no extra vendor
ingestrStandalone ingestion CLISource-available (FSL, converts to Apache 2.0)Anywhere a binary runs (cron, CI, Airflow, Kubernetes)Free for internal useFast, reliable loads from built-in sources with no server to run
Bruin CLI + Bruin CloudEnd-to-end pipelines (ingestion, transformation, checks, lineage)Bruin CLI Apache 2.0; Bruin Cloud commercialCLI anywhere; Bruin Cloud managed or in your VPCFree to run the CLI; Bruin Cloud is commercialCode-first teams who want ingestion, custom Python sources, transforms, checks, and lineage in one pipeline

Beyond moving rows

Loading is only the first step. What you need around it is where the tools differ most:

ToolLog-based CDCTransformationsData quality checksLineage / catalogPipelines as code
FivetranYesVia dbt (same company since 2026)Via dbt testsMetadata API and catalog integrationsAPI and Terraform; UI-first
Azure Data FactoryNative CDC for selected sourcesMapping data flows (Spark)Assert transformation in data flows; no dedicated DQ productVia Microsoft PurviewJSON with Git integration; visual-first
Informatica (IDMC)YesYesYes (separate data quality product)Yes (separate catalog product)GUI-first
AWS GlueNo native log-based CDC (pair with AWS DMS)PySpark / ScalaGlue Data QualityData CatalogYes (Spark code)
AirbyteYes (Postgres, MySQL, SQL Server, MongoDB; Oracle in Enterprise)No (pair with dbt)NoNoAPI, Terraform, low-code builder
MatillionYesYes (pushdown SQL)Basic assertion componentsTable- and column-level lineageVisual-first with Git integration
Snowflake nativeVia Openflow connectorsDynamic tables, tasks, stored proceduresData metric functionsSnowflake Horizon lineageSQL
ingestrYes, batch or streaming (Postgres, MongoDB, SQL Server into Snowflake; MySQL CDC not yet supported into Snowflake)No (load only)NoNoCLI flags, scriptable
Bruin CLI + Bruin CloudVia ingestr assetsSQL and Python assetsBuilt-in column and custom checksColumn-level lineageYes (YAML, SQL, Python in Git)

How to choose an ETL tool for Snowflake

Total cost, including Snowflake compute. Every tool here writes into Snowflake, so the warehouse bill moves with your sync frequency and load pattern. Row-based pricing (Fivetran) punishes high-churn tables, compute-based pricing (Glue, ADF data flows, Matillion) punishes inefficient jobs, and open-source tools shift the cost to the infrastructure and engineers you already have. Model a realistic month for your three largest tables before deciding. If warehouse compute is the bigger issue, tools like Greybeam's cost observability app help you see where Snowflake credits actually go.

Where the tool runs and what it can reach. Regulated teams often cannot let a SaaS vendor connect to production databases. ADF and Glue run in your cloud, Fivetran, Matillion, and Informatica offer agents in your network, Airbyte, ingestr, and the Bruin CLI can run entirely on your own infrastructure, and Bruin Cloud can run in your VPC. Read replicas, exports, and incremental loads avoid touching the primary database at all.

Connector coverage for your sources. Count the sources you have today and the ones you expect next year. Long-tail SaaS apps favor Fivetran or Airbyte, or a few lines of Python in a Bruin asset if you have engineers. Databases and a known set of SaaS tools are covered by almost everyone. Mainframes and ERPs still favor Informatica or Qlik.

CDC and freshness. Most analytics does not need sub-second data. Scheduled incremental loads every few minutes are simpler and cheaper. Use log-based CDC when you need hard deletes, have no reliable update timestamp, or genuinely need real-time data. Some tools, ingestr included, run the same CDC connector either on a schedule or as a continuous stream, so you can start with batch and switch later.

What happens after the load. If ingestion, transformation, testing, scheduling, and lineage live in five tools, a failure in one shows up as a wrong number in a dashboard three steps later. Suites (Informatica), platforms (Matillion, the Bruin CLI), and the merged Fivetran + dbt company all try to close that gap in different ways: GUI suites, visual pipelines, or code in Git.

Readiness for AI agents. AI agents that build or debug pipelines work best when the whole flow is readable text with lineage and tests they can run. Code-first tools (Glue, ingestr, the Bruin CLI, dbt) are easier for agents to work with than visual designers, and Matillion's Maia and Fivetran's AI features show the vendors know it.

Which Snowflake ETL tool should you pick?

  • You want zero maintenance and have the budget: Fivetran.
  • Your company runs on Azure or migrates from SSIS: Azure Data Factory, or Fabric Data Factory for new work.
  • You are a regulated enterprise with mainframe or ERP sources: Informatica or Qlik, and plan for the PowerCenter support timeline.
  • You are AWS-first and already write PySpark: AWS Glue, paired with DMS for CDC.
  • You want open-source connectors and can run infrastructure: Airbyte.
  • You want visual ELT that runs on Snowflake compute: Matillion.
  • You want nothing but Snowflake: COPY INTO, Snowpipe, and Openflow.
  • You want fast, reliable ingestion from built-in sources and already have a scheduler: ingestr on its own. Start with the replicate a database into Snowflake guide.
  • You need custom sources like internal APIs or niche SaaS tools: the Bruin CLI's Python materialization, which loads any DataFrame into Snowflake with merge, append, or incremental strategies.
  • You want ingestion, SQL and Python transformations, quality checks, and lineage in one Git-native pipeline, deployable in your own VPC: the open-source Bruin CLI, with Bruin Cloud for managed scheduling and observability.

FAQ

Fivetran. According to Crustdata's technographic data, charted by Greybeam, 12.7% of 52,771 Snowflake companies also run Fivetran, ahead of Azure Data Factory (7.9%), Informatica (6.5%), AWS Glue (2.7%), Airbyte (2.7%), and Matillion (1.5%).

What is the best ETL tool for Snowflake?

There is no single best tool; it depends on your cloud, sources, and team. Fivetran is the most hands-off for SaaS sources. Azure Data Factory and AWS Glue fit Azure-first and AWS-first companies. Informatica suits regulated enterprises with legacy sources. Airbyte and ingestr are code-first options you run yourself, and the Bruin CLI adds custom Python sources, transformations, quality checks, and lineage in one pipeline.

Fivetran vs Airbyte: which is better for Snowflake?

Fivetran is fully managed and priced on monthly active rows, so it costs little at low volume and more as tables grow or churn. Airbyte has a large connector catalog, is free to self-host under the Elastic License 2.0, and offers a usage-based cloud. Pick Fivetran to avoid running infrastructure, and Airbyte to avoid per-row pricing.

How do most companies load data into Snowflake?

There is no single dominant method. The six most common third-party tools cover at most about a third of Snowflake companies, since their shares add up to 34% before overlap. The rest use Snowflake's native loading (COPY INTO, Snowpipe, Snowpipe Streaming, Openflow), smaller vendors, open-source and code-first tools such as Airbyte, dlt, or ingestr, or custom scripts.

Fivetran grew up alongside Snowflake as the managed ELT half of the modern data stack. It loads raw data into Snowflake and leaves transformation to the warehouse, which matched Snowflake's design, and it was available through Snowflake Partner Connect. Its large catalog of maintained connectors means teams do not have to maintain syncs themselves. In June 2026 Fivetran completed its merger with dbt Labs.

Is Azure Data Factory good for loading data into Snowflake?

Yes, for Azure-first organizations. Snowflake runs on Azure, ADF has a native Snowflake V2 connector for copy activities and data flows, and it reaches on-prem sources through a self-hosted integration runtime. The trade-offs are Azure lock-in, visual pipelines that are harder to review as code, and a gradual shift of new Microsoft investment toward Fabric Data Factory.

What is the best open-source tool for loading data into Snowflake?

It depends on the job. Airbyte has the largest connector catalog and can be self-hosted. dlt suits teams that want ingestion written in Python. ingestr, Bruin's standalone ingestion CLI (source-available under the FSL), loads 130+ sources into Snowflake with a single command and supports incremental loads and log-based CDC. For custom sources, the open-source Bruin CLI's Python materialization loads any DataFrame into Snowflake.

Can I load data into Snowflake without a third-party ETL tool?

Yes. COPY INTO bulk-loads files from a stage, Snowpipe loads files automatically as they arrive in cloud storage, Snowpipe Streaming writes rows with low latency, and Snowflake Openflow provides managed connectors, including database CDC. You still need something to schedule jobs, transform data, and test the output.

What does Informatica's acquisition by Salesforce mean for Snowflake users?

Salesforce completed its acquisition of Informatica in November 2025, and standard support for PowerCenter 10.5 ended on March 31, 2026, with paid extended support for roughly one more year, into 2027. Snowflake teams still on PowerCenter need to choose between moving to Informatica's cloud platform (IDMC) and migrating to another tool.

Where does the 52,771 Snowflake companies number come from?

It comes from Crustdata, a Y Combinator (F24) company that provides company and people data through APIs. Crustdata detects the technologies a company uses from public web signals and job postings. Greybeam, a Snowflake cost-optimization company, used that data to chart which ETL tools Snowflake companies also run. The number is larger than Snowflake's reported customer count (14,554 as of July 2026) because it counts every company with a Snowflake signal, not billing accounts.

Sources

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.