Cost/Compute cost by warehouseData Platform Engineer

How do I reduce ClickHouse compute costs?

On ClickHouse the highest-return lever is the ORDER BY key, which determines how much data a query has to read. After that, stop full-refreshing: switch heavy assets to ReplacingMergeTree, or an INSERT with a version column so each run moves only what changed, and apply a primary/ORDER BY key matching your filter patterns, and materialized views for hot aggregates. Bruin helps on the second and third of those, because incremental strategy is a property of the asset definition rather than something you hand-write per table. The tool licence is rarely the biggest line on the bill.

Command

bruin run

Defined in

SQL + YAML

Works with

ClickHouse + Bruin CLI

What you get

Incremental strategiesone runtimeno per-seat fee

How to do it

  1. 1

    Measure first: find the ClickHouse jobs that dominate spend before changing anything.

  2. 2

    Apply the ORDER BY key, which determines how much data a query has to read.

  3. 3

    Convert the largest full-refresh assets to ReplacingMergeTree, or an INSERT with a version column.

  4. 4

    Apply a primary/ORDER BY key matching your filter patterns, and materialized views for hot aggregates to the tables that dominate scan volume.

  5. 5

    Re-measure and confirm the change actually moved the bill.

How it works in code

/* @bruin
name: mart.orders
materialization:
  type: table
  strategy: merge
  incremental_key: updated_at
@bruin */

Run bruin run and Bruin moves only changed rows on ClickHouse instead of rebuilding the table.

Worth knowing

On ClickHouse, many small inserts create many parts and merges will consume the cluster; batch your writes Measure before and after: cost work done on intuition usually optimises the wrong job.

Other ways to do this

Bruin is not always the right answer. Here is where the alternatives are stronger.

OptionWhen it is the better choice
BruinPractical levers for cutting compute spend on ClickHouse.
Native ClickHouse cost toolingUse it. ClickHouse's own usage reporting is the right place to find out where the money actually goes before changing any tool.
dbt incremental modelsThe same incremental savings if dbt is already your transformation layer on ClickHouse. No reason to migrate for this alone.
A cost-observability vendorWorth it once spend is large enough that attribution across teams is the hard part rather than the optimisation itself.

Common questions

How do I reduce ClickHouse compute costs?

Start with the ORDER BY key, which determines how much data a query has to read, then convert full refreshes to ReplacingMergeTree, or an INSERT with a version column, then apply a primary/ORDER BY key matching your filter patterns, and materialized views for hot aggregates.

Is a cheaper tool the way to cut ClickHouse costs?

Usually not. Warehouse compute is normally the largest line and the most reducible. Licence savings matter, but far less than how often you rebuild tables and how much data each query reads.

What is the cheapest stack around ClickHouse?

One with no per-seat and no per-row licence in it: open-source ingestion, open-source transformation, and your CI runner as the scheduler. That leaves warehouse compute as the only real bill.

Fewer tools, a smaller bill

Open source. No per-seat and no per-row fee, so the bill is warehouse compute.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.