How do I reduce Databricks compute costs?
On Databricks the highest-return lever is job clusters over all-purpose clusters, and Photon where it actually helps. After that, stop full-refreshing: switch heavy assets to MERGE INTO on a Delta table so each run moves only what changed, and apply OPTIMIZE and Z-ORDER on the columns you filter on. Bruin helps on the second and third of those, because incremental strategy is a property of the asset definition rather than something you hand-write per table. The tool licence is rarely the biggest line on the bill.
Command
bruin runDefined in
SQL + YAML
Works with
Databricks + Bruin CLI
What you get
How to do it
- 1
Measure first: find the Databricks jobs that dominate spend before changing anything.
- 2
Apply job clusters over all-purpose clusters, and Photon where it actually helps.
- 3
Convert the largest full-refresh assets to MERGE INTO on a Delta table.
- 4
Apply OPTIMIZE and Z-ORDER on the columns you filter on to the tables that dominate scan volume.
- 5
Re-measure and confirm the change actually moved the bill.
How it works in code
/* @bruin
name: mart.orders
materialization:
type: table
strategy: merge
incremental_key: updated_at
@bruin */Run bruin run and Bruin moves only changed rows on Databricks instead of rebuilding the table.
Worth knowing
On Databricks, an all-purpose cluster left warm for a nightly job costs far more than a job cluster that starts and stops Measure before and after: cost work done on intuition usually optimises the wrong job.
Other ways to do this
Bruin is not always the right answer. Here is where the alternatives are stronger.
| Option | When it is the better choice |
|---|---|
| Bruin | Practical levers for cutting compute spend on Databricks. |
| Native Databricks cost tooling | Use it. Databricks's own usage reporting is the right place to find out where the money actually goes before changing any tool. |
| dbt incremental models | The same incremental savings if dbt is already your transformation layer on Databricks. No reason to migrate for this alone. |
| A cost-observability vendor | Worth it once spend is large enough that attribution across teams is the hard part rather than the optimisation itself. |
Common questions
How do I reduce Databricks compute costs?
Start with job clusters over all-purpose clusters, and Photon where it actually helps, then convert full refreshes to MERGE INTO on a Delta table, then apply OPTIMIZE and Z-ORDER on the columns you filter on.
Is a cheaper tool the way to cut Databricks costs?
Usually not. Warehouse compute is normally the largest line and the most reducible. Licence savings matter, but far less than how often you rebuild tables and how much data each query reads.
What is the cheapest stack around Databricks?
One with no per-seat and no per-row licence in it: open-source ingestion, open-source transformation, and your CI runner as the scheduler. That leaves warehouse compute as the only real bill.
Fewer tools, a smaller bill
Open source. No per-seat and no per-row fee, so the bill is warehouse compute.