How do I check a schema diff before merging using GitHub Actions?
Add a job to .github/workflows/data.yml that installs the Bruin CLI and runs bruin data-diff. It compares the shape of the output against the current table so an unintended schema change is visible in review. Store warehouse credentials as GitHub Actions secrets, and use a concurrency group so two deploys can never write the same tables at once. Keep validation and deployment as separate jobs: validation should be fast and credential-free, deployment serialised and gated.
Command
bruin data-diffDefined in
YAML + CI
Works with
GitHub Actions + Bruin CLI
What you get
How to do it
- 1
Create or open .github/workflows/data.yml in your repository.
- 2
Install the Bruin CLI as a step: curl -LsSf https://getbruin.com/install/cli | sh.
- 3
Add bruin data-diff as the job's command.
- 4
Store warehouse credentials as GitHub Actions secrets, never in the repository.
- 5
Add a concurrency group so concurrent runs cannot write the same tables.
- 6
Make the job required for merge, so a failure actually blocks.
How it works in code
# .github/workflows/data.yml
- run: curl -LsSf https://getbruin.com/install/cli | sh
- run: bruin data-diffRun bruin data-diff and GitHub Actions fails the build when the pipeline does, before the change reaches production.
Worth knowing
A job that reports failures without failing the build gets ignored within weeks. Make it required. And give CI its own warehouse connection pointed at a scratch schema, because a pull request can come from anywhere and should never hold production credentials.
Other ways to do this
Bruin is not always the right answer. Here is where the alternatives are stronger.
| Option | When it is the better choice |
|---|---|
| Bruin | Wire bruin data-diff into GitHub Actions so pipeline changes get checked the same way application code does. |
| dbt in CI | The same pattern with dbt build and a state-based selector. Well documented, and the right choice if dbt is already your transformation layer. |
| A managed orchestrator | Tools like Dagster or Prefect Cloud handle scheduling and retries more richly than a CI runner, at the cost of another system to run. |
| Soda in CI | Soda's data-contract checks are the strongest documented option if formal contracts between teams are the main goal. |
Common questions
How do I run data pipelines in GitHub Actions?
Install the Bruin CLI in a job and run bruin data-diff. Validation needs no warehouse credentials; deployment reads them from GitHub Actions secrets.
How do I stop two GitHub Actions deploys running at once?
Use a concurrency group. Two pipeline runs writing the same tables concurrently is the hardest data bug to diagnose, and it is entirely preventable.
Should validation and deployment be separate GitHub Actions jobs?
Yes. Validation should run on every pull request, take seconds, and need no credentials. Deployment should run on merge, be serialised, and be gated. Combining them makes validation slow and deployment unsafe.
Related use cases
Check a schema diff before merging with GitLab CI
How do I check a schema diff before merging using GitLab CI?
Schema diff in CICheck a schema diff before merging with Bitbucket Pipelines
How do I check a schema diff before merging using Bitbucket Pipelines?
Schema diff in CICheck a schema diff before merging with Azure Pipelines
How do I check a schema diff before merging using Azure Pipelines?
Deploy data pipelines like software
Open source. Validate on every pull request, deploy on merge, roll back with git.