Technical
6 min read

What Is an MCP Server for Data Tools? How AI Agents Query Pipelines and Warehouses

An MCP server is a program that exposes a tool's capabilities, running a query, reading lineage, validating a pipeline, to any AI agent that speaks the Model Context Protocol, so Claude Code, Cursor, or ChatGPT can operate the tool directly. This explainer covers what an MCP server does for data work, what a good one exposes, and how Bruin's compares with the MCP servers from dbt, Snowflake, Databricks, and BigQuery.

What Is an MCP Server for Data Tools? How AI Agents Query Pipelines and Warehouses

An MCP server is a program that exposes a tool's capabilities to AI agents through the Model Context Protocol, an open standard, so that an agent such as Claude Code, Cursor, Codex, or ChatGPT can call the tool's functions directly rather than reconstructing them from documentation. For a data tool that means an agent can run a query against the warehouse, list the assets in a pipeline and what depends on what, read column-level lineage, validate a change, and run the checks, using the same commands a human engineer would. The server is the difference between an agent that writes plausible SQL and an agent that knows your schema, your dependencies, and whether its change broke anything. Bruin ships an MCP server inside the open-source CLI; dbt, Snowflake, Databricks, and Google BigQuery ship their own for their slice of the stack.

What the protocol does

MCP standardises three things: how a client discovers what tools a server offers, how it calls one with typed arguments, and how results come back. A server can also expose resources (documents the agent can read) and prompts (templates for common tasks). The consequence for data work is that the agent's knowledge stops being frozen at training time. It asks the server what tables exist now, what a column means now, and what depends on the table it is about to change.

Without a server, an agent editing a pipeline works from the files in the repository and its memory of SQL dialects. With one, it can:

Without MCPWith a data MCP server
Infers the schema from model filesQueries the warehouse for the actual columns and types
Guesses which tables depend on the one it is editingReads the dependency graph and column-level lineage
Writes SQL and asks you to run itRuns the query against a named connection and reads the result
Cannot tell whether the change is validRuns validation and the checks, and reads the failures
Reads documentation it was trained onReads the tool's current documentation through the server

What Bruin's server exposes

bruin mcp starts the server from the CLI. Connecting Claude Code is one command:

claude mcp add bruin -- bruin mcp

The agent then has, inside the repository it is editing:

  • Query. Run SQL against any connection the project defines, without the agent holding the credentials itself. bruin query --connection warehouse --query "select ..." is the underlying command.
  • Assets and dependencies. List the pipeline's assets, their types, and what each depends on.
  • Lineage. Column-level lineage for an asset, the same output as bruin lineage <asset> --full --output json, so the agent can see what a column change reaches.
  • Validate and run. bruin validate for the project, bruin run for a single asset or its checks, with the failures returned as text the agent can act on.
  • Documentation. Bruin's own docs, so the agent reads the current syntax for checks, materialisation, and connections rather than a remembered version.

The reason the server covers the whole pipeline rather than just SQL execution is that the useful agent workflows need all of it: read the failing check, read the lineage to find the cause, query the source to confirm, propose the fix, validate, run the checks, open a pull request. That loop is what agentic data engineering means in practice, and the walk-through is in how to build data pipelines with AI agents.

How the other servers differ

Each vendor's server covers its own layer. dbt's MCP exposes models, the semantic layer, and metric queries, so an agent can ask "what is revenue" and get the governed definition, but not the ingestion or the orchestration around it. Snowflake's exposes Cortex agents and SQL against Snowflake. Databricks' exposes Unity Catalog and Genie spaces. Google's MCP Toolbox for Databases exposes BigQuery, AlloyDB, and Cloud SQL queries. An agent working across a stack made of those tools connects to several servers and reasons across them. An agent working on a Bruin project connects to one, because the loads, models, checks, and lineage are one project. Which is right depends on how many tools the stack already has, and the tool-by-tool list is in which data tools have a built-in MCP server in 2026.

What a good server should not do

Hand the agent a raw database credential with no structure around it. An MCP server that only runs arbitrary SQL gives the agent power without context: it can drop a table but cannot tell what depended on it. The safer shape is the one above, where queries go through named connections the project controls, changes are validated before they run, and lineage is one call away, so the agent can check impact before it acts.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.