Quick answer: onboarding an AI data agent works a lot like hiring a data person. You build its context in four phases - orientation, onboard, internship, and hire - and it gets more access after each one. Every phase ends with a test you pass before moving on: a reviewed first draft, tested permissions, about 90% correct answers on questions you already know, and finally a feedback path that keeps the context current once real teams are using it.
This is the written version of the September Build Lab session on context for agents. It is tool-agnostic. I use Bruin for the examples because it is what I work in every day, and the same steps apply to dbt, Airflow, or whatever you run.
Why a hiring plan, not a documentation project
The usual advice is "write more docs", and it rarely works. Nobody can tell when they are done or which docs actually matter. And a blank page in front of a busy engineer stays blank.
A hiring plan gets around both problems, because you already know how to bring a new data person up to speed. You give them an overview, let them shadow someone, put them on real tasks with a reviewer, and only then let them work on their own. You would never hand a first-day hire write access to production and a link to a wiki.
An agent needs the same staged path for the same reason: it can make a production change long before it understands the system. Thinking of it as a hire also changes the question. "Is the documentation complete?" has no real answer. "Can the agent find, read, and use the context to do this job?" does, and you can test it.
| Phase | What it is | Access | Done when |
|---|---|---|---|
| 1 · Orientation | basic overview, simple questions | read the repo | a reviewed first draft, unknowns written down |
| 2 · Onboard | shadow real work, set the boundary | read the repo and logs, tools connected | tools and permissions tested and holding |
| 3 · Internship | real tasks under review | local and dev, PRs only | ~90% of known questions correct |
| 4 · Hire | works alone with a review path | team channels, scoped by use case | owners, audits, and a feedback loop in place |
One prerequisite comes before any of this. The agent needs a data model it can navigate, and if a senior data person cannot navigate your models, an agent cannot either. I cover that, along with the rule for what context to write at all, in The Best Context Is No Context.
Phase 1: Orientation
The first phase gives the agent a basic overview of the data, then checks whether it actually understood it.
Skip the blank page. Generate a first pass from what already exists and edit that instead. In Bruin it takes three commands, and every framework has an equivalent:
# Scaffold a project from a template, so you edit a real structure
bruin init empty my-pipeline
cd bruin
# Create metadata-only asset files from the tables in one schema
bruin import database --connection my-warehouse --schema analytics my-pipeline
# Draft descriptions, quality checks, and tags from real table statistics
bruin ai enhance my-pipeline
bruin import database reads the schema and writes an asset file per table, with column metadata filled in. bruin ai enhance then fetches real statistics (row counts, null counts, distinct values, min and max) and adds descriptions, checks, and tags. It needs a configured AI provider, validates what it wrote, and restores the original file if validation fails. The full walkthrough is in How to Build an AI Context Layer for Your Data Warehouse.
Be honest about what you get from this. It is a schema-informed starting point, and it is not business context yet. The import knows a column is called status and holds five distinct values. It has no idea that cancelled never counts as revenue.
Then ask the agent the kind of questions you would ask a new hire at the end of day one:
- What is this project for, and how do the pipelines relate?
- Why does this pipeline exist, and who uses its output?
- What does one row of this table represent?
- Who owns this table, and who do I ask when it looks wrong?
Review what comes back. When it guesses, write the guess down as an unknown instead of letting it stand.
Done when: you have a reviewed first draft of the context for one pipeline, and a short, explicit list of what is still unknown.
Phase 2: Onboard
Onboarding has two halves. The agent shadows real work, and you set the boundary it works inside.
Shadow real work
Give the agent the orientation context and watch how it explores. The questions it asks will tell you more than the answers it gives.
- Ask it to read the context and the data and come back with questions. Good questions show you where the context is still thin.
- Point it at the query logs, have it find the failing queries, and ask it to explain why each one failed.
- Repeat a task you already did (a small model change, a bug fix) and compare its pull request with yours.
- Ask it to explore the pipelines and name the most important tables. Check the reasoning, not just the list.
Do this one pipeline or one business area at a time, and keep a running list of the questions the agent should be able to answer. That list becomes your test set in the internship.
Set the operating boundary
All of this needs to be in place before the agent touches real work.
Start with instructions. A repo-level AGENTS.md states the conventions, what the agent may do on its own, and when it has to stop and ask. Add a context map to it: where each layer lives, and which source wins when two of them disagree.
Then connect the tools it needs - the CLI, MCP servers, APIs, and skills. In Bruin, bruin ai skills all installs the bundled skills and writes an AGENTS.md you can start from.
Access is defined by role and scope. One agent may need the full lineage; another only needs gold-tier reports for one business unit. Decide which agents can read, write, open a pull request, or only suggest a change.
Finally, test it. Check that the tool calls work and that the permissions hold. The easiest way is to ask it to do something it should not be able to do, and confirm that it cannot.
Here is a short excerpt of what that policy can look like in AGENTS.md:
## What you may do
- Read any asset, run log, and lineage in this repo.
- Run read-only queries with `bruin query`. Aggregate before sharing results.
- Run assets against the `dev` environment only.
## What needs a human
- Any change goes through a pull request. Never push to `main`.
- Never change a metric definition without the owner listed in the asset.
- If two sources disagree, stop and report both. Do not pick one.
Done when: the instructions are in the repo, the tools are connected, and you have tested that every permission holds.
Phase 3: Internship
Now the agent works on real problems. Access stays local and controlled, and you keep correcting it until it is reliably right.
The internship is a loop:
- Have the agent create or update an asset, or ask it a real question that someone on the team actually asked.
- Check the answer.
- If it is wrong, correct the context. Find the layer that misled it (or the one that was missing) and fix that, not just the answer.
- If it is right, enhance the context. Write down the business rule it had to infer, so next time the answer does not depend on luck.
- Repeat until it gets about 90% of your known questions right.
Step three is the one people skip. A plausible patch to one query closes today's bug. A corrected glossary entry or column description closes it for every agent and every person who reads that table next. So make it an instruction: when the agent is corrected, it updates the relevant context in the same task and the same pull request.
The known-question list from onboarding is your test set. Keep it in the repo next to the pipeline. The format is up to you. This is not a Bruin feature, it is just a file:
# evals/revenue.yml - questions with answers the team has already verified
- question: "What was net revenue for the first week of September?"
expect: "uses mart_revenue_daily, excludes refunds and cancelled orders"
- question: "Which table should I use for one row per order?"
expect: "stg_orders"
- question: "Who owns the MRR metric?"
expect: "Growth"
Raw correctness is only part of it. Watch whether the agent can move between layers - from a dashboard question to the semantic model to the asset to the owner - and whether the layers agree once it gets there. And watch whether it knows what it does not know. An agent that says "the definition of active customer is missing" is far more useful than one that picks a definition and sounds sure about it.
For the reviewing, pick a few superusers or domain experts. They correct answers and improve the source of truth as they go, which is work you want done anyway.
The 90% target is a practical exit check before wider rollout. It is not a claim that the agent is perfect, and it only means something if the question set reflects what people actually ask.
Done when: about 90% of the known questions come back correct, and every correction along the way left a change in the repo.
Phase 4: Hire
At this point the agent works on its own, with a review path behind it. Being hired means access plus accountability.
Put it where teams already work - Slack, Microsoft Teams, or another application - so questions and corrections happen in the open. Roll it out to one team and one use case at a time, and tell people what the agent can do, what it cannot, and how to report a mistake.
Then audit it. Sample the prompts, the context it used, the queries it ran, and the answers it gave; that sample is its performance review. Feedback that turns out to be useful becomes a tracked change, with an owner, an approval path, and a record of what changed.
This is where maintenance stops being a separate project and becomes part of the job. Code changes, business rules change, and a correction made in a Slack thread scrolls away unless something writes it back. The loops that keep context current (local review, CI, scheduled audits, and channel feedback) have their own post: How to Keep an AI Context Layer From Going Stale.
Done when: each use case has a named owner, a sampled audit on a schedule, and a way for feedback to become a reviewed change.
Why this is worth the effort
Written-down, curated context moves an agent's accuracy far more than piling on raw data. Anthropic's data team reported that without skills, Claude's accuracy on their analytics evals "didn't exceed 21%", and that adding skills got it "consistently above 95% in aggregate". In the same post, giving the agent grep access to thousands of historical SQL files moved accuracy "by less than a point in either direction".
That is one internal evaluation of a broader system, so I would read it as a direction rather than a benchmark for your stack. It is still the direction this plan bets on. Pointing the agent at everything looks like a shortcut and mostly is not. The curated context does the work, and the four phases are a way to produce it without starting from a blank page.
Common mistakes
- Skipping to hire. If you connect the agent to Slack on day one, before the internship, your users end up running the internship for you. They stop trusting it after the second wrong number.
- Boiling the warehouse. Onboarding every pipeline at once spreads review across too much context. Do one pipeline or one business area, then the next.
- Fixing the answer instead of the context. Any correction that does not change a file in the repo will come back.
- Testing on toy questions. Use questions people actually asked, with answers someone verified.
- Untested permissions. A boundary you did not try to cross is a boundary you hope holds.
Where to start
This week, map one pipeline. Run the orientation commands, review the first draft, write down five questions you already know the answers to, and see how many the agent gets right. Whatever that number is, it is your baseline, and the questions it missed are your to-do list.
If you want the hands-on version, the Build an AI Context Layer guide covers orientation step by step, and the Cloud AI Agent guide covers connecting an agent to Slack, browser chat, and dashboards.
FAQ
How do I onboard an AI agent to my data stack?
Treat it like hiring a data person, in four phases. Orientation: generate a first draft of the context and test basic understanding. Onboard: let the agent shadow real work, then set its instructions, tools, and permissions. Internship: give it real tasks under review and correct the context every time it is wrong, until it answers about 90% of a known question set correctly. Hire: connect it to one team and one use case at a time, with audits and a feedback path that turns corrections into tracked changes.
Where should I start building context for an AI data agent?
Pick one pipeline or one business area and leave the rest of the warehouse for later. Generate a first pass from what already exists - in Bruin, bruin import database creates asset files from your tables and bruin ai enhance drafts descriptions, checks, and tags from real table statistics - then review it and ask the agent a few basic questions about the project, the pipeline's purpose, table grain, and owners.
How do I know when an AI agent is ready for production use?
Look for evidence. Before wider rollout you want a reviewed first draft of the context, a written list of known unknowns, tool calls and permissions that were tested and held, about 90% correct answers on a set of questions you already know the answers to, and a feedback path with an owner. The 90% figure is an exit check for the internship phase. It is not a claim that the agent is perfect.
What permissions should an AI data agent have?
Broad read access and narrow, named write access. Decide per agent whether it may read, write, open a pull request, or only suggest a change, and scope its data by role - one agent may need the full lineage, another only the gold-tier reports for one business unit. Test those boundaries during onboarding, before anyone depends on them.
Does this onboarding plan only work with Bruin?
No. The phases are tool-agnostic. I use Bruin as the example because it keeps the context, the code, and the checks in one repository, but the same plan works with dbt, Airflow, and a warehouse, as long as the agent can read the context and reach your systems through a CLI, an API, or an MCP server.
Related Reading
- The Best Context Is No Context - what to write down for an agent, what to leave out, and why a conflict is worse than a gap.
- How to Keep an AI Context Layer From Going Stale - the feedback loops that keep context current after the agent is hired.
- How to Build an AI Context Layer for Your Data Warehouse - the orientation phase, command by command.
- What Is a Self-Healing Data Pipeline? - what a well-onboarded agent can do when a pipeline breaks.
- Is It Safe to Give AI Access to Your Data? - the security questions behind the access decisions in phase two.
