Opinion
8 min read

The Best Context Is No Context

What context to write for AI agents: reuse what code and logs already show, add only the meaning they cannot, and recycle conflicts, which are worse than gaps.

The Best Context Is No Context

The short version: most context written for AI agents is a copy of something the system already says. The copies drift, the agent finds two versions, and it picks one. So treat context like any other code and keep it small. Reuse what your code, configs, and logs already show, reduce what you add to the meaning they cannot show, and recycle anything that has gone stale or contradicts its source. The title sounds backwards, but the idea is simple: do not document what the system already states clearly.

This comes from the September Build Lab session on context for agents. It is also the rule that decides what you write in every phase of onboarding an agent like a new hire.


The instinct to write everything down

When an agent gets something wrong, the reflex is to add more context. Another paragraph in the README, a longer AGENTS.md, maybe a wiki page that explains the pipeline from scratch.

Some of that helps. Most of it restates facts the agent could already read, and every restated fact is a second copy that has to stay in sync with the first. Nobody keeps it in sync. Six months later the README says the pipeline runs every six hours, pipeline.yml says hourly, and the agent has to guess which one you meant.

The agent usually has plenty of text. What it is missing is meaning, and a clear answer to "which of these is true?"


Reuse: start with the context you already have

A lot of context is inherent. It exists because the system runs, and since it is what actually executes, it is the most accurate description you have.

Configs describe schedules, retries, dependencies, and where alerts go. Logs show what ran, what failed, and what the error said. Metadata exposes schemas, types, owners, tags, and lineage. The data itself shows which values actually appear and how tables relate. And the query shows how a table is built: which columns feed it, which joins and filters shape it, and where the business logic lives.

That holds whatever the stack is - Bruin, dbt, Airflow, BigQuery, or any mix of them. So before you write a sentence of context, ask whether a tool can already answer the question. If it can, point the agent at the tool, give it access, and keep the explanation out of the context layer.


Reduce: add only the missing meaning

Picture two circles. One is inherent context, the facts the system already exposes. The other is explicit context, which is everything a person wrote down. Where they overlap is waste: prose that repeats a fact the tools already show. Cut that down to the minimum viable information.

What does earn a line of explicit context:

  • Business definitions. "AOV" is a familiar acronym, but the agent still needs your definition of it: which orders count, and whether refunds, taxes, and shipping are excluded.
  • Decisions, like why de-duplication happens once in staging, so nobody adds a second pass downstream.
  • Ownership. A real person to escalate to means the agent asks instead of guessing.
  • Exceptions, such as "the spike on the first of the month is expected."
  • The why behind things. Anything a new engineer asks in their first week that the SQL cannot answer.

A gap prompts a question; an unclear definition produces a confident mistake. That difference is the reason explicit context is worth writing at all, and the reason it should be short.


Recycle: resolve conflicts and overlap

Recycle is the maintenance loop, and it runs again every time the code or the business changes:

  1. Identify where two sources overlap or disagree.
  2. Choose one source of truth for that fact.
  3. Remove the other copy, or rewrite it so it points at the source instead of competing with it.

This is not a one-time cleanup. Code changes, business rules change, and documentation goes stale without anyone noticing. Deleting a wrong page counts as a context improvement, and it is often a bigger one than writing a new page.


Why a conflict is worse than a gap

When context is missing, the agent hits a visible unknown. A well-instructed agent says so ("I could not find a definition of active customer"), you fill the gap, and you move on.

When two sources disagree, the agent has two plausible answers and nothing telling it which to trust. So it picks one and sounds sure about it, and the wrong answer looks justified because it came from your own documentation. Nobody notices until a number is wrong in a meeting.

Three rules keep this under control:

  • One fact, one home. Each fact has exactly one authoritative place, and everything else references it.
  • State the precedence. Where scopes overlap, narrower scope wins, and you write that rule down where the agent reads it. In Bruin, a column can extends a glossary term to inherit its type and description, and anything set explicitly on the asset overrides the glossary default.
  • Keep the open question visible. If a conflict cannot be resolved today, record it with an owner. Do not let the agent quietly pick a side.

This is one way to split the facts so each has a single home:

FactIts one home
Schedule, retries, alert channelspipeline config (pipeline.yml)
Why the pipeline exists, who consumes itpipeline README.md
What one row represents, owner, tierthe asset definition
How the table is builtthe query
Valid values, keys, invariantscolumn checks and custom checks
Shared terms like "order id"the glossary
Metric definitions like revenue or AOVthe semantic layer
Decisions, runbooks, project plansdocumentation, linked from the repo
Where each of the above lives, and what winsAGENTS.md and its context map

Shared meaning belongs in one layer

The facts behind most conflicts are business meanings: what "active customer", "net revenue", or "order id" mean. They show up in every pipeline, so they get redefined in every pipeline.

Give them one layer. In Bruin that is two files at the repository level. A glossary defines shared domains, entities, and attributes, so columns extend a definition instead of repeating it (the glossary is marked beta in the docs). A semantic layer defines reusable metrics and dimensions in YAML under semantic/:

# semantic/orders.yml
schema: v1
name: orders
label: Orders
description: One-time storefront order metrics.

source:
  table: shop_report.orders_enriched

dimensions:
  - name: order_date
    type: time
    granularities:
      month: date_trunc('month', order_date)
  - name: country
    type: string

metrics:
  - name: revenue_usd
    expression: sum(amount_usd)
  - name: order_count
    expression: count(distinct order_id)
  - name: average_order_value_usd
    expression: "{revenue_usd} / {order_count}"

If you want agents to build dashboards and charts, add a semantic layer. The agent then reuses revenue_usd instead of rebuilding revenue from raw columns on every question, slightly differently each time.


Fix the model before adding context

Context can explain the gaps around a sound system. It cannot repair a system nobody understands. If a senior data person cannot navigate your models, read the code, and understand the business, an agent cannot either. It will still write fluent SQL, but fluency is not understanding.

So before writing more context, check the model itself:

  • Add a semantic layer, so metrics, dimensions, entities, and relationships each have one reusable meaning.
  • Model with purpose. That means a stated grain, tiers that fit the team, readable SQL, and no logic duplicated across layers.
  • Get rid of stale data. Audit and delete unused tables, reports, and dashboards rather than making the agent search through them.
  • Delete outdated pages. A wrong page is worse than no page.

The evidence for less

The best published number I know of points the same way. Anthropic's data team reported that without skills, Claude's accuracy on their analytics evals "didn't exceed 21%", and adding skills got it "consistently above 95% in aggregate". Giving the agent grep access to thousands of historical dashboard, transformation, and notebook SQL files moved accuracy "by less than a point in either direction".

It is one internal evaluation of a broader system, and I would not treat it as a benchmark for your stack. Still, the signal is clear enough. More raw SQL barely moved accuracy, and the curated, written-down knowledge did the work. So I would aim for the least context that closes the real gaps, with every fact in one place.


FAQ

What does "the best context is no context" mean?

Do not write down what the system already states clearly. A table schema, a pipeline schedule, a query log, or the SQL itself is already context an agent can read. Extra prose only helps when it adds meaning the system cannot show on its own: business definitions, decisions, ownership, exceptions, and why a rule exists.

Why is conflicting context worse than missing context for an AI agent?

A missing definition leaves a visible gap the agent can report or ask about. Two conflicting definitions give it two plausible answers, so it picks one and sounds confident, and the wrong answer looks justified because it came from your own documentation. Give each fact one authoritative home and state a precedence rule where scopes overlap.

What context should I write for an AI data agent?

Only what closes a real gap: what a metric like AOV means in your business, which orders count as revenue, who owns a table, why a pipeline is shaped the way it is, and the exceptions people know but never wrote down. Leave schedules, column types, row counts, and dependencies to the configs and schemas that already hold them.

Where should each piece of context live?

One fact, one home. Run behaviour in the pipeline config, the reason a pipeline exists in its README, table grain and ownership in the asset definition, shared terms in a glossary, metric definitions in a semantic layer, valid values in checks, and decisions or runbooks in documentation that the repo links to. Everything else should point at that home rather than restate it.

Do I need a semantic layer for AI agents?

If you want agents to build dashboards and charts or answer metric questions, yes. A semantic layer gives the agent a reusable, reviewed definition of each metric and dimension, so it reuses revenue instead of reconstructing it from raw columns on every question, slightly differently each time.


Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.