Course overview/Governance as agent context3 of 5

Measure your context

Measure your context

Score an agent before and after context changes instead of treating documentation as an article of faith.

Measure the change, not the intention

Keep an untouched template for the baseline. Ask its instructor the eight supplied evaluation questions plus two verified questions you add. Score a fresh agent session. Then use a second fresh session with your descriptions, glossary, and rules, and score the same ten questions again.

Fresh sessions matter because a follow-up remembers the earlier reasoning. Exact scoring matters because a nearly-right answer often exposes a different unresolved assumption. The result may be flat or lower. That means the context did not address these failures, not that the exercise failed.

Your task

Add two query-verified questions to docs/eval/questions.md. Run both fresh sessions without exposing docs/eval/answers.md first, then write docs/eval/results.md with both scores, every question's result, timing, and the one context line you think caused the largest change.

Check your understanding

  • Why must the second run be a fresh session rather than a follow-up in the same conversation?
  • Your before score is 6 and your after score is 6. What have you learned?
  • Why does the scoring allow no partial credit?

Do it with your agent

Say next lesson, keep both sessions separate, then say review my work. The review should grade the recorded evidence, not whether the score improved.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.