Course overview/Governance as agent context3 of 5
Measure your context
Measure your context
Score an agent before and after context changes instead of treating documentation as an article of faith.
Measure the change, not the intention
Keep an untouched template for the baseline. Ask its instructor the eight supplied evaluation questions plus two verified questions you add. Score a fresh agent session. Then use a second fresh session with your descriptions, glossary, and rules, and score the same ten questions again.
Fresh sessions matter because a follow-up remembers the earlier reasoning. Exact scoring matters because a nearly-right answer often exposes a different unresolved assumption. The result may be flat or lower. That means the context did not address these failures, not that the exercise failed.
Your task
Add two query-verified questions to docs/eval/questions.md. Run both fresh sessions without exposing docs/eval/answers.md first, then write docs/eval/results.md with both scores, every question's result, timing, and the one context line you think caused the largest change.
Check your understanding
- Why must the second run be a fresh session rather than a follow-up in the same conversation?
- Your before score is 6 and your after score is 6. What have you learned?
- Why does the scoring allow no partial credit?
Do it with your agent
Say next lesson, keep both sessions separate, then say review my work. The review should grade the recorded evidence, not whether the score improved.