Build an evaluation set
For every test case, record:- the user intent and relevant context
- required and prohibited capability attributes
- acceptable capability IDs or a clear human relevance rubric
- whether an executable result should exist
- the capability revision used when the judgment was made