Skip to main content
Evaluate Search against the intents your product actually receives. A good test set includes straightforward requests, ambiguous requests, hard constraints, unavailable supply, and cases that should return no acceptable option.

Build an evaluation set

For every test case, record:
  • the user intent and relevant context
  • required and prohibited capability attributes
  • acceptable capability IDs or a clear human relevance rubric
  • whether an executable result should exist
  • the capability revision used when the judgment was made
Keep evaluation labels separate from production ranking. Review them when capabilities or requirements change.

Useful measures

Review failures

Classify failures as query ambiguity, missing capability metadata, incorrect filtering, stale availability, or ranking quality. Do not replace these labels with a fabricated confidence score. Run the same set before changing query construction, filters, or result presentation. For end-to-end quality, continue the selected result through an Act evaluation.