User task, expected outcome, constraints, and risk boundaries
AI SKILL
Agent Evaluation Rubric Builder
Turns product goals and task descriptions into reviewable dimensions, rubrics, and failure types.
Dimensions, scoring anchors, examples, and a human-review checklist
Best fit
Agent acceptance, regression testing, and product review
02 / WORKFLOW
Workflow
- 01
Identify task goals and non-negotiable constraints
- 02
Separate outcome, process, and risk quality
- 03
Write observable anchors for each level
- 04
Generate counterexamples for product review
03 / EXAMPLE
Sample input & output
Create constraint-adherence and fairness rubrics for a group restaurant agent.
Treat budget, distance, and opening status as hard constraints; explain commute disparity and option diversity as soft measures.
Limitations
- Does not replace domain experts in high-risk settings
- Rubrics require calibration with real failures
- Scores should not be compared across unrelated tasks

