AI SKILL

Agent Evaluation Rubric Builder

Turns product goals and task descriptions into reviewable dimensions, rubrics, and failure types.

INPUT

User task, expected outcome, constraints, and risk boundaries

OUTPUT

Dimensions, scoring anchors, examples, and a human-review checklist

01 / USE CASE

Best fit

Agent acceptance, regression testing, and product review

02 / WORKFLOW

Workflow

  1. 01

    Identify task goals and non-negotiable constraints

  2. 02

    Separate outcome, process, and risk quality

  3. 03

    Write observable anchors for each level

  4. 04

    Generate counterexamples for product review

03 / EXAMPLE

Sample input & output

SAMPLE INPUT

Create constraint-adherence and fairness rubrics for a group restaurant agent.

SAMPLE OUTPUT

Treat budget, distance, and opening status as hard constraints; explain commute disparity and option diversity as soft measures.

04 / LIMITS

Limitations

  • Does not replace domain experts in high-risk settings
  • Rubrics require calibration with real failures
  • Scores should not be compared across unrelated tasks