Human Feedback & Evaluation

Rubrics and Structured Evaluation

Rubrics and Structured Evaluation

A rubric is a scoring guide that defines what each quality level looks like for specific criteria. Rubrics transform subjective judgment into structured, reproducible evaluation. Most AIDASH tasks embed rubrics — explicitly in instructions or implicitly in the rating scales you use. Mastering rubric-based evaluation is core professional skill for AI evaluators.

Anatomy of a Rubric

A typical evaluation rubric includes:

  • Criteria: The dimensions being assessed (accuracy, relevance, safety, etc.)
  • Scale: Rating levels (1-5, pass/fail, binary choice, ranking)
  • Descriptors: What each level means for each criterion
  • Examples: Sample responses illustrating each level (when provided)

Example accuracy descriptor:

  • 5: All facts verified correct. No errors found.
  • 3: Mostly accurate with minor errors that do not change the core answer.
  • 1: Major factual errors that undermine the response's usefulness.

Why Rubrics Matter

Without rubrics, evaluators apply personal, inconsistent standards. One evaluator's "good" is another's "mediocre." Rubrics:

  • Align evaluators to shared standards
  • Make ratings auditable and defensible
  • Enable meaningful aggregation across many evaluators
  • Help identify which specific dimensions need improvement

Applying Rubrics Fairly

Rate each criterion independently. A response can be highly accurate but poorly written. Do not let one strong dimension inflate scores on others.

Use the full scale. If every response gets a 4, the rubric provides no discrimination. Reserve top scores for genuinely excellent work and low scores for clear failures.

Anchor to descriptors, not feelings. Check your rating against the written definition. If you rated accuracy a 5, can you defend that every fact is correct?

Note borderline cases. When a response falls between levels, pick the closest match and explain why in your comments.

Writing Rubric-Based Feedback

Strong rubric evaluations include:

  1. Rating for each criterion
  2. Evidence supporting each rating (quoted from the response)
  3. Specific improvement suggestions for failures
  4. Overall assessment synthesizing the criteria

"Weak" feedback: "Accuracy: 2. Not great."

"Strong" feedback: "Accuracy: 2. The response states that World War II ended in 1944 (line 3) and attributes the theory of relativity to Newton (line 7). Two major factual errors that would mislead a student using this as a study guide."

Calibration Tasks

Platforms often include hidden calibration items — responses with known correct ratings — to measure evaluator accuracy. Treat every task as if it could be calibration. Consistent, rubric-faithful work improves your quality score and unlocks access to higher-paying assignments over time.

Key Takeaways

  • Rubrics define criteria, scales, and descriptors for consistent evaluation
  • Rate each criterion independently using the full scale
  • Anchor ratings to written descriptors with quoted evidence
  • Specific, rubric-referenced feedback is far more valuable than vague scores