Glossary · Term

annotator agreement

← all terms

Definition

Plain language

How often two people scoring the same material give the same answer.

As stated in the literature

Inter-rater reliability, typically reported as Cohen's or Fleiss's kappa; low values indicate raters are effectively answering different questions and the metric is underspecified.

Also called: inter-rater agreement, inter-annotator agreement

Why it matters: Without it you can't tell whether a score reflects the system's behavior or just the opinions of whoever happened to do the scoring.

For example, if two reviewers each rate the same 100 chatbot replies as 'helpful' or 'not helpful' and disagree on a third of them, the agreement score will be low and the ratings can't be trusted as a single measure.

Related terms