Definition
Plain language
How often two people scoring the same material give the same answer.
As stated in the literature
Inter-rater reliability, typically reported as Cohen's or Fleiss's kappa; low values indicate raters are effectively answering different questions and the metric is underspecified.
Also called: inter-rater agreement, inter-annotator agreement
Why it matters: Without it you can't tell whether a score reflects the system's behavior or just the opinions of whoever happened to do the scoring.
For example, if two reviewers each rate the same 100 chatbot replies as 'helpful' or 'not helpful' and disagree on a third of them, the agreement score will be low and the ratings can't be trusted as a single measure.